Technology
Reliability of LLMs in Safety-Critical Requirements Engineering
Master's-thesis research measuring how LLM assistance affects correctness, efficiency and trust in safety-critical requirements engineering.
~15% · Efficiency gain
The challenge
Safety-critical systems live or die by their requirements: precise, unambiguous specifications that all downstream engineering depends on. As Large Language Models started entering engineering workflows, Fraunhofer IKS needed an evidence-based answer to a high-stakes question - do LLMs genuinely help engineers write better requirements, or do they introduce new failure modes such as over-reliance, subtle errors and misplaced trust?
Most published LLM evaluations stop at benchmark scores. They say little about how human engineers actually behave when a model sits in the loop, or how that behaviour shifts with the engineer's own experience. Answering that called for a controlled human study, not another leaderboard.
The approach
For my Master's thesis at Fraunhofer IKS I designed and ran an end-to-end study on the reliability of LLMs in safety-critical requirements engineering, carrying it from study design and ethics approval through to final data analysis.
I built and deployed a custom experimental platform in Streamlit that put participants through requirements-engineering tasks under controlled conditions, comparing LLM-assisted workflows against human-only ones. The platform captured each interaction so I could measure the LLM's impact on correctness, efficiency, user trust and learning rather than relying on self-reports.
Analysing the results showed a roughly 15% efficiency gain for experienced engineers working with LLM assistance - but the same assistance produced a clear over-reliance risk that grew as user expertise fell, with less-experienced participants more likely to accept flawed model output uncritically. The finding: the benefit of LLMs in safety-critical work is real, but conditional on user expertise.
The outcome
Results that moved the needle.
- ~15%
-
Efficiency gain
For experienced engineers using LLM assistance in requirements tasks
- Over-reliance
-
Key risk quantified
Risk rose as user expertise fell, strongest among less-experienced engineers
- 4
-
Dimensions measured
Correctness, efficiency, user trust and learning under controlled conditions
- End-to-end
-
Study delivered
From study design and ethics approval through to data analysis
Want similar results for your project?
Every project above started with a conversation. Let's figure out what yours needs.
Keep exploring
More projects.
Rail catenary pole placement, automated
Days of expert engineering work, reduced to seconds.
D-Risk - MedTech Marketplace with AI Company Profiling
A three-sided marketplace linking medtech startups with investors and specialist freelancers - matched through AI document profiling.
Integrated LoRa Sensor Monitoring & Analytics System
Real-time monitoring and analytics for LoRa environmental sensors - from raw ingestion to geospatial drought and growth insight.
Revenue Administration MCP Server
A retrieval-only MCP server answering Turkish tax and regulation questions strictly from official government sources - zero hallucinations.
DSGENAI - AI Safety Requirements Engineering Platform
Stabilising and modernising an AI-driven safety-requirements platform - Flask to Streamlit, GPT-5.1, and critical data-leak fixes.
Retrieval-Augmented Generation (RAG) Documentation Assistant
A Gemini-powered RAG assistant that turns PDF libraries into an accurate, source-grounded knowledge base.
LLM Prompt Optimization for Legal-Clause Classification
Prompt-engineering research that lifted legal-clause classification accuracy by up to 20% on Terms-of-Service documents.
Smart Contract Analysis with NLP
An NLP system that reviews business contracts and recommends changes to keep them compliant with company policy.
Football Player Potential Prediction Model
A machine-learning model that reads scouting attributes and predicts player talent level with roughly 85% accuracy.
Football Player Ranking System
An attribute-based scoring engine that ranks 750+ players and rewards standout attributes with a statistical bonus threshold.
Football Player Position Recommender System
A recommender that analyses 750+ players' attributes by position and surfaces players who would excel in an alternative role.
Automated Connecting-Flight Optimization
An automated system that recomputes viable connections across 1000+ flights the moment schedule times change.
Flight Passenger-Count Prediction
A scikit-learn regression model that predicts passenger counts per flight from operational flight features.
Turkish Image-Captioning Benchmark on MS COCO 2014
A large-scale Turkish image-captioning benchmark - 616,767 human-verified captions and five trained vision-language models.
ESG Diversity & Sentiment Solution - CFA Poland Hackathon
An ESG prototype scoring gender diversity and news sentiment - 2nd place among 44 teams from 28 universities.
Automated ESG Scoring System - HackBogazici
An automated ESG scoring engine built from four scraped data sources - 2nd place among 14 hackathon teams.