Technology
LLM Prompt Optimization for Legal-Clause Classification
Prompt-engineering research that lifted legal-clause classification accuracy by up to 20% on Terms-of-Service documents.
Up to 20% · Metric improvement
The challenge
Terms-of-Service documents are dense with clauses that carry very different legal weight, and classifying them accurately is the foundation for any tool that helps people understand what they are agreeing to. Off-the-shelf LLM prompts handled the task unevenly - strong on some clause types, weak on others - which limited how far the results could be trusted.
The research question was practical: how much can careful prompt design alone improve classification quality, without retraining or fine-tuning the underlying model?
The approach
At the Technical University of Munich I optimised LLM prompts to improve accuracy, precision, recall and F1 for classifying legal clauses in Terms-of-Service documents. Rather than changing the model, I focused on prompt-strategy optimisation and iterative feedback loops.
I ran the work as a measured, feedback-driven cycle: test a prompting strategy against the metrics, analyse where it failed, refine the prompt, and repeat. I evaluated across both binary and multi-label classification setups to make sure improvements generalised rather than overfitting one framing of the task.
The iterative approach delivered up to a 20% improvement in the core metrics across both binary and multi-label tasks, showing how much headroom disciplined prompt engineering can unlock before any model changes are needed.
The outcome
Results that moved the needle.
- Up to 20%
-
Metric improvement
Across accuracy, precision, recall and F1 on clause classification
- Binary + multi-label
-
Tasks optimised
Gains held across both classification setups
- Iterative
-
Prompt-strategy loop
Feedback-driven prompt refinement rather than model retraining
Want similar results for your project?
Every project above started with a conversation. Let's figure out what yours needs.
Keep exploring
More projects.
Rail catenary pole placement, automated
Days of expert engineering work, reduced to seconds.
D-Risk - MedTech Marketplace with AI Company Profiling
A three-sided marketplace linking medtech startups with investors and specialist freelancers - matched through AI document profiling.
Integrated LoRa Sensor Monitoring & Analytics System
Real-time monitoring and analytics for LoRa environmental sensors - from raw ingestion to geospatial drought and growth insight.
Revenue Administration MCP Server
A retrieval-only MCP server answering Turkish tax and regulation questions strictly from official government sources - zero hallucinations.
Reliability of LLMs in Safety-Critical Requirements Engineering
Master's-thesis research measuring how LLM assistance affects correctness, efficiency and trust in safety-critical requirements engineering.
DSGENAI - AI Safety Requirements Engineering Platform
Stabilising and modernising an AI-driven safety-requirements platform - Flask to Streamlit, GPT-5.1, and critical data-leak fixes.
Retrieval-Augmented Generation (RAG) Documentation Assistant
A Gemini-powered RAG assistant that turns PDF libraries into an accurate, source-grounded knowledge base.
Smart Contract Analysis with NLP
An NLP system that reviews business contracts and recommends changes to keep them compliant with company policy.
Football Player Potential Prediction Model
A machine-learning model that reads scouting attributes and predicts player talent level with roughly 85% accuracy.
Football Player Ranking System
An attribute-based scoring engine that ranks 750+ players and rewards standout attributes with a statistical bonus threshold.
Football Player Position Recommender System
A recommender that analyses 750+ players' attributes by position and surfaces players who would excel in an alternative role.
Automated Connecting-Flight Optimization
An automated system that recomputes viable connections across 1000+ flights the moment schedule times change.
Flight Passenger-Count Prediction
A scikit-learn regression model that predicts passenger counts per flight from operational flight features.
Turkish Image-Captioning Benchmark on MS COCO 2014
A large-scale Turkish image-captioning benchmark - 616,767 human-verified captions and five trained vision-language models.
ESG Diversity & Sentiment Solution - CFA Poland Hackathon
An ESG prototype scoring gender diversity and news sentiment - 2nd place among 44 teams from 28 universities.
Automated ESG Scoring System - HackBogazici
An automated ESG scoring engine built from four scraped data sources - 2nd place among 14 hackathon teams.