Technology
Turkish Image-Captioning Benchmark on MS COCO 2014
A large-scale Turkish image-captioning benchmark - 616,767 human-verified captions and five trained vision-language models.
616,767 · Turkish captions
The challenge
Progress in image captioning and machine translation depends on large, high-quality labelled datasets - and for Turkish, those barely existed. English benchmarks like MS COCO have driven years of research, but Turkish-language vision-language work had no comparable resource to train and evaluate against.
Closing that gap meant more than machine translation: it required a benchmark of human-verified Turkish captions large enough to support serious research.
The approach
At Istanbul Bilgi University I extended the MS COCO 2014 dataset into a benchmark for Turkish image-captioning and machine-translation research. The dataset adds 616,767 human-verified Turkish captions covering 123,287 images, giving the research community a reusable, quality-controlled resource in a language that previously lacked one.
To validate the benchmark, I trained five image-captioning models on it, spanning a range of architectures from CNN+LSTM models using a ResNet backbone through to Meshed-Memory Transformers. Comparing these models on the new data demonstrated that the benchmark supports meaningful evaluation across both classic and state-of-the-art approaches.
The outcome
Results that moved the needle.
- 616,767
-
Turkish captions
Human-verified, extending the MS COCO 2014 dataset
- 123,287
-
Images covered
A reusable benchmark for Turkish vision-language research
- 5
-
Models trained
From CNN+LSTM (ResNet) to Meshed-Memory Transformers
Want similar results for your project?
Every project above started with a conversation. Let's figure out what yours needs.
Keep exploring
More projects.
Rail catenary pole placement, automated
Days of expert engineering work, reduced to seconds.
D-Risk - MedTech Marketplace with AI Company Profiling
A three-sided marketplace linking medtech startups with investors and specialist freelancers - matched through AI document profiling.
Integrated LoRa Sensor Monitoring & Analytics System
Real-time monitoring and analytics for LoRa environmental sensors - from raw ingestion to geospatial drought and growth insight.
Revenue Administration MCP Server
A retrieval-only MCP server answering Turkish tax and regulation questions strictly from official government sources - zero hallucinations.
Reliability of LLMs in Safety-Critical Requirements Engineering
Master's-thesis research measuring how LLM assistance affects correctness, efficiency and trust in safety-critical requirements engineering.
DSGENAI - AI Safety Requirements Engineering Platform
Stabilising and modernising an AI-driven safety-requirements platform - Flask to Streamlit, GPT-5.1, and critical data-leak fixes.
Retrieval-Augmented Generation (RAG) Documentation Assistant
A Gemini-powered RAG assistant that turns PDF libraries into an accurate, source-grounded knowledge base.
LLM Prompt Optimization for Legal-Clause Classification
Prompt-engineering research that lifted legal-clause classification accuracy by up to 20% on Terms-of-Service documents.
Smart Contract Analysis with NLP
An NLP system that reviews business contracts and recommends changes to keep them compliant with company policy.
Football Player Potential Prediction Model
A machine-learning model that reads scouting attributes and predicts player talent level with roughly 85% accuracy.
Football Player Ranking System
An attribute-based scoring engine that ranks 750+ players and rewards standout attributes with a statistical bonus threshold.
Football Player Position Recommender System
A recommender that analyses 750+ players' attributes by position and surfaces players who would excel in an alternative role.
Automated Connecting-Flight Optimization
An automated system that recomputes viable connections across 1000+ flights the moment schedule times change.
Flight Passenger-Count Prediction
A scikit-learn regression model that predicts passenger counts per flight from operational flight features.
ESG Diversity & Sentiment Solution - CFA Poland Hackathon
An ESG prototype scoring gender diversity and news sentiment - 2nd place among 44 teams from 28 universities.
Automated ESG Scoring System - HackBogazici
An automated ESG scoring engine built from four scraped data sources - 2nd place among 14 hackathon teams.