Technology
Turkish Image-Captioning Benchmark on MS COCO 2014
A human-verified Turkish caption dataset covering all of MS COCO, plus five models trained on it - a new reference point for Turkish image captioning.
The challenge
Models that understand an image and describe it in a sentence need large, clean data to train on. For English, MS COCO has done that job for years - the field's shared measuring stick. For Turkish there was no equivalent.
The consequence: Turkish vision-language work could not be compared. Everyone built their own small dataset, and results had no common ground to be measured on.
Filling the gap by machine translation alone wouldn't do it either. Automatic translation produces captions that drift from the image or read badly in Turkish, and a benchmark built on those is a faulty measuring stick. What was needed was verified quality at scale.
The approach
As part of the research group at Istanbul Bilgi University, we turned MS COCO 2014 into a complete benchmark for Turkish image captioning research.
- Scale and human verification together. The dataset contains 616,767 Turkish captions covering 123,287 images, and the captions were human-verified - not just machine-translation output. Large enough for serious model training, reliable enough to serve as a measuring stick.
- We tested the dataset with models. To claim a benchmark works, you have to train on it. Five image-captioning models were trained and compared on the data, ranging from classical CNN+LSTM architectures with a ResNet backbone to the Meshed-Memory Transformer.
- Results surpassed the state of the art at the time. The trained models exceeded the best prior results in Turkish image captioning, demonstrating both the dataset's quality and the value of comparing architectures on it.
- Released openly. The dataset was published for the research community, so subsequent work can be measured on the same ground.
The work was published at ICECCME 2022. Co-authors: Sina Berk Golech, Saltuk Buğra Karacan, Elena Battini Sönmez and Hakan Ayral.
Stack: Python, PyTorch; CNN+LSTM (ResNet backbone), Meshed-Memory Transformer and other vision-language architectures; standard captioning evaluation metrics.
Want similar results for your project?
Every project above started with a conversation. Let's figure out what yours needs.
Keep exploring
More projects.
Rail catenary pole placement automation
Weeks of expert engineering work, reduced to seconds.
D-Risk - MedTech Marketplace with AI Company Profiling
A three-sided marketplace linking medtech startups with investors and specialist freelancers - matched through AI document profiling.
Integrated LoRa Sensor Monitoring & Analytics System
Turning raw LoRa telemetry from tree-mounted sensors into live dashboards that answer watering and growth questions - built in one week.
Revenue Administration MCP Server
An assistant that reads Turkish tax legislation from its official source at the moment you ask and answers with the article behind it - a system a certified acc
Reliability of LLMs in Safety-Critical Requirements Engineering
A controlled experiment measuring what an ungrounded, off-the-shelf chatbot contributes to safety-critical engineering.
DSGENAI - AI Safety Requirements Engineering Platform
Stabilising and modernising an AI-driven safety-requirements platform - Flask to Streamlit, GPT-5.1, and critical data-leak fixes.
Retrieval-Augmented Generation (RAG) Documentation Assistant
A RAG assistant that turns stacks of PDFs into a searchable knowledge base where every answer traces back to the source document.
LLM Prompt Optimization for Legal-Clause Classification
Comparative research that lifted F1 from 0.62 to 0.76 on Terms-of-Service clause classification through automatic prompt optimisation alone - without retraining
Smart Contract Analysis with NLP
An NLP system learns from Siemens' legal team's past contract revisions, flags the same clauses in a new contract, and proposes the edit that was made before.
Football Player Potential Prediction Model
A classification model predicting whether a player will be marked "highlighted" from 39 scout attribute scores - ROC-AUC 0.86 under 10-fold cross-validation.
Football Player Ranking System
A scoring engine that ranks players not by total score, but by how many attributes they exceed the statistical average for their own position.
Football Player Position Recommender System
A recommender that compares a player's attribute profile against the profiles of other positions and finds players could be more valuable in a different role.
Automated Connecting-Flight Optimization
A tool that recalculates every connection possible through the hub - day by day, with passenger volumes - when a single flight's time is shifted.
Flight Passenger-Count Prediction
A demand model predicting booked passengers on one-stop routes from a flight's own characteristics, selected by comparing eleven regression models.
ESG Diversity & Sentiment Solution - CFA Poland Hackathon
An ESG prototype scoring gender diversity and news sentiment - 2nd place among 44 teams from 28 universities.
Automated ESG Scoring System - HackBogazici
An automated ESG scoring engine built from four scraped data sources - 2nd place among 14 hackathon teams.