KVKK-RAG is a Turkish RAG (Retrieval-Augmented Generation) assistant that answers questions about the Personal Data Protection Law (KVKK) with sources. It runs over 373 official documents and uses hybrid search and cross-encoder re-ranking to measurably improve retrieval quality.
Problem and approach
KVKK legislation, board decisions and guidelines are scattered across documents, and keyword search often misses the right article. The project builds a RAG pipeline that chunks and indexes these documents, retrieves the relevant parts and generates answers with their sources.
Hybrid search and cross-encoder re-ranking
Candidate chunks are retrieved with hybrid search (keyword + vector); a cross-encoder then scores each candidate together with the query and re-ranks them, filtering out passages that are semantically close but irrelevant.
Measurement: 27 experiments, MRR +73%
Twenty-seven experimental configurations were compared against a hand-built gold dataset. In the best configuration MRR (Mean Reciprocal Rank) rose from 0.380 to 0.657, a 73% increase. Every change was chosen by measurement, not intuition.
API and deployment
The system exposes FastAPI and GraphQL APIs, and includes document generation and a web interface. It runs in GPU-enabled Docker containers.
Technologies
- Python
- FastAPI
- GraphQL
- Docker
- RAG
- Cross-encoder
- Hybrid Search