KODEKS is a legal assistant I am building for police officers in Bosnia and Herzegovina. It answers questions from laws and regulations, and only from them. The legal texts are split into chunks, embedded with OpenAI's text-embedding-3-small, and stored in Supabase with pgvector. When someone asks a question, I run a similarity search, send the relevant passages to Claude, and Claude answers with references to the source.
The architecture is described in the KODEKS case study, and this is the kind of system I build in AI integration work.
In testing, the pipeline worked well for most questions. For some, it returned nothing useful at all. No error, no crash, just an empty or evasive answer. This post is about why that happened, and what I changed.
The symptom
The pattern was not random. Questions that used the same wording as the legal text worked. Questions phrased the way a person would actually ask them, in everyday language, often came back empty.
That is a problem for a tool like this. Police officers do not ask questions in the language of the law. They describe a situation. If the assistant only works when you already know the exact legal phrasing, it is not much of an assistant.
Cause 1: the similarity threshold filtered out everything
The retrieval step did two things: it ranked chunks by cosine similarity to the question, and it dropped every chunk below a fixed similarity threshold. The threshold was meant to keep irrelevant passages out of the prompt.
For questions phrased close to the legal text, the relevant chunks scored well above the threshold. For questions phrased differently, the relevant chunks were still at the top of the ranking, but their scores were lower, and they fell just below the cutoff. The filter removed all of them, and the model received no context at all.
A simplified version of the search function looked like this:
-- Simplified sketch, not the production function
create or replace function match_chunks(
query_embedding vector(1536),
match_threshold float,
match_count int
)
returns table (id bigint, content text, source text, similarity float)
language sql stable
as $$
select id, content, source, 1 - (embedding <=> query_embedding) as similarity
from legal_chunks
where 1 - (embedding <=> query_embedding) > match_threshold
order by embedding <=> query_embedding
limit match_count;
$$;
The where clause is the whole story. A threshold is a single global number, but similarity scores depend heavily on how a question is phrased. One number cannot be right for every question.
Cause 2: the prompt had no plan for "no relevant context"
The second problem made the first one worse. The system prompt told Claude to answer from the provided passages and cite them. It said nothing about what to do when there were no passages, or when the passages did not actually answer the question.
So when retrieval came back empty, the model was in an undefined situation. Sometimes it answered vaguely. Sometimes it returned something close to nothing. In a legal context, the vague answer is the more dangerous of the two, because it can sound confident without being grounded in the text.
The fix
I treated the two causes separately.
1. Tune the threshold against real questions
Instead of picking a threshold by feel, I built a small test set of real questions, written the way people actually ask them, each paired with the passages that should answer it. Then I measured retrieval on its own: for each question, did the correct passages appear in the results, and at what threshold did they disappear?
This made the trade-off visible. A strict threshold keeps the context clean but loses everyday phrasing. A loose threshold catches more of the right passages but lets in noise. With the test set, I could choose a value based on evidence instead of guessing.
2. Add a top-k fallback
Even a well-tuned threshold will sometimes filter out everything. So when nothing passes the threshold, retrieval now falls back to the top few results by similarity, regardless of their score. The model then decides whether those passages are actually relevant, which is a judgment it is much better at than a fixed number.
// Simplified sketch of the retrieval step
const { data: matches } = await supabase.rpc("match_chunks", {
query_embedding: embedding,
match_threshold: THRESHOLD,
match_count: MAX_CHUNKS,
});
let context = matches ?? [];
let usedFallback = false;
if (context.length === 0) {
// Nothing passed the threshold: take the closest chunks anyway
// and let the model judge whether they answer the question.
const { data: nearest } = await supabase.rpc("match_chunks", {
query_embedding: embedding,
match_threshold: -1,
match_count: FALLBACK_CHUNKS,
});
context = nearest ?? [];
usedFallback = true;
}
3. Tell the model exactly what to do when context is missing
The system prompt now has explicit behaviour for the case where the passages do not answer the question: say what is missing, and do not guess. In practice that means the assistant tells the officer that the provided legal texts do not cover the question, or which part of the question is not covered, instead of filling the gap with general knowledge.
Answer only from the provided passages and cite the source for every claim. If the passages do not contain the answer, say clearly what information is missing. Do not guess and do not use knowledge outside the passages.
This instruction matters in both directions. It fixes the empty answers, and it also removes the risk of the model inventing law when retrieval is weak.
The lesson: evaluate retrieval separately from generation
The most useful thing I took from this is a way of debugging, not a specific setting.
When a RAG system gives a bad answer, it is tempting to look at the final answer and start changing the prompt. But a RAG pipeline has two independent parts, and they fail in different ways:
- Retrieval can fail to find the right passages, because of thresholds, chunking, or the embedding model.
- Generation can fail to use good passages well, or behave badly when the passages are missing.
If you only evaluate the final answer, you cannot tell which part failed. In my case, the prompt was not the root cause of the empty answers at all; retrieval was. But the prompt still needed fixing, because it had no defined behaviour for a situation that retrieval will always produce sometimes.
So now I evaluate them separately. For retrieval: for a set of real questions, do the right passages come back? For generation: given the right passages, or deliberately no passages, does the model answer correctly, cite its sources, and admit when something is missing?
That split turns "the assistant is sometimes bad" into specific, fixable problems. It is also the method I want to use before taking KODEKS any wider: measure answer quality on real questions first, then expand.
