Query layer
Semantic search
Data BetaAPI PlannedKeyword queries only find what you can already name, whereas vector search finds what you mean. Candidate position text is indexed by meaning, so a query for rural broadband expansion matches a stance that never uses any of those words.
The index
A Qdrant collection holding 1024-dimensional embeddings under cosine distance, produced by Voyage AI's voyage-3-large model.
| Property | Value |
|---|---|
| Embedding model | voyage-3-large, 1024 dimensions |
| Distance metric | Cosine |
| Index | HNSW, with indexed metadata so filters apply before the search runs |
| Point id | Derived from the node slug by content hash |
| Filterable fields | record_type, stance, issue_slug |
Two of those choices matter more than they might appear to.
- The model name is part of the collection name. Vectors written by one model can never be searched by another, so an upgrade surfaces immediately as a missing collection instead of quietly degrading result quality, and the match is checked before anything is written.
- Point ids come from the node slug rather than being assigned. Re-indexing therefore overwrites in place, which means there is no deduplication pass to run because duplicates cannot occur in the first place.
Querying
{
"query": "expanding broadband access in rural counties",
"filter": { "issue_slug": "issue:economic-development", "stance": "support" },
"limit": 10,
"min_confidence": 0.6
}{
"hits": [
{
"score": 0.8412,
"slug": "position:az-…-economic-development",
"candidacy_slug": "candidacy:az-…",
"stance": "support",
"summary": "Supports municipal fibre buildout in underserved…",
"citing_artifact_slug": "artifact:…",
"confidence": 0.7
}
]
}Filters apply before the search rather than as a pass over the results afterwards, so narrowing to a single issue never silently returns fewer rows than the limit you asked for.
Freshness
A search index is only as good as how stale it is allowed to become, and ours is measured in seconds rather than nightly rebuilds because re-indexing is triggered by each commit rather than by a schedule.
- On every commit. When a batch of changes commits cleanly, only the records whose text actually changed are re-indexed, and everything else reuses the vector already stored.
- Unchanged text costs nothing. A stored content hash is compared against the current text, so an edit that leaves the summary alone incurs no embedding cost at all.
- A nightly pass reconciles the rest. It removes records dropped from Postgres by comparing ids, scoped so that it can reach neither a live record nor a different record type.
- Every pass is auditable. Each run produces a record of what was added, re-indexed, reused and removed, listed by slug and posted for human review.
Hybrid retrieval
Semantic search on its own invents confident answers, while structured search on its own misses anything phrased differently. Natural-language answering therefore queries all three stores at once and merges the results by a fixed rule.
Four properties of that pipeline are worth stating outright, because they are what separates a retrieval system you can put in front of a customer from one you cannot.
- A failing store degrades the answer rather than the request. Any store that is down, slow or misconfigured is simply dropped from the merge, and the answer is composed from whatever did respond.
- The merge is fixed and repeatable. Each store's results are weighted by how much it is trusted, ordered by the resulting score and de-duplicated, with no clock or randomness involved, so the ordering is verified by test rather than assumed.
- The system is allowed to say it does not know. When nothing clears the confidence threshold it declines to answer instead of composing something fluent out of weak results.
- Citations are built from results, never written by the model. The source list is assembled from the ranked facts themselves, which means a fabricated citation is not merely unlikely but impossible.
What is indexed today
The retrieval machinery is production quality, while the body of text behind it is still deliberately narrow: candidate position statements, across the graph's issue taxonomy, and nothing else yet.
Every indexed position carries real text, so nothing in that count is an empty row. Further material such as candidate biographies, ballot measure text and full source documents is added by registering a new record type against the existing structure rather than by rebuilding the index, which is the honest answer to what else will become searchable.
What we do not claim. There is no second-stage reranking model today and no learned sparse retrieval, because the merge is a straightforward weighting that stays transparent and debuggable. We would rather say so than imply a component we have not built.