Applied AI/hybrid retrieval/grounded agent
Describe a mood. Get a film that exists.
CineMatch pairs a two-stage recommender with a grounded AI assistant. It searches 1,911 films and series by meaning, re-ranks with LambdaMART, and can only recommend titles it retrieved.
No account needed to try it. See how it works
See it in action
$ cinematch discover "mind-bending dream heist" hybrid retrieval over 1,911 titles 1 Inception cos 0.51 2 Money Heist cos 0.49 3 Play Dirty cos 0.48 live results, ranked by meaning, keywords, and title
CineMatchInside the assistant
One request, five steps, every one on the record.
01 · Plan
The model reads the request and chooses tools. Constraints like a series or under two hours become filters, so every pick meets them.
tool_call"a Korean thriller series, not too long"
search_catalog({
query: "tense thriller",
media_type: "tv",
language: "ko",
max_runtime: 60
})
02 · Retrieve
Hybrid search runs three rankers over the catalog: embeddings for meaning, full text for words, trigrams for typos. Reciprocal rank fusion merges them by rank, so no scores need calibrating.
search_titles_hybridsemantic
full text
title
fused: score = Σ 1 / (60 + rank)
03 · Ground
Every title a tool returns gets a short ref. The final answer may only cite refs the agent has seen. An invented title is rejected on the server before it reaches you, and the rejection is logged.
present_pickst1Mousetrap
t4Bloodhounds
t7Squid Game
t19a title no tool returned
3 grounded · 1 rejected
04 · Rank and explain
Your personal feed runs the two-stage recommender: pgvector retrieves 50 candidates near your taste, then a LightGBM LambdaMART model re-orders them. Its SHAP values say why each title ranked where it did.
lambdamart-v1why this pick
close to your taste+0.22highly rated+0.12recent release+0.06because you liked Arrival
05 · Audit and budget
Each run is logged with its tools, token cost, and a hashed prompt. Per-user limits and a global token budget are checked before any model call. Past the budget, the assistant answers from search alone.
assistant_runsstatus picks · prompt sha256:8afe…
tools search_catalog, find_similar
tokens 6,213 · 11.4 s · ungrounded 0
today's budget31%
Evaluation
Every number comes from an eval you can rerun.
NDCG@10, LambdaMART
Re-ranker on held-out users, up 14% on a popularity baseline.
p95 re-rank latency
Stage-two scoring of 50 candidates, measured locally.
P@10, hybrid search
Paraphrased descriptions scored by genre, 4.9x a random ranking.
MRR@10, title lookups
Exact titles and typos. Keyword-only search scores 0.01 on descriptions.
Agent eval cases passed
On gemini-3.5-flash-lite, the production model: constraints, named titles, taste, vague and off-topic requests, and prompt injection.
Median agent run
Tool calls, retrieval, and grounded picks end to end. Zero ungrounded picks shown; the server caught 1 attempt.
Production
Budgets, audits, and fallbacks on every request.
Streaming agent API
POST /assistant streams each tool call as a server-sent event, so the interface shows what the agent intends before the results arrive.
Grounded by construction
Picks are checked against tool results on the server. A title no tool returned is dropped and counted in the audit log.
Any model provider
One OpenAI-compatible client. Ollama on a laptop, a hosted free tier in production, switched by three environment variables.
Budgets that fail closed
Each run is counted against per-user and per-network limits before it starts, under a database lock. If that fails, no model call is made.
Audit trail
Every run records the model, prompt version, tool calls, and token cost. Prompts are stored only as SHA-256 hashes.
Evals in the loop
Retrieval and agent evals run against the real stack. Unit tests, type checks, and lint run in CI on every push.
The catalog
Find your next favorite.
One tap to sign in, no passwords. Your taste profile builds as you go.
Start matching













































