How it works
How CineMatch builds recommendations
A two-stage pipeline pairs vector similarity search with a learned ranking model, and a grounded assistant sits on top.
01
How recommendations work
The two-stage pipeline
Every recommendation request runs two stages. Vector search retrieves 50 candidates close to the user's taste, then a ranking model re-orders them using quality, popularity, and era signals.
02
The retrieval stage
Vector search with pgvector
Every title is embedded once with OpenAI's text-embedding-3-small model, from its title and overview, into a 1536-dimensional vector.
User preferences are encoded the same way, built from the embeddings of movies they have liked and watched, weighted by recency.
Finding candidates is a nearest-neighbor search: we use pgvector's HNSW index to find the 50 movies with the highest cosine similarity to the user's embedding. Titles the user already rated are excluded.
Embedding space
Try it yourself
Pick any title below to see its 5 nearest neighbors in embedding space. This calls the real pgvector index with live data.
Select a title above to see real-time vector similarity search in action
03
The ranking stage
Multi-signal re-ranking
Raw similarity is not enough. A movie can be close in embedding space but poorly rated, or popular but not to the user's taste. The ranking stage combines those signals into one score.
Scoring weights
How close the movie is to the user's taste in embedding space
TMDB community rating, normalized to a 0-1 scale
Log scale keeps blockbusters from dominating
Fraction of the movie's genres matching the user's preferences
Learned re-ranker: The weights above are the transparent linear baseline. Production serves a LambdaMART model (LightGBM) that directly optimizes NDCG and learns non-linear preferences the handcrafted weights cannot capture, such as a vote-average sweet spot or an era preference. It leads the offline eval below. The ranker service supports both models and routes between them per request.
04
Evaluation
Measuring recommendation quality
NDCG@10
Measures whether the most relevant movies appear at the top of the list, penalizing good recommendations buried at position 8 more than position 2.
MRR
How soon the first relevant title appears: the mean of 1 / rank of the first hit across users.
Hit Rate@10
Whether the top 10 contains at least one title the user rated highly.
| Model | NDCG@10 | MRR | Hit Rate |
|---|---|---|---|
| Popularity Baseline | 0.72 | 0.88 | 1.00 |
| Vector Retrieval Only | 0.80 | 0.94 | 1.00 |
| Linear Re-ranker | 0.80 | 0.95 | 1.00 |
05
The AI layer
A grounded, audited agent
On top of the recommender sits a tool-calling agent. It plans with a language model but answers only from the catalog, and every step it takes is streamed to the page and written to an audit log. The model is swappable: the same Go client talks to Ollama on a laptop or a hosted provider in production.
01 Request
Up to 12 turns, validated and length-capped. Quota and budget checked first; the check fails closed.
02 Plan
An OpenAI-compatible model chooses read-only tools. Constraints become filters.
03 Tools
search_catalog, find_similar, get_taste_profile, get_recommendations. Each result title gets a short ref.
04 Ground
present_picks may only cite refs a tool returned. Anything else is rejected and counted.
05 Guard
Replies that repeat the instructions are replaced before they are sent.
06 Record
Streamed as server-sent events, then written to the audit log with tokens, latency, and a hashed prompt.
Retrieval eval
15 title lookups (including typos) and 18 paraphrased descriptions that avoid genre words, over 1,843 titles. A description result counts as relevant when its TMDB genres match the target, a label none of the rankers read. A random ranking scores 0.15 P@10.
| Retrieval |
|---|
06
Cold start
What happens for new users
A new user has no likes yet, so there is no taste vector to search with. The feed switches over as soon as there is one:
No likes yet
Popular titles
The feed shows the most popular titles. The assistant still works, since search needs no history.
source: popular
One like or watch
Full pipeline
Each like or watch rebuilds the taste vector as a recency-weighted mean of liked titles' embeddings, and the two-stage pipeline takes over.
source: personalized
07
Tech stack
Built with
Go
API Backend
Small static binary on Cloud Run. The API streams assistant runs as server-sent events and gives every ranker, model, and database call a deadline.
Python FastAPI
Ranking Service
LightGBM and the eval pipeline are Python, so the ranker is too. Pydantic validates every request; re-ranking 50 candidates takes about 0.9 ms at p95.
Supabase + pgvector
Database & Vector Search
Vectors, full text, trigrams, and app data in one Postgres, with an HNSW index for kNN search.