guide
Why every “movies like this” list shows the same movies
Runs business operations and designed the app's UI. Brings the ideas for what to improve and where to grow next.
8 min read

Type a movie you love into almost any recommendation feature and you will meet the same small committee of movies. Ask for something like Arrival and you get Interstellar. Ask for something like Interstellar and you get Inception. Ask for something unlike either and one of them turns up anyway.
This looks like a content problem, or like laziness. It is neither. It is a property of the space these systems measure distance in, it has a name, and it was quantified in 2010.
Similarity is a distance, and that is the whole trick
A similar-movies feature does not read movies. It converts each one into a vector, which is a long list of numbers standing in for everything the system knows: genre and era, cast and crew overlap, the language people use when they write about it, how audiences behave around it.
Once every movie is a point, “similar” has a definition that a computer can act on. It becomes distance. The nearest points are the answer, and the question of what a movie is about never comes up.
Searching every pair exactly is not an option at catalogue scale, because the cost grows with the square of the number of movies. So production systems use approximate nearest-neighbour indexes. The most widely deployed of them, HNSW, was published by Malkov and Yashunin in 2016 and scales logarithmically rather than linearly, which is the difference between an answer arriving while you wait and an answer arriving tomorrow. It gives up a little accuracy for that, on purpose.
None of this is the problem. This part works.
The part nobody mentions
Here is the finding that explains your experience.
In 2010, Radovanović, Nanopoulos and Ivanović published a paper in the Journal of Machine Learning Research about what happens to nearest-neighbour lists as the number of dimensions rises. They counted how often each point appears in other points’ neighbour lists, and found that the distribution becomes badly skewed. A small number of points turn up as a neighbour of almost everything else.
They called those points hubs, and the effect hubness.
Read that again with movies in mind. It means a handful of titles are near Arrival, near The Godfather, near a Korean thriller from 1999 and near a comedy nobody has heard of. Not because those movies resemble each other. Because of where they sit in the space.
So when a similar-movies list feels generic, the generic titles are not being inserted by an editor with poor taste. They are the mathematically correct answer to a question that was never quite the question you were asking.
The hubs are load bearing, which is the uncomfortable part
You might expect hubness to be a flaw that better engineering removes. The opposite turned out to be true: the search machinery leans on it.
A 2024 paper looked at why HNSW’s hierarchy helps, and found that on high-dimensional data it mostly does not. A flat graph matched it on both latency and recall while using roughly 38 percent of the memory. What replaced the hierarchy was already in the graph. Queries spend most of their time travelling through a well-connected highway of hub nodes before dropping into a local neighbourhood at the end.
Above about 32 dimensions, the authors suggest, the hierarchy stops paying for itself. Below it, it still helps.
The practical reading is that hubs are not a bug in the index. They are how the index gets you across the space quickly. The same structure that makes similarity search fast also makes a few items disproportionately reachable, and a system that returns what it reaches will return them.
Why more dimensions makes it worse
The instinct when recommendations feel shallow is to reach for a richer representation. More signals, bigger vectors, a better model.
Hubness gets stronger with dimensionality. That was the original finding, not a caveat added later. A more expressive representation may be worth having for other reasons, and it will not on its own make the neighbour lists less skewed. You can attack the symptom directly, by penalising items that appear too often or by rescaling distances locally, and those techniques work. They are corrections applied after the fact rather than a property of the space.
Which points at the real issue. Distance answers “what resembles this”. It was never asked “what would this person enjoy”, and no amount of accuracy at the first question turns it into the second.
What actually changes the answer
Two things, and neither is a better distance metric.
Rank against a person rather than against a movie. A profile built from a run of your own reactions is a different input from a single title, and it does not sit where the hubs sit. This is what our own system does, described in how the Filmatic recommendation algorithm works.
Aim slightly away from the nearest result on purpose. There is a separate reason not to want the closest movie, which is that it tends to be the same movie in different clothing, and it is the argument behind how Double Feature works. Distance is a useful ingredient. It is a bad objective.
Getting more out of the tools you already use
None of this makes similar-movie features useless. It means reading them with the mechanism in mind.
Start from something specific rather than something famous. A widely known movie sits closer to the crowded middle of the space, which is where the hubs live. An unusual starting point produces an unusual list.
Say what you want carried over. A bare title cannot tell a system whether you cared about the pacing, the period, the register or the subject. Where a tool accepts more than a title, that is the part worth filling in.
Read past the top three. The first few results are the ones most likely to be hubs. The interesting answers are often four to ten places down, which is exactly where most people stop looking.
Treat a repeated title as information. If the same movie follows you across three unrelated searches, you have learned something about the index rather than about the movie.
Questions
Why do different movies return the same recommendations?
Because of a property of high-dimensional geometry called hubness. When movies are represented as vectors with hundreds of dimensions, the number of times each movie appears in other movies' nearest-neighbour lists becomes very unevenly distributed, and a few movies end up near almost everything. Those movies are returned for a wide range of inputs, which is why two unrelated starting points can produce overlapping lists.
How does a “more like this” feature actually work?
The movie is converted into a vector, a long list of numbers derived from its metadata, its cast and crew, the text written about it and how audiences behave around it. Similarity becomes distance between vectors, and the feature returns the closest ones. Searching millions of vectors exactly is too slow to do on demand, so production systems use approximate nearest-neighbour indexes, which trade a small amount of accuracy for a very large amount of speed.
Is the most similar movie the best movie to watch next?
Usually not, and for two separate reasons. The nearest movie is often the same movie in different clothing, which makes for a repetitive evening. And in a high-dimensional space the nearest movie is disproportionately likely to be a hub, meaning a title that sits near everything and is therefore not really about you. Distance answers “what resembles this” rather than “what would this person enjoy”.
Does adding more dimensions to the model fix it?
It makes it worse. Hubness increases with dimensionality, which is the finding that named the effect. More expressive representations are useful for other reasons, but they do not make nearest-neighbour lists less skewed on their own.
How do I get better answers out of a similar-movies tool?
Give it a specific input rather than a famous one, because a well-known movie sits closer to the centre of the space where the hubs live. Name what you actually want carried over, whether that is the pacing, the era or the subject, since a bare title cannot tell the system which of its properties mattered to you. And read past the first three results, which are the ones most likely to be hubs.
Keep reading

guide
How movie recommendation algorithms work: the four main types
Collaborative filtering, matrix factorisation, content-based filtering and hybrids. How each movie recommendation algorithm works and where it fails.

product
How Filmatic picks one movie for two people
Averaging two tastes gives you a movie neither person wants. Two-person mode scores every candidate against both profiles and ranks by the weaker of the two.

product
How Double Feature finds the movie to watch after yours
Pairing is not similarity. The nearest movie in the model is usually the worst thing to watch next, and the reason why explains how the whole feature is built.