Want something good to watch? First let's define "Good" - with statistical analysis - and derive the three (?) eternal dimensions of human morality.

Updated: Sep 6
Explore the dataset: https://somethinggoodtowatch.id-host.com/#/atlas
My wife and I can never find something we both want to watch. She's happy to binge on Netflix series or marvel blockbuster. I don't think I've seen a really good film since The Return of The King came out in 2003.
I feel like my morals are in conflict with most Hollywood movies. But what does that mean? Do movies have morals? Can we measure them? And if so, what can we learn about morality and ourselves.
The Moral Dimensions Hypothesis
Based on the success of the Big Five research, my hypothesis is this:
Films convey moral teaching
Moral teaching varies in a small number of dimensions
Those moral dimensions can be found statistically
Inspiration: The Big Five Personality Model.
One of the best established findings in psychology is the "Big Five" personality model. The "Big Five" are five stable dimensions on which someones personality can be measured. The interesting thing about the "Big Five" is that nobody chose the dimensions. They emerged by asking a massive bank of personality tests to 1000s of participants, then using statistical analysis to find out how the answers naturally clustered. What emerged was 5 dimensions:
Openness: Measures creativity, curiosity, and imagination.
Conscientiousness: Measures self-control, organization, and goal-directed behavior.
Extraversion: Measures sociability, energy, and assertiveness.
Agreeableness: Measures kindness, cooperation, and empathy.
Neuroticism: Measures emotional instability and sensitivity to stress
The Approach
The method is almost identical to how the Big Five model was established. But studying moral teaching in films instead of personality in human participants. The overall procedure is:
Generate a large bank of interesting moral questions
Find each film's position on each moral question (using an LLM)
Normalise for bias and confounding variables
Use statistical analysis (PCA) to find underlying factors.
Name the underlying factors, these are the moral dimensions.
Plot films on the moral dimensions, and see how they line up.
I now have a valid scientific excuse to boycot all films that don't have the exact moral dimensions as "The Lord of The Rings: The Return of The King (2003)".
1. Generate moral questions
A wide range of moral questions that films actually disagree on. I asked an LLM to analyse a film and create generic propositions that could apply throughout history:
Too Specific: "A lion prince must return to claim his father's kingdom."
Better: "Legitimate authority is inherited, and usurpation is wrong."
We started with ~3,000 questions, after removing duplicates we got down to ~300, after removing questions that practically all films agree on, we were left with 98.
2. Assign films to the moral positions using an LLM
Give the model the entire script/subtitles of the film, ask it to assign the film a position against the claim. e.g.
Using only the evidence given, categorise this films position against this moral claim. If the film does not clearly take a position on the question, respond with "Not Addressed".
Valid Answers:
"Affirm"
"Strongly Affirm"
"Deny"
"Strongly Deny"
"Not Addressed"
The Claim:
"Legitimate authority is inherited, and usurpation is wrong."
The Evidence:
♪ From the day we arrive
♪ On the planet
♪ And blinking
♪ Step into the sun
...Result:
Film: "The Lion King"
Proposition: "Legitimate authority is inherited, and usurpation is wrong."
Stance: "Strongly Affirmed"
Evidence: "Simba is 'the one true king' [Mufasa's ghost] and rightful heir; Scar's rule is illegitimate and disastrous."
That is one moral stance for one film. Repeat that for the 98 moral stances for every film in the database.
3. Normalise for talkativeness.
Normalise for the LLM being too agreeable by using matching and contradicting propositions e.g.
Affirm: "Legitimate authority is inherited, and usurpation is wrong."
Deny: "Legitimate authority is earned, not inherited."
This strategy is known in statistics as "Reverse Keying".
Also normalise for some films being more talkative than others by spreading a films "moral weight" amongst every proposition it had an answer for.
4. Extract the strongest dimensions with statistics.
Run Principal Component Analysis which extracts independent dimensions of covariance within the dataset. It starts off with every individual moral argument as a dimension and then plots films on that n-dimensional space (for us, 98 propositions = 98 dimensions at first). It then draws a best fit line through them to establish the strongest "principal component". Then, after flattening everything on that line, it draws a next best-fit line through the remaining data, flattens on that component, and so on. The below illustration shows PCS in two dimensions. Visualising this in 98 dimensions is an exercise left to the reader.
Once we've established the list of potential dimensions, we use a seperate statistical test called Horns Parallel analysis to see which dimensions are predictive enough to be worth talking about. Basically, we only want to talk about a dimension that can predict propositions measurably better than chance. Here's what we found with the film analysis: a multitude of potential dimensions, but only three that clearly beat random chance.

e.g. Two different questions that are consistently answered the same way A. Sexual desire is a natural, comic, and sometimes ridiculous force that pervades everyday life. B Political ideologies and their rallies are primarily objects of satire rather than serious commitment. Films that affirm one overwhelmingly affirm the other: r = +0.88 across the 53 films that took a position on both. |
5. Name the dimensions.
Ok, we have established that the way that movies answer moral questions can be split up into distinct dimensions, and that in our dataset of about 500 films, three of those dimensions beat random chance.
Now the question is, what are the names of those dimensions?
Dimension One
Starting with the first dimension, positive value is correlated with the following propositions:
1. The human capacity for love and forgiveness is greater than the capacity for hatred and vengeance.
2. Forgiveness and understanding can lead to healing and reconciliation.
3. Individuals and societies have the potential for redemption and change.
4. Forgiveness can heal and enable progress.And in the negative direction:
1. Past actions can define a person's future.
2. A person's happiness is dependent on external circumstances.
3. Revenge can be a valid motivation for action.
4. The past cannot be escaped and will eventually catch up with an individual.I can name this axis myself, I'd called it "Forgiveness vs Revenge", or even "Forgiveness vs Punishment".
Rather than naming it myself, I tasked another LLM with reviewing all of the correlated propositions and finding a suitable name. Running the task 5 times and taking a name that came up the most. here is what it came up with:
Deterministic pessimism vs Redemptive optimism
Can individuals and societies overcome past wrongs through forgiveness and change, or are they trapped by past actions and circumstances?
Not as nice as my names! But to be fair, it is trying to find a name that incorporates all 48 co-related moral questions, not just the top four.
Here's how the set of films films clustered on this axis. It looks like a skewed normal distribution with most film favouring mercy, but a long tail of films favouring punishment:

Moral Dimension | Punishment (deterministic pessimism) | Forgiveness (redemptive optimism) |
Moral Propositions | Past actions can define a person's future. A person's happiness is dependent on external circumstances. Revenge can be a valid motivation for action. | The human capacity for love and forgiveness is greater than the capacity for hatred and vengeance. Forgiveness and understanding can lead to healing and reconciliation. Individuals and societies have the potential for redemption and change. |
Most Extreme Films | Sweeney Todd: The Demon Barber of Fleet Street -0.57 Barry Lyndon -0.54 Requiem for a Dream -0.53 Whiplash -0.52 The Good, the Bad and the Ugly -0.51 | Zootopia +0.72 Mrs. Doubtfire +0.69 Love Actually +0.65 As Good as It Gets +0.65 Gravity +0.63 |
That seems believable to me. Also, the general shape of the distribution is encouraging, when I ran the analysis incorrectly I got lumpy or completely one sided distributions and film positions that didn't make sense. This axis seems to be something that films genuinely differ on.
The other dimensions:
Divine order vs Self-determination
What a name! Now we're getting into it. A normal distribution with Final Destination on one end, and 12 Years a Slave on the other end:

Moral Dimension | Divine Order | Self Determination |
Moral Qs | There is a right order that precedes individual choice. Legitimate authority is inherited, and usurpation is wrong. The bond between parents and children is stronger than the bond between siblings. | The pursuit of legacy can lead to self-destructive behavior. The pursuit of wealth can corrupt and lead to the downfall of individuals and societies. Individuals have the right to determine their own destiny. |
Strongest Films | Final Destination -0.35 The Lion King -0.34 The Exorcist -0.31 Home Alone -0.28 The Passion of the Christ -0.26 | Battleship Potemkin +0.61 Casino +0.52 The Good, the Bad and the Ugly +0.52 12 Years a Slave +0.51 Spartacus +0.49 |
Intrinsic worth vs Utilitarian sacrifice
A bit more complex, but still coherent. Could also be read as Absolute morality vs Relative morality. Are some actions wrong no matter what, or does it all depend on the outcome?

Moral Dimension | Intrinsic worth | Utilitarian sacrifice |
Moral Qs | Children have a duty to care for their aging parents. Professional success should not come at the expense of personal relationships. Some actions are inherently wrong regardless of their outcomes. | The value of a life is determined by its potential for future contributions. Violence can be justified in the defense of freedom. The pursuit of personal happiness can lead to fulfilling one's true potential. |
Strongest Films | All Quiet on the Western Front -0.57 Seven -0.54 Psycho -0.54 Mulholland Drive -0.52 The Shining -0.51 | Starship Troopers +0.48 Master and Commander: The Far Side of the World +0.45 Once Upon a Time in Hollywood +0.44 The Thing +0.43 Whiplash +0.39 |
Summary of the dimensions
In addition to the three strong dimensions, there are many more candidate dimensions. I set the rule that we only surface dimensions that can given the dataset is only ~300 questions across ~500 films, there are likely more dimensions worth talking about that haven't yet passed this threshold. I suspect that more than three meaningful dimensions would emerge with a wider corpus of films and moral questions. The dimensions:
Punishment vs Forgiveness
Should the guilty be punished like in Goodfellas, Sweeny Todd, and Saw. Or can we be redeemed and forgiven, like in Zootopia, Mrs. Doubtfire and Love Actually?
Divine Order vs Self Determination
Should we be masters of our own destiny, like in 12 Years a Slave, Scarface, and The Truman Show. Or should we follow the cosmic hand of fate like in The Lion King, Narnia, or Final Destination?
Intrinsic worth vs Utilitarian sacrifice
Does every life have sacred value like in "All Quiet on the Western Front" and "Westside Story", or can a man's value be measured in his contribution and sacrifice like in Starship Troopers, Master and Commander, and 300?
How do we know that these are the moral dimensions?
Selection of films
We only analysed about 500 relatively popular Hollywood films. These skew towards modern creators and audiences. I suspect if we analysed a wider range of stories including folk tales, bible stories, ancient epic poems, and more, we would have a much more complete picture of overall human morality.
Asking good questions
The bank of moral questions is another big factor, if we don't ask very good questions, we may never know what the underlying dimensions are. Questions were generated by an LLM, analysing the plot of a film and being asked to generate moral questions that apply to that film but could also apply across cultures and stories. We generated about ~3000 questions but found that most of them where almost literal duplicates, and of the remaining, most were answered the same way by 99% of films. We reduced this to 98 meaningful questions which films actually differ on.
When removing bad questions, the predictive strength of our model actually improved, because they add "noise" without differentiating between films.
An unbiased reviewer?
Asking a film reviewer what the moral of the film was will introduce their own bias. But even a synopsis would introduce the interpretation of the one writing it. In order to be as objective as possible, I decided to use only the content of the film itself as inputs to be ranked by an LLM.
Moral bias of the model is another potential confounder. Popular models like Claude have been explicitly trained in "Alignment" which is very similar to a moral stance. I tested a variety of models, including an unconstrained model with no alignment guardrails, dolphin-mistral-24b.
Dolphin did excel at asking good questions. The moral questions asked by Claude were mostly softballs, 80% of the time all films would agree. Dolphin's unconstrained questions were twice as likely to generate a useful response.
Good question - Dolphin (unaligned model) - Films genuinely differ:
"Legitimate authority is inherited, and usurpation is wrong."
Bad Question - Claude Opus - Films all take the same position:
"Wealth is the true measure of a person's worth"
LLM Cost vs Performance
But interpreting the moral lessons of an entire film based on its subtitles is not an easy task. I originally tried with Claude Opus 5 but the cost was around $2.00 per film. I tried downgrading to Claude Sonnet but it was not much cheaper. Analysing ~500 films could cost around $1000USD. I tried Grok, it was slightly cheaper but the failure rate increased dramatically. My unconstrained model "dolphin" was much cheaper but failed most of the time and would deliver very crude analysis including echoing back the example answers it from its own prompt about 30% of the time.
Then I found DeepSeek. It DeepSeek gave better results than Claude Sonnet for only 2.2% of the cost. That means the project was achievable for under $30 instead of over $1000 (!).
I actually did quite a bit of testing around this, the best result was found in getting the questions from dolphin and using deepseek to read the subtitles
Model | Unique Qs | Rate Contested | % Films failed | Answered per film | Cost/MToken |
dolphin-mistral-24b | 219 | 42% | 65% | 7 | $0.16 |
deepseek-chat | 298 | 24% | 8% | 43 | $0.11 |
grok-4 | 84 | 39% | 18% | 2 | $2.99 |
claude-sonnet-5 | 141 | 19% | 11% | 2 | $4.85 |
claude-opus-5 | 694 | — | — | — | $5.90 |
Reviewing the answers generated by deepseek, I found them to mostly quite coherent and similar to how I would answer. Although certain subtleties can be missed, for example in the lion king, it affirms that personal happiness leads to fulfilling one's true potential, but in that film, Simba has to give up on his personal happiness and return to fight Scar in order to fulfil his true potential as king. Analysing the responses, we estimate that this kind of mistake was made ~1% of the time and adds noise to the analysis but doesn't compromise the overall validity because of the weight of correctly interpreted questions.
Bad Interpretation (Affirms before reading the whole story):
Film: "The Lion King"
Proposition:The pursuit of personal happiness can lead to fulfilling one's true potential.
Stance: "Affirmed"
Evidence: "Simba finds happiness in Hakuna Matata, but later realizes his true potential by returning to his responsibility."
Good Interpretation:
Film: "The Lion King"
Proposition: There is a right order that precedes individual choice.
Stance: "Strongly Affirmed"
Evidence: "The 'circle of life' is presented as a given order that must be respected."
Another misread occurs when the LLM answers from its existing training/knowledge rather than the evidence contained within the subtitles. This is exactly the form of LLM bias that we want to avoid. It is remembering a plot point from The Lion King 2 and inserting it as evidence for something that is not supported by the subtitles for The Lion King (1994):
Bad Interpretation (Existing assumptions rather than evidence):
Film: "The Lion King"
Proposition: The 'us versus them' mentality can be overcome through understanding and cooperation.
Stance: "Affirmed"
Evidence: "Lionesses, hyenas, and others unite under Simba's rule; hyenas are defeated but the film shows cooperation among different species."
It's hard to estimate this category of error but one analysis is to compare howmany readings cite actual dialog from the subtitles. Claude cited dialog in 93% of its responses while deepseek only did so 73% of the time. This doesn't mean that the other 27% of readings are Hallucinated or incorrect but it does make it harder to verify. Here a correct reading summarises the film without citing evidence directly
Good Interpretation, but poor evidence:
Film: "Princess Mononoke"
Proposition: "Violence can be a natural response to conflict."
Stance: "Affirmed"
Evidence: Violence is presented as a natural response to conflict: boars attack humans, wolves retaliate, humans fight back. The film shows violence as a cycle,
Statistical results
Punishment vs Forgiveness | Divine Order vs Self Determination | Intrinsic worth vs Utilitarian sacrifice | |
% over null (Horn's parallel analysis result, % over p<0.05) | 267% | 24% | 25% |
Coherence ("Cronbach's alpha", correlation of Qs when used together) | 0.94 Very Strong | 0.79 Good | 0.77 Good |
Split-Half Replication. (can we re-derive the same dimensions from seperate halves of the data) | 0.94 null=0.32 p<0.05 = 0.45 | 0.60 null=0.17 p<0.05=0.28 | 0.45 null=0.06 p<0.05=0.13 |
Correlation (mean inter-item correlation of moral Qs) | +0.21 | +0.11 | +0.13 |
Summary
We analysed ~500 films for moral alignment, I found that the unaligned dolphin model was the best at asking interesting paired moral questions and that deepseek-chat was the best at evaluating a films position based on only its subtitles. Out of the statistical analysis, we found three strong dimensions which were independent, coherently named, and sorted films sensibly.
The first dimension is incredibly well established and predictive. The other two are good, but there are also latent dimensions that narrowly missed out on clearing the bar to be predictive. I believe that with more films that cover a wider moral landscape, or better questions that group films differently, further meaningful moral dimensions are likely to emerge.
I think that we can say based on these results that yes, films do convey moral teaching, and yes, we can measure this and visualise their positions on a small number of meaningful moral dimensions.
The dataset is live. Lookup how your favourite film scored: https://somethinggoodtowatch.id-host.com/#/atlas

Next Questions:
Can we predict which films someone will like using these moral factors?
Using an existing dataset of film reviewers, do individual users gravitate towards a moral position, or do they rate films by a different metric?
Can we measure the moral alignment of a religious or political movement based on the films they recommend?
Many groups publish recommended watching lists such as the most edifying Christian films of all time, and the
Can we also analyse stories, mythology, and other content with the same framework? What will we find?
Can we place the morality of ancient Greece by analysing poems like The Iliad and The Odyssey? What moral position is taught by Aesop's Fables? How do these compare with HollyWood? Do new dimensions emerge or can they be mapped to a position on the existing axes?




Comments