This started as an idle question Malcolm asked:
what % of the time, in fiction, does a character who says “trust me” turn out to betray that trust to some substantial degree? does it vary based on whether it’s “just trust me” or other variants? does it vary between movies and books? has it changed over time?
let’s make this into a research project.
So we did. Two corpora: 617 films (the Cornell Movie-Dialogs Corpus, median 1990s) and 171 novels (sampled from Project Gutenberg, median 1883). Every occurrence of the phrase was found by regex and then coded by an LLM judge that reads the entire work — who said it, what kind of speech act it was, whether the trust was betrayed and how, whether the audience already knew, how the person being asked reacted, and how much of the plot the whole thing carries.
That’s 581 coded occurrences. The answer to the literal question is 38% in film, 27% in books. But that number turns out to be the least interesting thing we found.
The single strongest pattern is about what the audience knows at the moment the line lands. We coded that separately: does the reader or viewer already know this person is lying, while the character being addressed doesn’t?
| What the audience already knows | Film: betrayed | Books: betrayed |
|---|---|---|
| Audience knows the speaker is untrustworthy | 83% | 92% |
| Strong cues, not yet confirmed | 75% | 57% |
| Audience knows no more than the addressee | 19% | 8% |
| Audience knows the speaker is sincere | 14% | 6% |
Look at that third row. Now narrow it to deliberate deception — cons and defections, excluding characters who meant it and simply failed. Across the 79 appeals in both corpora where the audience had no advance signal about the speaker, there was exactly one bad-faith betrayal. One. Across two media, a century apart.
So the phrase does not work the way the folk theory says. It isn’t a con artist’s tell that slips past you. It’s a confirmation device. Fiction tells you first — and if it hasn’t told you, the character almost certainly means it. Roughly a third of film appeals come from someone the film has already marked as a liar; in those, it’s overwhelmingly a lie. In the rest, it basically never is.
The failures are a different matter. Even with no advance signal, about 17% of sincere “trust me”s still end badly — the speaker meant every word and was outmatched anyway. Nick of Time (1995) is the purest case in the set: Gene’s plea to Krista is completely honest and completely accurate, and her trusting him gets her murdered.
Pooling all direct appeals: film 38% betrayal with 22% in bad faith; books 27% with 12%. Roughly half the bad faith in prose.
An honest confound. The books are ~1883 and the films are ~1990s, so medium and era are the same variable in this design. Nothing here separates them. The books might be gentler because they’re books, or because they’re Victorian. The fix is cheap and specific — code pre-1930 films, which exist in the OpenSubtitles archive back to 1902, against the same Gutenberg sample. Until someone does that, treat this row as unexplained.
Malcolm asked how often the person being asked actually refuses. More often than we guessed: only 43% of film appeals are simply granted. 23% are flatly refused, 22% granted under visible protest, the rest deflected or ignored.
The interesting part is whether that refusal is right. It isn’t. In film, characters who complied were betrayed 37% of the time; characters who refused were betrayed 41% of the time. Those are the same number. Fictional characters’ spoken suspicion carries essentially no information about whether they’re about to be betrayed.
Reservoir Dogs is the whole finding in one scene. Mr. White: “Joe, trust me on this, you’ve made a mistake.” Joe refuses — “You don’t know jack shit. I do.” Joe: “Larry, I’m askin’ you to trust me on this.” White refuses — “Don’t ask me that” — and shoots him. Joe was right, White was wrong, both refused, both die.
In Victorian novels it’s worse than useless. Characters who refuse face a 15% betrayal rate — well below the 27% base rate. Refuse someone in a 19th-century novel and you are probably wronging an honest person. There are 29 such cases. Dora Thorne does it twice in a row: Lillian is concealing only her sister’s secret, Lionel breaks the engagement over it, the narration pointedly notes he “saw no guilt on her fair brow”, and they’re married ten years later.
Prose has narration, so you can code something subtitles simply cannot: the addressee’s private reaction, as against what they say out loud.
| The listener’s inner state | Share | Betrayed |
|---|---|---|
| Inward feeling matches what they say | 42% | 19% |
| Narration gives no access | 36% | 30% |
| Agrees outwardly, privately suspicious | 17% | 48% |
| Resists outwardly, privately persuaded | 4% | 0% |
The unvoiced doubt tracks the betrayal; the voiced refusal doesn’t. Which is a lovely result and probably not the one it looks like.
The likelier reading. This is probably a fact about narration, not about people. The narrator mentions a flicker of private doubt precisely because the author is foreshadowing. Jack London gives the game away in a single clause: “You will first have to trust me while I go upstairs for my purse.” She saw the doubt flicker momentarily in his eyes — while her slipper is working the hidden police bell. So this may be a foreshadowing marker the prose emits, not evidence that fictional people have good instincts.
Of everything we measured, one variable really moves. A character who says “trust me” once in a film betrays 33% of the time, 18% in bad faith. A character who says it two or more times betrays 62% of the time, 43% in bad faith.
That’s across ten different speakers in ten different films, so it isn’t one movie skewing things — Delacroix in Bamboozled (con, con), Telly in Kids (con, con), Sam in Brazil (con, con, failure), against Lindsey in The Abyss (honoured, honoured) and Conrad in The Game (honoured, honoured). Protesting too much, literally and measurably. With n=21 it’s the first thing worth confirming at scale.
This one inverted my expectation. It sounds like the classic manipulator’s move — demand reliance while refusing to justify it. In fiction it’s the opposite. Characters who explicitly decline to explain themselves betray less: 29% in film against 40% for those who explain, and in books 19% against 31%, with bad faith nearly four times lower. Fiction makes the unexplaining character the honest one, because withholding is how you write someone who is right and constrained.
Some other texture. Throwaway uses are honest — a “trust me” that occupies a single beat is a con only 11% of the time, while one the writer builds a subplot on is 30%. Long fuses are dangerous: resolved in the same scene, 22% betrayal; resolved scenes later, 54%. And the phrase itself has drifted — the Victorian “you may trust me” construction is 31% of book uses but 20% of film, while the modern parenthetical hedge (“Trust me, you’ll want to take this call”) runs the other way, 26% of film against 17% of books.
Two favourites from the corpus. In Marie Corelli’s The Sorrows of Satan (1895), Satan himself, disguised as Prince Lucio Rimânez, says “Trust me!” to Geoffrey Tempest. Coded a con, high confidence. And of the four characters in the whole set whom the audience already knows to be untrustworthy and who nonetheless don’t betray, three are simply mid-redemption-arc (Dark City, Entrapment, Game 6) — and the fourth is a Lubitsch joke. In Trouble in Paradise (1932) a waiter begs Gaston not to reveal that he gossiped about a robbery. Gaston says “You can trust me.” He keeps that confidence perfectly. He is also the thief.
Worth being plain about the limits. The era/medium confound above is total. Every judgment is a single LLM pass with no second coder and no inter-rater reliability measured, so the headline rests on one judge’s calls. Several cells are small — the repetition result is n=21, and “just trust me” specifically appeared only 4 times, which is why the original question about phrase variants remains genuinely unanswered. The film transcripts are partial (only two-person exchanges survive in the Cornell corpus), so most film judgments leaned partly on the model’s own recall of the movie. And the categories coding what the audience knows are partly circular by construction — a known villain betrays by definition. The one row that isn’t circular is the row carrying the finding, which is the reason to take it seriously and also the reason to want it replicated.
Everything is public and openly licensed: github.com/malcolmocean/nntd.
Crucially, all 798 LLM judgments are cached in the repo as JSON, so you
don’t need an API key, a GPU, or the corpora to reproduce every table on this page.
Clone it and run compare.py and it works instantly, for free. The results are
plain TSV — open them in a spreadsheet and slice by phrase, decade, genre or outcome
however you like. Disagree with a judgment? Every one has a one-sentence justification
citing its evidence, so you can audit them.
The open questions we’d most like someone to take, roughly in order of value:
If you do something with it, Malcolm would love to hear about it — @Malcolm_Ocean.
Corpora: the Cornell Movie-Dialogs Corpus (Danescu-Niculescu-Mizil & Lee, ACL 2011) and Project Gutenberg. Both were chosen partly because they’re lawfully obtainable in bulk, which is not a given in this kind of research.