This is triage, not a full investigation — the goal is to find which parts of a fluent answer deserve your trust and which need evidence. When you want the whole method, how to research a word origin is the deeper guide. Here we only ask: where does generated etymology tend to break?
What a model is doing when it ‘remembers’ a word history
A language model does not look words up. It produces the most plausible continuation: the shape of a good etymology, trained on real ones. When a word’s history is famous — coffee, robot, sandwich — the model has likely absorbed the genuine account and reproduces it passably. When the history is obscure, the same fluent machinery fills the gap with whatever typically comes next.
Real reconstruction is possible and disciplined: the comparative method described in Essentials of Linguistics works from systematic correspondences between related languages. A model imitates the surface of that work — the chains, the italics, the dates — without performing it. Fluent output therefore tells you how a history should look, not that this one has been checked.
Error shape 1: the invented intermediate form
The most dangerous error is a form that looks right but does not exist. Between a known earlier word and the modern one, a generated chain may insert an intermediate — a Middle English or Old French shape that bridges the gap neatly. It fits the pattern; it may simply not be attested anywhere.
The check is cheap: search the claimed form in a dictionary that prints etymologies. If “Middle English cowfee” appears in no reference you can open, the step is unconfirmed — mark it and keep looking. A form that existed leaves traces: dictionaries collect them, and a real Middle English form tends to come with a century and a document attached. Call the form invented only when there is positive evidence of fabrication — for instance when the citation offered for it names a work that does not contain the form, or does not exist at all. The invented intermediates are also the hardest to spot in the wild because they are precisely plausible — built to match the real chain’s rhythm — which is why plausibility is not evidence.
Error shape 2: the smoothed chain
Real etymologies have edges — gaps where evidence is missing, and places where scholars disagree. A generated history tends to sand those edges off, presenting one clean line where the record actually forks or breaks. Merriam-Webster’s etymology discussion notes that “origin unknown” and competing accounts are ordinary parts of the record; a chain with no visible seams should make you more suspicious, not less.
Coffee is the standing example. The Online Etymology Dictionary’s coffee entry runs the last step through Dutch koffie; Merriam-Webster’s entry names two possible immediate sources — Italian caffè or Turkish kahve directly; the American Heritage Dictionary calls English coffee an alteration of Ottoman Turkish qahve influenced by Italian. A generated route that presents one untroubled line to English hides the disagreement — our coffee page keeps the three accounts separate so you can see what smoothing removes, and when dictionaries disagree explains why the forks happen.

- Attested form
- Gap shown honestly
- Later form
- Smooth line hides the gap
Error shape 3: confident dates with no attestation
Generated histories love round confidence: “first used in 1523,” “coined in the 1680s.” Real first-use evidence is narrower. A dictionary date is scoped — Merriam-Webster’s is the earliest record its editors found for a particular sense — and different references legitimately give different years for the same word, as our algorithm page shows with its 1690s and 1811 dates attached to different senses.
So check the date’s footing, not just the date. Which sense does it belong to? Does any reference give that year, and does the reference attach it to the same meaning? A precise year that no dictionary carries is not a detail — it is an unconfirmed claim with a number on it; mark it until a dated citation backs it or a source rules it out.
The two-source check that catches most errors
You do not need an archive to catch most of this. Take the generated chain one step at a time and open two independent references for the word — for English etymology, a general dictionary with etymologies plus a dedicated one such as the Online Etymology Dictionary will do. For each step ask one question: does this reference name the same link?
Keep the results in three columns — this original check sheet shows the shape:
| What the reading claims | What you open | How the check ends |
|---|---|---|
| A form or spelling | Two dictionary entries for the word | Confirmed if a source names it; unconfirmed if all are silent; invented only if evidence shows fabrication |
| A language link (“from X”) | Each reference’s etymology line | Confirmed if named; a real disagreement if they name different links |
| A date or “coined” claim | Each reference’s date line, with its sense label | Scoped to a sense; differing dates usually mean differing senses |
Three outcomes, three verdicts, and they stay the same whatever the claim. A source names the claim: confirmed, with the reference noted. Every source is silent: unconfirmed — mark it, do not trust it, and do not call it wrong on silence alone. Sources name different answers: a genuine disagreement, which is more interesting than a tidy chain — report both rather than merging them. Invented is a fourth verdict and it needs positive evidence of fabrication, not just missing citations. This check fails gracefully: even where you cannot resolve the answer, you end with an accurate map of what is supported.
Two cautions on sources. First, “independent” matters: two websites quoting the same dictionary are one source. Second, watch for lookalike SEO pages that restate an etymology without naming a reference at all — a claim with no source attached is itself a generated-looking answer. Prefer the dictionaries and teaching references linked in this guide’s source list.
When a generated reading is still worth running
The three shapes are the common failures, not an exhaustive list — a generated history can also mislabel a borrowing as shared ancestry, promote a folk tale to the status of record, or give an accurate account of a different word entirely. The two-source check still works in each case, because it never assumes the error is one of the three: it asks every claim to earn its place against a named reference.
None of this makes generated etymology useless — it makes it a question-generator. A Word Says Hi reading is fastest at proposing the structure: which languages might be involved, which forms to check, which meaning shifts to look for, and where it is uncertain. Run a word through the lookup, then apply the two-source check to each step it returns. The reading’s value is the quality of the questions it hands you; the checking is what turns them into answers.
Reserve full trust for what survives the check, and keep the uncertain parts visible when you share the story. The curated examples show what the output looks like when the checking has already been done — every step beside a named dictionary entry. That is the standard to hold a generated reading to.
Sources and further reading
- Merriam-Webster: etymology FAQ — how real dictionaries mark unknown origins and competing accounts.
- Online Etymology Dictionary: coffee — the Dutch-route account used in the worked example.
- Essentials of Linguistics: reconstructing the past — the comparative method a model imitates the surface of.
The three error shapes and the two-source check are original editorial guidance. Sources consulted 30 September 2026.
