words

900words · research · Danish

Where the numbers come from

Every figure on this site was counted, never estimated. This page is the full study: the three conversation corpora, the exact matching rules, the results, what they do not show, and every source. All the data is downloadable below.

87.2%

Equal-corpus mean, family-aware, across three frozen Danish conversation samples.

3independent conversation corpora
901,076recorded conversation words measured
85.7 – 89.0%family-aware coverage range
66.1 – 69.2%exact-surface-form range

the question

What can an honest “900 words” claim mean?

Plenty of courses promise a percentage. Very few say what was counted. This study asks one narrow question: if you select 900 Danish word forms by how often they actually occur in conversation, what share of the words in real recorded conversations do they cover?

Coverage means a word in a transcript matches a word in the inventory. It is the honest ceiling for what a vocabulary can touch. It is not comprehension, not listening ability, and not fluency. The claim boundary below spells out exactly what this study does and does not show.

the result

Measured coverage, corpus by corpus

The same frozen token rule and the same matching policy were applied to every sample. No pooling: each corpus keeps its own numerator and denominator.

Danish Gigaword · spont/train

639,485 tokens

85.7%

Spontaneous and pseudo-spontaneous two-person conversations. Family-aware 85.746%, exact form 66.069%.

CoRal v3 · conversation/test

95,380 tokens

89.0%

A frozen held-out conversational Danish sample. Family-aware 88.980%, exact form 69.223%.

Dideriksen et al. · Spontaneous

166,211 tokens

86.9%

The spontaneous subset of a 39-dyad Danish conversation archive, kept separate from its task dialogues. Family-aware 86.861%, exact form 68.833%.

The full comparison, including the inventories the headline does not apply to:

InventoryGigawordCoRalDideriksen
Current 900 app card headwords alone, family-aware14.319%14.754%13.939%
Expanded 1,158-form research inventory, family-aware85.788%89.093%86.917%
Frequency-selected 900-form research inventory, family-aware85.746%88.980%86.861%
Frequency-selected 900-form research inventory, exact surface form66.069%69.223%68.833%

The equal-corpus mean of the three highlighted figures is 87.196% (exact-form mean 68.042%). That mean is descriptive. It is not a claim that three samples represent all Danish conversation.

Why “family-aware” and “exact form” differ by twenty points

A learner who was taught købe (to buy) plus the grammar for its endings can be credited when a speaker says købte (bought). The family-aware measure gives that credit for taught inflections. The exact-form measure deliberately does not. The gap between the two is why any public number must name its matching rule, and why ours do.

the inventory

How the 900 were chosen

The candidate pool held 1,158 forms: the 900 card headwords of the app's course, 252 essential support forms (the small machinery of sentences), and six additions approved for the study because conversation is full of them: øh, jamen, ej, ligesom, ting and fald.

Calling 1,158 forms “900 words” would not be honest. So all 1,158 competed on equal terms: each corpus weighted equally, forms ranked by mean occurrences per million running tokens, and the 900 highest-ranked forms kept. No form received special protection.

Final composition of the selected 900

703 current card forms · 196 support forms · all 6 additions. The selection removes 197 card forms and 61 support forms. The exact removal list with every score is in the data files.

The selected 900 is a research inventory. It informs how the course evolves, but it does not automatically rewrite the playable app, and this page never presents a measurement of one inventory as a claim about another.

the words themselves

The most common Danish words

These lists feed the cards on the main page. A word must occur in all three corpora and pass a human check for automatic part-of-speech tagging errors. “Describing words” is deliberate wording: it is more honest than presenting every automatic adjective tag as a strict grammatical category.

Top five overall

  • detit / that
  • jegI
  • jayes
  • then / so
  • eram / is / are

Common nouns

  • gangtime / occasion
  • tingthing
  • tidtime
  • faldfall / case
  • åryear

Common verbs

  • værebe
  • havehave
  • sesee / look
  • trobelieve / think
  • sigesay / tell

Describing words

  • sådansuch / like that
  • okayokay
  • godgood
  • mangemany
  • alall / every

The single most common word, det, is about 6% of everything said. The glosses are display aids; Danish words often carry more than one sense. Raw rankings with per-corpus counts are in the data files.

the method

How the counting works

The whole audit holds one denominator rule and one matching policy fixed across every sample. Aggregate counts are retained; raw transcript text is not.

  1. Count every running Danish letter-token in the frozen transcript sample. This is the denominator, and it never changes between inventories.
  2. Match a token if its exact surface form, or its predicted Danish inflectional lemma, is in the inventory. Lemmatisation uses Danish Stanza.
  3. Apply exactly two approved spoken reductions, because Danes really say them: ha → have and ik → ikke.
  4. Give no credit for anything else. No compounds, no derivations, no semantic relatives, no words outside the inventory.

For the word rankings, each corpus gets equal weight regardless of size, so the largest sample cannot quietly dominate the result. Corpus revisions and inventory files are pinned by hash, which makes every number on this page reproducible.

the evidence set

Which corpora count, and which do not

A corpus enters the headline evidence only if it is genuinely conversational and its licence permits commercial use. Several attractive Danish corpora fail one of those tests, and honesty means listing them too.

SourceStatusReason
Danish Gigaword spont/trainIncludedFrozen spontaneous two-person conversations; commercially usable source pin.
CoRal v3 conversation/testIncludedFrozen held-out conversational sample; text handled only under the accepted access terms, then deleted.
Dideriksen et al. archiveIncluded, labelledCC BY 4.0. Its spontaneous subset is used; its collaborative-task dialogues are kept separate. Participants were largely students, so it is a sensitivity sample, not a census.
Nordic Dialect Corpus, DanPASS, SamtaleBank, NOMCOExcludedGood conversation data, but non-commercial licence terms rule them out of a commercial evidence set.
LANCHART, GaMMANot acquiredAccess-restricted, or not yet established as commercially reusable transcript sources.
Common Voice scripted Danish, Danish Parliament speechContext controls onlyCommercially usable, but read-aloud or formal parliamentary speech is not ordinary conversation.

Excluding a corpus you would love to use is what makes the included ones mean something.

the honest part

What this study does not show

Coverage is not comprehension.

This research does not show that a learner knows a word, recognises it in fast speech, understands its sense in context, can answer back, or has reached any CEFR level. Three samples, however carefully chosen, are not a census of all Danish conversation.

The narrowly defensible formulation, which every claim on this site is written to respect:

“In three frozen Danish conversational samples, a frequency-selected 900-form research inventory and its taught inflections matched 85.7 to 89.0% of running transcript tokens.”

It must not be shortened to “learn 900 words and understand 87% of Danish.” And it must not present the app's current 900 card headwords as if they were the selected research inventory: measured alone, the card headwords cover about 14% (the sentence machinery lives in the support forms, which is precisely why the selection competed both together).

the learning science

Why the game is built on connection puzzles

The coverage study says which words are worth learning. A separate body of memory research says something about how. Three findings, each replicated many times over the decades, shaped the game's design.

These studies are about memory in general, not about this app. We cite them as the design rationale they are. Measuring the game's own learning outcomes is future work, and until it is done we will not pretend otherwise.

the corpus sources

References

check our work

The data, downloadable

Everything below is aggregate data: word forms, counts, ranks and flags. No corpus transcript text was retained anywhere in the pipeline, so none is published here.

Citing this page:

900words (2026). The Danish coverage study: measured spoken-vocabulary coverage of a frequency-selected 900-form inventory. https://900words.app/research/. Published 23 August 2026, updated 27 August 2026.