Credits & data sources

← Back

Lexicon (lemma, POS, gender, IPA, English glosses)

Per-language dictionary data extracted from the English Wiktionary by Kaikki.org using Wiktextract (Tatu Ylonen et al.). French, English and German each use the full Kaikki dump for their language — French, English, German; Spanish ships a curated bootstrap covering top-frequency lemmas + irregular conjugations.

The German lexicon is additionally filtered by an OpenSubtitles-derived frequency list from hermitdave/FrequencyWords, licensed CC BY-SA 4.0.

Original content © Wiktionary contributors, licensed under CC BY-SA 3.0. Derivative use here remains under the same license.

Ylonen, T. (2022). Wiktextract: Wiktionary as Machine-Readable Structured Data. Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022).

On-demand definitions (Vocab page expand)

Per-lemma definitions and example sentences are fetched live from the English Wiktionary REST API. Content © Wiktionary contributors, CC BY-SA 3.0.

CEFR vocabulary leveling

Word lists tagged by CEFR proficiency level (A1 – C1). For French, FLELex (CENTAL, UCLouvain). For Spanish, a curated set drawn from the Plan Curricular del Instituto Cervantes inventory and SUBTLEX-ESP frequency data. For English, EFLLex (CENTAL, UCLouvain). For German, a curated set drawn from the Goethe-Institut Goethe-Zertifikat Wortlisten (A1 – B1), the Profile deutsch inventory, and DeReWo / SUBTLEX-DE frequency data. Used for educational purposes.

Pintard, A. & François, T. (2020). Beacco-FLELex: A graded lexical resource for French foreign learners. Proceedings of LREC 2020.
François, T., Gala, N., Watrin, P. & Fairon, C. (2014). FLELex: A graded lexical resource for French foreign learners. Proceedings of LREC 2014.

Sentence-aligned translation

When configured, sentence-by-sentence translations from the target language into your native language (English or Chinese) are produced by DeepL rather than the article-generation LLM. Stored alongside each article so the Write practice page can show the exact native-language sentence for the cloze you're filling.

Article generation

Reading passages are generated on demand by large language models (Google Gemma 4, OpenAI GPT-OSS, Anthropic Claude, etc.) selected by the user in Settings. The shared default uses Google AI Studio's free tier.

Pronunciation

Audio playback uses the browser's built-in Web Speech API. No audio data is shipped or stored — your browser's TTS voice for the target language synthesizes speech locally on each click.

Software

Mille-Mots itself

This app is open source under the MIT License. Source code on GitHub.

© 2026 Bingji Guo. Pull requests welcome.