Spike: pick the translation source for cross-language links #16

Open
opened 2026-09-26 13:58:32 +00:00 by belvedere · 0 comments
Collaborator

Objective: A written recommendation, with measured numbers, for where translations come
from. Time-box: one day. Output is a document, not code.

Candidates to compare (all offline-usable, no runtime API dependency):

  • WikDict (https://download.wikdict.com/dictionaries/sqlite/<release>/) — pre-built
    per-pair dictionaries derived from Wiktionary, released as <src>-<tgt>.sqlite3 for many
    pairs; licence CC-BY-SA. Verify the current release directory rather than trusting this note.
  • Kaikki.org / Wiktextract per-language JSONL
    (https://kaikki.org/dictionary/<Language>/kaikki.org-dictionary-<Language>.jsonl) —
    rich: senses, forms, translations, examples. Large (the English extract was ~3.3 GB on
    2026-09-26), so treat it as an offline batch job.

Measure for each: coverage of the top 1000 words of the target language; how well senses
disambiguate; licence and attribution obligations; total download size; how long a full import
takes; whether we can avoid holding a multi-GB file on the app server.

Steps:

  1. Download a bounded sample of each.
  2. Compute coverage against the frequency list for two languages.
  3. Write docs/decisions/0002-translation-source.md with a recommendation and the numbers.

Verify: the doc contains a coverage table and an explicit recommendation with rationale.

Pitfall: do this before building the importer — the schema details (sense_rank,
confidence) depend on which source wins.

**Objective:** A written recommendation, with measured numbers, for where translations come from. Time-box: one day. Output is a document, not code. **Candidates to compare** (all offline-usable, no runtime API dependency): - **WikDict** (`https://download.wikdict.com/dictionaries/sqlite/<release>/`) — pre-built per-pair dictionaries derived from Wiktionary, released as `<src>-<tgt>.sqlite3` for many pairs; licence CC-BY-SA. Verify the current release directory rather than trusting this note. - **Kaikki.org / Wiktextract** per-language JSONL (`https://kaikki.org/dictionary/<Language>/kaikki.org-dictionary-<Language>.jsonl`) — rich: senses, forms, translations, examples. Large (the English extract was ~3.3 GB on 2026-09-26), so treat it as an offline batch job. **Measure for each:** coverage of the top 1000 words of the target language; how well senses disambiguate; licence and attribution obligations; total download size; how long a full import takes; whether we can avoid holding a multi-GB file on the app server. **Steps:** 1. Download a bounded sample of each. 2. Compute coverage against the frequency list for two languages. 3. Write `docs/decisions/0002-translation-source.md` with a recommendation and the numbers. **Verify:** the doc contains a coverage table and an explicit recommendation with rationale. **Pitfall:** do this before building the importer — the schema details (`sense_rank`, `confidence`) depend on which source wins.
Sign in to join this conversation.
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
nickkeers/first-thousand-words#16
No description provided.