Lexicon.Word resource with frequency ranking #14

Open
opened 2026-09-26 13:58:32 +00:00 by belvedere · 0 comments
Collaborator

Objective: One row per surface form per language, carrying its corpus frequency rank.

Files:

  • Create: lib/first_thousand_words/lexicon/word.ex

Steps:

  1. Attributes: text, normalized (downcased, unaccented, for matching/typing checks),
    lemma, part_of_speech, frequency_rank (integer), frequency_count (integer).
  2. Relationships: belongs_to :language.
  3. Identity: unique on (language_id, text). Index on (language_id, frequency_rank).
  4. cast must not include language_id (AGENTS.md: programmatic fields are set explicitly).

Verify: an IEx query returns the top 20 English words ordered by rank.

Pitfall: the upstream corpus contains non-word noise tokens ('s, 't, single punctuation,
numbers). Filtering belongs in the importer, but the schema must tolerate whatever survives.

**Objective:** One row per surface form per language, carrying its corpus frequency rank. **Files:** - Create: `lib/first_thousand_words/lexicon/word.ex` **Steps:** 1. Attributes: `text`, `normalized` (downcased, unaccented, for matching/typing checks), `lemma`, `part_of_speech`, `frequency_rank` (integer), `frequency_count` (integer). 2. Relationships: `belongs_to :language`. 3. Identity: unique on `(language_id, text)`. Index on `(language_id, frequency_rank)`. 4. `cast` must not include `language_id` (AGENTS.md: programmatic fields are set explicitly). **Verify:** an IEx query returns the top 20 English words ordered by rank. **Pitfall:** the upstream corpus contains non-word noise tokens (`'s`, `'t`, single punctuation, numbers). Filtering belongs in the importer, but the schema must tolerate whatever survives.
Sign in to join this conversation.
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
nickkeers/first-thousand-words#14
No description provided.