Lexicon.Word resource with frequency ranking #14
Labels
No labels
area
auth
area
data
area
domain
area
infra
area
stats
area
study
area
tooling
area
ui
duplicate
future
kind
bug
kind
chore
kind
decision
kind
docs
kind
feature
kind
spike
kind
test
prio
blocker
risk
high
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
nickkeers/first-thousand-words#14
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Objective: One row per surface form per language, carrying its corpus frequency rank.
Files:
lib/first_thousand_words/lexicon/word.exSteps:
text,normalized(downcased, unaccented, for matching/typing checks),lemma,part_of_speech,frequency_rank(integer),frequency_count(integer).belongs_to :language.(language_id, text). Index on(language_id, frequency_rank).castmust not includelanguage_id(AGENTS.md: programmatic fields are set explicitly).Verify: an IEx query returns the top 20 English words ordered by rank.
Pitfall: the upstream corpus contains non-word noise tokens (
's,'t, single punctuation,numbers). Filtering belongs in the importer, but the schema must tolerate whatever survives.