mySpellChecker
syllableကျောင်းသာ→ကျောင်းသား

43 million speakers.Not a single proper writing tool.Let me put a stop to this.

Detect, correct, and understand Myanmar text. From syllable rules to AI semantics. Build your own dictionaries. Train your own models. One library that handles it all.

docs →github →
●●●terminal
$

Why it's hard

Sixteen reasons most spell checkers give up.

Most languages use spaces. Myanmar doesn’t.

English
I·love·Myanmar

Words separated by spaces. Easy to tokenize.

Myanmar
ကျွန်တော်မြန်မာနိုင်ငံကိုချစ်တယ်

No spaces. No word boundaries. Where does one word end?

Most spell checkers assume whitespace = word boundary. Myanmar text is a continuous stream of characters with no delimiters. You can’t even begin checking until you figure out where one word ends and the next begins. myspellchecker starts from syllables — the smallest reliable unit — and builds up.

So .. how?

You can't check what you can't split. So I split smaller — then built up, layer by layer, until nothing gets through.

LAYER 1 · Syllable Validation
22 structural rules, no dictionary needed
catches: Invalid syllable structures, medial ordering & compatibility, tone mark errors, virama & stacking issues, kinzi patterns, vowel exclusivity, diacritic uniqueness, Great Sa rules, particle typos, corruption detection, zero-width cleanup
↓
LAYER 2 · Word Validation
SymSpell O(1) lookup + neural reranker (20-feature ONNX MLP)
catches: Out-of-vocabulary words, compound word resolution, ambiguous segmentation, morphological root recovery, colloquial variant detection, phonetic similarity matching, Zawgyi detection & conversion
↓
LAYER 2.5 · Grammar Rules
POS tagging + 6 grammar checkers (YAML-driven)
catches: Register mixing (colloquial/formal), aspect marker errors, classifier-noun agreement, negation pattern mismatches, compound formation, merged word/particle detection, reduplication patterns
↓
LAYER 3 · Context & AI
10 validation strategies, N-grams, custom-trained AI models
catches: Homophones, confusable variants (semantic via MLM), broken compounds, question structure, POS sequence errors, n-gram improbability, tone ambiguity, orthography errors, semantic validation, NER name protection
// what's inside each layer?
▼

The full toolkit

One box. Every tool you'll ever need.

$ tree myspellchecker --strategy
myspellchecker/
[-]checking/
syllable-validation (22 rules)[+]
symspell-lookup + neural reranker[+]
grammar-checkers (6)[+]
pos-tagging (3 backends)[+]
ner (2 backends)[+]
ngram-context[+]
10 validation strategies[+]
[-]text-processing/
word-segmentation (3 engines)[+]
normalization[+]
zawgyi-detection & conversion[+]
phonetic-hashing[+]
morphology-analysis[+]
[-]dictionary-pipeline/
corpus-ingestion[+]
ngram-frequency-building[+]
enrichment[+]
incremental-builds[+]
providers (4)[+]
[-]ai-training/
semantic-model (train-model)[+]
neural-reranker[+]
onnx-export[+]
[-]api/
check, check_batch, check_async[+]
streaming-checker[+]
builder-pattern[+]
factory methods & presets[+]
[-]cli/
check[+]
build[+]
train-model[+]
segment, infer-pos, config[+]
[-]performance/
cython-extensions (11)[+]
openmp-parallelization[+]
connection-pooling & caching[+]
[-]i18n/
bilingual-errors (en/my)[+]
8 directories, 31 modules

Benchmarks

Tested on macOS Apple Silicon against 489 hand-annotated sentences. WORD-level validation with a 565MB production dictionary (601K words, 21K syllables, 2.2M bigrams, 808K trigrams, 504K fourgrams, 634K fivegrams) — built using myspellchecker's own dictionary pipeline and AI training tools.

Rules + SymSpell + grammar checkers + n-grams — no semantic model
0.0%F1 Score
0.0%Recall
0.0%Precision
0.0msMean Latency
COMPOSITE SCORE················································································0.9249
TOP-1 ACCURACY················································································85.2%
MRR················································································0.8731
P95 LATENCY················································································58.8ms
FALSE POSITIVES················································································10 / 489
FALSE NEGATIVES················································································25 / 470

Fast and accurate for most use cases. Catches 94.7% of errors at 11.8ms mean latency with better suggestion ranking. Dictionary and AI models are not bundled — you build them with your own data using the built-in pipeline.

Get started

$ pip install myspellcheckerRead the docs →
bash
$ pip install myspellchecker
# With AI inference (ONNX + tokenizers)
$ pip install "myspellchecker[ai]"
# With Transformer POS/NER backends
$ pip install "myspellchecker[transformers]"
# Full AI stack (ai + transformers)
$ pip install "myspellchecker[ai-full]"
# For dictionary building pipeline
$ pip install "myspellchecker[build]"
# For model training
$ pip install "myspellchecker[train]"
// see it in action
▼

What you can build

A spell checker that catches what others miss — from broken syllables to missing punctuation.

●●●
sarsit.app
Myanmar Spell Checker
ကျွနတော် မနက်ဖြန် ဆာမေးပွဲ ဖြေရမယ်
Built with myspellchecker

sarsit · စာစစ်

coming soon

sarsit is a spell checking product for Myanmar — web app, mobile, and browser extensions. Powered by myspellchecker under the hood. The library handles the hard parts: syllable segmentation, dictionary lookup, context validation, and grammar checking. sarsit wraps it in a user-friendly experience.