Skip to main content
Version: Next 🚧

restore_diacritics

Restore diacritics (thêm dấu) for Vietnamese text written without them.

Usage

from underthesea import restore_diacritics

text = "chung ta co the lam duoc"
restored = restore_diacritics(text)
print(restored)
# "chúng ta có thể làm được"

Function Signature

def restore_diacritics(text: str) -> str

Parameters

ParameterTypeDefaultDescription
textstrThe input text, with or without diacritics

Returns

TypeDescription
strThe text with Vietnamese diacritics restored

Examples

Basic Usage

from underthesea import restore_diacritics

restore_diacritics("toi yeu viet nam")
# "tôi yêu việt nam"

Context Awareness

Ambiguous syllables are resolved from their context using a syllable bigram model with Viterbi decoding:

restore_diacritics("chung ta co the lam duoc")
# "chúng ta có thể làm được"

restore_diacritics("Ha Noi la thu do cua Viet Nam")
# "Hà Nội là thủ đô của Việt Nam"

Mixed Content

Numbers, URLs, emails and syllables that already carry diacritics are kept unchanged:

restore_diacritics("toi co 2 con meo.")
# "tôi có 2 con mèo."

restore_diacritics("email cua toi la test@gmail.com")
# "email của tôi là test@gmail.com"

Notes

  • The first call lazily loads a ~1.5 MB n-gram model (once per process); subsequent calls take well under a millisecond per sentence.
  • This is the lightweight mode requested in issue #766: dictionary + n-gram statistics, no deep learning dependencies. A context-aware AI mode may be added later.