Kurdish-first is an engineering decision, not a slogan

7 min read

What it actually takes to tune speech and recognition for Sorani, Badini, and Iraqi Arabic — diacritics, code-switching, and the text normalisation nobody warns you about.

Translated-last is translated-badly

Most speech platforms treat Kurdish as an edge case bolted onto an English pipeline. We start from it. The text normalisation, the pronunciation rules, and the evaluation sets are written for Sorani first — Arabic and the rest ride on rails that already work.

The normalisation nobody warns you about

Real Kurdish text mixes scripts: Sorani in Arabic script, Kurdish in Latin, product names in either, digits in both. Numbers, dates, and abbreviations each need their own rules per script. Getting a price or a phone number read correctly is unglamorous work, and it is the difference between a demo and a product.

Ten languages of listening

Recognition covers ten languages — Sorani, Badini, Iraqi Arabic, Arabic, Turkish, English, Spanish, French, Persian, and Chinese — because speakers code-switch mid-sentence. The agent’s reply languages stay Kurdish and Arabic first, tuned until a caller in Erbil or Baghdad hears their own register, not a translated flat line.

How we keep ourselves honest

Every engine change runs against evaluation sets recorded by native speakers, in every supported language. If a tuning improves Arabic and quietly breaks Sorani, the build stops. The languages that built us do not regress for a benchmark.

Start with a script. Leave with a sealed file.

Free studio credits, no card, no commitment.

Create your account