Articles
Twenty-Two Million Speakers, Sixty-Two Million Words
Somali is listed in the world’s major language models and effectively absent from them. China faced a version of this question a century ago and answered it before it had any technology to protect. That decision is the most useful thing Somalia can learn from Beijing, and it costs nothing.

Mogadishu,(SONNA):When Facebook researchers built XLM-R, one of the multilingual language models that underpins a great deal of subsequent work in the field, they trained it on more than two terabytes of text drawn from the open internet and covered one hundred languages. Somali was among them. Somalia can therefore be said to be represented in one of the foundational multilingual models of the modern era, and officials who wish to say so are not lying.
The share of that training data allocated to Somali was approximately 0.4 gigabytes, around sixty-two million tokens.
Somali is spoken by more than twenty-two million people. To put the figure in proportion, sixty-two million tokens is roughly the volume of text a well-resourced newsroom produces in a few years. It is a rounding error against two terabytes. Somali was not excluded from the model. It was included at a scale that guarantees the model cannot do very much in it, which is a more subtle and more damaging condition, because it permits everyone involved to record the language as covered.
This is the actual state of Somali in artificial intelligence, and it is worth stating plainly because the alternative framings are both wrong. Somali is not absent from the technology. Nor is it adequately served by it. It occupies the position that specialists call low-resource, a term that sounds technical and describes something quite concrete: there is not enough written Somali in machine-readable form for a model to learn the language properly, and so every tool built on those models will work in Somali the way a foreigner speaks after six months of study.
What Somalis have already built
The assumption that follows naturally from this, and that I have heard in Mogadishu more than once, is that Somalia has not begun. That assumption is wrong, and correcting it is the most useful thing this article can do.
Somali researchers have been working on this problem for several years, in universities across the diaspora and increasingly at home, and the output is more substantial than the country realises. Shafie Abdi Mohamed and Muhidin A. Mohamed have published work on lexicon and rule-based lemmatisation for Somali, the unglamorous foundational task of reducing words to their base forms without which nothing else functions. Jama Hussein Mohamud, working through the African Master of Machine Intelligence programme, contributed to automatic speech recognition research covering Somali alongside Wolof and Ga. Badel and colleagues built a Somali information retrieval corpus. A team published SomBERTa this year, a transformer-based model for detecting fabricated news and toxic messages in Somali text, which is a matter of some interest to anyone who has watched Somali social media during an election season.
Most striking is recent work on Somali dialect identification. A study published last year assembled and annotated a dataset of 3,011 Somali text samples drawn from social media and formal documents, and trained models to distinguish between Maxaa Tiri and Maay. The reported classification accuracy exceeded ninety-nine per cent. That result matters for a reason beyond the technical: it demonstrates that when Somali data is carefully assembled by people who understand the language, the models work. The obstacle has never been that Somali is somehow unsuited to computation. The obstacle is that nobody has been paid to write it down at scale.
What is missing is not talent and not effort. It is coordination and it is a corpus. One survey of the field states the position bluntly: no monolingual corpus for Somali exists. Individual researchers assemble their own datasets for their own projects, publish, and move on. The next researcher starts again. That is the pattern that keeps a language low-resource indefinitely, and it is entirely solvable.
The comparison inside our own region
There is a second measurement worth making, and it is uncomfortable.
Kiswahili, the working language of the East African Community and the most widely spoken African language, appears in the major benchmark datasets that researchers use to test whether models actually work: FLORES-200, CCAligned, PanLex, and the No Language Left Behind corpus. Somali appears in some of these and is absent from others, including FLORES-200. The practical consequence is that a researcher can measure how well a model handles Kiswahili and cannot, with the same rigour, measure how well it handles Somali.
This has nothing to do with the relative worth of the languages and everything to do with institutional history. Kiswahili has had national language academies, a regional commission, standardised orthography enforced through school systems, and decades of published material accumulating in libraries and on the internet. Somali has had a written orthography only since 1972 and spent much of the period since 1991 without functioning state institutions of any kind.
The gap is therefore explicable. It is not acceptable, and it is not permanent. It is a measure of accumulated institutional work, which means it responds to institutional work.
Three institutions, eight weeks
Something began this year that has not been noticed outside the country, and it deserves recording.
On 1 June, the National Communications Authority, working with the Academy of Science, Culture and Literature and the Regional Somali Language Academy, launched a Somali cybersecurity terminology glossary. The purpose was to standardise Somali terms so that laws, regulations, academic materials and public awareness campaigns could be written in Somali rather than translated into it after the fact. Professor Abdalla Omar Mansur of the Language Academy stated the principle with precision: terminology is what allows a language to carry scientific and technical knowledge, and new terms must be built through a standardised process rather than assembled by literal translation.
On 16 July, the Somali National University launched the country’s first National Artificial Intelligence Centre, naming natural language processing among its declared research priorities. It named both the centre and the field itself in newly coined Somali rather than borrowing the English words, which is a small decision that reveals a large one.
In the same month, Somalia delivered its national statement in Kiswahili at the East African Kiswahili Commission conference in Bujumbura, and the theme under discussion was the impact of artificial intelligence on the preservation of indigenous languages.
Three institutions, working separately, arrived at versions of the same conclusion within eight weeks. A country cannot absorb a technology it can only discuss in another language. Whether this becomes a coordinated national effort or three unconnected events is a decision that has not yet been made, and it is the decision this article is written to influence.
Why China understands this argument better than most
China faced a version of this question in the late nineteenth and early twentieth centuries, when Chinese intellectuals debated openly whether the language could carry modern science at all. Some serious figures proposed abandoning Chinese characters entirely in favour of a romanised script, on the argument that the writing system itself obstructed technical education. That argument lost. Instead China standardised the spoken language, simplified the script, built systematic technical terminology, and made the deliberate choice to conduct scientific and engineering education in Chinese at every level from primary school through doctoral research.
The consequence is that a Chinese engineer today thinks about her field in her own vocabulary rather than translating in her head, and that scientific knowledge accumulates in Chinese rather than passing through a foreign language on its way in and out. That decision was made when China had no technological standing whatsoever to protect, which is precisely why it worked. It is one of the quieter foundations of everything that followed, and it is almost never mentioned in accounts of Chinese development that concentrate on factories and infrastructure.
Somalia is at the same decision point, at a comparably early stage, and appears to be making the same choice. That is worth saying to a Chinese audience, because it is a form of learning from China that costs nothing, requires no agreement and no financing, and reflects genuine understanding of how that country actually developed rather than admiration for its skylines.
There is also a practical dimension. Chinese laboratories have invested heavily in multilingual capability and have released open-weight models that developers across Africa have found handle African languages more capably than several better-known alternatives, a discovery generally made by testing rather than by being told. Those models can be downloaded and adapted by institutions that control their own infrastructure. For Somali language work, which will be done by Somali institutions on modest budgets, the difference between adapting a model you hold and renting access to one you do not is the difference between a research programme and a subscription.
What should actually be built
A national Somali corpus is the first and most important. Somalia has decades of broadcast archives, newspaper files, government records, published literature, poetry and religious scholarship sitting in physical and semi-digital form. Assembled, cleaned, licensed and released openly, this would be the single most valuable technical asset the country could create for its own language, and it would immediately serve every researcher currently building datasets from scratch. It is a digitisation and archiving project more than a computing one.
A speech corpus is the second, and for Somali it may matter more than text. This remains substantially an oral culture, a great deal of the most important Somali material exists as recorded speech rather than as writing, and a large share of the population would use a voice interface in preference to a typed one. Recording, transcribing and releasing thousands of hours of Somali speech across both major dialects is achievable with modest equipment and a great many willing participants.
Terminology work should continue and expand well beyond cybersecurity, into medicine, engineering, law, finance and computing, on the model the Language Academy has already established and using the standardised process Professor Mansur described.
Benchmarks must be built, because a language that cannot be measured cannot be improved. Somali needs its own evaluation datasets so that Somali institutions can test whether any given model actually works in Somali rather than accepting a vendor’s claim that the language is supported.
And the diaspora should be organised rather than admired. The Masakhane network across Africa produced translation datasets and benchmarks for more than thirty African languages by involving contributors who had no formal training in machine learning, on the principle that native speakers are the scarce resource and technical training is the abundant one. Somalia has a diaspora of millions, concentrated in countries with excellent universities, many of whom would contribute to a serious national language effort for nothing more than the knowledge that it existed and was competently run. What has been missing is an address to send the work to. The National Artificial Intelligence Centre could be that address.
The window
Language technology has a characteristic that most technology does not. It compounds, and it compounds in one direction. Every corpus assembled makes the next model better, every model makes the next tool more useful, every tool generates more text and speech that feeds the next corpus. Languages that enter that cycle move up and stay up. Languages that do not enter it fall further behind each year, not through anyone’s hostility but through simple accumulation elsewhere.
Somali is at the edge of that cycle now. The researchers exist, the institutions have been created, the terminology work has started, and the models available to a country with limited resources are better and cheaper than they have ever been. What is absent is the corpus and the coordination, and both are within Somalia’s own power to supply. Neither requires a foreign partner, a large budget or a functioning national grid.
Twenty-two million people speak this language. Sixty-two million tokens is what the world’s major models were given to learn it from. Somalia can change the second number without asking anyone’s permission, and until it does, every tool built anywhere in the world will speak Somali badly.
About the author
Abdiqani Abdullahi Ahmed is Senior Advisor for Communication and Analysis at Somalia's Ministry of Information, Culture and Tourism. He is Somalia's national focal point to the East African Kiswahili Commission, a juror for the IGAD Media Awards, and lead facilitator of the IGAD Youth Peace and Security Series. He writes here in a personal capacity.



