Skip to content
Back to Blog
NLP

Building NLP Pipelines for Swahili: Challenges and Solutions

Josephat Nyambura · Apr 2025 · 8 min read

Try running a standard sentiment classifier on a customer complaint written half in Swahili, half in English, with some Sheng mixed in for good measure — the kind of message that shows up constantly in East African customer support queues — and watch it fall apart. Most pretrained NLP models are trained overwhelmingly on English text, and Swahili, despite being spoken by well over 100 million people, is still a low-resource language in the NLP sense: fewer labeled datasets, fewer pretrained embeddings, less tooling. That gap is real, but it's not insurmountable, and the fixes aren't exotic — they just require treating code-switching as the norm rather than an edge case.

Why Swahili is still low-resource, despite the numbers

Speaker count and NLP resourcing don't track each other. Swahili has a huge number of speakers across East Africa, but comparatively little of that shows up as clean, labeled training data, and most general-purpose embeddings and tokenizers were built with English or other high-resource European languages as the primary case. Tools that work well out of the box for English quietly assume properties — spelling consistency, monolingual input, abundant labeled examples — that don't hold for Swahili in the real world.

Code-switching isn't an edge case — it's the default

Real customer messages in East Africa rarely stay in one language for a full sentence. Swahili, English, and Sheng mix within the same message, sometimes the same clause. Models trained on clean, monolingual corpora — the kind most public Swahili datasets are built from — fail on this in production, because the training distribution doesn't match the deployment distribution. Treating code-switched examples as noise to filter out of a training set is a common mistake, and it throws away a large share of the data that actually looks like what the model will see in production.

Morphology makes tokenization harder than it looks

Swahili is agglutinative — prefixes and suffixes stack onto a root to carry tense, subject, object, and negation all in a single word. A tokenizer built around English whitespace-and-punctuation conventions, or a subword vocabulary trained mostly on English text, segments Swahili words in ways that don't line up with anything linguistically meaningful. This shows up downstream as degraded performance that's easy to misdiagnose as a modeling problem when it's actually a tokenization problem.

Where usable labeled data actually comes from

Building a labeled dataset from scratch for a low-resource setting is slow if done naively. Active learning changes the economics: label the examples the model is currently most uncertain about first, rather than a random sample, and a small labeling budget goes much further. Weak supervision — bootstrapping initial labels from keyword rules or heuristics, then refining — is a reasonable way to get a first pass of training data quickly. And annotator disagreement, especially on code-switched or ambiguous text, is worth treating as information about genuine ambiguity rather than noise to be resolved by majority vote and forgotten.

What actually works: fine-tuning multilingual models

Starting from a multilingual pretrained model — something like mBERT or XLM-R — and fine-tuning it on Swahili and code-switched examples consistently outperforms training a Swahili-only model from scratch on the amount of local data most teams realistically have. The key detail is in the fine-tuning data itself: deliberately including code-switched examples in the fine-tuning set, rather than only clean monolingual Swahili, is what makes the resulting model behave sensibly on real production text instead of just benchmark text.

Evaluation has to reflect real text, not benchmarks

Most public Swahili NLP benchmarks are built from formal, monolingual sources like news articles, which look nothing like a customer support ticket or a WhatsApp complaint. A model that scores well on a news-based benchmark can still perform poorly in production. Building even a small, representative evaluation set from real, anonymized support messages is more useful than chasing a leaderboard number that doesn't reflect the actual deployment distribution.

Keep a human in the loop

For genuinely ambiguous, heavily code-switched, or Sheng-dense messages, the right move is routing to human review rather than forcing a confident-but-wrong automated classification. This isn't a workaround specific to Swahili — it's standard practice for production NLP in any low-resource setting — but it matters more here because the gap between benchmark performance and real-world performance tends to be wider than teams expect going in.

None of this is exotic. It's mostly a matter of designing the pipeline around how people actually write, instead of how a benchmark assumes they write.