OtuhoAI

Building an AI that speaks Otuho

AI assistants now work in hundreds of languages, but not in Otuho. We are changing that with the people who speak it: collecting, checking and protecting Otuho so that an open AI model can learn to understand, write and speak it.

You need a phone with WhatsApp. Contributions take a few minutes at a time.

“Ijiara bu eromo arabotie ihuo enne engiyetiere arabotie.”
“Let us dig together the way we eat together.”From Angata iko Inoro, by Cornelius Gulere, translated into Otuho by Peter Akalkal, African Storybook (CC BY 4.0)

Why it matters

A language that isn't in AI systems is left out of search, translation, education tools and voice assistants.

310,000
people speak Otuho as their first language (Ethnologue, 2017)
1,100+
languages that Meta's open speech model can transcribe. Otuho is not one of them.
200
languages in Meta's open translation model, NLLB. Otuho is not one of them.

Progress toward the first model

These are the first milestones in our data plan. The numbers update live as speakers contribute.

Checked translations0 of 5,000

English sentences translated into Otuho and approved by a second speaker.

Example conversations0 of 1,000

Questions and answers written in Otuho. These teach the model to hold a conversation.

Dialects represented0 of 7

So that the model learns every variety of Otuho, not just one.

How contributions become a model

  1. 1

    Speakers contribute

    Native speakers translate everyday sentences, write in their own words and draft example conversations.

  2. 2

    Speakers review

    A second speaker checks every entry. Only approved Otuho is used, so the model learns from good examples.

  3. 3

    We train an open model

    We adapt an existing open-source AI model to Otuho, together with published texts such as the Otuho Bible.

  4. 4

    Speakers test it

    The community judges the results, corrects mistakes, and those corrections become new training data.

The data belongs to the Otuho community

Contributors keep control of what they share. Names and phone numbers are never published. A community governance group decides how the data is used, following the CARE principles for Indigenous data. The code is open source.

Read how we handle your contributions