July 13, 2026

Before Training, the Data !

Written by:
Godwin Houdji


Building our linguistic dataset was a process of trial, refinement, and continuous learning. From selecting relevant phrases to ensuring accurate translations and clear recordings, every step came with its own challenges and adjustments. What started as a straightforward task quickly became an iterative process, shaped by real-world obstacles and practical solutions. This article takes you through our journey, detailing how we developed our corpus, translated it into Bamanankan, validated the data, and finally recorded high-quality audio.

Building the Corpus

🔗


The first step was to produce the textual data that would eventually be translated and vocalized.

To ensure the dataset was relevant, we based our phrases on real interactions with our AI system within the app. We analyzed conversations and extracted key topics to generate natural, contextually rich sentences.

Our goal was to create 10 hours of content, ensuring a diverse and comprehensive dataset.

French to Bamanankan Translation

🔗


We developed a custom platform to streamline the translation process while maintaining quality and efficiency. Translators see a sentence in French and enter the corresponding translation in Bamanankan. If a sentence is too difficult, they have the option to skip it and move on to the next one. To ensure consistency and prevent fatigue-related errors, we capped translations at 100 sentences per translator per day, after noticing that accuracy declined when translators worked on too many batches in a single session.

Figure: Screenshot of the translator's interface for adding translations.

Once a batch is submitted, it goes to a supervisor for validation. The supervisor randomly checks 25% of the sentences in each batch. If even one sentence is incorrect, the entire batch is rejected and sent back for correction with comments. This cycle repeats until the batch meets our quality standards. Our key approach is having native speakers handle translations, with rigorous expert validation.

Vocalizing the Translations

🔗


After translation, the next step was recording high-quality audio for each sentence.

Recording Process

Narrators see the translated Bamanankan sentence alongside the original French text to provide better context and ensure accurate pronunciation and intonation. Using our platform’s built-in recording feature, they capture their voice directly within the system. To help them produce high-quality recordings, we provided a detailed guide outlining best practices for sound clarity, volume consistency, and proper enunciation.

If an audio file does not meet our quality standards, the supervisor rejects the batch with comments, allowing the narrator to re-record and make necessary improvements. With the goal to deliver 10 hours of high-quality recorded speech from native speakers.


Learnings & Challenges

🔗

Our journey wasn't without obstacles. Here are some of the biggest lessons we learned:

Community Participation vs. Linguistic Expertise

Initially, we aimed for a community-driven approach. for both translation and vocalization. We launched an open call, inviting participants through a selection form that assessed their proficiency in reading and writing Bamanankan, their spoken dialect, and their gender to ensure balanced representation. However, despite this structured selection process, we encountered major quality issues as well as a lack of efficiency. Translations were inconsistent, and audio recordings often contained background noise or unnatural intonations. Another challenge was the slower-than-expected pace of work, which did not align with our production timeline.

These challenges made it clear that maintaining high-quality data required linguistic expertise rather than open participation. As a result, we shifted to working with a dedicated group of Bamanankan linguists to ensure the accuracy and consistency of both translations and recordings.

Fatigue Affects Quality

Initially, translators could complete as many sentences as they wanted per day. However, we observed a decline in accuracy after a certain threshold. Limiting batches to 100 sentences per day significantly improved consistency.

Context Matters for Natural Translations

Working with a low-resource language like Bamanankan, we found that intonation is crucial for meaning. The word ja, for example, can mean "to dry", "to petrify" or "the shade of a tree", depending on pronunciation and context. Without the original French text, narrators can often misinterpret meaning, leading to unnatural phrasing. Adding the French sentence during recording helped them adjust intonation and ensure accurate vocal delivery.

Validation Efficiency

The random validation method quickly highlights systematic errors, enabling prompt corrective action. We chose this method because a translator may start their day with high accuracy but produce lower-quality translations toward the end. A randomized selection ensures we capture a fair representation of the overall quality.

Interface Improvements

Translator feedback led to notable enhancements in our platform's design, making it more user-friendly and reducing errors.

For example, given the rigorous validation process, translators requested a feature to review their submitted work before sending it for supervisor validation. We implemented an interface where they can see all their translations in one place, edit them, and submit with confidence.

Recording Environment Impacts Sound Quality

Some early audio recordings had background noise or inconsistent volume. Providing a r helped standardize sound quality.


Final Thoughts

🔗


Data collection is an ongoing, iterative process. Every test, mistake, and adjustment has brought us closer to a high-quality dataset that will make AI-powered Bamanankan communication more natural and accessible.

🚀 This is just the beginning—our process will continue evolving as we refine our approach and learn even more!