Teaching AI to Speak Languages It Barely Knows: Inside CIPHER
Author: Gaurav Sarkar and Jay Gala | Date: October 1, 2026

Large language models can write essays, explain code, and translate between widely spoken languages. Ask them to use a language with very little digital text, however, and something strange can happen: the answer may look fluent while quietly drifting into a larger neighbouring language.
Our new paper, It Knows More Than It Can Say: Self-Evolving Prompts for Extremely Low-Resource Languages, explores a simple question: can an AI system study its own mistakes and build the language guide it needs?
The paper has been accepted at the ORACLE Workshop at EMNLP 2026. We are excited to share this work with the research community and continue improving it with broader evaluation and native-speaker review.
We developed CIPHER, a system that turns translation errors into linguistic rules, stores those rules in a structured prompt, and even changes the prompt's architecture when the existing structure is not enough. The goal is not to replace speakers or linguists. It is to make the expensive first draft of a language resource much easier to produce - so experts can spend more time reviewing and improving it.
Read the full research paper (PDF)
The hidden problem: fluent in the wrong language
We studied Tulu, a Dravidian language spoken by more than two million people, mainly in coastal Karnataka. Tulu has a rich spoken tradition but a relatively small digital footprint. That makes it difficult for modern AI systems to learn from the web at the scale available for languages such as English or Hindi.
Because Tulu and Kannada share geography, scripts, and parts of their linguistic history, a model asked to produce Tulu may fill gaps with Kannada words and forms. This is called vocabulary contamination. It is a particularly deceptive failure: the output does not look broken. To a non-speaker, it can look convincing.
That distinction matters for language technology. A confident answer in the wrong language is not meaningful inclusion. For learners, teachers, and communities, correctness includes respecting the boundary between related languages.
Prompting can help, but expert time does not scale
Earlier work showed that a carefully engineered, five-layer prompt could dramatically improve Tulu generation without changing the model's weights. The prompt gave the model an identity, a list of forms to avoid, grammar rules, examples, and a final self-check. It reduced contamination from 80% to 5% and raised measured grammar accuracy from 18% to 85%.
The result was impressive. The process was difficult to repeat. Three native speakers documented verb paradigms, the case system, and a detailed list of negative constraints through targeted elicitation, followed by several rounds of manual prompt design. Thousands of languages face a similar data shortage, and most do not have a dedicated team ready to build a custom prompt by hand.
This led to the central idea behind our work: a model may fail to produce the right sentence while still being able to recognise and explain how its attempt differs from a reference. CIPHER turns that gap into a learning signal.
How CIPHER teaches the prompt to evolve
CIPHER stands for Contrastive Iterative Prompt Heuristic Evolution for Resource-scarce languages. One language model plays several roles: translator, error analyst, rule writer, example generator, judge, and prompt architect.
In plain language, the system follows a repeating cycle:
- Try a translation. The model translates a sentence using its current structured prompt.
- Compare it with a reference. Instead of receiving only a score, the model examines the exact difference between its answer and a community-verified translation.
- Diagnose the error. It labels the problem as contamination, morphology, syntax, register, vocabulary, meaning, or a custom category, then records the wrong form, the correct form, and the likely rule.
- Update the right part of the prompt. A contamination error becomes a negative constraint. A grammatical error becomes a grammar rule. A vocabulary or meaning error can become a worked example.
- Test a better structure. An outer loop changes layer order, formatting, and token budgets. If recurring errors do not fit anywhere, it can propose an entirely new type of layer.
The system also generates targeted practice examples for its most frequent mistakes. Those examples enter the prompt only after passing a contamination check and a three-judge quality vote. Candidate prompts compete on three goals at once: better grammar, less contamination, and fewer prompt tokens.
The result is not a mysterious weight update. It is a compact document of rules, examples, and checks that a person can read and audit.
What happened in the Tulu experiments?
We used publicly released, community-verified Tulu resources, including 250 English-Tulu translation pairs, a basic vocabulary list, Kannada-to-Tulu constraints, a dictionary, and a short grammar sketch. No native speaker manually curated CIPHER's evolving prompt during the experiments.
Across the search, CIPHER mined 3,471 typed error records. About 36% were contamination errors; the rest captured recurring problems in morphology, syntax, vocabulary, and meaning. Instead of keeping thousands of isolated corrections, the system clustered repeated mistakes into reusable rules.
| Model and setup | Contamination before | Contamination after CIPHER | Other measured change |
|---|---|---|---|
| GPT-5, full evolution | 8% | 0% | Grammar accuracy rose from 28% to 44% |
| Llama-3.2-3B, reduced budget | 46.7% | 0% | The model's own grammar judging was unreliable at this scale |
| Qwen2.5-72B, paired comparison | 40% | 20% | Character-level translation similarity rose from .155 to .247 |
The clearest pattern was that contamination - a lexical problem that can be checked directly - improved across model sizes. Grammar gains depended much more on the model's ability to analyse language accurately.
The best prompt contained a layer nobody designed
One of the most interesting results came from the architecture search. The original prompt began with five familiar layers. During the final generation, CIPHER found repeated errors that did not fit cleanly into them and proposed a sixth: a Lexical Distinction and Contamination Filter.
This new layer stored confusable forms and decision rules, including distinctions between existential constructions. In a paired run against a candidate with the same parent but no discovered layer, the augmented prompt moved contamination from 24% to 0% and grammar accuracy from 16% to 44%. This is promising rather than conclusive - the inner optimisation paths also diverged - but the winning architecture was one no human had specified in advance.
The final prompt used 2,345 tokens. That is small enough to inspect, edit, and discuss with speakers, unlike knowledge hidden across billions of model parameters.
Why this matters to Indilingo
Indilingo exists to help people learn a language from the language they already know. Building that experience for widely represented languages is already hard. Building it for languages with little text, few benchmarks, and limited digital tooling is a much larger challenge.
CIPHER suggests a practical path for creating language-specific AI guidance from small, verified resources. A system can draft explicit rules and examples, reveal which forms it confuses, and package that knowledge for review. This could help teams explore support for more languages without pretending that an AI model is already an authority on them.
That last point is essential. Language belongs to its speakers. Machine-grown grammar can contain mistakes, flatten dialect differences, or amplify a bad example. The right workflow is not "AI replaces the expert." It is "AI prepares an auditable draft; speakers and linguists verify it."
What this study does not prove
We want to be clear about the limits of the evidence:
- The experiments cover one language, with only 15 to 25 evaluation pairs per run and a single seed.
- The same fixed development pairs guided improvement and evaluation. The results measure closed-loop knowledge extraction, not performance on a fully held-out test set.
- Grammar and fluency were judged by the model being improved. Those judgments were visibly unstable for the 3B model, so the paper relies more heavily on judge-free contamination and character-similarity metrics.
- Some mined rules may be noisy or wrong. Native-speaker review and held-out testing are the most important next steps.
- The expert-built prompt reported 85% grammar accuracy, compared with 44% in our best run. These figures come from different models and evaluation setups, so they are not directly comparable, but they show that expert knowledge remains the standard to aim for.
From expert-weeks to an expert-reviewable first draft
The full GPT-5 search took about 7.1 hours of wall-clock time. What it produced was more valuable than a higher score alone: a readable linguistic artifact that records what the model learned, where that knowledge belongs, and how it should be checked at generation time.
There is still a long road from this experiment to dependable language technology. We need held-out evaluation, native-speaker audits, stronger comparisons, and tests across many language families. But the direction is encouraging. An AI model that cannot reliably speak a language may still know enough to help build a better guide - if we give it references, a structure for learning from failure, and a process that keeps its claims open to human inspection.
Research by Gaurav Sarkar and Jay Gala.
Read It Knows More Than It Can Say: Self-Evolving Prompts for Extremely Low-Resource Languages (PDF)
Explore the original research
Open the complete six-page paper, including the CIPHER architecture, experiment tables, limitations, and references.

