
The Future of AI in Education: Why Multilingual Teaching Will Become the New Standard
The Future of AI in Education: Why Multilingual Teaching Will Become the New Standard

Article by
Milo
ESL Content Coordinator & Educator
ESL Content Coordinator & Educator
All Posts
Somewhere this week, a teacher will paste a progress report into a translation box, read the output, decide it looks right, and send it home. The Spanish will be clean. The grammar will be correct. The tone will be warm and professional. And one term in the middle of it, the term explaining why the child is being referred for assessment, will mean something the teacher never wrote.
That is not a worst-case scenario invented to make a point. It is close to the documented finding of a 2026 peer-reviewed study, and it explains why multilingual teaching is about to stop being a specialism and start being a baseline competency.
The usual story about AI in education is about lesson plans, marking, and getting evenings back. That story is true, and it is also too small. The bigger shift is that language stopped being a budget line. When a translation costs nothing and takes four seconds, the question is no longer whether a school can afford to communicate with a family. The question is whether anyone can confirm what was actually said.
University research teams reached this question years before schools did, because they had no way around it: a mistranslated source that ends up in a published paper is a mistake with a permanent address. Their working answers to how to use AI translation in education are worth borrowing, because the underlying problem is identical. Someone is relying on a translation they cannot personally check, in a document that carries consequences. What follows is what the evidence actually supports, and the workflow it points to.
Still grading everything by hand?
EMStudio is a free teaching management app — manage your classes, students, lessons, and more!
Learn More

Still grading everything by hand?
EMStudio is a free teaching management app — manage your classes, students, lessons, and more!
Learn More

Table of Contents
The multilingual classroom is not a forecast
Multilingual classrooms are already the norm rather than the future. According to the National Center for Education Statistics, in its Condition of Education reporting on English learners in public schools, English learners made up 10.6% of United States public school students in autumn 2021, roughly 5.3 million children, up from 9.4% a decade earlier. The distribution is wildly uneven: 0.8% in West Virginia, 20.2% in Texas.
The global picture is starker. UNESCO, in its Languages Matter guidance on multilingual education, estimates that 40% of people worldwide cannot access education in a language they speak and understand fluently. In some low and middle income countries that figure reaches 90%. More than a quarter of a billion learners are affected.
What makes this a workflow problem rather than a policy problem is the staffing gap underneath it. In the most recent national survey of school psychologists, only 12% of 1,308 practitioners identified as bilingual, and only 7% reported delivering services in a language other than English. That share has barely moved since 2015. The constraint has never been willingness. It is that the people writing the documents usually cannot read them once they are translated.
AI removed the cost barrier, not the accuracy barrier
Teachers have already adopted AI at scale. Gallup and the Walton Family Foundation surveyed 2,232 United States public school teachers and found that 60% had used AI tools for their work in the past year, and that those using them at least weekly saved an average of 5.9 hours a week. Across a school year that is roughly six weeks recovered.
The same survey found that only about one in five teachers works at a school with any AI policy at all. Capability arrived years before governance did, which is normal for classroom technology and usually harmless. Translation is where it stops being harmless.
Every other AI task leaves the teacher able to check the work. A generated rubric can be skimmed. A drafted email can be reread. A translated letter cannot, because the person who needs to verify it is the one person in the room who does not speak the language. Verification is exactly what gets skipped, and the output gives no reason to think it should not be.
The finding every teacher should know about
In a study by Shergill and colleagues, published in the peer-reviewed journal Contemporary School Psychology in June 2026, researchers took a 71-sentence psychoeducational report summary for a fictional third grader and translated it from English into Spanish three ways: with Google Translate, with ChatGPT-4o, and with a certified human translator holding more than two decades of experience in educational translation. Two bilingual graduate raters scored every sentence for fluency and accuracy separately, with disagreements settled by a bilingual school psychologist who did not know which system produced which text.
Two results matter for anyone sending documents home.
First, fluency showed no statistically significant difference across the three. All three read comparably well to a Spanish speaker. Second, the number of errors did differ significantly, and no method produced a perfect translation. Between 80% and 90% of sentences were free of any recorded error, which sounds reassuring until you look at where the remaining errors landed.
They landed on the technical vocabulary. Mistranslations clustered on terms like rapid automatic naming, comprehension knowledge, and phonological processing. In one case a spelling test was rendered as an ELA test. These are not decorative words. They are the words that determine what a parent believes about their child's needs, and they are precisely the words a monolingual reader has no way to spot-check.

Infographic 1. Fluency and accuracy were measured separately, and they did not track each other.
The result that surprised the researchers is worth stating plainly, along with its caveats. The professional human translator recorded the highest observed number of total errors.
The authors are careful about this, and so should we be: the study used one report, one translator, and a sample too small to detect moderate differences reliably. They describe their own findings as exploratory. The takeaway is not that machines have overtaken professionals. It is narrower and more useful.
A credential is not a guarantee, a clean-reading sentence is not a correct one, and the errors concentrate in the same place regardless of who or what produced the translation.
The practical version: fluency carries no error signal. The output that reads best is not the output most likely to be right, and nothing in the interface will tell you the difference.
One system's answer is not a second opinion
Here is the detail that changes daily practice. The three translations in that study did not make the same mistakes as each other. Google Translate logged 8 mistranslation errors, ChatGPT-4o logged 12, and the human translator 19. On grammatical errors the ordering reversed: 7 for ChatGPT-4o, 11 for Google Translate, 21 for the human translator. Different systems, different failure points, identical fluent surface.
That disagreement is not noise. For a teacher who cannot read the target language, it is the only free error signal available. Run the same paragraph through two independent systems and compare. Where both render a term the same way, the agreement is weak evidence but real evidence. Where they diverge, you have found the sentence that needs a human. You have not learned which version is correct, but you have located the risk, and locating the risk is most of the job.
The academic guidance mentioned at the top of this piece turns on the same distinction, and it is the one worth carrying into a staffroom. There is a hard line between using a translation to understand something and using it as a statement you will stand behind in public. The first tolerates error, because the cost of a small mistake is a moment of confusion. The second does not, because the output becomes the record. School communication has exactly the same split running through it, and most schools have never named it out loud.
Sort the document before you translate it
The most useful habit is triage, and it happens before anything is pasted anywhere. Ask one question of the document in front of you: what does it cost if a sentence in this is wrong? Low-stakes text can go straight through. Medium-stakes text should be cross-checked against a second system. High-stakes text needs a qualified human translator, full stop.

Infographic 2. A three-tier triage model for classroom documents.
The top tier is not a matter of preference. Under Title VI of the Civil Rights Act, United States districts must take reasonable steps to give families with limited English proficiency meaningful access to school communication, and federal special education rules require assessment information in the language most likely to yield accurate information. Machine output on its own is not a language-access strategy, and it is not a defence if a placement decision is later challenged.
There is a second question sitting underneath the first, and it arrives before translation quality does. Pasting a student's name, test scores, and diagnostic language into a consumer chatbot is a data-handling decision governed by FERPA. Education-licensed versions of the major platforms exist specifically for this, with data terms that consumer accounts do not carry. Check which account you are signed into before you check the translation.
Write it so it survives translation
The strongest lever a teacher controls is not the translation system. It is the English going into it. Documents written well above a twelfth-grade reading level, dense with jargon and idiom, translated badly across every system tested in that study. Clearer source text produces better translations everywhere, and it also produces better documents for the families who read English.
Cut jargon, or gloss it in plain language the first time it appears.
Short sentences. Active voice. One idea per sentence.
Never make a pronoun carry the meaning. In the study, the sentence about parents who had always had difficulty reading came back describing the child instead of the parents, turning a family history into a finding about the student.
Keep a running glossary of the thirty or so terms your school uses constantly, with an agreed translation for each, so the same concept does not arrive three different ways across three letters.
Spell out acronyms on first use, including the ones that feel universal.
List numbers, dates, names and test titles somewhere they can be checked without reading the prose. Those are the elements a monolingual teacher can verify directly in any language.
What multilingual by default actually looks like
The reason multilingual teaching becomes the standard is not that translation got good. It is that translation got cheap enough to move upstream. It stops being a request routed to the district office three weeks before a meeting, and becomes a property of how materials are built in the first place.
In practice that means four things, none of them technological:
Templates are written for translation from the start, not fixed afterwards.
The glossary belongs to the school, not to whichever bilingual staff member happens to be available.
Every document type carries a risk tier, decided once, so the judgement is not remade under time pressure by whoever is sending it.
Review is a scheduled step with a name attached, not a favour asked of a bilingual teaching assistant between lessons.
That last one matters more than it looks. The research literature is consistent that schools lean on bilingual staff and, worse, on family members, for work those people were never trained or paid to do. Making review an explicit, resourced step is what turns individual goodwill into something a school can rely on when the person who has always quietly handled it moves on.
The teachers who will look prepared three years from now are not the ones who picked the best translation system. They are the ones who decided in advance which documents were allowed to leave the building on a machine's word, and which were not.
Frequently asked questions
Is it acceptable for teachers to use AI translation for parent communication?
For low-stakes communication, yes. Newsletters, reminders and scheduling notes rarely carry consequences if a phrase is imperfect. For anything affecting eligibility, placement, consent or discipline, federal language-access obligations point to a qualified human translator. Treat machine output as a draft that a person still has to approve.
If a translation reads well, does that mean it is accurate?
No. That is the central finding of the 2026 study. Fluency and accuracy were scored separately, and fluency showed no significant difference across systems that produced significantly different numbers of errors. A translation can be grammatically flawless and still misstate a diagnosis, a score, or a recommendation.
How can a monolingual teacher check a translation they cannot read?
Run the same text through a second independent system and compare the two. Where they agree, confidence is reasonable. Where they diverge, flag that sentence for a bilingual colleague. Separately, verify numbers, dates, names and test titles directly, because those are readable in any language.
What should a school do before rolling out AI translation?
Three things. Classify document types by risk before anyone needs them. Choose a platform with education-appropriate data handling, given FERPA obligations around student records. Build a shared glossary of recurring terms. Only about one in five teachers currently works at a school with any AI policy, so a written one is still a genuine advantage.
Does AI translation reduce the need for bilingual staff?
It changes the work rather than removing it. AI absorbs volume in low-risk communication, which frees bilingual staff for the high-stakes review only a person can do. The constraint was never enthusiasm: national survey data shows only 12% of school psychologists surveyed identified as bilingual, a figure flat for a decade.
The multilingual classroom is not a forecast
Multilingual classrooms are already the norm rather than the future. According to the National Center for Education Statistics, in its Condition of Education reporting on English learners in public schools, English learners made up 10.6% of United States public school students in autumn 2021, roughly 5.3 million children, up from 9.4% a decade earlier. The distribution is wildly uneven: 0.8% in West Virginia, 20.2% in Texas.
The global picture is starker. UNESCO, in its Languages Matter guidance on multilingual education, estimates that 40% of people worldwide cannot access education in a language they speak and understand fluently. In some low and middle income countries that figure reaches 90%. More than a quarter of a billion learners are affected.
What makes this a workflow problem rather than a policy problem is the staffing gap underneath it. In the most recent national survey of school psychologists, only 12% of 1,308 practitioners identified as bilingual, and only 7% reported delivering services in a language other than English. That share has barely moved since 2015. The constraint has never been willingness. It is that the people writing the documents usually cannot read them once they are translated.
AI removed the cost barrier, not the accuracy barrier
Teachers have already adopted AI at scale. Gallup and the Walton Family Foundation surveyed 2,232 United States public school teachers and found that 60% had used AI tools for their work in the past year, and that those using them at least weekly saved an average of 5.9 hours a week. Across a school year that is roughly six weeks recovered.
The same survey found that only about one in five teachers works at a school with any AI policy at all. Capability arrived years before governance did, which is normal for classroom technology and usually harmless. Translation is where it stops being harmless.
Every other AI task leaves the teacher able to check the work. A generated rubric can be skimmed. A drafted email can be reread. A translated letter cannot, because the person who needs to verify it is the one person in the room who does not speak the language. Verification is exactly what gets skipped, and the output gives no reason to think it should not be.
The finding every teacher should know about
In a study by Shergill and colleagues, published in the peer-reviewed journal Contemporary School Psychology in June 2026, researchers took a 71-sentence psychoeducational report summary for a fictional third grader and translated it from English into Spanish three ways: with Google Translate, with ChatGPT-4o, and with a certified human translator holding more than two decades of experience in educational translation. Two bilingual graduate raters scored every sentence for fluency and accuracy separately, with disagreements settled by a bilingual school psychologist who did not know which system produced which text.
Two results matter for anyone sending documents home.
First, fluency showed no statistically significant difference across the three. All three read comparably well to a Spanish speaker. Second, the number of errors did differ significantly, and no method produced a perfect translation. Between 80% and 90% of sentences were free of any recorded error, which sounds reassuring until you look at where the remaining errors landed.
They landed on the technical vocabulary. Mistranslations clustered on terms like rapid automatic naming, comprehension knowledge, and phonological processing. In one case a spelling test was rendered as an ELA test. These are not decorative words. They are the words that determine what a parent believes about their child's needs, and they are precisely the words a monolingual reader has no way to spot-check.

Infographic 1. Fluency and accuracy were measured separately, and they did not track each other.
The result that surprised the researchers is worth stating plainly, along with its caveats. The professional human translator recorded the highest observed number of total errors.
The authors are careful about this, and so should we be: the study used one report, one translator, and a sample too small to detect moderate differences reliably. They describe their own findings as exploratory. The takeaway is not that machines have overtaken professionals. It is narrower and more useful.
A credential is not a guarantee, a clean-reading sentence is not a correct one, and the errors concentrate in the same place regardless of who or what produced the translation.
The practical version: fluency carries no error signal. The output that reads best is not the output most likely to be right, and nothing in the interface will tell you the difference.
One system's answer is not a second opinion
Here is the detail that changes daily practice. The three translations in that study did not make the same mistakes as each other. Google Translate logged 8 mistranslation errors, ChatGPT-4o logged 12, and the human translator 19. On grammatical errors the ordering reversed: 7 for ChatGPT-4o, 11 for Google Translate, 21 for the human translator. Different systems, different failure points, identical fluent surface.
That disagreement is not noise. For a teacher who cannot read the target language, it is the only free error signal available. Run the same paragraph through two independent systems and compare. Where both render a term the same way, the agreement is weak evidence but real evidence. Where they diverge, you have found the sentence that needs a human. You have not learned which version is correct, but you have located the risk, and locating the risk is most of the job.
The academic guidance mentioned at the top of this piece turns on the same distinction, and it is the one worth carrying into a staffroom. There is a hard line between using a translation to understand something and using it as a statement you will stand behind in public. The first tolerates error, because the cost of a small mistake is a moment of confusion. The second does not, because the output becomes the record. School communication has exactly the same split running through it, and most schools have never named it out loud.
Sort the document before you translate it
The most useful habit is triage, and it happens before anything is pasted anywhere. Ask one question of the document in front of you: what does it cost if a sentence in this is wrong? Low-stakes text can go straight through. Medium-stakes text should be cross-checked against a second system. High-stakes text needs a qualified human translator, full stop.

Infographic 2. A three-tier triage model for classroom documents.
The top tier is not a matter of preference. Under Title VI of the Civil Rights Act, United States districts must take reasonable steps to give families with limited English proficiency meaningful access to school communication, and federal special education rules require assessment information in the language most likely to yield accurate information. Machine output on its own is not a language-access strategy, and it is not a defence if a placement decision is later challenged.
There is a second question sitting underneath the first, and it arrives before translation quality does. Pasting a student's name, test scores, and diagnostic language into a consumer chatbot is a data-handling decision governed by FERPA. Education-licensed versions of the major platforms exist specifically for this, with data terms that consumer accounts do not carry. Check which account you are signed into before you check the translation.
Write it so it survives translation
The strongest lever a teacher controls is not the translation system. It is the English going into it. Documents written well above a twelfth-grade reading level, dense with jargon and idiom, translated badly across every system tested in that study. Clearer source text produces better translations everywhere, and it also produces better documents for the families who read English.
Cut jargon, or gloss it in plain language the first time it appears.
Short sentences. Active voice. One idea per sentence.
Never make a pronoun carry the meaning. In the study, the sentence about parents who had always had difficulty reading came back describing the child instead of the parents, turning a family history into a finding about the student.
Keep a running glossary of the thirty or so terms your school uses constantly, with an agreed translation for each, so the same concept does not arrive three different ways across three letters.
Spell out acronyms on first use, including the ones that feel universal.
List numbers, dates, names and test titles somewhere they can be checked without reading the prose. Those are the elements a monolingual teacher can verify directly in any language.
What multilingual by default actually looks like
The reason multilingual teaching becomes the standard is not that translation got good. It is that translation got cheap enough to move upstream. It stops being a request routed to the district office three weeks before a meeting, and becomes a property of how materials are built in the first place.
In practice that means four things, none of them technological:
Templates are written for translation from the start, not fixed afterwards.
The glossary belongs to the school, not to whichever bilingual staff member happens to be available.
Every document type carries a risk tier, decided once, so the judgement is not remade under time pressure by whoever is sending it.
Review is a scheduled step with a name attached, not a favour asked of a bilingual teaching assistant between lessons.
That last one matters more than it looks. The research literature is consistent that schools lean on bilingual staff and, worse, on family members, for work those people were never trained or paid to do. Making review an explicit, resourced step is what turns individual goodwill into something a school can rely on when the person who has always quietly handled it moves on.
The teachers who will look prepared three years from now are not the ones who picked the best translation system. They are the ones who decided in advance which documents were allowed to leave the building on a machine's word, and which were not.
Frequently asked questions
Is it acceptable for teachers to use AI translation for parent communication?
For low-stakes communication, yes. Newsletters, reminders and scheduling notes rarely carry consequences if a phrase is imperfect. For anything affecting eligibility, placement, consent or discipline, federal language-access obligations point to a qualified human translator. Treat machine output as a draft that a person still has to approve.
If a translation reads well, does that mean it is accurate?
No. That is the central finding of the 2026 study. Fluency and accuracy were scored separately, and fluency showed no significant difference across systems that produced significantly different numbers of errors. A translation can be grammatically flawless and still misstate a diagnosis, a score, or a recommendation.
How can a monolingual teacher check a translation they cannot read?
Run the same text through a second independent system and compare the two. Where they agree, confidence is reasonable. Where they diverge, flag that sentence for a bilingual colleague. Separately, verify numbers, dates, names and test titles directly, because those are readable in any language.
What should a school do before rolling out AI translation?
Three things. Classify document types by risk before anyone needs them. Choose a platform with education-appropriate data handling, given FERPA obligations around student records. Build a shared glossary of recurring terms. Only about one in five teachers currently works at a school with any AI policy, so a written one is still a genuine advantage.
Does AI translation reduce the need for bilingual staff?
It changes the work rather than removing it. AI absorbs volume in low-risk communication, which frees bilingual staff for the high-stakes review only a person can do. The constraint was never enthusiasm: national survey data shows only 12% of school psychologists surveyed identified as bilingual, a figure flat for a decade.
Still grading everything by hand?
EMStudio is a free teaching management app — manage your classes, students, lessons, and more!
Learn More

Still grading everything by hand?
EMStudio is a free teaching management app — manage your classes, students, lessons, and more!
Learn More

2026 Notion4Teachers. All Rights Reserved.
2026 Notion4Teachers. All Rights Reserved.
2026 Notion4Teachers. All Rights Reserved.








