Patients with limited English proficiency face persistent barriers to equitable healthcare. Large language models (LLMs) may expand access to translation services; however, their accuracy for Asian languages and clinical contexts remains understudied. Prior work suggests LLMs may produce clinically meaningful errors despite strong automated metric performance.
This study evaluates LLM accuracy in translating pediatric patient education materials into Asian languages, using American Academy of Pediatrics (AAP) translations as the reference standard. Source documents were originally in English, with AAP translations serving as gold-standard comparators. Languages included Chinese, Korean, Vietnamese, Arabic, Bengali, and Hmong. Twenty-four documents were selected based on availability across all languages and standardized for topic and reading level. Translations were generated using a consistent prompt template for ChatGPT and Gemini and default settings for Google Translate. Translation fidelity was assessed using corpus-level Bilingual Evaluation Understudy (BLEU) scores, selected for standardized comparison of n-gram overlap despite known limitations in capturing clinical nuance. Differences were evaluated using one-way ANOVA with post hoc comparisons. Google Translate demonstrated the highest fidelity (mean BLEU 0.817 ± 0.111), outperforming ChatGPT (0.133 ± 0.063) and Gemini (0.132 ± 0.061) in five of six languages; ChatGPT performed slightly better in Bengali. Differences were statistically significant (p = 0.0143), with large magnitude gaps suggesting meaningful discrepancies in lexical alignment.
These findings indicate that traditional neural machine translation currently outperforms LLMs for pediatric patient materials. Clinically, LLM-generated translations should be used cautiously and supplemented with human review, particularly for lower-resource languages. Limitations include reliance on BLEU and a lack of human evaluation. Future work should incorporate clinically grounded metrics and expand low-resource language representation.
This study evaluates LLM accuracy in translating pediatric patient education materials into Asian languages, using American Academy of Pediatrics (AAP) translations as the reference standard. Source documents were originally in English, with AAP translations serving as gold-standard comparators. Languages included Chinese, Korean, Vietnamese, Arabic, Bengali, and Hmong. Twenty-four documents were selected based on availability across all languages and standardized for topic and reading level. Translations were generated using a consistent prompt template for ChatGPT and Gemini and default settings for Google Translate. Translation fidelity was assessed using corpus-level Bilingual Evaluation Understudy (BLEU) scores, selected for standardized comparison of n-gram overlap despite known limitations in capturing clinical nuance. Differences were evaluated using one-way ANOVA with post hoc comparisons. Google Translate demonstrated the highest fidelity (mean BLEU 0.817 ± 0.111), outperforming ChatGPT (0.133 ± 0.063) and Gemini (0.132 ± 0.061) in five of six languages; ChatGPT performed slightly better in Bengali. Differences were statistically significant (p = 0.0143), with large magnitude gaps suggesting meaningful discrepancies in lexical alignment.
These findings indicate that traditional neural machine translation currently outperforms LLMs for pediatric patient materials. Clinically, LLM-generated translations should be used cautiously and supplemented with human review, particularly for lower-resource languages. Limitations include reliance on BLEU and a lack of human evaluation. Future work should incorporate clinically grounded metrics and expand low-resource language representation.