MIDDLE EAST - Patients do not always describe their symptoms in formal medical language. In Lebanon and across the Arab region, a health question may combine a local Arabic dialect with English or French terms, everyday expressions, and culturally familiar ways of describing pain or distress.
While a healthcare professional can ask follow-up questions and interpret these nuances, an AI tool may not always understand them as accurately.
This is the central challenge facing medical AI in the Arab region. The question is not simply whether these systems “speak Arabic,” but which forms of Arabic they understand and which patients must change how they communicate to be understood.
Arabic is spoken by more than 450 million people worldwide, yet only around 3% of online content is available in the language, according to UNESCO. This figure does not reveal how much Arabic material is used to train medical AI specifically, but it illustrates the broader shortage of digital Arabic resources.
Moreover, Arabic is not a uniform language. Modern Standard Arabic is mainly used in formal writing, education, and news, while patients are more likely to describe symptoms in Lebanese, Egyptian, Tunisian, Gulf, or other local dialects.
Many also mix Arabic with English or French, write Arabic using Latin letters and numbers, commonly known as Arabizi, or depend on voice messages because writing is difficult or inconvenient.
These differences matter in healthcare. A symptom that appears straightforward in formal Arabic may be expressed through colloquial language, metaphors or culturally familiar phrases.
A patient might describe pain as burning, heavy, or as pressure or “something eating” inside the body. Psychological distress may be communicated through headaches, exhaustion, chest tightness or bodily pain rather than words such as anxiety or depression.
An AI system may translate the individual words correctly while still missing what the person means.
Strong Answers, Under Controlled Conditions
Research suggests that AI can produce useful Arabic health information, particularly when questions are clearly written. In a 2025 study of standardized liver-health questions, Arabic answers generated by ChatGPT received an average accuracy score of 4.9 out of 6 and a comprehensibility score of 2.74 out of 3.
However, the questions were professionally translated, while the answers were assessed by medical specialists. This is considerably different from interpreting an unstructured description entered by a patient.
Other research reveals continuing gaps. In a comparison involving virology questions, ChatGPT-4 received an average quality score of 4.20 in Arabic, compared with 4.68 in English.
An earlier study involving 91 questions about liver cirrhosis found that only 24.2% of the Arabic answers were comprehensive. Around 13.2% were completely incorrect, while one-third were less accurate than the corresponding English answers. The study used an earlier version of ChatGPT, meaning the figures should not be treated as a measure of how every current AI system performs.
Performance may also vary between dialects. One comparative study found that ChatGPT-4 scored 3.53 out of 5 when responding to Jordanian Arabic and 3.20 in Tunisian Arabic. Responses in both dialects performed significantly worse than responses in English.
Voice-based tools face an additional challenge: they must first transcribe what the patient says before interpreting its medical meaning. A 2025 review of Arabic code-switched language technologies reported word-error rates ranging from 28% to 54% across speech datasets combining Arabic with other languages.
Although the voice-based tools were not used in a medical setting, it demonstrates the difficulty faced by automated systems in processing conversation when there is a mix of languages and dialects.
When Language Becomes an Inequality
The benefits may not be experienced equally. Patients who are comfortable using digital tools and familiar with medical terminology may find it easier to describe their symptoms clearly and assess the reliability of an answer.
Others, including older adults, people with limited health or digital literacy, rural communities, displaced populations, and those who communicate mainly through spoken dialects. May encounter additional barriers that the technology must be designed to address.
A Beirut outpatient study involving 587 adults found that 65.8% had inadequate or problematic functional health literacy. This highlights a serious concern: a health tool requiring carefully structured questions may work best for people already better equipped to navigate the healthcare system.
The risks are particularly important in mental, reproductive and sexual health. AI may give people a private space to ask questions they feel embarrassed or unable to raise elsewhere.
Among 525 young women surveyed in Lebanon, 43.8% reported using AI chatbots for menstrual concerns and 48.8% for mental health. Avoiding embarrassment was viewed as an advantage by 43.4%, yet 85.5% identified accuracy as a concern.
This captures AI’s promise and danger. A confidential tool could widen access to basic information where affordable professional care is limited. However, misunderstanding an indirect expression, emotional crisis or sensitive symptom could produce false reassurance, unnecessary fear or delayed treatment.
Designing AI Around Patients
These problems are not inevitable. Arabic medical AI could improve access if it is developed and tested using original Arabic health material, rather than relying mainly on information translated from English. Assessments should include different dialects, accents, literacy levels, Arabizi, voice messages and Arabic-English or Arabic-French communication.
Safe systems must be designed to recognize uncertainty. When a description is unclear, the appropriate response is not to guess but to ask simple follow-up questions and direct urgent or complicated cases to qualified professionals. Independent assessments should examine clinical accuracy, language performance, patient privacy, and whether certain groups consistently receive poorer answers.
Medical AI should complement healthcare workers, not replace them. This is consistent with World Health Organization guidance, which stresses the need for oversight, transparency, accountability, and inclusive design when generative AI is used in health.
Improving Arabic-language AI matters far beyond conversational chatbots. Artificial intelligence is increasingly used across the healthcare journey—from appointment booking and symptom assessment to patient communication and clinical decision support. Language limitations at any of these stages could lead to misunderstandings, delay appropriate care, or contribute to medical errors.
Ensuring that AI systems accurately understand Modern Standard Arabic, regional dialects, mixed-language communication, and everyday expressions should therefore be treated as a public health and patient safety priority, not simply as a technological improvement. AI can support more accessible and equitable healthcare, but only if it is designed around how patients communicate and used to complement, rather than replace, qualified healthcare professionals.