Multilingual Foundation Models
Large language and generative models designed to extend strong language understanding and generation beyond high-resource languages.
Multilingual & Low-Resource AI
Building reliable language and speech AI for underrepresented languages—from research to real-world systems.

I build multilingual language and speech AI that works beyond the small set of high-resource languages that dominate today's models. My research connects large language models, speech, data, evaluation, and scalable machine learning to make AI more reliable in linguistically diverse, real-world settings.
Arabic and African languages are a central proving ground for this work. Their dialect diversity, code-switching, cultural context, and uneven data availability expose limitations that conventional benchmarks often miss. I develop models, datasets, benchmarks, and evaluation frameworks that help close these gaps and expand who modern AI can serve.
My goal is not only to publish strong research, but to turn it into reusable resources, open benchmarks, and systems that enable broader research and real-world adoption. This work has appeared at ACL, EMNLP, EACL, INTERSPEECH, LREC, and ACM CSCW, and is advanced through collaborations connecting research communities across Canada, Africa, and the Arab world.
My research develops multilingual AI systems end to end—from foundation models and speech technologies to the data, evaluation, and infrastructure required to make them reliable in low-resource and culturally diverse settings.
Large language and generative models designed to extend strong language understanding and generation beyond high-resource languages.
Automatic speech recognition, NLP, and machine translation for Arabic, African, and other low-resource languages and varieties.
Datasets and rigorous evaluation frameworks that reveal linguistic, cultural, dialectal, and real-world capability gaps in modern AI.
Scalable training and evaluation pipelines, reasoning systems, and agentic methods that connect research advances to practical AI systems.
Community-driven projects that bring researchers and annotators together across the Arab world to build culturally grounded datasets, benchmarks, and evaluation resources.
A culturally inclusive and linguistically diverse Arabic instruction dataset for evaluating and developing Arabic LLMs.
A multimodal, culturally-aware Arabic instruction dataset designed for cultural understanding and visual question answering.
A multi-domain English↔Dialectal Arabic machine translation dataset and benchmark built around localized, conversational language.
Public research feature · CLEAR Global · June 2026
PALM · ACL 2025
Afrocentric-NLP · Global Top 100 Outstanding AI Projects
Industry research feature · AMD · YouTube · May 2023
Research · Speaking · Advisory
I welcome conversations around multilingual and low-resource AI, research collaborations, invited talks and panels, strategic advisory, and partnerships that move language and speech research into real-world impact.