Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
- 1Ministry of Defense, Saudi Arabia
- 2Ability Center, Saudi Arabia
- 3University of Rochester, USA

Abstract
Automated phoneme-level pronunciation assessment is vital for scalable speech therapy and language learning, yet validated tools for Arabic remain scarce. Harf-Speech is a modular system that scores Arabic pronunciation at the phoneme level on a clinical scale by combining an MSA phonetizer, a fine-tuned speech-to-phoneme model, Levenshtein alignment, and a blended scorer using longest common subsequence and edit-distance metrics.
The study fine-tunes three ASR architectures on Arabic phoneme data and benchmarks them against zero-shot multimodal models. The strongest model achieves an 8.92% phoneme error rate. Clinical validation by three certified speech-language pathologists shows a Pearson correlation of 0.791 and ICC(2,1) of 0.659 with mean expert scores, demonstrating clinically aligned and interpretable assessment comparable to inter-rater expert agreement.
Citation
Azad, A., Shanto, M. S. H., Hossain, M. S., Alwuqaysi, B., Boughorbel, S., Bokhari, Y., Aljouie, A., Sindi, A. O., & Hoque, E. (2026). Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment. Accepted at Interspeech 2026.
