Application of machine learning techniques to automate speech mispronunciation detection in children and adolescents

Authors

  • Nazila Ameli University of Alberta. Mike Petryk School of Dentistry.
  • Mahshid Nik Ravesh Autonomous Researcher
  • Luan Matheus Trindade Dalmazo
  • Karen Pollock Autonomous Researcher.
  • Daniel DeSantis Autonomous Researcher.
  • Manuel Lagravere University of Alberta. Mike Petryk School of Dentistry.
  • Hollis Lai University of Alberta. Mike Petryk School of Dentistry.

DOI:

https://doi.org/10.1590/1678-7765-2026-0078

Keywords:

Speech sound errors, Mispronunciation errors, Machine learning, Deep learning, Pediatrics

Abstract

Introduction  Speech sound disorders (SSDs) are common in children and can affect academic and psychosocial outcomes. Clinical identification relies on expert auditory-perceptual assessment, which is time-intensive and may vary across raters. Automated screening tools could support triage when access to specialists is limited. Objective  This study aimed to develop and evaluate a deep learning (DL) classifier for detecting mispronunciation errors in standardized pediatric speech recordings obtained using a structured fricative-focused word elicitation protocol, using expert-adjudicated labels as the reference standard. Methodology  In this cross-sectional study, we analyzed 100 participants (6–18 years) providing 1,800 standardized word recordings. Two expert speech-language pathologists (SLPs) labeled recordings as mispronunciation present vs. absent; disagreements were adjudicated by a third SLP. Inter- and intra-rater reliability were assessed. Audio was denoised and standardized to 16 kHz. A pretrained transformer speech model (WavLM Base+) was fine-tuned for recording-level binary classification. Class imbalance was addressed using augmentation, weighted sampling, and focal loss. Performance was assessed on a held-out test set at the recording level using accuracy, sensitivity, specificity, precision, F1-score, and area under the ROC curve (AUC). Results  The overall recording-level prevalence of speech sound errors was 16.15%, with the “th” category (/θ/ + /ð/) being the most frequently mispronounced phoneme group. No significant differences were observed by sex, age group (6–9, 10–13, 14–18 years), or malocclusion status (p > 0.05). Labeling showed strong agreement (Cohen’s κ = 0.84; intra-rater reliability 0.89–0.92). On the test set (253 recordings), the model achieved 90.9% accuracy, 77.8% sensitivity, 98.2% specificity, 95.9% precision, and AUC = 0.936. Conclusions  Fine-tuned pretrained speech representations demonstrated promising performance for screening pediatric mispronunciation errors under standardized recording conditions using expert labels. The high-specificity profile supports its use as a screening and triage decision-support tool, while future studies are needed to validate performance across broader clinical and real-world settings.

Downloads

Download data is not yet available.

References

1- American Speech-Language-Hearing Association. Speech sound disorders: articulation and phonology [Internet]. Rockville (MD): ASHA; 2014 [cited 2026 Jan 12]. Available from: https://www.asha.org/practice-portal/clinical-topics/articulation-and-phonology/

» https://www.asha.org/practice-portal/clinical-topics/articulation-and-phonology/

2- Hitchcock ER, Harel D, Byun TM. Social, emotional, and academic impact of residual speech errors in school-aged children: a survey study. Semin Speech Lang. 2015;36(4):283-94. doi: 10.1055/s-0035-1562911

» https://doi.org/10.1055/s-0035-1562911

3- Wren Y, Pagnamenta E, Peters TJ, Emond A, Northstone K, Miller LL, et al. Educational outcomes associated with persistent speech disorder. Int J Lang Commun Disord. 2021;56(2):299-312. doi: 10.1111/1460-6984.12599

» https://doi.org/10.1111/1460-6984.12599

4- Wren Y, Miller LL, Peters TJ, Emond A, Roulstone S. Prevalence and predictors of persistent speech sound disorder at eight years old: findings from a population cohort study. J Speech Lang Hear Res. 2016;59(4):647-73. doi: 10.1044/2015_JSLHR-S-14-0282

» https://doi.org/10.1044/2015_JSLHR-S-14-0282

5- McCormack J, McLeod S, Harrison LJ, McAllister L. The impact of speech impairment in early childhood: investigating parents' and speech-language pathologists' perspectives using the ICF-CY. J Commun Disord. 2010;43(5):378-96. doi: 10.1016/j.jcomdis.2010.04.009

» https://doi.org/10.1016/j.jcomdis.2010.04.009

6- Namasivayam AK, Coleman D, O'Dwyer A, van Lieshout P. Speech sound disorders in children: an articulatory phonology perspective. Front Psychol. 2020;10:2998. doi: 10.3389/fpsyg.2019.02998

» https://doi.org/10.3389/fpsyg.2019.02998

7- World Health Organization. International classification of functioning, disability and health (ICF). Geneva: World Health Organization; 2001.

8- Korkalainen J, McCabe P, Smidt A, Morgan C. Motor speech interventions for children with cerebral palsy: a systematic review. J Speech Lang Hear Res. 2023;66(1):110-25. doi: 10.1044/2022_JSLHR-22-00375

» https://doi.org/10.1044/2022_JSLHR-22-00375

9- Tashkandi NE, AlDosary R, Zamandar H, Alalwan M, Alwothainani M, Aljoaid H, et al. The relationship between malocclusion and speech patterns: a cross-sectional study. BMC Oral Health. 2025;25(1):65. doi: 10.1186/s12903-025-05437-0

» https://doi.org/10.1186/s12903-025-05437-0

10- Palakolanu SV, Dodda KK, Yelchuru SH, Kurapati J. Comparison of speech defects in different types of malocclusion. Cureus. 2024;16(6):e62290. doi: 10.7759/cureus.62290

» https://doi.org/10.7759/cureus.62290

11- Mohammed SA, Saloom JE, Obaid DH, Alhuwaizi AF. Malocclusion traits and speech disorders. Med J Babylon. 2023;20(4):661-4. doi: 10.4103/MJBL.MJBL_570_23

» https://doi.org/10.4103/MJBL.MJBL_570_23

12- Aprile M, Verdecchia A, Dettori C, Spinas E. Malocclusion and its relationship with sound speech disorders in deciduous and mixed dentition: a scoping review. Dent J (Basel). 2025;13(1):27. doi: 10.3390/dj13010027

» https://doi.org/10.3390/dj13010027

13- Assaf DC, Knorst JK, Busanello-Stella AR, Ferrazzo VA, Berwig LC, Ardenghi TM, et al. Association between malocclusion, tongue position and speech distortion in mixed-dentition schoolchildren: an epidemiological study. J Appl Oral Sci. 2021;29:e20201005. doi: 10.1590/1678-7757-2020-1005

» https://doi.org/10.1590/1678-7757-2020-1005

14- Sahad MG, Nahás AC, Scavone-Junior H, Jabur LB, Guedes-Pinto E. Vertical interincisal trespass assessment in children with speech disorders. Braz Oral Res. 2008;22(3):247-51. doi: 10.1590/S1806-83242008000300010

15- Kalia G, Tandon S, Bhupali NR, Rathore A, Mathur R, Rathore K. Speech evaluation in children with missing anterior teeth and after prosthetic rehabilitation with fixed functional space maintainer. J Indian Soc Pedod Prev Dent. 2018;36(4):391-5. doi: 10.4103/JISPPD.JISPPD_221_18

» https://doi.org/10.4103/JISPPD.JISPPD_221_18

16- Hyde AC, Moriarty L, Morgan AG, Elsharkasi LM, Deery C. Speech and the dental interface. Dent Update. 2018;45(9):795-803. doi: 10.12968/denu.2018.45.9.795

» https://doi.org/10.12968/denu.2018.45.9.795

17- Pernon M, Assal F, Kodrasi I, Laganaro M. Perceptual classification of motor speech disorders: the role of severity, speech task, and listener's expertise. J Speech Lang Hear Res. 2022;65(8):2727-47. doi: 10.1044/2022_JSLHR-21-00519

» https://doi.org/10.1044/2022_JSLHR-21-00519

18- Bates S, Titterington J, Child Speech Disorder Research Network. Good practice guidelines for the analysis of child speech. 2nd ed. [Internet]. Ulster University; 2021 [cited 2026 Jan 12]. Available from: https://www.rcslt.org/wp-content/uploads/2019/11/guidelines-for-analysis-of-child-speech-data.pdf

» https://www.rcslt.org/wp-content/uploads/2019/11/guidelines-for-analysis-of-child-speech-data.pdf

19- Munson B, Johnson JM, Edwards J. The role of experience in the perception of phonetic detail in children's speech: a comparison between speech-language pathologists and clinically untrained listeners. Am J Speech Lang Pathol. 2012;21(2):124-39. doi: 10.1044/1058-0360(2011/11-0009)

» https://doi.org/10.1044/1058-0360(2011/11-0009)

20- Skahan SM, Watson M, Lof GL. Speech-language pathologists' assessment practices for children with suspected speech sound disorders: results of a national survey. Am J Speech Lang Pathol. 2007;16(3):246-59. doi: 10.1044/1058-0360(2007/029)

» https://doi.org/10.1044/1058-0360(2007/029)

21- Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44-56. doi: 10.1038/s41591-018-0300-7

» https://doi.org/10.1038/s41591-018-0300-7

22- Kim DH, Jeong JW, Kang D, Ahn T, Hong Y, Im Y, et al. Usefulness of automatic speech recognition assessment of children with speech sound disorders: validation study. J Med Internet Res. 2025;27:e60520. doi: 10.2196/60520

» https://doi.org/10.2196/60520

23- Al-Nasheri A, Muhammad G, Alsulaiman M, Ali Z, Malki KH, Mesallam TA, et al. Voice pathology detection and classification using auto-correlation and entropy features in different frequency regions. IEEE Access. 2018;6:6961-74. doi: 10.1109/ACCESS.2017.2696056

» https://doi.org/10.1109/ACCESS.2017.2696056

24- Naqvi Y, Gupta V. Functional voice disorders. In: StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing; 2026 [cited 2026 Jan 12]. Available from: https://www.ncbi.nlm.nih.gov/books/NBK563182/

» https://www.ncbi.nlm.nih.gov/books/NBK563182/

25- Proffit WR, Fields HW, Larson BE, Sarver DM. Contemporary orthodontics. 6th ed. St. Louis: Elsevier; 2019. p. 192-5.

26- McLeod S, Baker E. Children's speech: an evidence-based approach to assessment and intervention. Boston: Pearson; 2017.

27- Dodd B, Hua Z, Crosbie S, Holm A, Ozanne A. Diagnostic evaluation of articulation and phonology (DEAP). London: The Psychological Corporation; 2002.

28- Chen S, Wang C, Chen Z, Wu Y, Liu S, Chen Z, et al. WavLM: large-scale self-supervised pre-training for full stack speech processing. IEEE J Sel Top Signal Process. 2022;16(6):1505-18. doi: 10.1109/JSTSP.2022.3188113

29- Microsoft. WavLM-Base-Plus [Internet]. Hugging Face; 2021 [cited 2026 Jan 12]. Available from: https://huggingface.co/microsoft/wavlm-base-plus

» https://huggingface.co/microsoft/wavlm-base-plus

30- Chen S, Wang C, Chen Z, Wu Y, Liu S, Chen Z, et al. WavLM [Internet]. Hugging Face; 2021 [cited 2026 Jan 12]. Available from: https://huggingface.co/docs/transformers/model_doc/wavlm

» https://huggingface.co/docs/transformers/model_doc/wavlm

31- Feltner C, Wallace IF, Nowell SW, Orr CJ, Raffa B, Middleton JC, et al. Screening for speech and language delay and disorders in children 5 years or younger: evidence report and systematic review for the US preventive services task force. JAMA. 2024;331(4):335-51. doi: 10.1001/jama.2023.24647

» https://doi.org/10.1001/jama.2023.24647

32- McGill N, McLeod S, Crowe KM, Wang C, Hopf SC. Waiting lists and prioritization of children for services: speech-language pathologists' perspectives. J Commun Disord. 2021;91:106099. doi: 10.1016/j.jcomdis.2021.106099

» https://doi.org/10.1016/j.jcomdis.2021.106099

33- Rvachew S, Rafaat S. Report on benchmark wait times for pediatric speech sound disorders. Can J Speech Lang Pathol Audiol. 2014;38(1):82-96.

34- Sugden E, Cleland J. Using ultrasound tongue imaging to support the phonetic transcription of childhood speech sound disorders. Clin Linguist Phon. 2022;36(12):1047-66. doi: 10.1080/02699206.2021.2003433

» https://doi.org/10.1080/02699206.2021.2003433

35- Northcutt CG, Jiang L, Chuang IL. Confident learning: estimating uncertainty in dataset labels. J Artif Intell Res. 2021;70:1373-411. doi: 10.1613/jair.1.12125

» https://doi.org/10.1613/jair.1.12125

36- Song H, Kim M, Park D, Shin Y, Lee JG. Learning from noisy labels with deep neural networks: a survey. IEEE Trans Neural Netw Learn Syst. 2022;34(11):8135-53. doi: 10.1109/TNNLS.2022.3152527

» https://doi.org/10.1109/TNNLS.2022.3152527

37- Ahn T, Hong Y, Im Y, Kim DH, Kang D, Jeong JW, et al. Automatic speech recognition (ASR) for the diagnosis of pronunciation of speech sound disorders in Korean children. Clin Linguist Phon. 2025;39(10):913-26. doi: 10.1080/02699206.2024.2387609

» https://doi.org/10.1080/02699206.2024.2387609

38- Shahin M, Zafar U, Ahmed B. The automatic detection of speech disorders in children: challenges, opportunities, and preliminary results. IEEE J Sel Top Signal Process. 2020;14(2):400-12. doi: 10.1109/JSTSP.2019.2959393

39- Suthar K, Yousefi Zowj F, Speights Atkins M, He QP. Feature engineering and machine learning for computer-assisted screening of children with speech disorders. PLOS Digit Health. 2022;1(5):e0000041. doi: 10.1371/journal.pdig.0000041

» https://doi.org/10.1371/journal.pdig.0000041

40- Ng SI, Ng CW, Wang J, Lee T. Automatic detection of speech sound disorder in child speech using posterior-based speaker representations. In: Interspeech 2022; 2022 Sep 18-22; Incheon, Republic of Korea. p. 2853-7. doi: 10.21437/Interspeech.2022-935

» https://doi.org/10.21437/Interspeech.2022-935

41- Luan Y, Speights M, Dozier G, Seals C. Automatic speech disorder detection (ASDD) system with self-supervised representation of children's speech. In: Wei J, Margetis G, Degen H, Ntoa S, editors. HCI International 2025 - late breaking papers. Cham: Springer; 2026. (Lecture Notes in Computer Science; vol 16346). doi: 10.1007/978-3-032-13187-4_13

» https://doi.org/10.1007/978-3-032-13187-4_13

42- Tbaishat M, Al-Shafei E, Odeh E. The role of AI in the diagnosis of speech and language disorders: a systematic mapping study. Digit Health. 2025;11:20552076251379769. doi: 10.1177/20552076251379769

» https://doi.org/10.1177/20552076251379769

43- Hsu TY, Li CA, Wu TY, Lee HY. Model extraction attack against self-supervised speech models. arXiv v2 [Preprint]. 2023 [cited 2026 July 29]. doi: 10.48550/arXiv.2211.16044

» https://doi.org/10.48550/arXiv.2211.16044

44- Patel T, Scharenborg O. Improving end-to-end models for children's speech recognition. Appl Sci (Basel). 2024;14(6):2353. doi: 10.3390/app14062353

» https://doi.org/10.3390/app14062353

45- McLeod S, Crowe K. Children's consonant acquisition in 27 languages: a cross-linguistic review. Am J Speech Lang Pathol. 2018;27(4):1546-71. doi: 10.1044/2018_AJSLP-17-0100

» https://doi.org/10.1044/2018_AJSLP-17-0100

46- Shriberg LD, Austin D, Lewis BA, McSweeny JL, Wilson DL. The percentage of consonants correct (PCC) metric: extensions and reliability data. J Speech Lang Hear Res. 1997;40(4):708-22. doi: 10.1044/jslhr.4004.708

» https://doi.org/10.1044/jslhr.4004.708

47- Lin TY, Goyal P, Girshick R, He K, Dollár P. Focal loss for dense object detection. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 2017. p. 2980-8.

48- Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi: 10.1371/journal.pone.0118432

» https://doi.org/10.1371/journal.pone.0118432

49- Panda PK, Ghosh S. Ethical use of AI in infectious diagnostic decision and therapeutic stewardship. IDCases. 2025;42:e02356. doi: 10.1016/j.idcr.2025.e02356

» https://doi.org/10.1016/j.idcr.2025.e02356

50- Wang CH, Tay J, Wu CY, Wu MC, Su PI, Fang YD, et al. External validation and comparison of statistical and machine learning-based models in predicting outcomes following out-of-hospital cardiac arrest: a multicenter retrospective analysis. J Am Heart Assoc. 2024;13(20):e037088. doi: 10.1161/JAHA.124.037088

» https://doi.org/10.1161/JAHA.124.037088

Downloads

Published

2026-09-08

Issue

Section

Original Articles

How to Cite

Ameli, N., Ravesh, M. N., Dalmazo, L. M. T., Pollock, K., DeSantis, D., Lagravere, M., & Lai, H. (2026). Application of machine learning techniques to automate speech mispronunciation detection in children and adolescents. Journal of Applied Oral Science, 34, e20260078. https://doi.org/10.1590/1678-7765-2026-0078