135931 20260217205549.0 doi 10.1016/j.jvoice.2024.03.001 sideral 138888 ART-2024-138888 eng Vidal, Jazmin Automatic voice disorder detection from a practical perspective 2024 Access copy available to the general public Unrestricted Voice disorders, such as dysphonia, are common among the general population. These pathologies often remain untreated until they reach a high level of severity. Assisting the detection of voice disorders could facilitate early diagnosis and subsequent treatment. In this study, we address the practical aspects of automatic voice disorders detection (AVDD). In real-world scenarios, data annotated for voice disorders is usually scarce due to various challenges involved in the collection and annotation of such data. However, some relatively large datasets are available for a reduced number of domains. In this context, we propose the use of a combination of out-of-domain and in-domain data for training a deep neural network-based AVDD system, and offer guidance on the minimum amount of in-domain data required to achieve acceptable performance. Further, we propose the use of a cost-based metric, the normalized expected cost (EC), to evaluate performance of AVDD systems in a way that closely reflects the needs of the application. As an added benefit, optimal decisions for the EC can be made in a principled way given by Bayes decision theory. Finally, we argue that for medical applications like AVDD, the categorical decisions need to be accompanied by interpretable scores that reflect the confidence of the system. Even very accurate models often produce scores that are not suited for interpretation. Here, we show that such models can be easily improved by adding a calibration stage-trained with just a few minutes of in-domain data. The outputs of the resulting calibrated system can then better support practitioners in their decision-making process. info:eu-repo/grantAgreement/ES/AEI/PID2021-126061OB-C44 info:eu-repo/grantAgreement/ES/DGA/T36-23R info:eu-repo/grantAgreement/EC/H2020/101007666/EU/Exchanges for SPEech ReseArch aNd TechnOlogies/ESPERANTO This project has received funding from the European Union’s Horizon 2020 research and innovation program under grant agreement No H2020 101007666-ESPERANTO info:eu-repo/semantics/openAccess All rights reserved http://www.europeana.eu/rights/rr-f/ 2.4 2024 AUDIOLOGY & SPEECH-LANGUAGE PATHOLOGY 6 / 35 = 0.171 2024 Q1 T1 OTORHINOLARYNGOLOGY 11 / 67 = 0.164 2024 Q1 T1 0.573 2024 LPN and LVN 2024 Q2 Speech and Hearing 2024 Q2 Otorhinolaryngology 2024 Q2 4.2 2024 info:eu-repo/semantics/article info:eu-repo/semantics/acceptedVersion Ribas, Dayana Universidad de Zaragoza (orcid)0000-0003-3813-4998 Bonomi, Cyntia Lleida, Eduardo Universidad de Zaragoza (orcid)0000-0001-9137-4013 Ferrer, Luciana Ortega, Alfonso Universidad de Zaragoza (orcid)0000-0002-3886-7748 5008 800 Universidad de Zaragoza Dpto. Ingeniería Electrón.Com. Área Teoría Señal y Comunicac. (2024), [16 pp.] J. voice JOURNAL OF VOICE 0892-1997 732864 http://zaguan.unizar.es/record/135931/files/texto_completo.pdf Postprint 2244007 http://zaguan.unizar.es/record/135931/files/texto_completo.jpg?subformat=icon icon Postprint oai:zaguan.unizar.es:135931 articulos driver 2026-02-17-20:39:43 ARTICLE