Tied hidden factors in neural networks for end-To-end speaker recognition

Miguel, A.; Ortega, A.; Lleida, E.; Llombart, J.

doi:10.21437/Interspeech.2017-1314

Tied hidden factors in neural networks for end-To-end speaker recognition

Miguel, A. (Universidad de Zaragoza) ; Llombart, J. (Universidad de Zaragoza) ; Ortega, A. (Universidad de Zaragoza) ; Lleida, E. (Universidad de Zaragoza)

Resumen: In this paper we propose a method to model speaker and session variability and able to generate likelihood ratios using neural networks in an end-To-end phrase dependent speaker verification system. As in Joint Factor Analysis, the model uses tied hidden variables to model speaker and session variability and a MAP adaptation of some of the parameters of the model. In the training procedure our method jointly estimates the network parameters and the values of the speaker and channel hidden variables. This is done in a two-step backpropagation algorithm, first the network weights and factor loading matrices are updated and then the hidden variables, whose gradients are calculated by aggregating the corresponding speaker or session frames, since these hidden variables are tied. The last layer of the network is defined as a linear regression probabilistic model whose inputs are the previous layer outputs. This choice has the advantage that it produces likelihoods and additionally it can be adapted during the enrolment using MAP without the need of a gradient optimization. The decisions are made based on the ratio of the output likelihoods of two neural network models, speaker adapted and universal background model. The method was evaluated on the RSR2015 database.
Idioma: Inglés
DOI: 10.21437/Interspeech.2017-1314
Año: 2017
Publicado en: Interspeech (USB) 2017-August (2017), 2819-2823
ISSN: 2308-457X
Originalmente disponible en: Texto completo de la revista

Financiación: info:eu-repo/grantAgreement/EC/FP7/610986/EU/IRIS: Towards Natural Interaction and Communication/IRIS
Financiación: info:eu-repo/grantAgreement/ES/MINEC0/TIN2014-54288-C4-2-R
Tipo y forma: Artículo (PostPrint)
Área (Departamento): Área Teoría Señal y Comunicac. (Dpto. Ingeniería Electrón.Com.)

Derechos reservados por el editor de la revista

Exportado de SIDERAL (2021-02-08-17:42:18)

Enlace permanente:

Visitas y descargas

Este artículo se encuentra en las siguientes colecciones:
Artículos

Volver a la búsqueda

Registro creado el 2018-08-16, última modificación el 2021-02-08

Postprint:
PDF

Valore este documento:

(Sin ninguna reseña)

Añadir a una carpeta personal
Exportar como BibTeX, MARC, MARCXML, DC, EndNote, NLM, RefWorks

Repositorio Institucional de Documentos

Tied hidden factors in neural networks for end-To-end speaker recognition