Mostrar el registro sencillo del ítem

dc.creatorKoumparoulis A., Potamianos G.en
dc.date.accessioned2023-01-31T08:45:24Z
dc.date.available2023-01-31T08:45:24Z
dc.date.issued2019
dc.identifier10.1109/SLT.2018.8639698
dc.identifier.isbn9781538643341
dc.identifier.urihttp://hdl.handle.net/11615/75302
dc.description.abstractRecently, visual-only and audio-visual speech recognition have made significant progress thanks to deep-learning based, trainable visual front-ends (VFEs), with most research focusing on frontal or near-frontal face videos. In this paper, we seek to expand the applicability of VFEs targeted on frontal face views to non-frontal ones, without making assumptions on the VFE type, and allowing systems trained on frontal-view data to be applied on mismatched, non-frontal videos. For this purpose, we adapt the 'pix2pix' model, recently proposed for image translation tasks, to transform non-frontal speaker mouth regions to frontal, employing a convolutional neural network architecture, which we call 'view2view'. We develop our approach on the OuluVS2 multiview lipreading dataset, allowing training of four such networks that map views at predefined non-frontal angles (up to profile) to frontal ones, which we subsequently feed to a frontal-view VFE. We compare the 'view2view' network against a baseline that performs linear cross-view regression at the VFE space. Results on visual-only, as well as audio-visual automatic speech recognition over multiple acoustic noise conditions, demonstrate that the 'view2view' significantly outperforms the baseline, narrowing the performance gap from an ideal, matched scenario of view-specific systems. Improvements are retained when the approach is coupled with an automatic view estimator. © 2018 IEEE.en
dc.language.isoenen
dc.source2018 IEEE Spoken Language Technology Workshop, SLT 2018 - Proceedingsen
dc.source.urihttps://www.scopus.com/inward/record.uri?eid=2-s2.0-85063087085&doi=10.1109%2fSLT.2018.8639698&partnerID=40&md5=7403ea83fab455685cf3ae73e8f09864
dc.subjectAcoustic noiseen
dc.subjectAudio acousticsen
dc.subjectAudio systemsen
dc.subjectDeep learningen
dc.subjectNetwork architectureen
dc.subjectNeural networksen
dc.subjectSpeech analysisen
dc.subjectAudio visual speech recognitionen
dc.subjectAutomatic speech recognitionen
dc.subjectConvolutional neural networken
dc.subjectImage translationen
dc.subjectLipreadingen
dc.subjectNoise conditionsen
dc.subjectPerformance gapsen
dc.subjectregressionen
dc.subjectSpeech recognitionen
dc.subjectInstitute of Electrical and Electronics Engineers Inc.en
dc.titleDeep View2View Mapping for View-Invariant Lipreadingen
dc.typeconferenceItemen


Ficheros en el ítem

FicherosTamañoFormatoVer

No hay ficheros asociados a este ítem.

Este ítem aparece en la(s) siguiente(s) colección(ones)

Mostrar el registro sencillo del ítem