Mostrar el registro sencillo del ítem
Deep View2View Mapping for View-Invariant Lipreading
| dc.creator | Koumparoulis A., Potamianos G. | en |
| dc.date.accessioned | 2023-01-31T08:45:24Z | |
| dc.date.available | 2023-01-31T08:45:24Z | |
| dc.date.issued | 2019 | |
| dc.identifier | 10.1109/SLT.2018.8639698 | |
| dc.identifier.isbn | 9781538643341 | |
| dc.identifier.uri | http://hdl.handle.net/11615/75302 | |
| dc.description.abstract | Recently, visual-only and audio-visual speech recognition have made significant progress thanks to deep-learning based, trainable visual front-ends (VFEs), with most research focusing on frontal or near-frontal face videos. In this paper, we seek to expand the applicability of VFEs targeted on frontal face views to non-frontal ones, without making assumptions on the VFE type, and allowing systems trained on frontal-view data to be applied on mismatched, non-frontal videos. For this purpose, we adapt the 'pix2pix' model, recently proposed for image translation tasks, to transform non-frontal speaker mouth regions to frontal, employing a convolutional neural network architecture, which we call 'view2view'. We develop our approach on the OuluVS2 multiview lipreading dataset, allowing training of four such networks that map views at predefined non-frontal angles (up to profile) to frontal ones, which we subsequently feed to a frontal-view VFE. We compare the 'view2view' network against a baseline that performs linear cross-view regression at the VFE space. Results on visual-only, as well as audio-visual automatic speech recognition over multiple acoustic noise conditions, demonstrate that the 'view2view' significantly outperforms the baseline, narrowing the performance gap from an ideal, matched scenario of view-specific systems. Improvements are retained when the approach is coupled with an automatic view estimator. © 2018 IEEE. | en |
| dc.language.iso | en | en |
| dc.source | 2018 IEEE Spoken Language Technology Workshop, SLT 2018 - Proceedings | en |
| dc.source.uri | https://www.scopus.com/inward/record.uri?eid=2-s2.0-85063087085&doi=10.1109%2fSLT.2018.8639698&partnerID=40&md5=7403ea83fab455685cf3ae73e8f09864 | |
| dc.subject | Acoustic noise | en |
| dc.subject | Audio acoustics | en |
| dc.subject | Audio systems | en |
| dc.subject | Deep learning | en |
| dc.subject | Network architecture | en |
| dc.subject | Neural networks | en |
| dc.subject | Speech analysis | en |
| dc.subject | Audio visual speech recognition | en |
| dc.subject | Automatic speech recognition | en |
| dc.subject | Convolutional neural network | en |
| dc.subject | Image translation | en |
| dc.subject | Lipreading | en |
| dc.subject | Noise conditions | en |
| dc.subject | Performance gaps | en |
| dc.subject | regression | en |
| dc.subject | Speech recognition | en |
| dc.subject | Institute of Electrical and Electronics Engineers Inc. | en |
| dc.title | Deep View2View Mapping for View-Invariant Lipreading | en |
| dc.type | conferenceItem | en |
Ficheros en el ítem
| Ficheros | Tamaño | Formato | Ver |
|---|---|---|---|
|
No hay ficheros asociados a este ítem. |
|||