• English
    • Ελληνικά
    • Deutsch
    • français
    • italiano
    • español
  • español 
    • English
    • Ελληνικά
    • Deutsch
    • français
    • italiano
    • español
  • Login
Ver ítem 
  •   DSpace Principal
  • Επιστημονικές Δημοσιεύσεις Μελών ΠΘ (ΕΔΠΘ)
  • Δημοσιεύσεις σε περιοδικά, συνέδρια, κεφάλαια βιβλίων κλπ.
  • Ver ítem
  •   DSpace Principal
  • Επιστημονικές Δημοσιεύσεις Μελών ΠΘ (ΕΔΠΘ)
  • Δημοσιεύσεις σε περιοδικά, συνέδρια, κεφάλαια βιβλίων κλπ.
  • Ver ítem
JavaScript is disabled for your browser. Some features of this site may not work without it.
Todo DSpace
  • Comunidades & Colecciones
  • Por fecha de publicación
  • Autores
  • Títulos
  • Materias

A non-linguistic approach for human emotion recognition from speech

Thumbnail
Autor
Spyrou E., Vernikos I., Nikopoulou R., Mylonas P.
Fecha
2019
Language
en
DOI
10.1109/IISA.2018.8633644
Materia
Acoustic properties
Classification (of information)
Human computer interaction
Linguistics
Students
Bag-of-visual-words
Emotion classification
Human emotion recognition
Linguistic approach
Middle school students
Speech information
Visual representations
Visual vocabularies
Speech recognition
Institute of Electrical and Electronics Engineers Inc.
Mostrar el registro completo del ítem
Resumen
One of the most important issues in several aspects of human-computer interaction is the understanding of the users' emotional state. In several applications such as monitoring of humans in assistive living environments, or assessing students' affective state during a course, it is imperative to use an unobtrusive method, so as to avoid discomforting or distracting the user. Thus, one should opt for approaches that use either visual or audio sensors which may observe users without any kind of direct contact. In this work, our goal is to recognize the emotional state of humans using only the non-linguistic aspect of speech information, i.e., the acoustic properties of speech. Therefore, we propose an emotion classification that is based on the bag-of-visual words model that has been previously applied in many computer vision tasks. A given audio segment is transformed to a spectrogram, i.e., a visual representation of its spectrum. From this representation we first extract SURF features and using a previously constructed visual vocabulary, we quantize them into a set of visual words. Then a histogram is constructed per image; These feature vectors are used to train SVM classifiers. Our approach is evaluated using a) 3 publicly available datasets that contain speech from different languages and b) a custom dataset that has been constructed during a real-life classroom experiments, involving middle-school students. ©2018 IEEE
URI
http://hdl.handle.net/11615/79346
Colecciones
  • Δημοσιεύσεις σε περιοδικά, συνέδρια, κεφάλαια βιβλίων κλπ. [19743]
htmlmap 

 

Listar

Todo DSpaceComunidades & ColeccionesPor fecha de publicaciónAutoresTítulosMateriasEsta colecciónPor fecha de publicaciónAutoresTítulosMaterias

Mi cuenta

AccederRegistro
Help Contact
DepositionAboutHelpContacto
Choose LanguageTodo DSpace
EnglishΕλληνικά
htmlmap