Research on multimodal student speech emotion recognitionmethods for precise education
摘要:
本文面向国家“精准育人”战略需求,针对课堂场景中学生面部表情变化细微,导致现有基于计算机视觉的情感识别方法面临识别困难的现实瓶颈,创新性地从学生言语维度出发,提出了一种融合多维度声学特征与文本特征的多模态学生言语情感识别方法,并构建了多模态感知神经网络.其中,在该网络的声学编码器中设计了注意力统计池化层,能够有效提取学生言语中的关键帧级别特征和序列级特征.实验结果表明,所提出的方法中的不同模块均能有效提升情感识别性能,且在识别效果上显著优于现有单模态方法.该方法有效融合了学生言语中的声学和文本的信息,突破了单模态感知的局限性,能够为“精准育人”实践提供可靠的技术支持,切实提高育人成效.
This study addressed the national strategic requirement of precision education by focusing on the subtle changes in students’facial expresions in classroom, which pose significant challenges to existing computer vision-based emotion recognition methods, This study innovatively approach the issue from the dimension of student speech, proposing a multimodal student speech emotion recognition method that integrates multidimensional acoustic and textual features, and constructing a mulimodal perception neural network. Specifically, an attention statistical pooling layer is designed in the acoustic encoder of this network, effectively extracting key frame-level and sequence-level features from student spech. Experimental results demonstrate that the different modules within the proposed method significanty enhance emotion recognition perormance, outper-forming existing unimodal methods in recognition effectivenes. This method effectively integrates acoustic and textual information from student speech, overcoming the limitations of unimodal perception, and provides reliable technical support for the practice of precision education, thereby substantially improving educational outcomes.
作者:
刘振中,张夷楠,王翠萍
Liu Zhenzhong, Zhang Yinan, Wang Cuiping
机构地区:
洛阳师范学院心理健康教育与咨询中心
引用本文:
刘振中,张夷楠,王翠萍。面向精准育人的多模态学生言语情感识别方法研究[J].河南师范大学学报(自然科学),2026,54(4): 98-105.(Liu Zhenzhong,Zhang Yinan, Wang Cuiping.Research on multimodal student speech emotion recognition methods for precise education[J].Journal of Henan Normal University (Natural Science Edition) 2026,54(4) :98-105.DOI:10.16366/j.cnki.1000-2367.2025.10.30.0002.)
基金:
国家社会科学基金教育学青年项目;河南省软科学研究计划项目
关键词:
精准育人;学生言语情感;声学特征;多模态情感识别
precise education; students speech emotion; acoustic features; multimodal emotion recognition
分类号:
O413


