Voice evaluation method and device, electronic equipment and storage medium

By combining multimodal fusion processing of speech audio and facial video data, the scoring bias problem of existing speech assessment schemes has been solved, achieving a more accurate and comprehensive oral speech assessment.

CN122417072APending Publication Date: 2026-07-17HEFEI IFLYTEK TOYCLOUD TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI IFLYTEK TOYCLOUD TECH
Filing Date
2026-04-14
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing speech assessment solutions cannot comprehensively, objectively, and accurately reflect users' true spoken language assessment level, and are difficult to simulate the comprehensive reasoning and judgment process of human experts.

Method used

By acquiring speech audio data and facial video data, acoustic feature extraction, speech-to-text consistency analysis, and pronunciation action feature extraction are performed, followed by multimodal fusion processing to generate a comprehensive pronunciation score.

Benefits of technology

It significantly improves the comprehensiveness and accuracy of speech evaluation results, reduces scoring bias caused by single-modal evaluation, and provides objective and reliable evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122417072A_ABST
    Figure CN122417072A_ABST
Patent Text Reader

Abstract

本发明提供一种语音测评方法、装置、电子设备及存储介质,涉及数据处理技术领域,包括如下步骤:获取待测评样本数据;其中,待测评样本数据包括语音音频数据以及与语音音频数据同步采集的面部视频数据;基于语音音频数据进行声学特征提取和声学评分处理,获得第一测评数据;基于语音音频数据进行语音转写,并基于转写结果与参考文本进行文本一致性分析和语义分析,获得第二测评数据;基于面部视频数据和语音音频数据,进行发音动作特征提取和发音动作合理性评估,获得第三测评数据;基于第一测评数据、第二测评数据和第三测评数据进行多模态融合处理,获得待测评样本数据的目标测评数据;其中,目标测评数据至少包括多模态综合发音评分。
Need to check novelty before this filing date? Find Prior Art