Voice evaluation method and device, electronic equipment and storage medium
By combining multimodal fusion processing of speech audio and facial video data, the scoring bias problem of existing speech assessment schemes has been solved, achieving a more accurate and comprehensive oral speech assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI IFLYTEK TOYCLOUD TECH
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-17
AI Technical Summary
Existing speech assessment solutions cannot comprehensively, objectively, and accurately reflect users' true spoken language assessment level, and are difficult to simulate the comprehensive reasoning and judgment process of human experts.
By acquiring speech audio data and facial video data, acoustic feature extraction, speech-to-text consistency analysis, and pronunciation action feature extraction are performed, followed by multimodal fusion processing to generate a comprehensive pronunciation score.
It significantly improves the comprehensiveness and accuracy of speech evaluation results, reduces scoring bias caused by single-modal evaluation, and provides objective and reliable evaluation results.
Smart Images

Figure CN122417072A_ABST