The speech emotion recognition system and methodology utilize a multi-branch machine learning model with independent two-way contrast learning and pseudo-labeling in the emotion space.

VN126757APending Publication Date: 2026-07-01TAP OAN CONG NGHIEP VIEN THONG QUAN OI
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
VN · VN
Patent Type
Applications
Current Assignee / Owner
TAP OAN CONG NGHIEP VIEN THONG QUAN OI
Filing Date
2026-05-29
Publication Date
2026-07-01

Smart Images

  • Figure VN1202604444_0
    Figure VN1202604444_0
Patent Text Reader

Abstract

This invention proposes a system and method for speech emotion recognition using a multi-branch machine learning model with independent two-way contrast learning and pseudo-labeling in the emotion space. The method is implemented in four steps: step 1 builds a multi-branch machine learning model including: speech encoder, attentional synthesis block, activation level encoder, emotion polarity encoder, activation level decoder, emotion polarity decoder, and emotion classification branch.Step 2 involves automatically assigning pseudo-arousalvalence labels from discrete emotion classification labels based on Russell's two-dimensional emotion space model; Step 3 involves training the model using the total loss function ℒ𝑡𝑜𝑡𝑎𝑙 = α1ℒ𝑐𝑙𝑎𝑠𝑠 + α2ℒ𝑟𝑒𝑔 + α3ℒ𝑐𝑜𝑛𝑡, where ℒ𝑐𝑙𝑎𝑠𝑠 optimizes the classification branch, ℒ𝑟𝑒𝑔 optimizes the Decoder branches based on regression error, and ℒ𝑐𝑜𝑛𝑡 applies independent contrast learning on each dimension (arousal and valence); Step 4 performs emotion recognition by feeding speech signals through the trained model and selecting the emotion label with the highest probability.
Need to check novelty before this filing date? Find Prior Art