The speech emotion recognition system and methodology utilize a multi-branch machine learning model with independent two-way contrast learning and pseudo-labeling in the emotion space.
VN126757APending Publication Date: 2026-07-01TAP OAN CONG NGHIEP VIEN THONG QUAN OI
0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- VN · VN
- Patent Type
- Applications
- Current Assignee / Owner
- TAP OAN CONG NGHIEP VIEN THONG QUAN OI
- Filing Date
- 2026-05-29
- Publication Date
- 2026-07-01
Smart Images

Figure VN1202604444_0
Abstract
This invention proposes a system and method for speech emotion recognition using a multi-branch machine learning model with independent two-way contrast learning and pseudo-labeling in the emotion space. The method is implemented in four steps: step 1 builds a multi-branch machine learning model including: speech encoder, attentional synthesis block, activation level encoder, emotion polarity encoder, activation level decoder, emotion polarity decoder, and emotion classification branch.Step 2 involves automatically assigning pseudo-arousalvalence labels from discrete emotion classification labels based on Russell's two-dimensional emotion space model; Step 3 involves training the model using the total loss function ℒ𝑡𝑜𝑡𝑎𝑙 = α1ℒ𝑐𝑙𝑎𝑠𝑠 + α2ℒ𝑟𝑒𝑔 + α3ℒ𝑐𝑜𝑛𝑡, where ℒ𝑐𝑙𝑎𝑠𝑠 optimizes the classification branch, ℒ𝑟𝑒𝑔 optimizes the Decoder branches based on regression error, and ℒ𝑐𝑜𝑛𝑡 applies independent contrast learning on each dimension (arousal and valence); Step 4 performs emotion recognition by feeding speech signals through the trained model and selecting the emotion label with the highest probability.
Need to check novelty before this filing date? Find Prior Art