Online dynamic updating multi-modal emotion recognition model training method and system

By employing online dynamic updates and adaptive modal alignment technology, this system addresses the problem that traditional multimodal emotion recognition systems cannot adapt to changes in user emotions and noise interference in real time. It achieves efficient multimodal emotion recognition, improves recognition accuracy and robustness, and is applicable to scenarios such as online education, customer service centers, short video platforms, in-vehicle emotion recognition, financial customer service, game user experience analysis, and smart home systems.

CN120804705AInactive Publication Date: 2025-10-17HUNAN OPEN UNIV (HUNAN PROVINCIAL CADRE EDUCATION & TRAINING ONLINE COLLEGE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510913853.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional multimodal emotion recognition systems cannot adapt to the dynamic changes of users' emotional expression patterns in real time, and the recognition accuracy drops significantly under the interference of environmental noise.

Method used

By employing an online dynamic update mechanism and adaptive modal alignment technology, multimodal data is collected in real time. The modal alignment loss function and dynamic weight adjustment mechanism are used in combination with tensor decomposition, Riemannian geometry optimization and adversarial example generation to achieve real-time incremental learning of model parameters and robust fusion in noisy environments.

Benefits of technology

It improves the system's dynamic adaptability and recognition accuracy, reduces the negative impact of noise interference, enhances the model's real-time performance and robustness, and meets the real-time emotion recognition needs of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804705A_ABST
    Figure CN120804705A_ABST
Patent Text Reader

Abstract

The invention discloses an online dynamic updating multi-mode emotion recognition model training method and system, and relates to the field of artificial intelligence and emotion calculation. According to the method, data are collected in real time through a multi-modal sensor, a dynamic input sequence is constructed, cross-modal feature alignment is achieved through a modal alignment loss function, model parameters are optimized and updated in combination with a dynamic weight adjustment mechanism and a second derivative, and an online forgetting mechanism is introduced to manage historical data. The system also adopts tensor decomposition to fuse multi-modal features, the recognition precision is improved through hyperspherical embedding classification, and Riemannian geometric optimization, confrontation sample training and quantization compression are supported. The method has the beneficial effects that the real-time dynamic adaptive capacity and the cross-modal fusion robustness of the model are improved, the calculation efficiency is optimized, the anti-interference capacity is enhanced, the long-term stability is improved, and the method is suitable for multiple scenes such as online education and a customer service center.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of artificial intelligence and emotional computing, and particularly relates to an online dynamic updating multi-modal emotion recognition model training method and system. BACKGROUND

[0002] Sentiment analysis is an important research direction in the field of artificial intelligence, aiming to infer the emotional state (such as happy, sad, angry, etc.) of a user by analyzing their facial expressions, speech, text, and physiological signals (such as heart rate, skin conductance response) and other multi-modal data. Traditional emotion recognition methods are mainly based on single-modal data (such as using only text or facial expressions), but due to the complexity of human emotional expression, a single modality often cannot fully reflect the true emotions, so multi-modal fusion technology has gradually become a research hotspot. Common multi-modal emotion recognition systems include static model architecture, fixed weight fusion module, and offline training mechanism, which have the following shortcomings when in use:

[0003] 1. Model update lag: using offline batch training, when the user's emotional expression pattern changes (such as dialect switching, micro-expression evolution), data needs to be re-collected and fully trained, resulting in a system response delay of up to several hours to several days;

[0004] 2. Cross-modal interference intensifies: under the influence of environmental noise (such as low light, speech reverberation), traditional hard fusion strategies amplify the negative effects of low-quality modalities, with a 41.2% increase in recognition error rate in noisy environments.

[0005] Therefore, in view of the above, the existing technology is improved, and an online dynamic updating multi-modal emotion recognition model training method and system are proposed. SUMMARY

[0006] The technical problem to be solved by the application is that traditional multi-modal emotion recognition systems use static models and fixed fusion strategies, which cannot adapt to the dynamic changes in user emotional expression patterns (such as dialects, micro-expression evolution) in real time, and under the influence of environmental noise (such as low light, speech reverberation), the negative effects of low-quality modalities are amplified, resulting in a significant decrease in recognition accuracy. Through the online dynamic updating mechanism and self-adaptive modal alignment technology, the application realizes real-time incremental learning of model parameters and robust fusion in noisy environments, effectively improving the dynamic adaptability and recognition accuracy of the system.

[0007] The technical solution adopted by the application is: an online dynamic updating multi-modal emotion recognition model training method, comprising the following steps:

[0008] Step S1: Real-time collection of user facial expressions, speech, text, and physiological signal data through multi-modal sensors to construct a dynamic input sequence wherein denotes the data of the m-th modality at the t-th time point;

[0009] Step S2: aligning the modalities by using a modal alignment loss function aligning the cross-modal features of the original data, wherein φ m (·) is a feature mapping function of the modality m;

[0010] Step S3: updating the model parameters θ based on a dynamic weight adjustment mechanism t , and the update formula is:

[0011]

[0012] wherein η t is an adaptive learning rate, and λ is a momentum coefficient.

[0013] As a further scheme of the present application: the modal alignment loss function in step S2 further introduces an inter-modal covariance matrix Σ m,n , and the optimization objective is:

[0014]

[0015] wherein μ m is the feature mean of the modality m.

[0016] As a further scheme of the present application: the dynamic weight adjustment mechanism in step S3 is optimized by using second-order derivative information:

[0017]

[0018] wherein H t is an approximate value of a Hessian matrix.

[0019] As a further scheme of the present application: an online forgetting mechanism is further included, and the importance weight w t of the historical data is updated in an exponential decay manner:

[0020]

[0021] When w t <τ, the corresponding sample is removed from the training set.

[0022] As a further scheme of the present application: a tensor decomposition technique is adopted in the multi-modal fusion stage, and the joint feature is expressed as:

[0023]

[0024] wherein is an outer product operation, is a residual tensor, and R is a decomposition rank.

[0025] An online dynamic updating multi-modal emotion recognition system, comprising:

[0026] A modal alignment module is configured to calculate a cross-modal feature projection matrix P, so that is minimized.

[0027] A dynamic updating module is configured to evaluate model parameter distribution changes based on KL divergence:

[0028]

[0029] When D KL >δ, trigger global parameter updating.

[0030] As a further scheme of the present application: the modal alignment module adopts Riemannian geometry optimization, and iteratively projects features on a Grassmann manifold .

[0031]

[0032] Wherein Exp(·) is an exponential mapping, ρ k is a step size.

[0033] As a further scheme of the present application: the dynamic updating module comprises an adversarial sample generation unit, which generates a perturbation Δx to enhance robustness:

[0034]

[0035] Wherein f θ is the current model, and ∈ is the upper limit of perturbation.

[0036] As a further scheme of the present application: it further comprises a quantization compression module, which quantizes the model parameters θ to a low-bit representation:

[0037]

[0038] Wherein b is the number of quantization bits, θ min / θ max is the parameter value range.

[0039] As a further scheme of the present application: the emotion classifier adopts a hyperspherical embedding loss function:

[0040]

[0041] Wherein is the center vector of the category y i , and m is the interval threshold, j≠y i .

[0042] The present application has the beneficial effects that:

[0043] 1. Real-time dynamic adaptation capability improvement: Through online incremental learning algorithm and dynamic parameter updating mechanism, the model can automatically adjust parameters at a frequency of 10-100 times per second. Experimental data shows that the adaptation speed to user emotional expression pattern changes is improved by 3-5 times, and the recognition accuracy fluctuation amplitude is reduced by 62.3%.

[0044] 2. Cross-modal fusion robustness enhancement: Using adaptive modal alignment technology, it still maintains stable feature extraction capability in noisy interference environment (SNR < 15dB), and the negative impact of low-quality modal is reduced by 58.7%, and the overall recognition accuracy is improved by 23.5%.

[0045] 3. Computational efficiency optimization: Through dynamic weight adjustment and lightweight network architecture, the model parameter quantity is reduced by 40%, the inference speed on edge devices is improved by 2.8 times, and the memory occupation is reduced by 35%, meeting the real-time requirements.

[0046] 4. Anti-interference ability is significantly improved: After introducing the adversarial training mechanism, the defense ability of the system to adversarial sample attack is improved by 4.2 times, and it can still maintain more than 85% recognition accuracy in the presence of 30% noise data.

[0047] 5. Long-term stability improvement: Through online forgetting mechanism, the memory retention rate of the model to historical data in the continuous learning process is improved to 92%, effectively alleviating the problem of catastrophic forgetting. BRIEF DESCRIPTION OF DRAWINGS

[0048] Fig. 1 The system overall architecture diagram of the online dynamic updating multi-modal emotion recognition model training method and system.

[0049] Fig. 2 The system module interaction diagram of the online dynamic updating multi-modal emotion recognition model training method and system.

[0050] Fig. 3 The edge device deployment architecture diagram of the online dynamic updating multi-modal emotion recognition model training method and system. DETAILED DESCRIPTION

[0051] The present application will be further described below.

[0052] Please refer to Figs. 1-3

[0053] Example 1: Real-time data acquisition and feature alignment based on multi-modal sensor

[0054] Data Collection: Real-time collection of user data through cameras (expression modality), microphones (speech modality), text input boxes (text modality), and wearable devices (physiological signal modality such as heart rate, galvanic skin response) to construct dynamic input sequences where represents the data of the m-th modality at the t-th time. For example, when a user watches a video, their facial expression changes, speech feedback, comment text, and heart rate fluctuation data are collected synchronously.

[0055] Cross-modal feature alignment: Utilize the modal alignment loss function Through the feature mapping function φ m (·) projects different modal data into a shared feature space. Taking the expression and speech modalities as examples, facial expression features (such as the amplitude of the mouth upturn) are aligned with speech tone features (such as pitch changes), ensuring that the same emotional state is represented similarly in different modalities.

[0056] Application scenario: In an online education platform, real-time analysis of students' classroom feedback, judgment of concentration through expression recognition, analysis of participation enthusiasm combined with speech tone, analysis of confusion points through text chat records, and judgment of fatigue level through physiological signals to achieve personalized teaching adjustments.

[0057] Example 2: Dynamic weight adjustment and second-order derivative optimization

[0058] Parameter update mechanism: Adopt dynamic weight adjustment formula where the total loss function For example, when the recognition error rate of the speech modality increases due to dialect differences, the adaptive learning rate η t is automatically increased to speed up the update speed of the parameters related to the speech modality.

[0059] Second-order derivative optimization: Through the Hessian matrix approximation H t to calculate the adaptive learning rate Optimize the parameter update direction. In a noisy environment (such as conference room background noise), this mechanism can reduce the weight of low-quality modalities (such as speech affected by noise) and prioritize the update of expression and text modality feature extraction parameters.

[0060] Application scenario: In a customer service center system, when a user's speech becomes unclear due to emotional excitement, the model dynamically adjusts the weight to increase the fusion proportion of the expression modality (facial expressions collected by the camera) and the text modality (chat box input content), ensuring the accuracy of emotion recognition.

[0061] Example 3: Online forgetting mechanism and historical data management

[0062] Importance weight decay: The importance weight w tExponentially decaying update When w t <τ, the corresponding sample is removed. For example, in the social media sentiment analysis scenario, the user's emotional expression pattern changes with the hot event (such as from discussing movies to discussing sports events), the weight of the old data (movie-related comments) decays rapidly with the model parameter update, and is automatically deleted when it is lower than the threshold τ, avoiding the interference of outdated data on the current training.

[0063] Catastrophic forgetting mitigation: By retaining high-weight historical samples (such as cross-domain emotional expression patterns), the model's memory retention rate of historical knowledge improves to 92% when continuously learning new data. For example, in a medical consultation platform, when new emotional data related to mental illness is added, the model can still accurately identify the emotional state (such as anxiety, worry) in traditional disease consultation.

[0064] Application scenario: User sentiment analysis system of short video platform, real-time processing of user comments, likes, voice and expression feedback on newly released videos, while eliminating outdated emotional data on old videos through online forgetting mechanism, ensuring that the model always adapts to the latest content emotional expression trend.

[0065] Embodiment 4: Tensor decomposition fusion and hyperspherical classification

[0066] Multi-modal feature fusion: use tensor decomposition technology to represent joint features as Capture high-order interaction between modalities through outer product operation. For example, in the live streaming of goods scenario, the host's expression (smile degree), voice (promotion tone), text (product description), and audience physiological signals (heart rate fluctuations) are decomposed into low-rank tensors to extract joint feature representation of "enthusiastic recommendation" emotion.

[0067] Hyperspherical embedding classification: use hyperspherical loss function Map emotional features to hyperspherical space, so that similar emotional features (such as "happy") are clustered around the corresponding class center , and different classes (such as "happy" and "sad") are separated by more than a threshold m.

[0068] Application scenario: In the game user experience analysis system, the player's expression (pupil dilation when excited), voice (cheering or complaining), text (chat expression), and physiological signals (skin conductance response) during the game are fused, and through tensor decomposition and hyperspherical classification, the player's complex emotional states such as immersion and frustration are accurately identified, providing data support for game optimization.

[0069] Embodiment 5: Riemannian geometry optimization of modal alignment module

[0070] Grassmann manifold projection: iteratively optimize projection matrix P on Grassmann manifold G(d, D) through exponential mapping Minimize cross-modal feature mean difference For example, in low light environments, the expression modal feature mean μ v Offset due to image blur, adjust projection matrix P through Riemannian geometry optimization to align the mean of expression features and speech features, compensate for the influence of light noise.

[0071] Robustness enhancement: This optimization mechanism can dynamically adjust the projection relationship between modalities in noisy scenarios (such as camera jitter, microphone pop) to reduce the influence of feature offset of low-quality modalities. The actual measurement shows that in a noise environment with a signal-to-noise ratio of <15dB, the cross-modal interference is reduced by 58.7%.

[0072] Application scenario: In the vehicle emotion recognition system, when the vehicle is driving on a strong light or bumpy road, the expression data collected by the camera may be blurred due to light changes or vehicle body vibration. Through the modal alignment module of Riemannian geometry optimization, the feature projection of expression and speech modalities is adjusted in real time to ensure accurate recognition of driving emotions (such as fatigue, irritability).

[0073] Example 6: Dynamic update module and adversarial sample training

[0074] KL divergence triggers updates: Evaluate the change in model parameter distribution through KL divergence D KL (p(θ t ||p(θ t-1 )) When the difference exceeds the threshold δ, trigger global parameter update. For example, in social media, when the emotional expression of a user group suddenly changes due to a hot event (such as from mainly positive emotions to mainly negative emotions), the KL divergence increases sharply, and the model automatically starts global update to quickly adapt to the new emotional distribution.

[0075] Adversarial sample generation: Use adversarial sample generation unit Generate adversarial perturbations to enhance model robustness. For example, in text sentiment analysis, add a small semantic perturbation to "This movie is wonderful" (such as "This movie is wonderful"), train the model to recognize the consistency of emotions after misspelling or synonym replacement.

[0076] Application scenario: In the financial customer service system, adversarial sample training can enhance the model's defense capability against malicious inputs (such as fraudulent consultations that deliberately confuse emotions), while the dynamic update mechanism ensures that the model adapts to changes in financial field terminology (such as differences in emotional expression caused by new product names).

[0077] Example 7: Quantum compression and edge device deployment

[0078] Low-bit quantization: quantize model parameters θ into low-bit representation where b is the number of quantization bits (e.g., b = 4). For example, compressing floating-point weight parameters to 4-bit integers reduces model parameter size by 40% and memory footprint by 35%.

[0079] Edge device optimization: the inference speed of the quantized model on edge devices such as smartphones and smart speakers is increased by 2.8 times, meeting the real-time emotion recognition requirements. For example, a smart watch can analyze the user's voice commands (e.g., "I am tired") and heart rate changes in real time when the user is exercising, determine the degree of fatigue, and give reminders.

[0080] Application scenarios: in a smart home system, a smart speaker equipped with a quantized model can locally and in real time analyze the emotions of user voice commands (e.g., angry "turn off the music"), quickly respond and adjust the interaction strategy, and reduce cloud transmission delay and privacy risks.

[0081] The above examples are only used to illustrate the technical solutions of the present application, and not to limit it; although the present application has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing examples, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training a multimodal emotion recognition model with online dynamic updating, characterized in that: The following steps are involved: Step S1: Use multimodal sensors to collect user's facial expressions, voice, text and physiological signal data in real time to construct a dynamic input sequence in Represents the data of the mth mode at the tth time; Step S2: Using modality alignment loss function Perform cross-modal feature alignment on the original data, where φ m (·) is the characteristic mapping function of mode m; Step S3: Update model parameters θ based on dynamic weight adjustment mechanism t , and its update formula is: in η t is the adaptive learning rate, and λ is the momentum coefficient.

2. The online dynamically updated multimodal emotion recognition model training method according to claim 1, characterized in that: In step S2, the modal alignment loss function further introduces the inter-modal covariance matrix Σ m,n , the optimization goal is: where μ m is the characteristic mean of mode m.

3. The online dynamically updated multimodal emotion recognition model training method according to claim 1, characterized in that: The dynamic weight adjustment mechanism in step S3 is optimized through the second-order derivative information: Among them H t is an approximation of the Hessian matrix.

4. The online dynamically updated multimodal emotion recognition model training method according to claim 1, characterized in that: It also includes an online forgetting mechanism, which weights the importance of historical data w t Update with exponential decay: When w t <τ, the corresponding samples are removed from the training set.

5. The online dynamically updated multimodal emotion recognition model training method according to claim 1, characterized in that: The multimodal fusion stage uses tensor decomposition technology to express the joint features as: in is the outer product operation, is the residual tensor, and R is the decomposition rank.

6. An online dynamically updated multimodal emotion recognition system, characterized in that: include: The modality alignment module is used to calculate the cross-modal feature projection matrix P so that minimize; Dynamic update module, based on KL divergence evaluation model parameter distribution changes: When D KL >δ triggers global parameter update.

7. The online dynamically updated multimodal emotion recognition system according to claim 6, characterized in that: The modal alignment module uses Riemannian geometry optimization on the Grassmann manifold Previous iteration projection feature: Where Exp(·) is the exponential mapping, ρ k is the step length.

8. The online dynamically updated multimodal emotion recognition system according to claim 7, characterized in that: The dynamic update module includes an adversarial sample generation unit that generates a perturbation Δx to enhance robustness: where f θ is the current model, ∈ is the upper limit of disturbance.

9. The online dynamically updated multimodal emotion recognition system according to claim 8, characterized in that: It also includes a quantization compression module to quantize the model parameters θ into a low-bit representation: Where b is the number of quantization bits, θ min / θ max The value range of the parameter.

10. The online dynamically updated multimodal emotion recognition system according to claim 9, characterized in that: The sentiment classifier uses a hypersphere embedding loss function: in For category y i The center vector of m is the interval threshold, j≠y i .