Brain-like multi-mode emotion recognition network, brain-like multi-mode emotion recognition method and emotion robot

Through the brain-like multimodal emotion recognition network, multimodal emotion information is converted into pulse sequences, a continuous spectrum emotion representation space is constructed, and the pulse encoding of mixed emotions is generated is solved, which solves the problem of roughening and insufficient prediction of emotional representation in the existing technology, and achieves efficient and accurate emotion recognition and prediction.

CN120492978APending Publication Date: 2025-08-15SHENZHEN YIYUANZHEN TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510632616.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, in the emotion recognition, there are problems such as roughening of emotion representation, difficulty in expressing mixed emotions, insufficient prediction of dynamic emotion evolution, insufficient computational efficiency and biological interpretability, and limited multimodal fusion effect.

Method used

The brain-like multimodal emotion recognition network is adopted to convert multimodal emotion information into pulse sequences, build a continuous spectrum emotion representation space, generate pulse encoding of mixed emotions, establish an emotional state transition probability model, and realize the prediction and processing of the emotional gradient process through the pulse neural network to capture subtle changes in emotions.

Benefits of technology

It improves the accuracy of emotion recognition, enhances the ability to represent mixed emotions, improves the dynamic prediction performance of emotions, optimizes the efficiency of computing resources, enhances biological interpretability, and improves situational adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492978A_ABST
    Figure CN120492978A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a brain-like multi-mode emotion recognition network, a brain-like multi-mode emotion recognition method and an emotion robot, and the brain-like multi-mode emotion recognition method comprises the steps: converting multi-mode emotion information into a pulse sequence; on the basis of the obtained pulse sequence, constructing a continuous pedigree emotion representation space, and realizing continuous representation of an emotional state; a continuous pedigree emotion representation space is utilized to generate pulse codes of mixed emotions, and a complex emotional state is represented; according to the pulse codes of the mixed emotions, establishing an emotional state conversion probability model, and describing a conversion relation between the emotions; based on an emotional state transition probability model, realizing prediction and processing of an emotional gradual change process, and capturing subtle emotional changes; according to the method, the working principle of human brain neurons can be used for reference, fine representation and accurate prediction of the human emotional state are achieved, and meanwhile high calculation efficiency and biological interpretability are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a brain-like multimodal emotion recognition network, a recognition method, and an emotion robot. Background Art

[0002] With the rapid development of artificial intelligence (AI), affective computing, as a key research area in human-computer interaction, has attracted widespread attention from both academia and industry. As a core component of affective computing, emotion recognition technology has significant application value in areas such as improving human-computer interaction experiences, intelligent education systems, and mental health monitoring.

[0003] At present, mainstream emotion recognition technologies are mainly based on traditional deep learning models. These methods have achieved certain results in single-modality emotion recognition tasks, but they still face the following key technical challenges in practical applications: the problem of coarsening emotion representation; difficulty in expressing mixed emotions; insufficient prediction of the dynamic evolution of emotions; insufficient computational efficiency and biological interpretability; and limited multimodal fusion effects. In addition, existing technologies fail to fully draw on the neural mechanisms of the human brain to process emotional information when processing emotional information, resulting in poor performance of the model when facing complex and dynamically changing emotional scenes.

[0004] Therefore, a new emotion recognition method is needed that can draw on the working principles of human brain neurons to achieve detailed representation and accurate prediction of human emotional states, while having high computational efficiency and biological interpretability. Summary of the Invention

[0005] The present invention provides a brain-like multimodal emotion recognition network, recognition method and emotion robot to solve the technical problems in related technologies such as coarsening of emotion representation, difficulty in expressing mixed emotions, insufficient prediction of dynamic evolution of emotions, insufficient computational efficiency and biological interpretability, and limited multimodal fusion effect.

[0006] The present invention provides a brain-like multimodal emotion recognition method, comprising: Convert multimodal emotional information into pulse trains; Based on the obtained pulse sequence, a continuous spectrum emotion representation space is constructed to achieve continuous representation of emotional states; Using the continuous spectrum emotion representation space, pulse codes of mixed emotions are generated to represent complex emotional states; According to the pulse coding of mixed emotions, an emotional state transition probability model is established to describe the transition relationship between emotions; Based on the emotional state transition probability model, the gradual change of emotions can be predicted and processed to capture subtle changes in emotions.

[0007] In a preferred embodiment, the method for converting multimodal emotion information into a pulse sequence includes: Through the multimodal feature extraction network, emotional features of facial expressions, speech and text modalities are extracted respectively; Map the extracted emotional features to the pulse emission rate and construct a mapping function from features to pulse rate; The pulse train is generated based on the random threshold reset model, and the membrane potential update formula is: ; in, express The neuronal membrane potential at that moment, express The neuron membrane potential at each moment, when the membrane potential exceeds the threshold, the neuron emits a pulse and resets, otherwise the membrane potential is updated according to the input and attenuation, express Input current at all times, is the membrane potential attenuation coefficient, is the reset potential value after the neuron fires, The firing threshold for the neuron; The temporal change information of emotion is encoded into the pulse time interval to form a time-coded pulse sequence.

[0008] In a preferred embodiment, the method for constructing a continuous spectrum emotion representation space includes: Construct a nonlinear manifold based on emotion perception characteristics, including three basic dimensions: emotion arousal, emotion valence, and emotion dominance; Introducing emotion mixing degree and emotion change rate as derived features; The manifold is constructed using a weighted locality preserving projection method by solving the optimization problem: ; satisfy: ; in, Represents the projection matrix Find the minimum value, is the Laplace matrix, is a diagonal weight matrix, is the projection matrix, is the identity matrix, represents the sum of the diagonal elements of a matrix; A parameterized cubic B-spline curve is used to represent the emotion gradient trajectory, which satisfies the emotion continuity constraint, physiological constraint and psychological constraint.

[0009] In a preferred embodiment, the method for generating pulse codes of mixed emotions includes: Represent the mixed emotion as a weighted sum of emotion basis vectors: ; in, A vector representation of mixed emotions, For the The mixed weights of the basic emotions, For the The feature vectors of basic emotions, Express emotions and emotions The interaction strength coefficient between Express emotions and emotions The interaction feature vector of is the total number of basic emotions defined in the system; Establish a dynamic equilibrium model of mixed emotions to describe the evolution of mixed emotions over time; Encode mixed emotional states into multi-channel pulse trains, with different channels corresponding to different emotional components; Express the intensity of emotions and the relationship between emotions through pulse frequency and phase modulation.

[0010] In a preferred embodiment, the method for establishing an emotional state transition probability model includes: Construct an emotional state transition matrix to represent the probability of transitioning from one emotional state to another; Introducing contextual factors to dynamically adjust the conversion probability according to the situation; Build a conditional probability model based on variational autoencoders to learn the latent distribution of emotion transitions; Utilize a sequence-to-sequence learning framework to predict the distribution of future emotional state sequences under given conditions.

[0011] In a preferred embodiment, the method for predicting and processing the gradual change of emotions includes: Design a phase gradient modulation algorithm to encode the rate of change of emotion by modulating the phase of the pulse sequence; Use an emotion trajectory smoothing interpolation algorithm to generate natural transitions between discrete emotion states; Build a predictive model of the gradual change of emotions, and predict future emotional changes based on historical emotional trajectories and current situations; An adaptive threshold mechanism is introduced to detect emotion change inflection points and emotion conversion events.

[0012] In a preferred embodiment, the multimodal collaborative fusion step is also included: Design a modal fusion mechanism based on spiking neural networks, including inter-modal suppression / enhancement mechanisms, modal reliability assessment, and modal adaptive weighting; Construct a multimodal loss and inconsistency handling mechanism, including modality reconstruction, zero-shot modality transfer, inconsistency detection, and conflict mediation; Implement impulse response decoding based on membrane potential accumulation and decode the fused pulse train into emotional state representation; Inter-modal collaborative reinforcement learning is introduced to enhance the semantic consistency between modalities through contrastive learning and mutual information maximization.

[0013] In a preferred embodiment, a brain-inspired multimodal emotion recognition network is used to implement a brain-inspired multimodal emotion recognition method, and is applied to an intelligent education system, including: Build a system architecture that includes the front-end data collection layer, edge computing layer, and cloud analysis layer; Achieve real-time recognition and prediction of students' emotional states, supporting higher accuracy than traditional methods; Analyze the correlation between emotional trajectories and learning behaviors, and discover the association between emotional patterns and learning outcomes; Generate adaptive teaching strategies based on emotion recognition results to improve learning outcomes.

[0014] In a preferred embodiment, the system is also applied to a mental health monitoring and emotion regulation system, including: Collect multimodal data from users through smartphones and wearable devices; Identify abnormal emotional patterns and predict the trend of emotional deterioration, and provide early warning of emotional abnormalities; Generate personalized emotion regulation suggestions based on emotion state analysis and transition probability modeling; When used in clinical practice, it significantly improves the treatment effect for patients with depression and anxiety.

[0015] In a preferred embodiment, a brain-like multimodal emotion robot is used to execute a brain-like multimodal emotion recognition method and is applied to an emotion interaction robot to achieve natural and coherent human-computer emotion interaction.

[0016] The beneficial effects of the present invention are: Refinement of emotional expression: Pulse coding and continuous spectrum representation are used to improve recognition accuracy.

[0017] Enhanced mixed emotion representation capability: The accuracy of mixed emotion recognition is improved through continuous spectrum emotion representation space and mixed emotion pulse coding.

[0018] Improved performance in predicting dynamic emotions: Based on an emotional state transition probability model and a forward-looking pulse generation algorithm, this improves the accuracy of predicting emotional transitions and shortens the response time to sudden emotional fluctuations.

[0019] Computing resource efficiency optimization: The pulse neural network computing mode adopted reduces energy consumption and hardware implementation area under the same performance indicators.

[0020] Enhanced biological interpretability: The pulse coding method and emotional state transition model are highly similar to the working principles of human brain neurons, which enhances the biological interpretability of the model and helps to understand the neurobiological basis of human emotional processing.

[0021] Improved situational adaptability: Through the attention modulation algorithm and emotion mutation detection and response system, the system can dynamically adjust its attention to emotional components according to different situations, improving the adaptability of emotion recognition in complex social scenarios and enhancing cross-scenario generalization performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flow chart of the brain-like multimodal emotion recognition method of the present invention. DETAILED DESCRIPTION

[0023] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0024] At least one embodiment of the present invention discloses a brain-like multimodal emotion recognition method, such as Figure 1 As shown, the following steps are included: Step 1: Convert multimodal emotion information into pulse sequences; The specific implementation steps are as follows: Step 1.1, multimodal emotional information acquisition and preprocessing; Acquire multimodal emotional information, including facial expression video data, speech audio data, and physiological signal data, and perform normalization and noise reduction on this data. For video data, extract facial key points and micro-expression variation features; for audio data, extract acoustic features such as pitch, intensity, and speaking rate; and for physiological signals, extract biometric features such as heart rate variability, electrodermal conductivity, and electromyography.

[0025] Step 1.2, emotional feature pulse frequency mapping; Construct an emotion intensity pulse frequency mapping function to map the emotion intensity value to the corresponding pulse frequency. The mapping function is defined as: ; in, represents the pulse frequency after mapping, Indicates the emotional intensity value, , and Represent the minimum and maximum pulse frequencies, and are the power exponential parameter and the semi-saturation parameter respectively. The mapping function has nonlinear characteristics and can more finely distinguish the emotional states of different intensity levels.

[0026] Step 1.3, encoding of spatiotemporal patterns of emotion categories; A spatiotemporal coding model is used to represent different emotion categories, mapping each basic emotion category to a specific spatiotemporal distribution pattern of pulses. The coding equation is: ; in, Indicates the spatiotemporal coding patterns of emotion categories, Indicates the number of pulses, Indicates the The weight of the pulse, Indicates the The time point of the pulse, Represents the unit pulse function. Different emotion categories form unique encoding patterns through the time distribution and weight distribution of pulses.

[0027] Step 1.4, multimodal pulse train generation; Based on the above mapping and encoding, a pulse sequence representing emotional information is generated. The mathematical expression of the pulse sequence is: ; in, is the pulse sequence generated, represents the unit impulse function, For the The time point of the pulse occurrence, The emotional coding information carried by the pulse.

[0028] Step 2: Based on the obtained pulse sequence, a continuous spectrum emotion representation space is constructed to achieve continuous representation of emotional states; The specific implementation steps are as follows: Step 2.1, determination of basic emotion dimensions; Based on psychological research, we identify a set of basic emotional dimensions as the basis vectors of the emotional representation space. We select emotional dimensions with a psychological basis, including but not limited to pairs such as joy-sadness, anger-fear, and surprise-disgust, to form the basic framework for emotional representation.

[0029] Step 2.2, continuous emotion representation space construction; Construct a multi-dimensional continuous emotion representation space and represent each emotional state as a point or region in the space. The representation space is defined as: ; in, 、 、 Respectively represent 、 、 The proportion of basic emotions in mixed emotions, 、 、 Indicates the 、 、 The intensity value of the basic emotion, is the number of basic emotions.

[0030] Step 2.3, emotional state density distribution construction; For each point in the emotion representation space , define the probability density function of emotional state: ; in, Indicates the point in the representation space The probability density of the emotional state at Indicates the most likely emotional state point, express and The Euclidean distance between is the standard deviation parameter of the distribution, is the normalization constant, Represents the exponential function.

[0031] Step 2.4, emotion vector embedding learning; Using self-supervised learning methods, we learn the embedding representation of emotion vectors from multimodal emotion data, so that semantically similar emotional states have a small distance in the representation space. The learning objective function is: ; in, represents the triplet loss function, Indicates the maximum value, Indicates emotional state and The distance measure between represents the set of positive sample pairs, Represents Dissimilar emotional states, is the boundary parameter.

[0032] Step 3: Using the continuous spectrum emotion representation space, generate pulse codes of mixed emotions to represent complex emotional states; The specific implementation steps are as follows: Step 3.1, basic emotion pulse template generation; A characteristic pulse template is generated for each basic emotion category as the basis for mixed emotion encoding. basic emotions, the pulse template is defined as: ; in, Indicates the intensity No. Pulse templates of basic emotions, is the amplitude modulation function related to the intensity of emotion, is the number of pulses in the template, Indicates the The weight of the pulse, Indicates the The time point of the pulse, represents the unit impulse function.

[0033] Step 3.2, modeling the interaction between emotions; Construct an interaction model between emotions to capture the interactive effects of different emotional components when they exist simultaneously. and emotions The interaction between them defines the interaction function: ; in, Express emotions and emotions The interaction strength between and Respectively represent the proportion of the two emotions in the mixed emotions, and Represents the intensity of two emotions, is the interaction coefficient, is the intensity-dependent modulation function.

[0034] Step 3.3, construction of mixed emotion pulse coding function; Based on the basic emotion pulse template and the emotion interaction model, the pulse coding function of mixed emotions is constructed: ; in, Indicates mixed emotional states In time The pulse code at A mixed emotional state, Represents the number of basic emotions, Indicates the The proportion of basic emotions, Indicates strength, The impulse generating function representing a single emotion, Express emotions and emotions The interaction encoding function between .

[0035] Step 3.4, attention modulation algorithm integration; The attention modulation algorithm is introduced to dynamically adjust the weights of different emotional components according to the current task and environmental factors, thereby enhancing the contextual adaptability of the encoding. The mathematical expression of attention modulation is: ; in, Indicates the situation Mixed emotional state In time The pulse code after attention modulation at Indicates the situation Next pair The attention weight of the emotion, Indicates the situation Lower emotions and emotions The attention weights of the interactions, A mixed emotional state, Represents the number of basic emotions, Indicates the The proportion of basic emotions, Indicates strength, The impulse generating function representing a single emotion, Express emotions and emotions The interaction encoding function between .

[0036] Step 4: Based on the pulse coding of mixed emotions, an emotional state transition probability model is established to describe the transition relationship between emotions; The specific implementation steps are as follows: Step 4.1, emotional state sequence extraction; Extract the emotional state sequence at discrete time points from the continuous time pulse code. By decoding the pulse sequence , get the emotional state of the time window , forming a sequence of emotional states ,in, 、 、 Respectively represent 、 、 an emotional state, is the sequence length, is the time window width.

[0037] Step 4.2, recursive spiking neural network construction; A recursive spiking neural network (RSNN) model is constructed to extract temporal features and dependencies from the emotional state sequence. The neuron dynamic equation of RSNN is: ; If and only if and ; in, represents the rate of change of membrane potential over time, Indicates the neurons at time The membrane potential, is the membrane potential time constant, From neurons to neurons The synaptic weights, Represents neurons No. Pulse emission time, is the postsynaptic potential kernel function, is the external input current, Represents neurons In time The pulse output, represents the unit impulse function, Represents neurons No. Pulse emission time, The firing threshold is when the membrane potential exceeds this threshold, the neuron fires a pulse. The refractory period is the recovery time after a neuron fires a pulse, during which the neuron cannot fire a pulse again.

[0038] The recursive spiking neural network in this embodiment adopts a three-layer structure: input layer, hidden layer, and output layer. The input layer receives the pulse encoding of the emotional state sequence, the hidden layer is composed of multiple recursively connected spiking neurons, including feedback connections to capture temporal dependencies, and the output layer generates a pulse sequence that represents the probability of emotional state transitions. The specific implementation method is as follows: Network structure: The input layer contains Spiking neurons correspond to the dimensions of emotional state; The hidden layer contains spiking neurons with recurrent connections; The output layer contains Spiking neurons correspond to possible emotional state transitions.

[0039] Synaptic connections: The connection weights from the input layer to the hidden layer Training is performed through supervised learning after random initialization; Recurrent connection weights within the hidden layer Use FORCE learning algorithm for optimization; The connection weights from the hidden layer to the output layer Adjustment is performed through the back-propagation algorithm.

[0040] Learning algorithm: The network is trained using a remote supervised learning algorithm (RSM) based on error backpropagation, combined with spike timing-dependent backpropagation (STDBP) to adjust synaptic weights, enabling the network to accurately predict transitions in emotional states.

[0041] When applied in facial micro-expression recognition scenarios, the recurrent spiking neural network is able to capture subtle temporal dynamic changes in micro-expression sequences.

[0042] Step 4.3, Markov switching model construction; Based on the hidden state representation of the recurrent spiking neural network, a high-order Markov transition model is constructed to capture the conditional probability distribution of emotional state changes. The transition probability model is defined as: ; in, Indicates a given past The conditional probability of the emotional state at the next time step is given by the emotional state at the time step. 、 、 Represents the time step 、 、 emotional state, Indicates the length of the historical time window, Represents the time step The predicted emotional state, represents a parameterized neural network model, are model parameters, represents all possible emotional states at the next time step, Represents the exponential function.

[0043] Step 4.4, emotional inertia and conversion impedance model; The emotional inertia and conversion impedance model is introduced to express the dynamic characteristics of emotional state changes. and The conversion between , defines the conversion impedance function: ; in, Indicates emotional state Convert to The impedance, Indicates emotional state and The Euclidean distance in the representation space, and Represent emotional states and Middle The proportion of basic emotions, and Represent emotional states and The corresponding intensity, 、 、 are the weight parameters for controlling the Euclidean distance term, the emotion proportion difference term, and the emotion intensity difference term, respectively.

[0044] The relationship between conversion impedance and conversion probability is: ; in, Indicates emotional state Shifting to an emotional state The conditional probability of Indicates emotional state Convert to The impedance, is the temperature parameter, Indicates a proportional relationship, Represents the exponential function.

[0045] Step 5: Based on the emotional state transition probability model, the gradual emotional change process is predicted and processed to capture subtle changes in emotions; The specific implementation steps are as follows: Step 5.1, phase gradient modulation algorithm construction; Construct a phase gradient modulation algorithm to encode the rate of change of emotions by adjusting the phase characteristics of neural pulses, and achieve a fine expression of the gradual change of emotions. , and its phase function is defined as: ; in, Indicates time The phase value at is the initial phase, Indicates the The time derivative of the ratio of the emotional components, Represents the number of basic emotions, is the weight coefficient. Through this phase modulation, the system can encode the changing dynamics of emotions in the temporal structure of the pulse sequence, thereby distinguishing static emotional states from dynamic changing processes.

[0046] The phase gradient modulation algorithm in this embodiment adopts a multi-channel phase encoding method to achieve high-precision encoding of the emotion change rate by modulating the phase characteristics of different frequency carriers. The specific implementation method is as follows: Multi-frequency phase encoding: Select Carrier frequency (Typical value is , frequency range is 20-100Hz), a set of carrier waves is assigned to each basic emotion, and the rate of change of emotion is expressed by modulating the phase of these carrier waves, where 、 、 Respectively represent 、 、 carrier frequencies, represents the number of carrier frequencies. The complete phase function expands to: ; in, Indicates the The phase function of the carrier, is the initial phase, For emotions Carrier The weight coefficient of Represents the number of basic emotions, Indicates the The time derivative of the ratio of the emotional components, Indicates the number of carrier frequencies.

[0047] Phase modulation is achieved through a voltage-controlled oscillator (VCO) model. The output signal is: ; in, For the Modulation output of each channel, is the amplitude, which controls the signal strength, is the carrier frequency, is a sine function, is the fundamental phase of the carrier, Phase modulation caused by emotional changes.

[0048] Multi-channel fusion: The modulated outputs of the channels are weighted and fused to obtain the final phase modulated pulse sequence: ; in, is the fused phase modulated pulse sequence, For the The weight coefficient of each channel, Indicates the number of carrier frequencies, For the Modulation output of each channel.

[0049] When applied in the emotion gradient analysis scenario, the phase gradient modulation algorithm can accurately capture subtle changes in emotion intensity.

[0050] Step 5.2, prospective pulse generation algorithm; Based on the emotional state transition probability model, a forward-looking pulse generation algorithm is constructed to predict the emotional state changes in the future time period. emotional state Its historical status , predicting future time emotional state , and generate the corresponding predictive pulse sequence, the prediction equation is: ; in, Expressing the future The predictive pulse sequence, Pulse code representing the current emotional state, Indicates that in the known current and historical Under the condition of an emotional state, the future moment Emotional state The conditional probability of 、 、 Represents the time step 、 、 emotional state, Indicates the length of the historical time window, is the prediction function, is the forecast time span.

[0051] Step 5.3, emotion trajectory smoothing interpolation algorithm; Construct an emotion trajectory smooth interpolation algorithm to continuously interpolate between the emotion states at discrete time points to achieve a smooth expression of the emotion gradual change process. and emotional state and , define time The interpolation state at : ; in, Indicates time The interpolated emotional state at Indicates at a point in time The emotional state, is the time normalization factor, Indicates at a point in time The emotional state, represents the difference vector between two adjacent emotional states, is a smooth function.

[0052] The emotion trajectory smoothing interpolation algorithm in this embodiment uses an adaptive nonlinear spline interpolation method, combined with the emotion physics model, to achieve a detailed description of the emotion gradual change process. The specific implementation method is as follows: Adaptive control point selection: for discrete state points on the emotion trajectory , adaptively select control points according to the changing characteristics of emotional state, where, 、 、 Represents the emotional trajectory 、 、 A point in time, Indicates the total number of time points in the emotion trajectory. Increase the density of control points in areas where emotions change dramatically, and reduce the density of control points in areas where emotions change slowly. The control points are selected based on the change rate threshold. : ; in, is the selected control point set, 、 Indicates at a point in time 、 The emotional state, Represents the Euclidean distance between adjacent emotional states, which is used to measure the intensity of emotional changes. is the basic sampling interval, is the rate of change threshold, Represents an index Can be Condition for divisibility.

[0053] Physically constrained smooth function: Based on the physical characteristics of the emotional state, a smooth function that satisfies dynamic constraints is constructed: ; in, represents a smooth function that satisfies the dynamic constraints, is the normalized time, is the emotional state conversion impedance, is the impedance influence coefficient, The cubic Hermite interpolation polynomial based on is used to ensure the continuity of the interpolation curve at the control points. is the impedance modulation term, which makes the high impedance emotion conversion present nonlinear characteristics.

[0054] Multi-resolution trajectory generation: A pyramid structure is used for multi-resolution emotion trajectory generation, which first interpolates at the coarsest level and then gradually refines it: ; in, Indicates the Level-resolution emotional trajectories, Indicates the emotional trajectory of the previous level, For the The basis functions of level , is the weight coefficient, For the The number of basis functions at each level is set. Through coarse-to-fine layered interpolation, the smoothness of the global trajectory and the accuracy of local details are maintained.

[0055] In the application of movie sentiment analysis, this algorithm can construct a continuous emotional trajectory and accurately depict the audience's emotional changes during the viewing process.

[0056] Step 5.4, emotional change detection and response system; Build an emotion mutation detection and response system to identify sudden changes in the emotional state sequence and quickly adjust the prediction model. , define the sentiment change rate: ; in, Indicates a time point The rate of change of emotions, Indicates the previous time point The emotional state vector, represents the Euclidean distance between two emotional state vectors, Represents the time interval between two adjacent time points. 、 、 Represents the emotional trajectory 、 、 A point in time, Represents the total number of time points of the emotion trajectory.

[0057] when Exceeding the preset threshold When the system detects a sudden change in sentiment, it triggers the following response actions: Adjust the forecast time window, shorten the forecast time scale, and improve the accuracy of short-term forecasts; Recalculate the transition probability matrix to improve adaptability to new emotional dynamics; Increase the weight of recent emotional states and weaken the influence of historical states.

[0058] Through this adaptive approach, the system can quickly respond to sudden changes in emotional state and adjust its prediction strategy to maintain accurate tracking of emotional dynamics.

[0059] Real-world application examples of this implementation: Student emotion tracking and personalized teaching system in intelligent education: This embodiment of the brain-inspired multimodal emotion recognition method is applied to intelligent education scenarios, building a student emotion tracking and personalized teaching system. This system uses cameras, microphones, and wearable devices to collect multimodal data such as students' facial expressions, voice, and physiological signals. Combined with the emotion recognition method of this embodiment, it accurately identifies and predicts students' emotional states during learning, thereby providing personalized teaching strategies and content adjustments.

[0060] System architecture and data collection: The system adopts a distributed architecture, including the front-end collection layer, edge computing layer and cloud analysis layer: Front-end acquisition layer: High-definition cameras (30fps, 1080p) are installed in the classroom to capture students' facial expressions; a microphone array (sampling rate 48kHz) collects voice signals; and smart bracelets (sampling rate 100Hz) worn by students can optionally collect physiological signals such as heart rate and skin conductivity.

[0061] Edge computing layer: Edge computing devices (8-core CPU, 16GB RAM, dedicated neural network acceleration chip) are deployed in classrooms to run a simplified version of the pulse neural network to handle real-time emotion recognition tasks with latency controlled within 50ms.

[0062] Cloud-based analysis layer: Deploys a complete brain-inspired multimodal emotion recognition system to perform deep analysis and emotion prediction tasks, generate personalized teaching suggestions, and continuously optimize the model based on new data.

[0063] Real-time emotion recognition and prediction examples: Taking mathematics classes as an example, the system's effectiveness in identifying and predicting students' emotions in actual applications is shown in the following table: Table 1: Comparison of students’ emotion recognition and prediction effects in mathematics classroom;

[0064] When the system detects a student's trend of shifting from "concentration" to "confusion" and "frustration", it can predict this change 1.5-2.1 seconds in advance, giving the teaching system enough time to adjust the teaching strategy.

[0065] Emotional trajectory analysis and learning behavior association: The system uses an emotion trajectory smoothing interpolation algorithm to generate a curve of students' emotions throughout the course, and associates it with learning behavior and learning outcomes. The following table shows the relationship between different emotion trajectory patterns and learning outcomes: Table 2: Association between emotional trajectory patterns and learning outcomes;

[0066] Based on emotional trajectory pattern analysis, the system triggers personalized intervention at the appropriate time. For example, for students with "fluctuating downward type", when a "confusion-frustration" transition trend is detected, targeted concept explanations and encouragement are provided, successfully transforming 78% of such emotional trajectories into "U-shaped" or "fluctuating types", thereby improving learning outcomes.

[0067] Teaching strategy adaptation and effect verification: The system automatically adjusts teaching strategies and learning content based on emotion recognition and prediction results. The following table shows the effectiveness of the system in real classrooms: Table 3: Verification of the effect of emotion perception teaching intervention;

[0068] In a one-semester controlled experiment, the average grades of students in the experimental class that received assisted teaching from this system increased by 17.3%, and their learning enthusiasm increased by 32.5%.

[0069] Mental health monitoring and emotion regulation system: The brain-like multimodal emotion recognition method of this embodiment is applied to the field of mental health, and a set of emotion monitoring and regulation systems are constructed, which can be used for emotion monitoring, early warning and intervention in patients with anxiety and depression.

[0070] System architecture and data collection: The system adopts a "end-cloud" architecture, including: Mobile terminal layer: User smartphones (collecting voice, text, and facial expressions) and smart wristbands (collecting physiological and behavioral data such as heart rate variability, skin conduction, and activity level). Mobile terminals run a lightweight spiking neural network model, performing basic sentiment analysis every 10 minutes.

[0071] Cloud analysis layer: Deploy a complete brain-like multimodal emotion recognition system to perform deep emotion analysis, emotion trajectory generation, and abnormal warning tasks, and generate personalized intervention recommendations.

[0072] While protecting user privacy, the system collects the following data: voice emotion features (original speech is not stored), text emotion content (processed using local encryption), facial expression feature points (original images are not stored), physiological signal features, and daily activity patterns.

[0073] Examples of abnormal emotion detection and warning: The system uses a continuous spectrum of emotion representation and a probabilistic model of emotional state transitions to identify abnormal emotional patterns and predict trends of emotional deterioration. The following table shows the system's performance in detecting abnormal emotions in clinical applications: Table 4: Emotional anomaly detection and warning effects;

[0074] The system can detect a trend of worsening depression 1.5 days in advance and warn of anxiety attacks 4.3 hours in advance, giving medical teams and patients enough time to take preventive measures.

[0075] Emotion Regulation Intervention Strategies and Effects: The system generates personalized emotion regulation suggestions based on emotional state analysis and transition probability modeling. The following table shows the application scenarios and effects of different intervention strategies: Table 5: Emotion regulation intervention strategies and effects;

[0076] The system generates a continuous emotion curve based on the emotion trajectory smoothing interpolation algorithm, triggering intervention at the optimal time to improve the intervention effect.

[0077] Long-term effect verification and clinical application: In clinical trials, the experimental group using this system for emotion monitoring and intervention achieved significantly better treatment results compared to the control group (which only received conventional treatment): Table 6: Clinical application effect of the system;

[0078] The system has the following advantages in practical applications: Objective quantification of emotional states: Through continuous spectrum emotion representation, it provides detailed quantitative data of emotional states to assist clinicians in evaluating disease conditions and treatment effects.

[0079] Early warning of mood changes: By modeling the probability of mood state transitions, early warning of mood deterioration can be achieved, making intervention measures more forward-looking.

[0080] Personalized intervention plan: Based on the emotion change analysis of the phase gradient modulation algorithm, the system can identify each patient's emotion change pattern and adjustment needs, and generate personalized treatment recommendations.

[0081] Improved treatment compliance: Through timely emotional feedback and effective intervention suggestions, patient treatment compliance increased from 65.3% in the control group to 82.7% in the experimental group.

[0082] Emotional continuum representation method: In this embodiment, the emotion continuum representation method converts discrete basic emotion categories into a continuous emotion space representation, which can more accurately capture the gradual change and mixing of emotions. The method includes the following steps: Nonlinear manifold construction based on emotion perception features: This method first constructs a nonlinear manifold based on emotion perception features and uses a variety of technical means to capture the complex relationships between emotions.

[0083] Emotion Perception Feature Definition: Emotional perception characteristics include the following three dimensions: Emotional arousal ( ): Indicates the degree of emotional activation, with a value range of [0,1].

[0084] Emotional valence ( ): Indicates the positive or negative nature of emotions, with a value range of [-1,1].

[0085] Emotional dominance ( ): Indicates the sense of control over emotions, with a value range of [0,1].

[0086] Based on these three basic dimensions, derived features are further defined: Emotional Mixture ( ): The degree of mixing of different emotional components is calculated as follows: ; in, Indicates the degree of emotional mixing, Indicates the intensity value of basic emotions, Indicates the maximum value, Indicates the average value.

[0087] Mood change rate ( ): The rate of change of emotional state over time, calculated as: ; in, represents the rate of change of emotion, Indicates time emotional state, represents the emotional state vector at the previous time point, Represents the time interval between two adjacent time points. Indicates the number of sampling points in the time series.

[0088] Manifold construction method: The manifold construction adopts the Weighted Locality Preserving Projections (WLPP) method, which is implemented by the following steps: Construct a neighbor graph: For each point in the sentiment feature space , find its k nearest neighbor point set .

[0089] Calculate the weight matrix: for neighboring points , calculate the weight : ; in, represents the weight between two points, Indicates a point and point The square of the Euclidean distance in feature space, represents the bandwidth parameter of the Gaussian kernel, represents the exponential function, is the sentiment similarity, defined as: ; in, is the sentiment similarity, 、 、 Represent points valence vector, arousal and dominance of 、 、 Represent points valence vector, arousal and dominance of Represents the cosine function.

[0090] Construct the Laplacian matrix: ; in, represents the Laplace matrix, is a diagonal matrix, represents the weight matrix.

[0091] Manifold Embedding: Solving Optimization Problems: ,satisfy ; This problem is equivalent to solving the generalized eigenvalue problem: ; in, Represents the projection matrix Find the minimum value, is the Laplace matrix, is a diagonal weight matrix, is the projection matrix, is the identity matrix, represents the sum of the diagonal elements of the matrix, represents the eigenvalue in the generalized eigenvalue problem, represents the generalized eigenvector.

[0092] Detailed modeling of emotion gradient trajectories: Gradient trajectory representation: The emotion gradient trajectory is represented by a parameterized cubic B-spline curve: ; in, For time emotional state, is the control point, for Order B-spline basis function, Control Point Determined by interpolation of emotional key points, the following constraints are met: is the number of control points: Emotional continuity constraints: Continuously differentiable at all time points; Physiological constraints: The rate of change of emotions is limited by physiological constraints. ,in, is the maximum rate of change of emotion; Psychological constraints: The acceleration of emotional changes is limited by psychological constraints. ,in, is the maximum acceleration of emotion change.

[0093] Keypoint adaptive selection: In order to accurately capture emotional changes, this method proposes a key point adaptive selection algorithm: Initial key point selection: In the emotion sequence, according to the emotion change rate Select local extreme points as initial key points.

[0094] Keypoint optimization: For each initially selected keypoint set , calculate the error between the trajectory model constructed using this set and the actual emotion sequence : ; in, represents the error between the trajectory model and the actual emotion sequence, represents the total number of evaluation time points, The index of the time point. Indicates at a point in time The emotional state vector predicted by the B-spline curve model, Indicates at a point in time The actual observed emotional state vector; If the error exceeds the threshold , then the key point is increased; if the error is less than half of the threshold, the key point can be reduced.

[0095] Determination of the optimal key point set: Use a dynamic programming algorithm to find a set that uses the least key points while meeting the accuracy requirements.

[0096] Quantitative expression of mixed emotions: This method proposes a quantitative expression model for mixed emotions, which can accurately describe the complex state of multiple emotions coexisting.

[0097] Mathematical model of mixed emotions: The mixed emotion is represented as the weighted sum of emotion basis vectors: ; in, A mixed emotional state, is the basic emotion vector, is the total number of basic emotions, is the mixing weight, satisfying and ; In order to express the mutual inhibition or enhancement relationship between emotions, the emotion interaction matrix is introduced , the modified mixed model is: ; in, represents the final mixed emotional state, 、 Separate 、 The weight of the basic emotions, is the total number of basic emotions, Express emotions and emotions The interaction strength, Represents the new emotion component generated by the interaction.

[0098] Mixed Emotions Dynamic Balance: Mixed emotions have a dynamic equilibrium process over time, which can be expressed as: ; in, Express emotions The rate of change of the weight over time, For emotions Emotions The conversion rate, For emotions The decay rate, External stimuli affect emotions The activation strength, 、 Respectively indicate time Time Emotion 、 The weight or intensity coefficient of The total number of basic emotions represented; By solving this set of differential equations, the evolution of mixed emotions over time can be obtained.

[0099] Implementation effect verification: The emotion continuum representation method of this embodiment was verified on multiple test sets, and the results are shown in the following table: Table 7: Performance comparison of emotion continuum representation methods on different datasets;

[0100] As can be seen from the table, this method is significantly superior to traditional discrete emotion classification methods in terms of indicators such as emotion gradient accuracy, mixed emotion recognition accuracy, and emotion state prediction accuracy.

[0101] Emotional multimodal collaborative fusion method: Multimodal sentiment feature extraction: This method first extracts emotional features from each modality, including the following main modalities: Facial expression modality feature extraction: Facial expression feature extraction includes two levels: Low-level features: 68 key points on the face and their dynamic change characteristics; High-level features: facial action unit (AU) activation intensity and combination patterns.

[0102] The specific implementation steps are as follows: Extracting facial region feature maps using an improved densely connected convolutional network ; Apply spatial attention mechanism to highlight key areas of facial expressions: ; in, represents the spatial attention weight matrix, represents the sigmoid activation function, represents the learnable weight matrix, represents the facial region feature map, represents the facial feature map after spatial attention weighting, represents the Hadamard product; Calculate facial action unit activation strength vector : ; in, represents the facial action unit activation strength vector, represents a multilayer perceptron network, Represents the facial feature map after spatial attention weighting; Extract dynamic change features of facial key points : ; in, Indicates the dynamic change characteristics of facial key points, 、 Respectively indicate time hour, Time The location of a key point.

[0103] Speech modality feature extraction: Speech feature extraction is divided into two parts: acoustic features and semantic features: Acoustic feature extraction: Basic features: pitch (F0), energy, formant frequency (F1-F4), harmonic-to-noise ratio (HNR), etc. Statistical characteristics: mean, standard deviation, skewness, kurtosis and other statistics; Temporal features: An autoregressive model is used to capture the temporal changes of acoustic features.

[0104] Speech emotion expression pattern recognition: Design a multi-scale 1D-CNN network to extract acoustic patterns at different time scales; Apply the self-attention mechanism and calculate the attention weight: ; in, represents the time step, represents the softmax function, represents the attention weight matrix, represents the hyperbolic tangent activation function, represents the hidden state transformation matrix, Represents the time step The hidden state vector of represents the bias term; Weighted fusion obtains speech emotion features: ; in, represents the final fused speech emotion feature vector, represents all time steps, Represents the time step The attention weight, Represents the time step The hidden state vector of .

[0105] Text modality feature extraction: Text features include semantic features and sentiment word features: Semantic feature extraction: Extract contextual word embeddings using a pre-trained language model; Design a bidirectional GRU network to capture long-distance semantic dependencies; A hierarchical attention mechanism is introduced to calculate attention weights at the word level and sentence level respectively.

[0106] Sentiment word feature extraction: Construct a sentiment dictionary and assign sentiment intensity and sentiment category to each sentiment word; Calculate the sentiment word density and sentiment word distribution of the text; The relative position encoding of sentiment words is introduced to indicate the relative position of sentiment words in a sentence.

[0107] Fusion of semantic and sentiment word features: ; in, represents the final fused text feature vector, represents the semantic feature vector extracted from the text, Represents the sentiment word feature vector extracted from the text, is the adaptive weight.

[0108] Pulse coded modality fusion network: Modal characteristic pulse conversion: Convert each modal feature into a pulse sequence as follows: Feature to pulse rate mapping: ; in, is modal Features, is modal The pulse rate, Represents the mapping function from features to pulse rates.

[0109] Pulse generation: The pulse train is generated using a Poisson process: ; in, Indicates modality In time The pulse sequence, represents a Poisson random process, Indicates modality In time Pulse rate; At the same time, a random threshold reset model is introduced to enhance biological interpretability: ; in, express The neuronal membrane potential at that moment, express The neuron membrane potential at each moment, when the membrane potential exceeds the threshold, the neuron emits a pulse and resets, otherwise the membrane potential is updated according to the input and attenuation, express Input current at all times, is the membrane potential attenuation coefficient, is the reset potential value after the neuron fires, is the firing threshold of the neuron.

[0110] Temporal coding: Encoding the temporal change information of emotions into the pulse time interval: ; in, Indicates modality The pulse time interval, represents the time encoding function, Indicates modality The temporal variation characteristics of .

[0111] Multimodal pulse co-processing: Design a pulse co-processing network, including the following core components: Intermodal suppression / enhancement mechanism: ; in, Indicates modality In time The processed output pulse, Indicates modality In time The original pulse input, Indicates modality Modal The influence weight of represents the propagation delay, Indicates modality In time Pulse input, Indicates that the modal To Modal propagation delay.

[0112] Modal reliability assessment: ; in, Indicates modality The reliability score, represents the reliability evaluation function, represents modal consistency, represents the signal-to-noise ratio, Indicates timing stability.

[0113] Modal Adaptive Weighting: ; in, Indicates modality The weight coefficient in the final fusion, 、 Represents the mode 、 The reliability score, is the temperature parameter, Represents the exponential function.

[0114] Impulse response decoding: Decode the fused spike train into an emotional state representation: Membrane potential accumulation: ; in, Indicates time The accumulated membrane potential at time Indicates modality The weight coefficient of is the membrane time constant, Indicates modality In time The output pulse after processing at any moment, represents an exponential decay function.

[0115] Emotional state decoding: ; in, Indicates time The emotional state representation obtained by moment decoding, represents the emotional state decoder, represents the cumulative membrane potential, Represents all modes at time The output pulse set of The decoder adopts a spatiotemporal attention mechanism to focus on important time points and modalities.

[0116] Multimodal missing and inconsistent processing: Modality missing processing: Modal reconstruction: Designing complementary reconstruction networks between modalities: ; in, Represents the reconstructed mode The feature representation of Indicates the reconstruction mode The neural network model, Indicates that the mode is removed All other modal feature sets except ; The training objectives include reconstruction loss and feature distribution matching: ; in, represents the reconstruction loss function, represents the square of the Euclidean distance between the original feature and the reconstructed feature, represents the KL divergence between the original feature distribution and the reconstructed feature distribution, 、 Represent the original modal features and reconstructed modal features The probability distribution of .

[0117] Zero-shot modality transfer: Using a shared semantic space to achieve migration from a known modality to an unknown modality: ; in, represents the predicted feature representation of the unknown (missing) modality, is the migration function, represents the feature representation of the known (available) modality, It’s contextual information.

[0118] Modal inconsistency handling: Inconsistency detection: Calculate the consistency score of emotion representation between modalities: ; in, represents the consistency score of emotion representation between modalities, and They are modal and emotional representation, Represents the cosine function.

[0119] Conflict Mediation: When a conflict in emotional representation between modalities is detected, the system will mediate the conflict based on the reliability of each modality and the contextual situation; Introducing the Bayesian framework, calculate the posterior probability of each emotion hypothesis: ; in, Represents given all modal features Conditions, emotional state The posterior probability of Indicates emotional state Under the condition, the current modal feature set is observed The likelihood probability, Indicates emotional state The prior probability of Indicates a proportional relationship; The emotional state with the highest posterior probability is selected as the final output.

[0120] Inter-modal collaborative reinforcement learning: Design comparative learning objectives to bring different modal representations of the same emotional state closer together and push different emotional states further apart; The mutual information maximization principle is introduced to enhance the semantic consistency between modalities.

[0121] Implementation effect verification: The emotional multimodal collaborative fusion method of this embodiment was verified on multiple multimodal emotional datasets, and the results are shown in the following table: Table 8: Performance comparison of emotion multimodal collaborative fusion methods on different datasets;

[0122] As can be seen from the table, this method has significant improvements over traditional feature-level fusion and decision-level fusion methods. In particular, the performance improvement is more obvious in the cases of modality loss and modality inconsistency, reaching 50.3% and 63.0% respectively.

[0123] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A brain-like multimodal emotion recognition method, characterized in that: The following steps are involved: Convert multimodal emotional information into pulse trains; Based on the obtained pulse sequence, a continuous spectrum emotion representation space is constructed to achieve continuous representation of emotional states; Using the continuous spectrum emotion representation space, pulse codes of mixed emotions are generated to represent complex emotional states; According to the pulse coding of mixed emotions, an emotional state transition probability model is established to describe the transition relationship between emotions; Based on the emotional state transition probability model, the gradual change of emotions can be predicted and processed to capture subtle changes in emotions.

2. The brain-like multimodal emotion recognition method according to claim 1, characterized in that: The method for converting multimodal emotion information into a pulse sequence comprises: Through the multimodal feature extraction network, emotional features of facial expressions, speech and text modalities are extracted respectively; Map the extracted emotional features to the pulse emission rate and construct a mapping function from features to pulse rate; The pulse train is generated based on the random threshold reset model, and the membrane potential update formula is: ; in, express The neuronal membrane potential at that moment, express The neuron membrane potential at each moment, when the membrane potential exceeds the threshold, the neuron emits a pulse and resets, otherwise the membrane potential is updated according to the input and attenuation, express Input current at all times, is the membrane potential attenuation coefficient, is the reset potential value after the neuron fires, The firing threshold for the neuron; The temporal change information of emotion is encoded into the pulse time interval to form a time-coded pulse sequence.

3. The brain-like multimodal emotion recognition method according to claim 1, characterized in that: The method for constructing a continuous spectrum emotion representation space includes: Construct a nonlinear manifold based on emotion perception characteristics, including three basic dimensions: emotion arousal, emotion valence, and emotion dominance; Introducing emotion mixing degree and emotion change rate as derived features; The manifold is constructed using a weighted locality preserving projection method by solving the optimization problem: ; satisfy: ; in, Represents the projection matrix Find the minimum value, is the Laplace matrix, is a diagonal weight matrix, is the projection matrix, is the identity matrix, represents the sum of the diagonal elements of a matrix; A parameterized cubic B-spline curve is used to represent the emotion gradient trajectory, which satisfies the emotion continuity constraint, physiological constraint and psychological constraint.

4. The brain-like multimodal emotion recognition method according to claim 1, characterized in that: The method for generating pulse codes of mixed emotions comprises: Represent the mixed emotion as a weighted sum of emotion basis vectors: ; in, A vector representation of mixed emotions, For the The mixed weights of the basic emotions, For the The feature vectors of basic emotions, Express emotions and emotions The interaction strength coefficient between Express emotions and emotions The interaction feature vector of is the total number of basic emotions defined in the system; Establish a dynamic equilibrium model of mixed emotions to describe the evolution of mixed emotions over time; Encode mixed emotional states into multi-channel pulse trains, with different channels corresponding to different emotional components; Express the intensity of emotions and the relationship between emotions through pulse frequency and phase modulation.

5. The brain-like multimodal emotion recognition method according to claim 1, characterized in that: The method for establishing the emotional state transition probability model comprises: Construct an emotional state transition matrix to represent the probability of transitioning from one emotional state to another; Introducing contextual factors to dynamically adjust the conversion probability according to the situation; Build a conditional probability model based on variational autoencoders to learn the latent distribution of emotion transitions; Utilize a sequence-to-sequence learning framework to predict the distribution of future emotional state sequences under given conditions.

6. The brain-inspired multimodal emotion recognition method according to claim 1, characterized in that: The method for predicting and processing the gradual change of emotions includes: Design a phase gradient modulation algorithm to encode the rate of change of emotion by modulating the phase of the pulse sequence; Use an emotion trajectory smoothing interpolation algorithm to generate natural transitions between discrete emotion states; Build a predictive model of the gradual change of emotions, and predict future emotional changes based on historical emotional trajectories and current situations; An adaptive threshold mechanism is introduced to detect emotion change inflection points and emotion conversion events.

7. The brain-inspired multimodal emotion recognition method according to claim 1, characterized in that: It also includes multimodal collaborative fusion steps: Design a modal fusion mechanism based on spiking neural networks, including inter-modal suppression or enhancement mechanisms, modal reliability assessment, and modal adaptive weighting; Construct a multimodal loss and inconsistency handling mechanism, including modality reconstruction, zero-shot modality transfer, inconsistency detection, and conflict mediation; Implement impulse response decoding based on membrane potential accumulation and decode the fused pulse train into emotional state representation; Inter-modal collaborative reinforcement learning is introduced to enhance the semantic consistency between modalities through contrastive learning and mutual information maximization.

8. A brain-inspired multimodal emotion recognition network, configured to execute the brain-inspired multimodal emotion recognition method according to any one of claims 1 to 7, characterized in that: Applied to intelligent education systems, including: Build a system architecture that includes the front-end data collection layer, edge computing layer, and cloud analysis layer; Achieve real-time recognition and prediction of students' emotional states, supporting higher accuracy than traditional methods; Analyze the correlation between emotional trajectories and learning behaviors, and discover the association between emotional patterns and learning outcomes; Generate adaptive teaching strategies based on emotion recognition results to improve learning outcomes.

9. The brain-inspired multimodal emotion recognition network according to claim 8, characterized in that: It is also used in mental health monitoring and emotion regulation systems, including: Collect multimodal data from users through smartphones and wearable devices; Identify abnormal emotional patterns and predict the trend of emotional deterioration, and provide early warning of emotional abnormalities; Generate personalized emotion regulation suggestions based on emotion state analysis and transition probability modeling; When used in clinical practice, it significantly improves the treatment effect for patients with depression and anxiety.

10. A brain-inspired multimodal emotion robot, configured to execute the brain-inspired multimodal emotion recognition method according to any one of claims 1 to 7, characterized in that: Applied to emotional interaction robots to achieve natural and coherent human-computer emotional interaction.

Citation Information

Cited By

  • Asynchronous multi-mode emotion recognition method and device, equipment and medium

    CN120873762A

  • An asynchronous multi-modal emotion recognition method, device, equipment and medium

    CN120873762B

  • Speech synthesis method and system with emotion recognition capability

    CN121214907A