Emotion understanding enhancement method for multi-mode emotion interaction data fusion

By using a multimodal mutual intuition algorithm and an emotion state evolution network, modal weights and emotion recognition strategies are dynamically adjusted, solving the problems of instability and inaccuracy in emotion understanding in multimodal human-computer interaction. This enables continuous modeling of emotion changes and consistent responses, improving the naturalness and adaptability of human-computer interaction.

CN121834665APending Publication Date: 2026-04-10CHENGDU MINGTU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing multimodal human-computer interaction systems struggle to dynamically adjust modal importance in emotion understanding, leading to unstable and inaccurate emotion understanding results and a lack of continuous modeling and effective response mechanisms for user emotion changes.

Method used

By introducing a multimodal mutual perception algorithm and an emotion state evolution network, the weights of different modalities are dynamically adjusted to construct the user's emotion evolution trajectory, and the emotion recognition and task response strategies are adjusted in combination with interaction context information.

Benefits of technology

It improves the accuracy and stability of multimodal emotional information expression, enhances the naturalness and adaptability of human-computer interaction, and can more accurately reflect the trend of user emotional changes and generate consistent responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834665A_ABST
    Figure CN121834665A_ABST
Patent Text Reader

Abstract

The invention discloses an emotional understanding enhancement method based on multi-modal emotional interaction data fusion, and the method comprises the steps: collecting voice, text, image and other multi-modal data of a user in an interaction process, the multi-modal data comprises voice, text and image, and constructing a multi-modal time sequence feature flow; on the basis of the feature flow, dynamic weight adjustment of the features is achieved through a multi-modal mutual inductance algorithm, and real-time fused emotion features are obtained; constructing an emotional state evolution network by using the fused emotional features, and outputting a user emotional evolution trajectory; dynamically adjusting a classification boundary of an emotion recognition module in combination with the emotion evolution trajectory and interaction context information, and updating a user emotion state; according to the emotional state and the emotional expression intensity of the user, generating final task response content; according to the method, the problems of inflexible fusion and inaccurate recognition in multi-modal emotion recognition are solved, accurate judgment of the emotion state and response generation facing user requirements are realized, and more natural emotion interaction experience is provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to a method for enhancing emotion understanding through multimodal emotion interaction data fusion. Background Technology

[0002] With the continuous development of artificial intelligence technology, human-computer interaction systems based on voice, text, and visual information are widely used in scenarios such as intelligent customer service, virtual assistants, companion-style interaction, and intelligent education. Compared with traditional single-modal interaction methods, multimodal interaction systems can comprehensively utilize multi-source information from users during the interaction process, and conduct a more comprehensive analysis of user intentions and emotional states, thereby improving the naturalness and accuracy of the system's response.

[0003] Existing human-computer interaction systems achieve the recognition and understanding of users' emotional states by extracting and fusing features from multimodal data. However, in practical applications, multimodal emotion expression exhibits significant dynamism and instability, with different modalities contributing considerably to emotion expression at different stages of the interaction. Current multimodal feature fusion methods often rely on fixed weights or attention results obtained from a single computation, making it difficult to continuously adjust the importance of each modality as the interaction process changes, thus affecting the stability and accuracy of emotion understanding results.

[0004] Furthermore, existing emotion recognition methods mostly focus on determining the emotional state at the current moment, treating emotions as discrete, static labels, and lacking systematic modeling of how user emotions change over time. This approach fails to reflect the evolutionary trend of user emotions and is not conducive to interactive systems predicting and responding continuously to emotional changes.

[0005] Emotion recognition results are often output independently and are difficult to dynamically influence subsequent decision-making or task response modules. The lack of an effective linkage mechanism between emotion recognition and task response makes it difficult for the system-generated response content to maintain consistency with the user's true emotional state in terms of emotional expression.

[0006] Therefore, there is an urgent need for an emotion understanding method for multimodal interaction that can dynamically integrate and continuously model user emotions during multimodal interaction, and adaptively adjust emotion recognition and task response strategies in combination with interaction context information, thereby improving the naturalness and emotional consistency of human-computer interaction. Summary of the Invention

[0007] The purpose of this invention is to provide a method for enhancing emotion understanding through multimodal emotion interaction data fusion. By introducing a multimodal weight adaptive evolution mechanism based on mutual intuition energy, the contribution of different modal emotion information can be dynamically adjusted, thereby improving the stability and accuracy of emotion understanding results in complex interaction scenarios.

[0008] The objective of this invention is achieved through the following technical solution: A method for enhancing sentiment understanding through multimodal sentiment interaction data fusion includes the following steps: Step S1: Collect multimodal data of the user during the interaction process. The multimodal data includes voice, text and images, and construct a multimodal temporal feature stream; Step S2: Based on this feature flow, the multimodal mutual intuition algorithm is used to realize the dynamic weight adjustment of the features and obtain the real-time fused sentiment features; Step S3: Utilize the fused emotional features to construct an emotional state evolution network and output the user's emotional evolution trajectory; Step S4: Combine the emotion evolution trajectory with the interaction context information to dynamically adjust the classification boundary of the emotion recognition module and update the user's emotional state; Step S5: Generate the final task response content based on the user's emotional state and the intensity of their emotional expression.

[0009] Furthermore, step S1 specifically includes: Step S101: Asynchronously acquire multimodal data, including speech signal sequences (S), image frame sequences (P), and text input (Q); Step S102: Convert S into a frame sequence, where the length of each frame is within (1 / 2) of the given length. min -ι max Within the range of ms, the frame shift is (r' min -r' max Within the ms range, there are a total of m frames; the P sampling frame rate is... Each frame is h×b×v per second, where h represents the length, b represents the width, and v represents the number of channels. There are a total of n frames. Q is generated by keyboard input, and each q records its corresponding start and end time. There are a total of k q. S=[s1, s2, s3, …, s m ] P=[p1, p2, p3, …, p n ] Q = [q1, q2, q3, …, q] k ] Step S103: Detect pitch changes and short pauses in the obtained speech frames to obtain the corresponding trigger times t. s The obtained image frames are used to detect key point displacements and region changes, and the corresponding trigger time t is obtained. p The obtained text fragments are then analyzed for emotion words, negation words, and conjunctions to determine the corresponding trigger times t. q ; Step S104: For each trigger point t i Define a symmetric window [t]i - Δ, t i + Δ], extract the corresponding frames for this time period from S, P, and Q respectively; Step S105: Encode the local segments of each modality into a fixed-dimensional vector S i P i Q i Using an attention network, the instantaneous attention coefficient 'a' is obtained. i =( , , Construct an event container F and sort it by time; F= Where a represents the attention coefficient, m represents the set of meta-information for event i, and N is the number of events.

[0010] Furthermore, step S2 specifically includes: Step S201: Establish an online processing state for the sparse sequence F, based on the attention coefficient a. i Initialization involves pre-estimating the initial weights for each dimension of each modal characteristic vector. Where clip(•) and norm(•) represent clipping the vector to the 0~1 region and normalization function, respectively, β is the modal attention coefficient, and 1 d Let be a vector of length d consisting entirely of 1s; Step S202: For each sampling unit in F, with As the initial state, a multimodal mutual inductance algorithm is used to perform weight self-adjustment evolution locally to obtain the weight of each mode in the final output. , , ; Step S203: Check whether the deviation τ of the weight result exceeds [the specified value]. The current weights are rolled back to the stable weight configuration of the previous time step; otherwise, the original weights are retained, resulting in the corrected weights. , , ); Where γ represents the ratio of the current weight to the initial weight. To prevent the denominator from being 0, M∈{S, P, Q}; Step S204: Multiply the three-modal features dimension by dimension according to their weights. The results are summed and a fused output is obtained through a fixed nonlinear transformation. This output is then bound to a timestamp t. By repeatedly applying this process to each event frame, a continuous emotional sequence feature Z is obtained. Z i =tanh( ) Z={(t1, Z1), (t2, Z2), …, (t n Z n )} Furthermore, step S202 specifically includes: First, construct a mapped response vector R for each mode. S R P R Q The interaction between the three modes is calculated, and the difference is regarded as the internal mutual inductance potential D. SP D SQ D PQ The mutual inductance coupling term E1, representing the three sets of modes, was calculated: R S =f(w S ⊙S i ) R P =f(w P ⊙P i ) R Q =f(w Q ⊙Q i ) D SP = D SQ = D PQ = E1=D SP + D SQ + D PQ Where f(•) is a fixed mapping function, and ⊙ denotes the dimension-wise multiplication of vectors. The squared norm of a vector is denoted by 2. Secondly, amplitude constraint energy E2 and anchoring inertia energy E3 are introduced to ensure that the weight evolution results converge to a reasonable state. E2=||w S ⊙S i ||1 + ||w P ⊙P i ||1 + ||w Q ⊙Qi ||1 E3= Finally, E1, E2, and E3 will be coupled to obtain the optimal equilibrium state of the three modes in the feature representation space: E = λ1×E1 + λ2×E2 + λ3×E3. The dynamic change of the weights follows a downward process along the energy gradient direction. w (t+1) =w (t) - When the gradient approaches 0, it enters a local stable state, and no significant weight changes occur. At this point, the optimal balanced configuration of the output multimodal in this sequence unit is reached. , , ).

[0011] Furthermore, step S3 specifically includes: Step S301: Based on the sentiment sequence features Z, use the similarity space mapping algorithm to obtain the representation Z of the sentiment state at each time point in two-dimensional space. 2D ={(x1, y1), (x2, y2), …, (x n , y n )}; Step S302: Capture Z using a convolutional neural network 2D The relationships between nodes are analyzed, and the representation of each node is updated through information propagation to construct an emotion state evolution network J. H (0) =Z 2D H (L) = Relu( H (L-2) W (L-2) ) J = (V, A) = {(h i A ij )| h i ∈H (L) A ij ∈{0, 1}} Where A is the adjacency matrix. Represents the adjacency matrix obtained after self-looping and normalization, ReLU(•) is the activation function, W is the learnable weight matrix, and H... (L) For a graph convolution with L layers, V is a set of nodes containing the sentiment state h of each node. i A ij Indicates whether there is an edge between nodes; Step S303: Using the Emotional State Evolution Network algorithm, based on the current node state of J and related historical data, output the predicted emotional state Y for the next time point; Step S304: Associate each time point with the corresponding emotional state to generate a continuous emotional evolution trajectory. =[Y0, ​​Y1, …, Y T ].

[0012] Furthermore, step S301 specifically includes: First, the similarity between each pair of data points and its neighbors is calculated using a Gaussian distribution, thus obtaining the conditional probability z. j|i Then calculate the symmetric probability u of each pair of nodes. ij Then initialize the low-dimensional representation Y = {y1, y2, …, y}. n In two-dimensional space, calculate each pair of y i and y j Low-dimensional space similarity u j|I The objective function U is obtained by measuring the difference between the probability distributions z and u in high-dimensional and low-dimensional regions using relative fitness. z j|i = z ij = U= Among them, ||·|| 2 This represents the Euclidean distance, where N is the total number of sample points. k is the number of neighbors considered; Calculate the objective function U with respect to y i The partial derivatives yield a lower-dimensional representation of Y. in, Control the step size; continuously iterate the above steps and update y. i This continues until the objective function U converges.

[0013] Furthermore, step S303 specifically includes: (1) Obtain the current state h t =H (L) and historical state D={D t-1 D t-2, …, D t-k}, where D represents the state information in the past l time steps; (2) For each historical state ej Calculate its fusion feature a j ; e j = Among them, h t-j and D t-j These represent the time steps. Current and historical status information; (3) By analyzing the current state h t The current state is updated by weighted summation of fused features. ; (4) Obtain the fused feature representation through matrix operations, and obtain the final prediction result Y by relying on the activation function. t+1 .

[0014] H (L+1) =Relu( H (L) W ’ ) Y t+1 =g(H (L+1) ) Furthermore, step S4 specifically includes: Step S401: Define the emotion category I = [I1, I1, …, I n Based on the user's emotional evolution trajectory, the frequency f=[f] of each emotional category within a fixed time period is obtained. I1 , f I2 , …, f In ], where f In = ; Step S402: Extract the number of user interactions j and the dialogue topic preference g=[g1, g1, …, g] from the interaction context. m The user feature vector X=[j, g, I] is obtained by concatenating the emotional category frequencies. Step S403: Employ a dynamic adjustment mechanism to adjust the emotion classification boundaries, obtaining the updated user emotion state O = [O new (I1), O new (I2), …, O new (I n )).

[0015] Furthermore, step S403 specifically includes: Set an initial emotion classification threshold O0(I) for each emotion category. n ), calculate the mean μ of the emotional frequency.f and standard deviation σ f This indicates the overall level and volatility of user sentiment: μ f = σ f = The emotion classification boundary is updated based on a dynamic adjustment mechanism to reflect changes in emotion. The formula for the latest emotion classification boundary is: O new (I n )=O0(I n ) + ν×(f In - μ f ) +κ×O bad (f In ) -ε×|j - | O bad (f In )=ρ×(f In - O0(I n )) 2 O new (I n =max(0, min(1, O) new (I n ))) Where ν is the learning rate, the degree of influence of the adjustment frequency on the boundary, κ is the regularization system, and O bad As a penalty term, ε is used to adjust for deviations in the number of user interactions, ρ and These are the adjustment parameters for the number of user interactions and the average number of past user interactions, respectively.

[0016] Furthermore, step S5 specifically includes: Step S501: Calculate the emotional state consistency coefficient C based on the user's latest emotional state O to characterize the concentration of the user's emotions; C=1- Step S502: Based on the emotional state consistency coefficient C, perform consistency adjustment on each emotional component in the user's emotional state vector to obtain the emotional expression intensity vector R; R(I k )=C×O new (I k ) Step S503: Perform reduction processing on the emotion expression intensity vector R to determine the overall emotion expression intensity, and adjust the emotional expression level of the basic task G0 response based on the overall emotion expression intensity to generate the final task response G: in, , This refers to the response form obtained by enhancing the emotional expression of G0 while keeping the semantics of the task unchanged.

[0017] The beneficial effects of this invention include: (1) By unifying the temporal organization and feature representation of multimodal interactive data such as speech, text and images, data from different sources can maintain consistency in the time dimension, thereby providing more complete and comparable input features for subsequent sentiment analysis and improving the accuracy of multimodal sentiment information expression; (2) Introducing a dynamic weight adjustment mechanism in the process of multimodal feature fusion can adaptively adjust the influence of different modal features in the fusion result according to the emotional changes during the interaction process, avoid the bias caused by fixed weights, and thus improve the ability of fused emotional features to depict the user's real emotions. (3) By continuously modeling the fused emotional features, we can construct an evolutionary representation of user emotions over time, so that emotional states are no longer limited to discrete judgments at a single point in time, which is conducive to more accurately reflecting the changing trends of user emotions. (4) By combining the results of emotion evolution with the context of interaction, the classification boundary used in the emotion recognition process is dynamically adjusted, thereby reducing the impact of short-term emotional fluctuations on the recognition results and improving the stability and consistency of emotion recognition results in continuous interaction. (5) When generating task responses, an emotion expression intensity regulation mechanism is introduced, so that the system can adaptively adjust the degree of emotion expression in the response content while maintaining the semantic consistency of the task, thereby avoiding the problem of excessive or insufficient emotion expression and improving the naturalness of the human-computer interaction process. (6) By linking multimodal emotion fusion, emotion evolution modeling, emotion recognition and task response regulation, the system can adapt to the emotional changes of different users and different interaction scenarios, thereby improving the adaptability and application generalization of the emotional interaction system in complex application environments.

[0018] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein: Figure 1 This is a schematic diagram of the overall workflow of the present invention; Figure 2 This is a multimodal timing feature map of the present invention; Figure 3 This is a flowchart of the multimodal weighted self-adjusting fusion processing of the present invention; Figure 4 This is a schematic diagram of the emotional state evolution network of the present invention; Figure 5 This is a data graph showing the evolution trajectory of user emotions in this invention. Detailed Implementation

[0020] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the preferred embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0021] This invention provides a method for enhancing emotion understanding through multimodal emotion interaction data fusion, such as... Figure 1 As shown, it includes the following steps: Step S1: Collect multimodal data of the user during the interaction process. The multimodal data includes voice, text and images, and construct a multimodal temporal feature stream; Step S2: Based on this feature flow, the multimodal mutual intuition algorithm is used to realize the dynamic weight adjustment of the features and obtain the real-time fused sentiment features; Step S3: Utilize the fused emotional features to construct an emotional state evolution network and output the user's emotional evolution trajectory; Step S4: Combine the emotion evolution trajectory with the interaction context information to dynamically adjust the classification boundary of the emotion recognition module and update the user's emotional state; Step S5: Generate the final task response content based on the user's emotional state and the intensity of their emotional expression.

[0022] The specific steps of the above method will be further explained below through a specific embodiment.

[0023] In this embodiment, step S1 specifically includes the following steps: Step S101: Asynchronously acquire multimodal data, including speech signal sequence (S), image frame sequence (P), and text input (Q); In this embodiment, the voice signal is acquired at a sampling rate of 16 kHz, the image frames are acquired at a frame rate of 20 frames / second, and the text input is recorded once each time the user completes input confirmation; Step S102: Convert S into a frame sequence, where the length of each frame is within (1 / 2) of the given length. min -ι max Within the range of ms, the frame shift is (r' min -r' max Within the ms range, there are a total of m frames; the P sampling frame rate is... Each frame is h×b×v per second, where h represents the length, b represents the width, and v represents the number of channels. There are a total of n frames. Q is generated by keyboard input, and each q records its corresponding start and end time. There are a total of k q. S=[s1, s2, s3, …, s m ] P=[p1, p2, p3, …, p n ] Q = [q1, q2, q3, …, q] k ] In this embodiment, for the speech signal sequence S, under the condition of a sampling rate of 16 kHz, the continuous speech signal is framed according to the method of 25 milliseconds per frame and 10 milliseconds interval between adjacent frames; in this 15-second interaction, a total of 1498 speech frames are obtained, each speech frame corresponds to a speech segment of fixed length, thus forming a speech frame sequence: S=[s1, s2, s3, …, s 1498 ] For image data, the video stream was periodically sampled at a frame rate of 20 frames per second. A total of 300 frames of image data were acquired within 15 seconds. Each frame had a size of 224×224×3, representing the image's length, width, and number of channels, respectively, thus forming an image frame sequence. P=[p1, p2, p3, …, p 300 ] For text input data, when a user completes an input operation via the keyboard or input interface, the corresponding text content is recorded as a text segment, along with the start and end times of that text segment. In this interaction, a total of 8 text segments are recorded, thus forming a text input sequence: Q = [q1, q2, q3, …, q8] Step S103: Detect pitch changes and short pauses in the obtained speech frames to obtain the corresponding trigger times t. s The obtained image frames are used to detect key point displacements and region changes, and the corresponding trigger time t is obtained. p The obtained text fragments are then analyzed for emotion words, negation words, and conjunctions to determine the corresponding trigger times t. qMultimodal temporal feature maps, such as Figure 2 As shown; The pitch and energy changes between adjacent speech frames were calculated sequentially. A speech trigger event was identified when the pitch change between two adjacent frames exceeded 30Hz, or when a low-energy interval lasting longer than 200 milliseconds appeared in consecutive speech frames. The corresponding time of this speech frame was recorded as the speech trigger time. During this 15-second interaction, a total of 5 speech trigger time points were detected, namely: t s ={2.4s, 4.8s, 7.2s, 9.6s, 12.9s} The system compares and analyzes adjacent image frames in the image stream, calculating the overall pixel change ratio between two adjacent frames. When the pixel change ratio between two adjacent frames exceeds 15%, it is determined as an image trigger event, and the time corresponding to the current image frame is taken as the image trigger time. During this 15-second interaction, a total of 5 image trigger time points were detected, namely: t p ={2.5s, 4.9s, 7.0s, 9.8s, 13.0s} Content detection is performed on each text segment in the input text sequence. When the text contains at least one of the following: sentiment words, negation words, or conjunctions, the text segment is determined to be a text trigger event, and its start time is taken as the corresponding text trigger time. In this interaction, a total of two text trigger time points were detected: t q ={5.0s, 10.1s} Step S104: For each trigger point t i Define a symmetric window [t] i - Δ, t i + Δ], extract the corresponding frames for this time period from S, P, and Q respectively; With a window half-width Δ = 0.5s, the voice trigger frames are: S1 = [1.9s, 2.9s], S2 = [4.3s, 5.3s], S3 = [6.7s, 7.7s], S4 = [9.1s, 10.1s], S5 = [12.4s, 13.4s]; the image trigger frames are: P1 = [2.0s, 3.0s], P2 = [4.4s, 5.4s], P3 = [6.5s, 7.5s], P4 = [9.3s, 10.3s], P5 = [12.5s, 13.5s]; and the image trigger frames are: Q1 = [4.5s, 5.5s], Q2 = [9.6s, 10.6s]. Step S105: Encode the local segments of each modality into a fixed-dimensional vector S i P i Q i Using an attention network, the instantaneous attention coefficient 'a' is obtained. i = ( , , Construct an event container F and sort it by time; F= Where a represents the attention coefficient, m represents the set of meta-information for event i, and N is the number of events.

[0024] When the fixed dimension vector is 4 and t1=2.4s, the speech coding vectors are S1=[0.63, 0.48, 0.22,0.55], P1=[0.45, 0.55, 0.40, 0.37], Q1=[0.25, 0.125, 0.125, 0.8], and a1=(0.38, 0.36, 0.26). Therefore, F1=(t1=2.4s, S1, P1, Q1, a1, m1).

[0025] The specific steps of step S2 are as follows: Step S201: Establish an online processing state for the sparse sequence F, based on the attention coefficient a. i Initialization involves pre-estimating the initial weights for each dimension of each modal characteristic vector. Where clip(•) and norm(•) represent clipping the vector to the 0~1 region and normalization function, respectively, β is the modal attention coefficient, and 1 d Let be a vector of length d consisting entirely of 1s; Choosing β=0.5, at this point, norm(S1)=[0.335, 0.255, 0.117, 0.293], norm(P1)=[0.254, 0.311, 0.226, 0.209], and norm(Q1)=[0.192, 0.096, 0.096, 0.616]. Therefore, the initial weights... =[0.357, 0.317, 0.248, 0.337] =[0.317, 0.336, 0.293, 0.285] =[0.261,0.213, 0.213, 0.439]; Step S202: For each sampling unit in F, with As the initial state, a multimodal mutual inductance algorithm is used to perform weight self-adjustment evolution locally to obtain the weight of each mode in the final output. , , ; First, construct a mapped response vector R for each mode. S R P R Q The interaction between the three modes is calculated, and the difference is regarded as the internal mutual inductance potential D. SP D SQ D PQ The mutual inductance coupling term E1, representing the three sets of modes, was calculated: R S =f(w S ⊙S i ) R P =f(w P ⊙P i ) R Q =f(w Q ⊙Q i ) D SP = D SQ = D PQ = E1=D SP + D SQ + D PQ Where f(•) is a fixed mapping function, and ⊙ denotes the dimension-wise multiplication of vectors. The squared norm of a vector is denoted by 2. Secondly, amplitude constraint energy E2 and anchoring inertia energy E3 are introduced to ensure that the weight evolution results converge to a reasonable state. E2=||w S ⊙S i ||1 + ||w P ⊙P i ||1 + ||w Q ⊙Q i ||1 E3= Finally, E1, E2, and E3 will be coupled to obtain the optimal equilibrium state of the three modes in the feature representation space: E = λ1×E1 + λ2×E2 + λ3×E3. The dynamic change of the weights follows a downward process along the energy gradient direction. w (t+1) =w (t) - When the gradient approaches 0, it enters a local stable state, and no significant weight changes occur. At this point, the optimal balanced configuration of the output multimodal in this sequence unit is reached. , , ).

[0026] Taking the simplest mapping function f(x)=x, we get R S =[0.225, 0.152, 0.055, 0.185], R P =[0.143, 0.185, 0.117, 0.105], R Q =[0.065, 0.027, 0.027, 0.351]; At this time, D SP =0.014, D SQ =0.062, D PQ =0.082, then E1=0.158, E2=1.637; Under the constraint of the energy gradient, the weights undergo a small adjustment. = [0.357, 0.317, 0.248, 0.337] =[0.317, 0.336, 0.293, 0.285] =[0.261, 0.213, 0.213, 0.439], at this time E3=0.0011 Step S203: Check whether the deviation τ of the weight result exceeds [the specified value]. The current weights are rolled back to the stable weight configuration of the previous time step; otherwise, the original weights are retained, resulting in the corrected weights. , , ); Where γ represents the ratio of the current weight to the initial weight. To prevent the denominator from being 0, M∈{S, P, Q}; Step S204: Multiply the three-modal features dimension by dimension according to their weights. The values ​​are summed and a fused output is obtained through a fixed nonlinear transformation. A timestamp t is then bound to the sum. By continuously repeating this process for each event frame, a continuous emotional sequence feature Z is obtained. The entire multimodal weighted self-adjusting fusion processing flow is as follows: Figure 3 As shown: Z i =tanh( ) Z={(t1, Z1), (t2, Z2), …, (t n Z n )} Based on the above data, we obtain Z1=[0.410, 0.355, 0.194, 0.557], and using the same calculation method, we obtain Z2=[0.425, 0.368, 0.212, 0.540], Z3=[0.398, 0.351, 0.198, 0.563], Z4=[0.412, 0.359, 0.205, 0.552], and Z5=[0.409, 0.362, 0.200, 0.558]. Then Z={(t1=2.4s, [0.410, 0.355, 0.194, 0.557]), (t2=4.8s, [0.425, 0.368,0.212, 0.540]), (t3=7.2s, [0.398, 0.351, 0.198, 0.563]), (t4=9.6s, [0.412,0.359, 0.205, 0.552]), (t5=12.9s, [0.409, 0.362, 0.200, 0.558])}.

[0027] The specific steps of step S3 are as follows: Step S301: Based on the sentiment sequence features Z, use the similarity space mapping algorithm to obtain the representation Z of the sentiment state at each time point in two-dimensional space. 2D ={(x1, y1), (x2, y2), …, (x n , y n )}; First, the similarity between each pair of data points and its neighbors is calculated using a Gaussian distribution, thus obtaining the conditional probability z. j|i Then calculate the symmetric probability u of each pair of nodes. ij Then initialize the low-dimensional representation Y = {y1, y2, …, y}. n In two-dimensional space, calculate each pair of y i and y j Low-dimensional space similarity u j|IThe objective function U is obtained by measuring the difference between the probability distributions z and u in high-dimensional and low-dimensional regions using relative fitness. z j|i = z ij = U= Among them, ||·|| 2 This represents the Euclidean distance, where N is the total number of sample points. k is the number of neighbors considered; Calculate the objective function U with respect to y i The partial derivatives yield a lower-dimensional representation of Y. in, Control the step size; continuously iterate the above steps and update y. i This continues until the objective function U converges.

[0028] This embodiment is set as follows: If the neighbor k=2, and Y is initialized randomly in [0, 1], then Z is obtained. 2D ={(0.42,0.35), (0.44,0.36), (0.40,0.37), (0.41,0.355), (0.405,0.36)}; Step S302: Capture Z using a convolutional neural network 2D The relationships between nodes are established, and the representation of each node is updated through information propagation to construct an emotion state evolution network J. A schematic diagram of the emotion state evolution network is shown below. Figure 4 As shown; H (0) =Z 2D H (L) =Relu( H (L-2) W (L-2) ) J =(V, A) ={(h i A ij )| h i ∈H (L) A ij ∈{0, 1}} Where A is the adjacency matrix. Represents the adjacency matrix obtained after self-looping and normalization, ReLU(•) is the activation function, W is the learnable weight matrix, and H...(L) For a graph convolution with L layers, V is a set of nodes containing the sentiment state h of each node. i A ij Indicates whether there is an edge between nodes; Step S303: Using the Emotional State Evolution Network algorithm, based on the current node state of J and related historical data, output the predicted emotional state Y for the next time point; (1) Obtain the current state h t =H (L) and historical state D={D t-1 D t-2, …, D t-k}, where D represents the state information in the past l time steps; (2) For each historical state e j Calculate its fusion feature a j ; e j = Among them, h t-j and D t-j These represent the time steps. Current and historical status information; (3) By analyzing the current state h t The current state is updated by weighted summation of fused features. ; (4) Obtain the fused feature representation through matrix operations, and obtain the final prediction result Y by relying on the activation function. t+1 .

[0029] H (L+1) =Relu( H (L) W ’ ) Y t+1 =g(H (L+1) ) In this embodiment, the historical step size is 2, and the calculated values ​​are: Y1=[0.425, 0.365, 0.21, 0.545], Y2=[0.402, 0.355, 0.202, 0.558], Y3=[0.415, 0.362, 0.207, 0.552], Y4=[0.410, 0.364, 0.203, 0.556]. Step S304: Associate each time point with the corresponding emotional state to generate a continuous emotional evolution trajectory. =[Y0, ​​Y1, …, Y T ],like Figure 5 As shown.

[0030] ={Y0=[0.41,0.36,0.20,0.55], Y1=[0.425, 0.365, 0.21, 0.545], Y2=[0.402, 0.355, 0.202, 0.558], Y3=[0.415, 0.362, 0.207, 0.552], Y4=[0.410,0.364, 0.203, 0.556]} The specific steps of step S4 are as follows: Step S401: Define the emotion category I = [I1, I1, …, I n Based on the user's emotional evolution trajectory, the frequency f=[f] of each emotional category within a fixed time period is obtained. I1 , f I2 , …, f In ], where f In = ; Here, the emotion category is defined as four levels, from low to high: emotional stability, mild excitement, moderate excitement, and high excitement, I∈[0,1], and equal interval thresholds are used for classification. At this time, f=[0.25,0.50,0.25,0]; Step S402: Extract the number of user interactions j and the dialogue topic preference g=[g1, g1, …, g] from the interaction context. m The user feature vector X=[j, g, I] is obtained by concatenating the emotional category frequencies. In this embodiment, the number of user interactions during this period is 8, the user's historical average number of interactions is 6, and the dialogue topic is set as [task-oriented] X = [8, 0.45, 0.30, 0.25, 0.25, 0.50, 0.25, 0]; Step S403: Employ a dynamic adjustment mechanism to adjust the emotion classification boundaries, obtaining the updated user emotion state O = [O new (I1), O new (I2), …, O new (I n )]; Set an initial emotion classification threshold O0(I) for each emotion category. n ), calculate the mean μ of the emotional frequency. f and standard deviation σ f This indicates the overall level and volatility of user sentiment: μ f = σ f = The emotion classification boundary is updated based on a dynamic adjustment mechanism to reflect changes in emotion. The formula for the latest emotion classification boundary is: O new (I n )=O0(I n ) + ν×(f In - μ f ) +κ×O bad (f In ) -ε×|j - | O bad (f In )=ρ×(f In - O0(I n )) 2 O new (I n =max(0, min(1, O) new (I n ))) Where ν is the learning rate, the degree of influence of the adjustment frequency on the boundary, κ is the regularization system, and O bad As a penalty term, ε is used to adjust for deviations in the number of user interactions, ρ and These are the adjustment parameters for the number of user interactions and the average number of past user interactions, respectively.

[0031] μ was calculated f =0.25, σ f =0.177, the updated user emotion state vector is O=[0.138,0.458,0.633,0.875], the user's current emotion is mainly concentrated in the range of mild to moderate excitement.

[0032] The specific steps of step S5 are as follows: Step S501: Calculate the emotional state consistency coefficient C based on the user's latest emotional state O to characterize the concentration of the user's emotions; C=1- The calculation yielded C=0.85.

[0033] Step S502: Based on the emotional state consistency coefficient C, perform consistency adjustment on each emotional component in the user's emotional state vector to obtain the emotional expression intensity vector R; R(I k )=C×Onew (I k ) R=[0.117, 0.389, 0.538, 0.744] Step S503: Perform reduction processing on the emotion expression intensity vector R to determine the overall emotion expression intensity, and adjust the emotional expression level of the basic task G0 response based on the overall emotion expression intensity to generate the final task response G: in, , This refers to the response form obtained by enhancing the emotional expression of G0 while keeping the semantics of the task unchanged.

[0034] This indicates that the user is slightly agitated at this moment. Based on the set response, only a few reassuring or explanatory sentences need to be added. The basic task G0 response content is set as: "The system has received your request and is processing it." At this time, the task response is "The system has received your request. Please wait patiently. It is being processed." It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable storage medium. The method can be implemented using standard programming techniques, including a non-transitory computer-readable storage medium configured with a computer program within the computer program, wherein the storage medium is configured such that the computer operates in a specific and predefined manner—according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Furthermore, for this purpose, the program can run on a programmed application-specific integrated circuit (ASIC).

[0035] Furthermore, the procedures described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The procedures described herein (or variations and / or combinations thereof) may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors. The computer program comprises a plurality of instructions executable by one or more processors.

[0036] Furthermore, the method can be implemented in any suitable type of computing platform, including but not limited to personal computers, minicomputers, mainframes, workstations, networked or distributed computing environments, standalone or integrated computer platforms, or in communication with charged particle tools or other imaging devices, etc. Aspects of the invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it is readable by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein. Furthermore, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. The invention described herein includes these and other different types of non-transitory computer-readable storage media when such media comprises instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor. When programmed according to the methods and techniques described in the invention, the invention also includes the computer itself.

[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multi-modal affective interaction data fusion based affective understanding enhancement method, characterized in that: The method comprises the following steps: Step S1: collecting multi-modal data of the user in the interaction process, the multi-modal data comprising voice, text and images, and constructing a multi-modal time sequence feature flow; Step S2: based on the feature flow, using a multi-modal mutual sensing algorithm to realize dynamic weight adjustment of the features, and obtaining real-time fusion emotional features; Step S3: using the fused emotional features to construct an emotional state evolution network, and outputting an emotional evolution trajectory of the user; Step S4: combining the emotional evolution trajectory and interaction context information, dynamically adjusting the classification boundary of the emotional recognition module, and updating the emotional state of the user; Step S5: generating a final task response content according to the emotional state of the user and the emotional expression intensity.

2. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 1, wherein: The step S1 specifically comprises: Step S101: asynchronously collecting multi-modal data, including a voice signal sequence (S), an image frame sequence (P) and text input (Q); Step S102: converting S into a frame sequence, each frame having a length in the range of (r min -ι max ) ms, the frame shift being in the range of (r min -r’ max ) ms, a total of m frames; the P sampling frame rate being h×b×v per second, each frame having a size of h×b×v, h representing length, b representing width, and v representing the number of channels, a total of n frames; Q being formed by keyboard input to record the start and end times corresponding to each q, a total of k. S=[s1, s2, s3, …, s m ] P=[p1, p2, p3, …, p n ] Q = [q1, q2, q3, …, q k ] Step S103: Detecting the pitch mutation and short pause of the obtained voice frame to obtain the corresponding trigger time t s ; Detecting the key point displacement and area change of the obtained image frame to obtain the corresponding trigger time t p ; Detecting the emotional vocabulary, negative words and conjunction words of the obtained text segment to obtain the corresponding trigger time t q ; Step S104: for each trigger point t i , define a symmetric window [t i - Δ, t i + Δ] and extract the corresponding frames from S, P, Q for this period; Step S105: encode each local segment of each modality into a fixed-dimension vector S i , P i , Q i , using an attention network to obtain instantaneous attention coefficients a i = ( , , ), build an event container F and sort by time; F = Wherein, a represents an attention coefficient, m represents a meta information set of event i, and N is the number of events.

3. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 1, wherein: The step S2 specifically comprises: Step S201: Establishing an online processing state for the sparse sequence F, in accordance with the attention coefficient a i Initializing each dimension of each modal characteristic vector is given an initial weight estimate in advance where clip(•) and norm(•) denote the clipping of the vector to the 0~1 region and the normalization function, respectively, β is the modal attention coefficient, 1 d is an all-one vector of length d; Step S202: for each sampling unit in F, to As the initial state, the weight self-tuning evolution is performed locally and internally using a multi-modal mutual inductance algorithm to obtain the weight of each mode in the final output , , ; Step S203: Check whether the weight result deviation τ exceeds , the current weight is rolled back to the stable weight configuration of the previous time step, otherwise, the original weight is directly retained, and finally the effective weight is obtained , , ); wherein γ represents the proportion of the current weight and the initial weight, , to prevent the denominator from being 0, M ∈ {S, P, Q}; Step S204: multiply the three-modal features by weight dimension by dimension and add, get the fusion output through fixed nonlinear transformation, and bind the timestamp t, get the continuous emotion sequence feature Z by repeating each event frame repeatedly. Z i =tanh( ) Z = {(t1, Z1), (t2, Z2),..., (t n , Z n )}.

4. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 3, wherein: The step S202 multi-modal mutual sensing algorithm specifically comprises: First, construct a mapped response vector R for each modality S , R P , R Q , calculate the interaction among the three modalities, and regard the difference as the internal mutual potential energy D SP , D SQ , D PQ , calculate the mutual coupling term E1 between the three groups of modalities R S =f(w S ⊙S i ) R P = f(w P ⊙P i ) R Q = f(w Q ⊙Q i ) D SP = D SQ = D PQ = E1 = D SP + D SQ + D PQ where f(•) is a fixed mapping function, and represents the element-wise multiplication of vectors, denotes the square of the two-norm of a vector; Secondly, the amplitude constraint energy E2 and the anchoring inertial energy E3 are introduced to ensure that the weight evolution result converges to a reasonable state; E2=||w S ⊙S i ||1 + ||w P ⊙P i ||1 + ||w Q ⊙Q i ||1 E3= Finally, E1, E2 and E3 are coupled to obtain the optimal balance state E of the three modalities in the feature representation space E = λ1 × E1 + λ2 × E2 + λ3 × E3, and the dynamic change of the weight follows the downward process along the energy gradient direction; w (t+1) =w (t) - When the gradient tends to 0, the local stable state is entered, and no obvious weight change occurs, and at this time the output multi-modal is in the optimal balanced configuration in the sequence unit , , ).

5. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 1, wherein: The step S3 specifically comprises: Step S301: According to the emotion sequence feature Z, using a similarity space mapping algorithm, the representation of the emotion state of each time point in a two-dimensional space Z is obtained 2D ={(x1, y1), (x2, y2), …, (x n , y n )}; Step S302: Capture Z using a convolutional neural network 2D and update the representation of each node through information propagation, build the sentiment state evolution network J: H (0) =Z 2D H (L) = Relu( H (L-2) W (L-2) ) J = (V, A) = {(h i , A ij ) | h i ∈ H (L) , A ij ∈ {0, 1}} wherein A is an adjacency matrix, denotes the adjacency matrix obtained by self-loop and normalization, Relu(•) is an activation function, W is a learnable weight matrix, H (L) is the graph convolution through L layers, V is a set of nodes, containing the sentiment state h i of each node, ij denotes whether there is an edge between nodes; Step S303: using an emotional state evolution network algorithm to output a prediction Y of the emotional state at the next time point based on the current node state of J and related historical data; Step S304: associate each time point with the corresponding emotional state, and generate a continuous emotional evolution trajectory =[Y0, Y1,…,Y T ].

6. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 5, wherein: The step S301 similarity space mapping algorithm specifically comprises: First, the similarity of each pair of data points and its adjacent points is calculated according to the Gaussian distribution, and the conditional probability z is obtained j|i Then, the symmetric probability u of each pair of nodes is calculated ij ; Then, the low-dimensional representation Y={y1, y2, …, y n} is initialized in the two-dimensional space, and the low-dimensional space similarity u i of each pair of y j and y j|I is calculated; the difference between the probability distribution z and u between the high-dimensional and low-dimensional is measured by the relative fitness, and the objective function U is obtained; z j|i = z ij = U= where || · || denotes the Euclidean distance, N is the total number of sample points, 2 where || · || denotes the Euclidean distance, N is the total number of sample points, k is the number of neighbors considered; Compute the partial derivative of the objective function U with respect to y i to obtain a lower-dimensional representation Y: wherein, , control the step size; continue iterating the above steps, updating y i , until the objective function U converges.

7. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 5, wherein: The step S303 emotional state evolution network algorithm specifically comprises: (1) Obtain the current state h t = H (L) and the history state D = {D t-1 , D t-2, …, D t-k}, D represents the state information in the past l time steps; (2) for each historical state e j compute its fusion feature a j ; e j = where h t-j and D t-j respectively represent the current state and historical state information at time step . (3) update the current state h t and fusion feature weighted sum update the current state ; (4) The fused feature expression is obtained through matrix operation, and the final prediction result Y is obtained by relying on an activation function t+1 : H (L+1) = Relu( H (L) W ’ ) Y t+1 = g(H (L+1) ).

8. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 1, wherein: The step S4 specifically comprises: Step S401: define the emotion category I=[I1, I1, …, I n ] , get the frequency of each emotion category in a fixed time period f=[f I1 , f I2 , …, f In ] according to the user emotion evolution track, wherein f In = ; Step S402: extract the number of user interactions j and the dialogue topic preference g=[g1, g1, …, g m ] in the interaction context, and combine the emotion category frequency to splice the user feature vector X=[j, g, I]; Step S403: Employ a dynamic adjustment mechanism to adjust the emotion classification boundaries, obtaining the updated user emotion state O=[O new (I1), O new (I2), …, O new (I n )).

9. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 8, wherein: The step S403 dynamic adjustment mechanism specifically comprises: Set an initial emotion classification threshold O0(I) for each emotion category. n ), calculate the mean μ of the emotional frequency. f and standard deviation σ f This indicates the overall level and volatility of user sentiment: μ f = σ f = According to the dynamic adjustment mechanism, the emotional classification boundary is updated to reflect the change of the emotion, and the formula of the latest emotional classification boundary is: O new (I n )=O0(I n ) + ν×(f In - μ f ) +κ×O bad (f In ) -ε×|j - | O bad (f In )=ρ×(f In - O0(I n )) 2 O new (I n )=max(0, min(1, O new (I n ))) where v is the learning rate, adjusts the influence of the frequency on the boundary, K is the regularization system, O bad is the penalty term, e is used to adjust the deviation of the number of user interactions, p and are the adjustment parameters of the number of user interactions and the mean of the past number of user interactions, respectively.

10. The multi-modal affective interaction data fusion based affective understanding enhancement method of claim 1, wherein: The step S5 specifically comprises: Step S501: calculating an emotional state consistency coefficient C according to the latest emotional state O of the user, to depict the concentration degree of the user's emotion; C=1- Step S502: performing consistency adjustment on each emotional component in the emotional state vector of the user according to the emotional state consistency coefficient C, to obtain an emotional expression intensity vector R; R(I k )=C×O new (I k ) Step S503: performing reduction processing on the emotional expression intensity vector R to determine the overall emotional expression intensity, and adjusting the emotional expression degree of the basic task G0 response based on the overall emotional expression intensity, to generate a final task response G: wherein, , is the response form obtained by enhancing the emotional expression of G0 while keeping the task semantics unchanged.