AI interaction method and device based on emotion recognition, electronic equipment and storage medium

Through multimodal data fusion and emotional recognition, a personalized AI response strategy is generated, which solves the problem of poor AI interaction effects in the existing technology, and achieves a more accurate and in line with user preferences.

CN120372487AActive Publication Date: 2025-07-25FIBOCOM WIRELESS

Patent Information

Application Number
CN202510876979.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the prior art, the AI interaction effect based on emotion recognition is poor, mainly due to emotional recognition of single modal data, resulting in poor interaction effect.

Method used

By obtaining multimodal data of the target user, using a pre-trained emotional interaction engine for feature fusion, combining emotion recognition and context analysis, a personalized AI response strategy is generated, and the target emotional relief content is determined from the emotional relief database to generate response content in line with user preferences.

Benefits of technology

It improves the accuracy of emotional recognition results, so that the generated response content can meet the personal preferences of the target users, thereby improving the AI interaction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372487A_ABST
    Figure CN120372487A_ABST
Patent Text Reader

Abstract

The invention relates to an AI interaction method and device based on emotion recognition, electronic equipment and a storage medium. The method comprises the steps of obtaining multi-modal data of a target user; performing feature fusion on the multi-modal data by using a pre-trained emotion interaction engine to obtain feature fusion data; performing emotion recognition and context analysis on the feature fusion data to obtain an emotion recognition result and comprehensive representation of the current state of the user; based on the emotion recognition result and the comprehensive representation of the current state of the user, generating an AI response strategy corresponding to the target user, and determining target emotion relieving content from a pre-constructed emotion relieving database; and based on the AI response strategy, the feature fusion data, the emotion recognition result and the target emotion relief content, generating response content, and outputting the response content. In this way, the accuracy of the emotion recognition result can be improved, the generated response content can meet the personal preference of the target user, and then the purpose of improving the AI interaction effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to an AI interaction method, device, electronic device, and storage medium based on emotion recognition. Background Art

[0002] With the rapid development of artificial intelligence technology, artificial intelligence (AI) interaction technology based on emotion recognition can recognize users' emotions and provide positive emotional value for users, so it has been widely used in many fields, such as mental health support, customer service, education counseling, etc.

[0003] However, in the prior art, traditional large language models are usually used to perform emotion recognition on single-modal data. For example, models such as Generative Pre-trained Transformer (abbreviated as GPT) are used to perform emotion recognition on the text or speech input by users and then implement simple question-and-answer dialogues. Therefore, there is a problem of poor AI interaction effect. Therefore, how to improve the AI interaction effect has become an urgent technical problem to be solved. Summary of the Invention

[0004] This application provides an AI interaction method, device, electronic device, and storage medium based on emotion recognition to solve the problem of poor AI interaction effect in the prior art.

[0005] In a first aspect, an embodiment of this application provides an AI interaction method based on emotion recognition, and the method includes: Obtain multi-modal data of a target user; Use a pre-trained emotion interaction engine to perform feature fusion on the multi-modal data to obtain feature fusion data; Perform emotion recognition and context analysis on the feature fusion data respectively to obtain an emotion recognition result and a comprehensive representation of the user's current state; Based on the emotion recognition result and the comprehensive representation of the user's current state, generate an AI response strategy corresponding to the target user, and determine target emotion relief content from a pre-constructed emotion relief database, where multiple emotion relief contents are stored in the emotion relief database, and each emotion relief content corresponds to an attribute label one by one; Generate a response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion relief content, and output the response content.

[0006] Optionally, the feature fusion of the multimodal data by using the pre-trained emotion interaction engine to obtain feature fusion data includes: Preprocessing the multimodal data by using the pre-trained emotion interaction engine to obtain feature data corresponding to each modality data, where the multimodal data includes at least one of image data, voice data, touch data, heart rate data, and motion data; Performing first feature fusion on the feature data corresponding to each modality data by using a graph attention network to obtain an intermediate feature vector corresponding to each modality data, where the intermediate feature vector is used to reflect the correlation between its own modality data and other modality data; Performing second feature fusion on the intermediate feature vector by using a deep learning model to obtain the feature fusion data, where the feature fusion data is used to characterize the temporal features of each modality data.

[0007] Optionally, the emotion recognition and context analysis of the feature fusion data respectively to obtain an emotion recognition result and a comprehensive representation of the user's current state includes: Performing emotion recognition on the feature fusion data by using an emotion recognition model to obtain the emotion recognition result, where the emotion recognition result is used to characterize the emotion category and emotion intensity of the target user; Obtaining the historical interaction data of the target user, and performing context analysis based on the historical interaction data and the feature fusion data to generate a comprehensive representation of the user's current state, where the comprehensive representation of the user's current state is used to characterize the emotion category, emotion intensity, physiological state, and behavior state of the target user.

[0008] Optionally, generating an AI response strategy corresponding to the target user based on the emotion recognition result and the comprehensive representation of the user's current state, and determining target emotion relief content from a pre-constructed emotion relief database includes: Constructing a causal graph based on the emotion recognition result and the comprehensive representation of the user's current state, where the causal graph is used to characterize the causal relationship between the user state and the response action, and the user state includes the emotion recognition result and the comprehensive representation of the user's current state; Determining the AI response strategy corresponding to the target user based on the causal graph and a preset reward function, where the AI response strategy is used to characterize one or more response actions; Obtaining the historical interaction data of the target user, and determining the historical preference vector of the target user based on the historical interaction data of the target user. Determining the target emotion relief content from the emotion relief database based on the historical preference vector of the target user and the comprehensive representation of the user's current state.

[0009] Optionally, generating a response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content, and outputting the response content includes: When the AI response strategy includes text response, generating a target response text based on the feature fusion data and the emotion recognition result; When the AI response strategy includes voice response, generating a target response text based on the feature fusion data and the emotion recognition result, and converting the target response text into voice to obtain a target response voice; When the AI response strategy includes multimedia content response, obtaining target multimedia content based on the target emotion alleviation content, where the type of the target multimedia content includes at least one of music, video, and picture; When the AI response strategy includes physical feedback response, generating a target physical feedback based on the emotion recognition result, where the target physical feedback includes at least one of vibration of a tactile vibration motor and mechanical actions performed by a flexible mechanical component.

[0010] Optionally, after generating a response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content, and outputting the response content, the method further includes: Collecting feedback made by the target user based on the response content to obtain multi-modal feedback data; Using the feedback data to adaptively optimize the model parameters of the emotion interaction engine.

[0011] Optionally, before using a pre-trained emotion interaction engine to perform feature fusion on the multi-modal data to obtain feature fusion data, the method further includes: Training the initial model parameters of the emotion interaction engine using meta-learning, where the initial model parameters are learned based on multiple learning tasks of a single user; Training the global model parameters of the emotion interaction engine using federated learning to obtain the emotion interaction engine, where the global model parameters are aggregated based on the initial model parameters of multiple users.

[0012] In a second aspect, an AI interaction device based on emotion recognition provided by an embodiment of the present application includes: An acquisition module, configured to acquire multi-modal data of a target user; A feature fusion module, configured to perform feature fusion on the multi-modal data by using a pre-trained emotion interaction engine to obtain feature fusion data; An analysis module, configured to perform emotion recognition and context analysis on the feature fusion data respectively to obtain an emotion recognition result and a comprehensive representation of the user's current state; A determination module, configured to generate an AI response strategy corresponding to the target user based on the emotion recognition result and the comprehensive representation of the user's current state, and determine target emotion alleviation content from a pre-constructed emotion alleviation database, where multiple emotion alleviation contents are stored in the emotion alleviation database, and each emotion alleviation content corresponds to an attribute label one by one; An output module, configured to generate a response content based on the AI response strategy, the feature fusion data, the emotion recognition result and the target emotion alleviation content, and output the response content.

[0013] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory and a communication bus, where the processor, the communication interface and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is configured to implement the AI interaction method based on emotion recognition according to any embodiment of the first aspect when executing the program stored on the memory.

[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the AI interaction method based on emotion recognition according to any embodiment of the first aspect.

[0015] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: The method provided by the embodiment of this application obtains multimodal data of a target user; uses a pre-trained emotion interaction engine to perform feature fusion on the multimodal data to obtain feature fusion data; performs emotion recognition and context analysis on the feature fusion data respectively to obtain an emotion recognition result and a comprehensive representation of the user's current state; generates an AI response strategy corresponding to the target user based on the emotion recognition result and the comprehensive representation of the user's current state, and determines target emotion alleviation content from a pre-constructed emotion alleviation database, where multiple emotion alleviation contents are stored in the emotion alleviation database, and each emotion alleviation content corresponds to an attribute label one by one; generates a response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content, and outputs the response content. Through the above method, the pre-trained emotion interaction engine can be used to fuse and perform emotion recognition on the multimodal data of the target user, and an AI response strategy that meets the personalization of the target user can be generated through the emotion recognition result and the comprehensive representation of the user's current state, thereby improving the accuracy of the emotion recognition result and enabling the generated response content to meet the personal preferences of the target user, and further achieving the purpose of improving the AI interaction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart of an AI interaction method based on emotion recognition provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of another AI interaction method based on emotion recognition provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of an AI interaction device based on emotion recognition provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application.

[0020] Refer to Figure 1 , Figure 1 which is a schematic flowchart of an AI interaction method based on emotion recognition provided by an embodiment of this application. As Figure 1 shown, the AI interaction method based on emotion recognition may include the following steps: Step S101: Obtain multimodal data of a target user.

[0021] It should be noted that the AI interaction method based on emotion recognition provided by the embodiments of this application can be applied alone to a user terminal (such as a smart toy, an educational terminal, a medical device, etc.), or can be jointly applied to a user terminal and a cloud platform. The embodiments of this application do not make specific limitations. When applied alone to a user terminal, the user terminal can collect the multimodal data of the target user, and then perform feature fusion and emotion recognition on the multimodal data of the target user to obtain AI response content for the target user; when jointly applied to a user terminal and a cloud platform, the user terminal can collect the multimodal data of the target user and send it to the cloud platform, and then the cloud platform performs feature fusion and emotion recognition on the multimodal data of the target user to obtain AI response content and return it to the user terminal.

[0022] Specifically, the above target user can be any user who needs to perform AI interaction. The above multimodal data can be user data collected through multi-source sensors, which may include, but are not limited to, image data, voice data, touch data, heart rate data, and motion data, etc. Among them, the image data here can be collected through a camera, the voice data can be collected through a microphone, the touch data can be collected through a flexible electronic skin, the heart rate data can be collected through a Photoplethysmogram Sensor (abbreviated as PPG sensor), and the motion data can be collected through an Inertial Measurement Unit (abbreviated as IMU).

[0023] Step S102: Use a pre-trained emotion interaction engine to perform feature fusion on the multimodal data to obtain feature fusion data.

[0024] Specifically, the above-mentioned emotion interaction engine is pre-trained. The emotion interaction engine may include a multi-modal feature fusion module, an emotion recognition module, a context analysis module, an AI response strategy generation module, a response content generation module, etc. Among them, the multi-modal feature fusion module is mainly used for fusing features of multi-modal data to obtain feature fusion data. The emotion recognition module is mainly used for performing emotion recognition on the feature fusion data to obtain an emotion recognition result. The context analysis module is mainly used for performing context analysis on the feature fusion data to obtain a comprehensive representation of the user's current state. The AI response strategy generation module is mainly used for generating an AI response strategy corresponding to the target user based on the emotion recognition result and the comprehensive representation of the user's current state. The response content generation module is mainly used for generating response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content. The above-mentioned multi-modal feature fusion module may include a Graph Attention Network (GAT for short) and a deep learning model (such as a Transformer model, etc.). That is, the GAT and the deep learning model can be used to fuse features of multi-modal data to obtain feature fusion data. Among them, the GAT generates a unified feature representation by analyzing the correlation between various modal data (such as the co-variation of facial expressions and heart rate). The deep learning model is used to process the dynamic changes of data (such as the change of speech intonation over time) to ensure capturing the temporal features of emotions, etc.

[0025] Step S103: Respectively perform emotion recognition and context analysis on the feature fusion data to obtain an emotion recognition result and a comprehensive representation of the user's current state.

[0026] Specifically, the above-mentioned emotion recognition module can be implemented by a Bidirectional Long Short-Term Memory (BiLSTM for short), a Convolutional Neural Networks - Long Short-Term Memory networks (CNN-LSTM for short), or other networks. The above-mentioned emotion recognition result is used to characterize the emotion probability distribution of the target user, such as , where represents the probability that the predicted user is currently in the i-th emotion category, satisfying .

[0027] The above context analysis is used to generate a comprehensive representation of the user's current state by combining the user's historical interaction data with the user's real-time physiological state and behavioral state. The comprehensive representation of the user's current state may include representations in dimensions such as emotion categories (such as "anxiety", etc.), emotional intensity (levels 1-5, etc.), physiological state (such as heart rate change rate, etc.), and behavioral state (such as activity level).

[0028] Step S104: Based on the emotion recognition result and the comprehensive representation of the user's current state, generate an AI response strategy corresponding to the target user, and determine the target emotion alleviation content from a pre-constructed emotion alleviation database, where multiple emotion alleviation contents are stored in the emotion alleviation database, and each emotion alleviation content corresponds to an attribute label one by one.

[0029] Specifically, the above AI response strategy generation module can be implemented using causal reinforcement learning. The above AI response strategy is used to represent one or more response actions that conform to the user's preferences, such as response actions like playing music and nodding. The above response content generation module may include a Bidirectional Encoder Representations from Transformers (abbreviated as BERT) and an end-to-end text-to-speech synthesis model Tacotron, etc. Among them, the BERT model is used to generate the target response text, and the Tacotron model is used for the target response speech. During the construction phase of the above emotion alleviation database, it is necessary to collect multimedia data from the Internet , including emotion alleviation contents such as soothing music, inspiring pictures, healing videos, and physiological feedback samples (such as heart rate changes). Each emotion alleviation content is labeled with an attribute label , where is the emotion category (such as "anxiety") of the i-th emotion alleviation content, is the emotional intensity (levels 1-5) of the i-th emotion alleviation content, is the physiological effect (such as heart rate change rate) of the i-th emotion alleviation content, is the behavioral characteristic (such as activity level) of the i-th emotion alleviation content. The above emotion alleviation database can be stored in a NoSQL database (such as MongoDB, etc.) on the cloud platform to support fast retrieval. The above target emotion alleviation content may refer to the emotion alleviation content in the emotion alleviation database that has the highest degree of matching with the emotion recognition result and the comprehensive representation of the user's current state of the target user.

[0030] Step S105: Based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content, generate a response content and output the response content.

[0031] Specifically, the above response content can be in the form of text, voice, multimedia, physical feedback, etc. When outputting the response content, it can be responded according to the form of the response content. For example, for multimedia in the form of pictures, videos or text, etc., it is output through the display interface; for multimedia in the form of music or voice, etc., it is output through the speaker; for physical feedback, it is output through the vibration of the tactile vibration motor, the movement of the flexible mechanical component, etc.

[0032] Through the above method, the pre-trained emotion interaction engine can be used to fuse and recognize the emotions of the multimodal data of the target user, and through the comprehensive representation of the emotion recognition result and the user's current state, an AI response strategy that conforms to the personalization of the target user can be generated, thereby improving the accuracy of the emotion recognition result, and enabling the generated response content to meet the personal preferences of the target user, and further achieving the purpose of improving the AI interaction effect.

[0033] In an optional embodiment, the above step S102, using the pre-trained emotion interaction engine to perform feature fusion on the multimodal data to obtain feature fusion data, includes: Using the pre-trained emotion interaction engine to preprocess the multimodal data to obtain the feature data corresponding to each modal data, where the multimodal data includes at least one of image data, voice data, touch data, heart rate data, and motion data; Using a graph attention network to perform first feature fusion on the feature data corresponding to each modal data to obtain the intermediate feature vector corresponding to each modal data, where the intermediate feature vector is used to reflect the correlation between its own modal data and other modal data; Using a deep learning model to perform second feature fusion on the intermediate feature vector to obtain feature fusion data, where the feature fusion data is used to characterize the temporal features of each modal data.

[0034] Specifically, when using the multimodal feature fusion module in the pre-trained emotion interaction engine to perform feature fusion on the multimodal data, the multimodal data can be preprocessed first to obtain the feature data corresponding to each modal data, then a graph attention network is used to perform first feature fusion on the feature data corresponding to each modal data to obtain the intermediate feature vector corresponding to each modal data, and then a deep learning model is used to perform second feature fusion on the intermediate feature vector to obtain feature fusion data. When preprocessing the image data, the original image can be adjusted to a fixed resolution (such as 224×224 pixels, etc.), and then normalized using the mean and standard deviation of the image pixel values (usually calculated based on the training dataset). Specifically, it can be represented by the following formula: ; where, represents the original image data at time t (this original image data is usually the user's facial image captured by a camera, including expression information such as frowning and smiling). represents the original image data the image feature vector after preprocessing, represents adjusting the original image to a fixed resolution of 224×224 pixels. Of course, it can also be other resolutions to unify the image size, facilitate model processing, and ensure input consistency. represents the mean of the image pixel values, represents the standard deviation of the adjusted image pixel values. Subtracting the mean in the formula can make the pixel values centered around 0 and reduce the data distribution deviation. Dividing by the standard deviation in the formula can make the pixel value scales consistent and improve the stability of model training.

[0035] When preprocessing speech data, the original speech signal can be denoised, and then Mel-Frequency Cepstral Coefficients are extracted from the denoised speech signal to obtain a speech feature matrix. Specifically, it can be expressed by the following formula: ; where, represents the original speech signal at time t (this original speech signal can be collected by a microphone, and it contains speech features such as intonation and speech rate, which can reflect the user's emotion). represents the original speech signal the speech feature matrix after preprocessing. The dimension of the speech feature matrix is , and the speech feature matrix indicates that each frame of speech contains 13 cepstral coefficients, represents the number of frames of the speech segment. represents denoising the original speech signal to remove background noise, improve the quality of the speech signal, and ensure the accuracy of the extracted features. represents the extraction process of Mel-Frequency Cepstral Coefficients, which is used to convert the denoised speech signal into a set of features, capture the frequency and time-domain characteristics of the speech, and is suitable for emotion analysis.

[0036] When preprocessing touch data, the original touch data can be normalized to obtain a touch feature vector. Specifically, it can be expressed by the following formula: ; where, Represents the original touch data at time t (this original touch data can be the touch force collected by a flexible electronic skin sensor, which can reflect user interaction behaviors such as gentle strokes and hard presses, providing emotional cues). Represents the original touch data The touch feature vector after preprocessing. Norm( ) represents performing normalization on the original touch data to scale the touch force to the range of 0 to 1, eliminating the dimension difference for easier model processing. Represents the maximum value in the original touch data which is used as the denominator for normalization to ensure that the touch data divided by the maximum value is in the range of 0 to 1.

[0037] When preprocessing heart rate data, the original heart rate signal can be filtered, and then the heart rate variability can be calculated based on the filtered heart rate signal to obtain the heart rate feature vector. It can be specifically expressed by the following formula: ; where represents the original heart rate signal at time t (this original heart rate signal can be collected by a PPG sensor, which can reflect the user's physiological state, such as heart rate, and is used for emotional analysis). Represents the heart rate signal after filtering the original heart rate signal for calculating the heart rate variability (Heart Rate Variability, abbreviated as HRV). Represents applying a band - pass filter (abbreviated as ) to the original heart rate signal The frequency range of this band - pass filter can be from 0.5 Hz to 5 Hz. In this way, low - frequency drift and high - frequency noise can be removed, and the main signals related to the heart rate can be retained. Represents the heart rate variability, which measures the degree of fluctuation of the heart rate signal and can reflect the user's emotional state. Represents the calculation formula of HRV, and the calculation result is the standard deviation based on the filtered heart rate signal. Its function is to quantify the heart rate fluctuation. The larger the standard deviation, the higher the HRV, reflecting emotional changes. Represents the number of sampling points of the heart rate signal. Represents the mean value of the filtered heart rate signal .

[0038] When preprocessing motion data, the original IMU data can be subjected to a fast Fourier transform to obtain the IMU feature vector. It can be specifically expressed by the following formula: ; Among them, represents the original IMU data at time t, including acceleration and angular velocity. It reflects the user's motion state, such as stationary or jittering, etc., and provides behavioral characteristics. represents applying the Fast Fourier Transform (abbreviated as ) to the original IMU data to convert the time-domain signal into a frequency-domain signal and extract the frequency characteristics of the motion, such as the jitter frequency. represents the original IMU data The IMU feature vector after preprocessing, which contains information such as the jitter frequency. represents the sum of the squares of the accelerations on the x, y, and z axes in the original IMU data. In this way, the accelerations of the three axes can be integrated to calculate the total motion intensity, which is used to extract the jitter frequency. represents the absolute value of the spectral amplitude of the total acceleration, which can reflect the frequency distribution of the motion. The larger the amplitude, the more significant the motion at that frequency. argmax represents the frequency index of taking the maximum value, so that the frequency with the largest amplitude in the spectrum can be found as the main jitter frequency. represents the extracted main jitter frequency, which is used as an IMU feature vector to reflect the user's motion pattern, such as high-frequency jitter indicating nervousness, etc.

[0039] When performing the first feature fusion on the feature data corresponding to each modality data using the graph attention network to obtain the intermediate feature vector corresponding to each modality data, the preprocessed can be input into the attention network, so that the attention network can be used to obtain the correlation between each modality data. Among them, GAT updates the node features through the attention mechanism, and its formula is as follows: ; Among them, is the i-th modality feature (such as expression, etc.), is the j-th modality feature (such as heart rate, etc.), is and the attention weight between them, and the numerical range is from 0 to 1. This determines the contribution size of the update to . The larger the value of , the greater the influence of on . . W represents the learnable weight matrix, and σ is the ReLU activation function, which is used to perform a non-linear transformation on the fused features, enhance the model's expression ability, and highlight important features. represents the i-th modality feature The updated feature vector at the l+1-th (next) layer of the GAT. It is the new feature representation output by the GAT, which integrates modal features information of itself and other relevant modalities for subsequent sentiment recognition. Denote the feature vector of the j-th modal feature at the l-th (current) layer of the GAT. Denote the i-th modal feature the set of neighbor modalities of. Denote the scaling factor, which is used to prevent the similarity value from being too large or too small.

[0040] Using the above method, the node features of each layer of the GAT can be updated, and finally the intermediate feature vectors corresponding to each modality data can be obtained , that is, the node features of the last layer of the GAT.

[0041] Next, input the intermediate feature vectors corresponding to each modality data into a deep learning model (such as a Transformer model) to perform a second feature fusion on the intermediate feature vectors to obtain feature fusion data. Specifically, it can be expressed by the following formula: ; where, denote the feature fusion data output by the model, which is the feature representation after being processed by the Transformer, containing the temporal information of multi-modal data (such as the change of emotion over time, etc.) for subsequent tasks. denote the intermediate feature vectors output to the model, which contains the comprehensive information of modalities such as expression, speech, heart rate, touch, movement, etc., reflecting the current state of the user. Q, K, and V are the query, key, and value respectively, , , , , and are learnable weight matrices respectively. denote the normalization factor.

[0042] In the above way, the graph attention network can be used to perform the first feature fusion on the feature data corresponding to each modality data, dynamically integrate each modality through the attention mechanism to generate a unified feature representation, and then perform a second feature fusion on the intermediate feature vectors through a deep learning model to extract the temporal features of each modality data to obtain the feature fusion data, which is convenient for subsequent sentiment classification and response content generation based on the feature fusion data.

[0043] In an optional embodiment, in the above step S103, sentiment recognition and context analysis are respectively performed on the feature fusion data to obtain a sentiment recognition result and a comprehensive representation of the user's current state, including: Using a sentiment recognition model to perform sentiment recognition on the feature fusion data to obtain a sentiment recognition result, where the sentiment recognition result is used to characterize the sentiment category and sentiment intensity of the target user; Obtain the historical interaction data of the target user, and perform context analysis based on the historical interaction data and the feature fusion data to generate a comprehensive representation of the user's current state, where the comprehensive representation of the user's current state is used to characterize the sentiment category, sentiment intensity, physiological state, and behavioral state of the target user.

[0044] Specifically, a bidirectional long short-term memory network BiLSTM can be used as the sentiment recognition model to perform sentiment recognition on the feature fusion data to obtain a sentiment recognition result, that is, to obtain the sentiment category and sentiment intensity of the target user. Among them, the processing flow of the bidirectional long short-term memory network BiLSTM is as follows: ; ; ; Among them, represents the feature fusion data, which can be generated by a graph attention network (GAT) and a Transformer processing multi-source data (such as expressions, voices, heart rates, touches, etc.). It is the input of sentiment recognition, contains cross-modal comprehensive information, and the dimension depends on the model design. and respectively represent the hidden states of the bidirectional long short-term memory network BiLSTM at the previous moment (forward) and the next moment (backward) at time step t. BiLSTM is bidirectional, is used for the forward LSTM to capture the context from the start of the sequence to time step t-1; is used for the backward LSTM to capture the context from the end of the sequence to time step t+1. These hidden states can help the model understand the context dependencies in the time series. and respectively represent the output hidden states of the forward LSTM and the backward LSTM at time step t. is calculated based on and to reflect the information from the past to the present; is calculated based on and to reflect the information from the future to the present. and combine to provide a complete temporal context. A concatenated vector representing the forward and backward hidden states, with a dimension of and twice the dimension. The concatenated vector combines bidirectional context information and is used for subsequent sentiment classification. Represents the weight matrix of the output layer, which maps the concatenated hidden state to the dimension of the sentiment category (i.e., k-dimensional, where k is the number of sentiment categories). Represents the bias vector of the output layer, with a dimension of k, which is used to adjust the result of the linear transformation and enhance the expressive power of the model. Represents the sentiment classification output at time step t, which is a k-dimensional vector. After being normalized by the softmax function, it represents a probability distribution, that is , where is the probability of the i-th sentiment category, satisfying . Represents the sentiment probability distribution, which is a k-dimensional vector .

[0045] In addition, the sentiment interaction engine can also generate a comprehensive representation of the user's current state based on the historical interaction data of the target user and real-time physiological and behavioral characteristics (such as heart rate change rate, activity level, etc.). This comprehensive representation of the user's current state contains information corresponding to the labeled attributes in the database (such as sentiment category, sentiment intensity, physiological state, and behavioral state, etc.). In this way, it is convenient to determine the target emotion relief content from the emotion relief database based on the comprehensive representation of the user's current state.

[0046] In an optional embodiment, the above step S104, generating an AI response strategy corresponding to the target user based on the sentiment recognition result and the comprehensive representation of the user's current state, and determining the target emotion relief content from the pre-constructed emotion relief database, includes: Based on the sentiment recognition result and the comprehensive representation of the user's current state, construct a causal graph, where the causal graph is used to characterize the causal relationship between the user state and the response action, and the user state includes the sentiment recognition result and the comprehensive representation of the user's current state; Based on the causal graph and the preset reward function, determine the AI response strategy corresponding to the target user, where the AI response strategy is used to characterize one or more response actions; Obtain the historical interaction data of the target user, and based on the historical interaction data of the target user, determine the historical preference vector of the target user. Based on the historical preference vector of the target user and the comprehensive representation of the user's current state, determine the target emotion relief content from the emotion relief database.

[0047] Specifically, when generating the AI response strategy corresponding to the target user based on the comprehensive representation of the emotion recognition result and the user's current state, a causal graph can be constructed based on the comprehensive representation of the emotion recognition result and the user's current state, and the AI response strategy corresponding to the target user can be determined based on the causal graph and a preset reward function. The causal graph here can be expressed as G=(V,E), where G represents the causal graph, V represents the node set of the causal graph, and each node can include emotion state , fused features , actions and other variables. E represents the edge set of the causal graph, which is used to describe the causal relationship between variables. For example, emotion state affects action or action affects the emotion at the next moment . The causal graph is used to model the causal association between the emotion state and the response action, improving the interpretability of the strategy. The reward function here can be expressed as , where represents the state vector at time step t, which is composed of the current emotion state and the feature fusion data , that is . is the current emotion category extracted from (usually the category with the highest probability or a certain representation of the probability distribution). provides multi-modal context information, enhancing the richness of the state. is the input data for reinforcement learning, which is used to determine the action. represents the action at time step t, such as "play music", "display a healing video", "perform haptic feedback", etc. represents the reward function for performing action in state . represents the expected value of the emotion state at the next moment after performing action in state . is similar to , which is the emotion probability distribution at the next moment. This part measures the immediate impact of the action on the user's emotion. represents the long-term effect difference based on the Q value, is the weight coefficient, which is used to balance the immediate and long-term rewards. is the reward Q value of the current action, is the reward Q value of other alternative actions. This part optimizes the strategy by comparing the long-term value of the actions. represents the state-action pair The reward Q-value measures the long-term cumulative reward for performing action a in state s. , is the concatenated vector of state and action, is the learnable weight matrix. The Q-value is used to evaluate the quality of actions and optimize the long-term interaction effect.

[0048] When obtaining the target emotion-relieving content, the historical interaction data of the target user can be obtained, and based on the historical interaction data of the target user, the historical preference vector of the target user can be determined. Based on the comprehensive representation of the historical preference vector of the target user and the user's current state, the target emotion-relieving content can be determined from the emotion-relieving database. Among them, determining the target emotion-relieving content from the emotion-relieving database can be achieved through the following formula: ; Among them, C represents the extracted target emotion-relieving content (such as soothing music, inspiring pictures, healing videos, etc.). D represents the emotion-relieving database, which stores multimedia content annotated with attributes such as emotion category, emotion intensity, physiological state, and behavior state. S represents the comprehensive representation of the user's current state. represents the similarity between the comprehensive representation S of the user's current state and the emotion-relieving content c in the emotion-relieving database. represents the historical preference vector of the target user, reflecting the user's past interaction preferences (such as preferring a certain type of music, etc.). represents the matching degree between the historical preference vector and the emotion-relieving content c, usually calculated by the inner product of vectors. represents the weight parameter, which is used to balance the similarity score and the matching degree score.

[0049] In this way, by constructing a causal graph and combining the reward function to optimize the AI response strategy, the interaction strategy can be dynamically optimized, and the optimal response can be selected according to the user feedback (such as heart rate change, touch feedback), rather than relying on static mapping. And causal reinforcement learning evaluates the potential effects of different actions through counterfactual reasoning, such as "what would the effect be if another response was selected". This ensures that the system preferentially selects actions that have a long-term contribution to the improvement of the user's emotion (such as continuously relieving anxiety). The optimized reward is based on the Q-value, combined with the historical preference vector , ensuring that the strategy takes into account the user's historical interaction data, thereby improving the long-term user satisfaction. In addition, since the causal graph G = (V, E) explicitly models the causal relationship between the state S and the action A, such as "does playing a healing video directly lead to a decrease in heart rate", this makes the interaction strategy more interpretable, facilitating the analysis of which responses are effective for emotion improvement, and is superior to the "black box" method of direct mapping.

[0050] In an optional embodiment, the above step S105, based on the AI response strategy, feature fusion data, emotion recognition result, and target emotion alleviation content, generates a response content and outputs the response content, including: When the AI response strategy includes text response, based on the feature fusion data and emotion recognition result, generate a target response text; When the AI response strategy includes voice response, based on the feature fusion data and emotion recognition result, generate a target response text, and perform voice conversion on the target response text to obtain a target response voice; When the AI response strategy includes multimedia content response, based on the target emotion alleviation content, obtain target multimedia content, where the type of the target multimedia content includes at least one of music, video, and picture; When the AI response strategy includes physical feedback response, based on the emotion recognition result, generate a target physical feedback, where the target physical feedback includes at least one of the vibration of a tactile vibration motor and the mechanical action performed by a flexible mechanical component.

[0051] Specifically, when the AI response strategy includes text response, a target response text can be generated based on the feature fusion data and emotion recognition result. The implementation process is as follows: ; Among them, represents the generated target response text, which is the output of the BERT model. is a natural language text for language interaction with the user (such as comforting words or guiding statements). represents the current emotional state at time step t, usually the highest probability emotional category (such as "happy", "anxious") or the probability distribution itself extracted from the emotion probability distribution . is used to guide the BERT model to generate text content that matches the user's emotion. represents the feature fusion data, which is generated after the graph attention network and Transformer process multi-source data (such as expressions, voices, heart rates, touches, movements, etc.). provides rich context information to help the BERT model generate more accurate and personalized text. represents the t-th word in the generated target response text. During the generation process of the BERT model, is the next word predicted according to the previous text , emotional state and feature fusion data . represents given the previous text , emotional state and feature fusion data Under the condition of to generate the probability of a word. The BERT model selects the most appropriate word by maximizing this probability to generate the complete text T.

[0052] When the AI response strategy includes a voice response, based on the feature fusion data and the sentiment recognition result, a target response text can be generated, and the target response text can be converted into speech to obtain a target response voice. The implementation process is as follows: ; where represents the target response voice synthesized by the Tacotron2 model, which is the output of converting the target response text T into audio. is a two-dimensional matrix with dimensions , where represents the time step of the target response voice, that is, the total number of frames of the target response voice, and its value depends on the length of the target response text T and the sampling rate of speech synthesis. represents the feature dimension of each frame of the target response voice, usually the dimension of the Mel spectrogram output by Tacotron2. describes the spectral characteristics of the voice and is used for the subsequent vocoder to generate waveforms.

[0053] When the AI response strategy includes a multimedia content response, based on the target emotion alleviation content, the target multimedia content can be obtained. The implementation process has been described in detail in the above embodiments and will not be elaborated here.

[0054] When the AI response strategy includes a physical feedback response, based on the sentiment recognition result, a target physical feedback can be generated. The implementation process is as follows: ; ; where represents the vibration frequency of the haptic vibration motor (in Hz) for physical feedback, which reflects the intensity of the user's emotional state. represents the scaling coefficient of the vibration frequency, and the value of k can be 10Hz / unit. is used to convert the quantified value of the emotional state into the actual vibration frequency. represents the quantified value of the emotional state , usually a scalar calculated based on the emotion category or probability distribution. reflects the intensity of the emotion or the characteristics of a specific emotion, and determines the magnitude of the vibration frequency. Represents the action angle (in degrees) of the flexible mechanical component, used for physical feedback (such as nodding or arm movement). Represents the emotional state Symbol (positive or negative), used to determine the action direction. For example, a positive emotion (such as "happy") may correspond to a positive action (such as nodding), and a negative emotion (such as "sad") corresponds to a reverse action (such as lowering the head). Represents the emotional state Intensity, defined as the maximum value of the emotional probability distribution P( ), that is, Intensity( ) = max(P( ))). It is a scalar that reflects the confidence or intensity of the current emotional category and is used to adjust the action amplitude.

[0055] In the above way, various forms of response content can be obtained, making the AI response content richer. In particular, a tactile vibration motor is added to simulate the heartbeat rhythm and cooperate with the flexible mechanical component to achieve action feedback, significantly enhancing the immersion and emotional expression ability of physical interaction compared with traditional hardware designs.

[0056] In an optional embodiment, after the above step S105, based on the AI response strategy, feature fusion data, emotion recognition result, and target emotion relief content, generate response content and output the response content, the method further includes: Collect the feedback made by the target user based on the response content to obtain multi-modal feedback data; Use the feedback data to adaptively optimize the model parameters of the emotion interaction engine.

[0057] Specifically, after outputting the response content, it is also possible to collect the feedback made by the target user based on the response content to obtain multi-modal feedback data, such as new voice data, new heart rate data, etc. fed back by the target user. Then use the feedback data to adaptively optimize the model parameters of the emotion interaction engine. The specific process is as follows: ; Among them, Represents the optimization objective function, used to evaluate and optimize the performance of the model parameters of. Integrates the contributions of three aspects: emotion recognition accuracy, voice feedback, and heart rate feedback, and is calculated by weighted summation of the three. Represents the model parameters of the emotion interaction engine, including the weights of networks such as GAT, Transformer, BiLSTM, etc. The optimization objective is to improve the model performance by adjusting . Represents the accuracy of emotion recognition, which measures the consistency between the predicted emotion output y of the model (i.e., the category with the highest probability in P(e)) and the true emotion label of. The calculation formula is the number of correctly predicted samples divided by the total number of samples, and the weight is 0.4, indicating that accuracy is the main factor for optimization. y represents the predicted emotion output of the model, usually the highest probability emotion category selected from , which is generated by the BiLSTM model and used to compare with the true emotion label . Represents the true emotion label of the user, usually obtained through annotation or external reference (such as user self-report). Represents the average value of the voice feedback of all users , which reflects the overall satisfaction of users with the AI response, with a weight of 0.3, and N is the total number of users. Represents the normalized value of the heart rate change feedback , usually the result after standardizing or taking the absolute value of . Quantifies the intensity of physiological feedback, with a weight of 0.3, reflecting the contribution of heart rate changes to model optimization.

[0058] In the above way, the emotion interaction engine can be continuously optimized according to user feedback, making the emotion interaction engine more and more in line with user preferences.

[0059] In an optional embodiment, before the above step S102, using the pre-trained emotion interaction engine to perform feature fusion on multimodal data to obtain feature fusion data, the method further includes: Using meta-learning to train the initial model parameters of the emotion interaction engine, where the initial model parameters are learned based on multiple learning tasks of a single user; Using federated learning to train the global model parameters of the emotion interaction engine to obtain the emotion interaction engine, where the global model parameters are aggregated based on the initial model parameters of multiple users.

[0060] Specifically, when training to obtain the emotion interaction engine, meta-learning can be used to train the initial model parameters of the emotion interaction engine, and then federated learning can be used to train the global model parameters of the emotion interaction engine to obtain the emotion interaction engine. Here, the meta-learning can be the Model-Agnostic Meta-Learning (abbreviated as MAML) algorithm, which can be expressed by the following formula: ; Among them, represents the original model parameters of the kth user, denotes the model parameters after the k-th user's learning. denotes the learning rate. denotes gradient calculation. denotes the o-th task, such as the user's sentiment recognition. denotes the task 's loss function. The federated learning (abbreviated as FL) here can be expressed by the following formula: ; Among them, is the user weight, denotes the global model parameters, denotes the model parameters after the k-th user's learning, and N denotes the total number of users.

[0061] In this way, meta-learning can achieve fast cross-user adaptation, and combined with federated learning, it optimizes the global model through local training and cloud aggregation, avoiding centralized data upload, improving cross-scenario adaptability compared with traditional solutions and ensuring user privacy.

[0062] In an optional embodiment, the AI interaction method based on sentiment recognition provided by the embodiments of the present application can be applied to a sentiment interaction system, which can include five core modules: a sensor module, a processing module, an output module, a cloud connection module, and a cloud platform. Among them, the sensor module integrates a high-definition camera (for facial expression collection), a microphone (for voice collection), a touch sensor (for interaction action collection), and newly adds a PPG sensor (for heart rate monitoring), a flexible electronic skin (for touch force perception), and an IMU (for motion data collection) to achieve multi-dimensional data input. The processing module uses an embedded AI processor to run the sentiment AI interaction engine, fuses multi-modal features through a graph attention network and a Transformer, and combines meta-learning and federated learning to improve the model adaptability and privacy protection ability. The output module includes a speaker (for voice response), a display screen (for displaying emotion-relieving content), and newly adds a tactile vibration motor and a flexible mechanical component to provide physical feedback and action expression, enhancing the interaction experience. The cloud connection module communicates with the cloud platform through the network, accesses the emotion-relieving database and aggregates the model parameters to achieve dynamic optimization. The hardware connection relationship is: the data collected by the sensor module is transmitted to the processing module, and after the processing module generates a real-time response, it is presented through the output module, and at the same time, it interacts with the cloud platform for the optimization engine. This architecture significantly improves the sentiment recognition accuracy, interaction intelligence, and cross-scenario adaptability through multi-source data fusion and causal reinforcement learning, and is applicable to fields such as intelligent toys, education, and medical care, with technological leadership and application expansion potential. Its AI interaction process based on sentiment recognition is as Figure 2As shown, it may include the following steps: Step S201, construct an emotion alleviation database and an emotional interaction engine.

[0063] Step S202, collect and fuse multi-source and multi-modal data.

[0064] Step S203, generate an AI response strategy and target emotion alleviation content based on causal reasoning.

[0065] Step S204, output multi-sensory real-time AI responses.

[0066] Step S205, optimize the adaptive model based on multi-source feedback.

[0067] This application significantly improves emotion recognition and AI interaction performance through multi-source data fusion, integrating multiple algorithms and hardware designs. It enhances the accuracy and robustness of emotion recognition, enabling stable operation in complex environments. The model can quickly adapt to new users, with enhanced cross-scenario adaptability. The interaction strategy is more intelligent, improving user satisfaction and having better long-term effects. The addition of flexible skin and tactile feedback enhances the immersion, providing a richer user experience. Privacy is protected. This application achieves high precision, intelligence, and excellent experience in fields such as smart toys and education, with significant practical value.

[0068] See Figure 3 , Figure 3 which is a schematic structural diagram of an AI interaction device based on emotion recognition provided by an embodiment of this application. As Figure 3 shown, the AI interaction device 300 based on emotion recognition includes: An acquisition module 301, configured to acquire multi-modal data of a target user; A feature fusion module 302, configured to perform feature fusion on the multi-modal data by using a pre-trained emotional interaction engine to obtain feature fusion data; An analysis module 303, configured to perform emotion recognition and context analysis on the feature fusion data respectively to obtain an emotion recognition result and a comprehensive representation of the user's current state; A determination module 304, configured to generate an AI response strategy corresponding to the target user based on the comprehensive representation of the emotion recognition result and the user's current state, and determine target emotion alleviation content from a pre-constructed emotion alleviation database, where multiple emotion alleviation contents are stored in the emotion alleviation database, and each emotion alleviation content corresponds to an attribute label one by one; An output module 305, configured to generate a response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content, and output the response content.

[0069] Further, the feature fusion module 302 includes: A preprocessing sub-module for preprocessing multi-modal data using a pre-trained emotion interaction engine to obtain feature data corresponding to each modality of data, where the multi-modal data includes at least one of image data, speech data, touch data, heart rate data, and motion data; A first fusion sub-module for performing first feature fusion on the feature data corresponding to each modality of data using a graph attention network to obtain an intermediate feature vector corresponding to each modality of data, where the intermediate feature vector is used to reflect the correlation between its own modality of data and other modalities of data; A second fusion sub-module for performing second feature fusion on the intermediate feature vector using a deep learning model to obtain feature fusion data, where the feature fusion data is used to characterize the temporal features of each modality of data.

[0070] Further, the analysis module 303 includes: An emotion recognition sub-module for performing emotion recognition on the feature fusion data using an emotion recognition model to obtain an emotion recognition result, where the emotion recognition result is used to characterize the emotion category and emotion intensity of the target user; An analysis sub-module for obtaining the historical interaction data of the target user and performing context analysis based on the historical interaction data and the feature fusion data to generate a comprehensive representation of the user's current state, where the comprehensive representation of the user's current state is used to characterize the emotion category, emotion intensity, physiological state, and behavior state of the target user.

[0071] Further, the determination module 304 includes: A construction sub-module for constructing a causal graph based on the emotion recognition result and the comprehensive representation of the user's current state, where the causal graph is used to characterize the causal relationship between the user state and the response action, and the user state includes the emotion recognition result and the comprehensive representation of the user's current state; A first determination sub-module for determining the AI response strategy corresponding to the target user based on the causal graph and a preset reward function, where the AI response strategy is used to characterize one or more response actions; A second determination sub-module for obtaining the historical interaction data of the target user and determining the historical preference vector of the target user based on the historical interaction data of the target user, and determining the target emotion alleviation content from the emotion alleviation database based on the historical preference vector of the target user and the comprehensive representation of the user's current state.

[0072] Further, the output module 305 includes: A first generation sub-module for generating a target response text based on the feature fusion data and the emotion recognition result when the AI response strategy includes a text response; A second generation sub-module, configured to generate a target response text based on the feature fusion data and the emotion recognition result when the AI response strategy includes a voice response, and perform voice conversion on the target response text to obtain a target response voice; A third generation sub-module, configured to obtain a target multimedia content based on the target emotion alleviation content when the AI response strategy includes a multimedia content response, where the type of the target multimedia content includes at least one of music, video, and picture; A fourth generation sub-module, configured to generate a target physical feedback based on the emotion recognition result when the AI response strategy includes a physical feedback response, where the target physical feedback includes at least one of the vibration of a tactile vibration motor and the execution of a mechanical action by a flexible mechanical component.

[0073] Further, the AI interaction device 300 based on emotion recognition further includes: An acquisition module, configured to acquire the feedback made by the target user based on the response content to obtain multi-modal feedback data; An optimization module, configured to adaptively optimize the model parameters of the emotion interaction engine by using the feedback data.

[0074] Further, the AI interaction device 300 based on emotion recognition further includes: A first training module, configured to train the initial model parameters of the emotion interaction engine by using meta-learning, where the initial model parameters are learned based on multiple learning tasks of a single user; A second training module, configured to train the global model parameters of the emotion interaction engine by using federated learning to obtain the emotion interaction engine, where the global model parameters are aggregated based on the initial model parameters of multiple users.

[0075] It should be noted that the AI interaction device 300 based on emotion recognition can implement the steps of the AI interaction method based on emotion recognition provided in any of the foregoing method embodiments, and can achieve the same technical effects, which will not be elaborated herein one by one.

[0076] As Figure 4 shown, an embodiment of the present application further provides an electronic device, including a processor 411, a communication interface 412, a memory 413, and a communication bus 414, where the processor 411, the communication interface 412, and the memory 413 complete communication with each other through the communication bus 414, The memory 413 is used to store a computer program; In an embodiment of the present application, when the processor 411 executes the program stored on the memory 413, it implements the AI interaction method based on emotion recognition provided in any of the foregoing method embodiments.

[0077] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the AI interaction method based on emotion recognition provided by any one of the foregoing method embodiments.

[0078] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0079] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. An AI interaction method based on emotion recognition, characterized in that, The method includes: Obtain multimodal data of the target user; Use a pre-trained emotion interaction engine to perform feature fusion on the multimodal data to obtain feature fusion data; Perform emotion recognition and context analysis on the feature fusion data respectively to obtain an emotion recognition result and a comprehensive representation of the user's current state; Based on the emotion recognition result and the comprehensive representation of the user's current state, generate an AI response strategy corresponding to the target user, and determine target emotion alleviation content from a pre-constructed emotion alleviation database, where multiple emotion alleviation contents are stored in the emotion alleviation database, and each emotion alleviation content corresponds to an attribute label one by one; Generate a response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content, and output the response content.

2. The AI interaction method based on emotion recognition according to claim 1, characterized in that, The using a pre-trained emotion interaction engine to perform feature fusion on the multimodal data to obtain feature fusion data includes: Use a pre-trained emotion interaction engine to preprocess the multimodal data to obtain feature data corresponding to each modality data, where the multimodal data includes at least one of image data, voice data, touch data, heart rate data, and motion data; Use a graph attention network to perform first feature fusion on the feature data corresponding to each modality data to obtain an intermediate feature vector corresponding to each modality data, where the intermediate feature vector is used to reflect the correlation between its own modality data and other modality data; Use a deep learning model to perform second feature fusion on the intermediate feature vector to obtain the feature fusion data, where the feature fusion data is used to characterize the temporal features of each modality data.

3. The AI interaction method based on emotion recognition according to claim 1, wherein, The performing emotion recognition and context analysis on the feature fusion data respectively to obtain an emotion recognition result and a comprehensive representation of the user's current state includes: Use an emotion recognition model to perform emotion recognition on the feature fusion data to obtain the emotion recognition result, where the emotion recognition result is used to characterize the emotion category and emotion intensity of the target user; Obtain the historical interaction data of the target user, and perform context analysis based on the historical interaction data and the feature fusion data to generate a comprehensive representation of the user's current state, where the comprehensive representation of the user's current state is used to characterize the emotion category, emotion intensity, physiological state, and behavior state of the target user.

4. The AI interaction method based on emotion recognition according to claim 1, characterized in that, The based on the emotion recognition result and the comprehensive representation of the user's current state, generate an AI response strategy corresponding to the target user, and determine target emotion alleviation content from a pre-constructed emotion alleviation database includes: Based on the emotion recognition result and the comprehensive representation of the user's current state, construct a causal graph, where the causal graph is used to characterize the causal relationship between the user state and the response action, and the user state includes the emotion recognition result and the comprehensive representation of the user's current state; Based on the causal graph and a preset reward function, determine the AI response strategy corresponding to the target user, where the AI response strategy is used to represent one or more response actions; Obtain the historical interaction data of the target user, and based on the historical interaction data of the target user, determine the historical preference vector of the target user. Based on the comprehensive representation of the historical preference vector of the target user and the current state of the user, determine the target emotion alleviation content from the emotion alleviation database.

5. The AI interaction method based on emotion recognition according to claim 1, wherein The generating and outputting the response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content includes: When the AI response strategy includes a text response, generate a target response text based on the feature fusion data and the emotion recognition result; When the AI response strategy includes a voice response, generate a target response text based on the feature fusion data and the emotion recognition result, and perform voice conversion on the target response text to obtain a target response voice; When the AI response strategy includes a multimedia content response, obtain target multimedia content based on the target emotion alleviation content, where the type of the target multimedia content includes at least one of music, video, and picture; When the AI response strategy includes a physical feedback response, generate a target physical feedback based on the emotion recognition result, where the target physical feedback includes at least one of the vibration of a tactile vibration motor and the execution of a mechanical action by a flexible mechanical component.

6. The AI interaction method based on emotion recognition according to claim 1, wherein After generating and outputting the response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content, the method further includes: Collect the feedback made by the target user based on the response content to obtain multi-modal feedback data; Use the feedback data to adaptively optimize the model parameters of the emotion interaction engine.

7. The AI interaction method based on emotion recognition according to claim 1, wherein Before using the pre-trained emotion interaction engine to perform feature fusion on the multi-modal data to obtain feature fusion data, the method further includes: Use meta-learning to train the initial model parameters of the emotion interaction engine, where the initial model parameters are learned based on multiple learning tasks of a single user; Use federated learning to train the global model parameters of the emotion interaction engine to obtain the emotion interaction engine, where the global model parameters are aggregated based on the initial model parameters of multiple users.

8. An AI interaction device based on emotion recognition, characterized in that, The device includes: An acquisition module, configured to acquire multi-modal data of a target user; A feature fusion module, configured to use a pre-trained emotion interaction engine to perform feature fusion on the multi-modal data to obtain feature fusion data; An analysis module, configured to perform emotion recognition and context analysis on the feature fusion data respectively to obtain an emotion recognition result and a comprehensive representation of the current state of the user; A determination module, configured to generate an AI response strategy corresponding to the target user based on a comprehensive representation of the emotion recognition result and the current state of the user, and determine target emotion alleviation content from a pre-constructed emotion alleviation database, where multiple emotion alleviation contents are stored in the emotion alleviation database, and each emotion alleviation content corresponds to an attribute label one by one; An output module, configured to generate a response content based on the AI response strategy, the feature fusion data, the emotion recognition result, and the target emotion alleviation content, and output the response content.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is configured to implement the AI interaction method based on emotion recognition according to any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the AI interaction method based on emotion recognition according to any one of claims 1-7.

Citation Information

Patent Citations

  • LUI and GUI method and system based on emotion module

    CN119441439A

  • AI interaction method and system based on emotion recognition

    CN119476488A

  • Multi-modal digital employee reception system and method based on edge calculation

    CN119693750A

  • Method for user recognition and emotion monitoring based on smart headset

    US20220188392A1

Cited By

  • Language interaction method and device based on multi-perception ability, equipment and medium

    CN117093893A