Emotional interaction agent interaction method based on scene and modal double attention mechanism
By employing an emotional interaction agent interaction method based on scene and modality dual attention mechanism, multimodal interaction data and scene context data are comprehensively collected, and multimodal feature extraction and correlation analysis are performed. This solves the problem of insufficient accuracy in user state recognition by the emotional interaction agent, and achieves more accurate and natural human-computer interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 广州云趣信息科技有限公司
- Filing Date
- 2026-04-20
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies for emotional interaction agents lack sufficient accuracy in recognizing user states, resulting in incomplete and inaccurate emotional understanding. This makes it difficult to accurately identify user states in complex interaction scenarios, affecting the smoothness of interaction and user experience.
An emotional interaction agent interaction method based on scene and modality dual attention mechanism is adopted. By comprehensively collecting multimodal interaction data and scene context data, multimodal feature extraction and correlation analysis are performed, and comprehensive emotional features and scene features are integrated to generate an adaptive response scheme.
It improves the accuracy and rationality of user state recognition, enhances the precision, consistency and naturalness of human-computer interaction, and ensures that intelligent agent interaction is more in line with users' real emotions and scenario needs.
Smart Images

Figure CN122064233A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, specifically to an emotional interaction intelligent agent interaction method based on a scene and modality dual attention mechanism. Background Technology
[0002] As human-computer interaction technology evolves towards greater intelligence and naturalness, the application scenarios for emotional interactive intelligent agents continue to expand. Existing technologies are gradually developing from single-modal interaction to multi-modal fusion interaction, and from static response to dynamic scene adaptation. Traditional methods generally employ multi-modal feature extraction and basic attention mechanisms to process interactive data such as speech, text, and micro-expressions, while introducing scene context modeling to improve interaction adaptability.
[0003] However, existing technologies only analyze emotional features of a single modality or simple fusion, resulting in incomplete and inaccurate emotional understanding. Scene context data is often processed in isolation or used only as auxiliary information, making it difficult to truly reflect the dynamic state of users in changing situations. This makes the emotional interaction agent less accurate in recognizing user states when facing complex interaction scenarios, and easily leads to problems such as emotional misjudgment, mechanical response, or disconnect from user needs, affecting the smoothness of interaction and user experience. Summary of the Invention
[0004] This invention provides an interaction method for emotional interaction agents based on a scene and modality dual attention mechanism, aiming to solve the technical problem of insufficient accuracy in recognizing user states in existing emotional interaction agents.
[0005] In view of the above problems, this invention provides an interaction method for emotion-based interactive intelligent agents based on a scene and modality dual attention mechanism, including: Acquire interaction data and scene context data during the interaction between the emotional interaction intelligent agent and the user, and obtain the interaction dataset and scene dataset; Multimodal feature extraction is performed on the interactive dataset to obtain demand features and sentiment features, and correlation analysis is performed on the sentiment features to obtain comprehensive sentiment features; The scene context data is subjected to feature extraction and analysis to obtain scene features; By integrating the comprehensive emotional features and the scene features, user status recognition is performed to obtain user status data; Based on the user status data and demand data, a response plan for the intelligent agent is obtained and the intelligent agent interacts.
[0006] One or more technical solutions provided in this invention have at least the following technical effects or advantages: This invention provides an emotional interaction agent interaction method based on a dual attention mechanism of scene and modality. By comprehensively collecting interaction data and scene context data, it ensures complete acquisition of interaction and scene information. Through multimodal feature extraction and emotional feature correlation analysis, single-modal biases can be eliminated to obtain accurate and reliable comprehensive emotional features. Scene context feature extraction and analysis enable precise filtering of effective scene information. By integrating comprehensive emotional features and scene features for user state recognition, the accuracy and rationality of user state judgment are improved. Finally, by combining user state and demand data, an adaptive response scheme is generated, making the agent interaction more aligned with the user's true emotions and scene needs, thus improving the accuracy, coherence, and naturalness of human-computer emotional interaction. Attached Figure Description
[0007] Figure 1 This is a flowchart illustrating the interaction method of the emotion-interactive intelligent agent based on scene and modality dual attention mechanism provided in an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the process of obtaining comprehensive emotional features in the emotional interaction agent interaction method based on scene and modality dual attention mechanism provided in the embodiments of the present invention. Detailed Implementation
[0008] This invention provides an interaction method for emotional interaction agents based on a scene and modality dual attention mechanism, which is used to address the technical problem of insufficient accuracy in recognizing user states by emotional interaction agents in the prior art.
[0009] Examples, such as Figure 1 As shown, this invention provides an interaction method for an emotion-based interactive intelligent agent based on a scene and modality dual attention mechanism, the method comprising: S100: Acquire interaction data and scene context data during the interaction process between the emotional interaction agent and the user, and acquire the interaction dataset and scene dataset.
[0010] In this embodiment of the invention, interaction data and scene context data during the interaction between the emotional interaction agent and the user are acquired to obtain an interaction dataset and a scene dataset. The interaction effect of the emotional interaction agent depends on a high-quality raw data foundation. This step comprehensively acquires multimodal interaction raw data and historical interaction scene semantics through a standardized, multi-dimensional collection and aggregation process, constructing a standardized interaction dataset and scene dataset, providing a complete, coherent, and high-quality raw data foundation for subsequent multimodal feature extraction and scene feature analysis.
[0011] Step S100 in the method provided in this embodiment of the invention includes: Collect multimodal raw interaction information during the interaction process between the user and the emotional interaction intelligent agent, wherein the multimodal raw interaction information includes voice interaction information, text interaction information and micro-expression imaging information; Semantic extraction is performed on the original multimodal interaction information, and the original interaction information is extracted into semantic vectors and integrated to form an interaction dataset; Extract the semantics of previous interactions before this interaction as the semantic context information of the interaction, collect and organize them into scene context data and form a scene dataset.
[0012] First, multimodal raw interaction information is collected during the interaction between the user and the emotional interaction agent. This multimodal raw interaction information includes voice interaction information, text interaction information, and micro-expression imaging information. Multimodal raw interaction information refers to the various types of unprocessed raw interaction signals generated during the interaction between the user and the emotional interaction agent, covering three modalities: voice interaction information, text interaction information, and micro-expression imaging information. It is the original source for extracting emotional features and interaction needs.
[0013] Voice interaction information refers to the raw voice signals generated when a user communicates with the intelligent agent via voice, including raw features such as voice content, tone of voice, and speech rate. Text interaction information refers to the raw text content generated when a user communicates with the intelligent agent via text input. Micro-expression imaging information refers to the raw image data formed by capturing the user's dynamic facial micro-expressions through the image acquisition module on the intelligent agent, including visual features such as subtle textures of facial muscle movements and facial postures, such as frowning and downturned corners of the mouth.
[0014] Specifically, the emotional interaction agent synchronously triggers three types of acquisition units through a built-in multimodal acquisition module: the voice acquisition unit captures the user's voice signal in real time and generates raw voice interaction information; the text acquisition unit receives the user's text input and generates raw text interaction information; and the micro-expression imaging acquisition unit captures the user's facial dynamics at a fixed frame rate through a high-definition camera and generates raw micro-expression imaging information. The three types of acquisition units start and stop synchronously to ensure that the three types of raw information are completely aligned in the time dimension and avoid time misalignment between modalities.
[0015] For example, when a user speaks to the emotional interaction agent, saying, "I was criticized by my boss at work today, and the plan was rejected even after I revised it three times. I'm in a really bad mood," the voice acquisition unit of the emotional interaction agent captures the original voice signal, which is the voice interaction information. At the same time, the user inputs the same text content via the keyboard, and the text acquisition unit obtains the original text, which is the text interaction information. The agent's micro-expression camera simultaneously captures the user's facial dynamics: furrowed brows, downturned corners of the mouth, and slightly red eyes, generating the original micro-expression imaging information.
[0016] Secondly, semantic extraction is performed on the original multimodal interaction information, extracting semantic vectors from each original interaction and integrating them to form an interaction dataset. Semantic extraction refers to the process of semantic parsing and feature transformation of the original interaction information, converting unstructured raw signals into structured semantic features that can be recognized and computed by computers. A semantic vector is a structured numerical vector generated after semantic extraction, used to represent the semantic connotation of the original interaction information. The value of each dimension in the vector corresponds to the weight of the semantic feature; for example, the higher the value of a certain dimension, the more significant the semantic feature. The interaction dataset is a set of semantic vectors generated after semantic extraction of the original multimodal interaction information, and serves as the direct input data for subsequent multimodal feature extraction.
[0017] For example, a speech semantic extraction model is built based on the Wav2Vec2 pre-trained model to analyze the speech content "being criticized by the boss at work, the plan was revised three times, feeling bad". Combining the acoustic features of low tone and slow speech rate, a 512-dimensional speech semantic vector V1 is generated. This vector fully captures the semantic connotation, emotional tendency and demand features in the speech signal. A text semantic extraction model is built based on the BERT pre-trained model to analyze the text content "being criticized by the boss at work, the plan was revised three times and still rejected, feeling particularly bad", generating a 768-dimensional text semantic vector V2. This vector accurately represents the semantic information, emotional state and demand tendency of the text content. A micro-expression semantic extraction model is built based on the ResNet pre-trained model to identify micro-expression images of furrowed brows and downturned corners of the mouth, generating a 256-dimensional micro-expression semantic vector V3. This vector effectively captures the emotional semantics and potential demands corresponding to facial micro-expressions. Integrating V1, V2, and V3, an interactive dataset D1={V1,V2,V3} is obtained, completing the construction of the interactive dataset.
[0018] Finally, the historical interaction semantics preceding the current interaction are extracted as interaction semantic context information, collected and organized into scene context data, and form a scene dataset. Interaction semantic context information refers to the semantic information generated during the user's historical interactions with the agent. It reflects the user's emotional habits and needs, serving as the basis for scene perception. The scene dataset is a collection formed by aggregating interaction semantic context information, used to characterize the scene background of the user's current interaction, such as historical emotional tendencies and historical needs preferences.
[0019] Specifically, from the agent's historical interaction database, historical interaction semantics related to the current interaction semantics are retrieved and extracted, with priority given to historical semantics that match the current interaction topic and sentiment. The extracted historical interaction semantics are standardized and organized to remove redundant and repetitive semantics, forming interaction semantic context information. The interaction semantic context information is integrated according to time order and semantic relevance to form a scene dataset, ensuring that the semantic information in the scene dataset is consistent with the current interaction scene.
[0020] For example, extract historical interaction semantics: In the agent's historical interaction database, the user's previous interaction semantics are stored, such as "After being criticized by my boss, I would be sad for a long time and hope someone could listen to me," which is denoted as interaction semantic context information S1; "I was criticized by my boss at work, the plan was not approved after three revisions, I was in a bad mood, and I wanted to confide in someone and get comfort," which is denoted as S2 after removing redundancy; S1 and S2 are standardized into scene context data and integrated to obtain scene dataset D2={S1,S2,……}, thus completing the construction of the scene dataset.
[0021] In this embodiment of the invention, three emotional interaction modalities—voice, text, and micro-expressions—are fully covered to ensure the integrity of the original interaction information. The semantic standardization and structural transformation of the original interaction information are realized, and the generated interaction dataset has the characteristics of dimensional alignment and temporal consistency, providing high-quality input for subsequent multimodal feature extraction. Historical interaction semantics connected with the current scene are collected to form a complete scene dataset, providing a basis for scene perception for subsequent scene feature analysis and user state recognition.
[0022] S200: Perform multimodal feature extraction on the interactive dataset to obtain demand features and sentiment features, and perform correlation analysis on the sentiment features to obtain comprehensive sentiment features.
[0023] In this embodiment of the invention, multimodal feature extraction is performed on the interactive dataset to obtain demand features and emotional features, and correlation analysis is performed on the emotional features to obtain comprehensive emotional features. In S100, an interactive dataset containing three types of semantic vectors—voice, text, and micro-expressions—has been obtained. However, the original semantic vectors mainly encode the content information expressed by the user and have not explicitly separated the user's inner needs and emotional state. Furthermore, different modalities often reflect the same emotion with different intensities: sometimes voice tone is more realistic than text, and sometimes micro-expressions are more subtle than voice. Simply fusing or averaging these modalities weakens the contribution of strongly correlated modalities and may even be interfered with by noisy modalities. Therefore, S200 aims to extract demand features and emotional features from the semantic vectors using a specialized modal feature extraction model, and then automatically select the modality that best represents the user's true emotions through multimodal correlation analysis, generating comprehensive emotional features to provide reliable and focused emotional input for subsequent user state recognition.
[0024] like Figure 2 As shown, step S200 in the method provided in this embodiment of the invention includes: Construct a modal feature extraction model, input the interaction dataset into the modal feature extraction model, and obtain demand features and sentiment features; The correlation analysis is performed on the aforementioned demand features and the aforementioned emotional features to obtain the voice-text correlation degree, voice-expression correlation degree, and text-expression correlation degree; Based on the speech-text correlation, speech-expression correlation, and text-expression correlation, modality consistency analysis is performed, and the sentiment feature corresponding to the modality with the highest correlation is determined as the comprehensive sentiment feature.
[0025] First, a modal feature extraction model is constructed. The interactive dataset is input into the modal feature extraction model to obtain demand features and sentiment features. The modal feature extraction model is a deep learning model used to parse features from multimodal semantic vectors and output demand features and sentiment features.
[0026] The construction of the modal feature extraction model includes: Obtain a sample interaction dataset, wherein the sample interaction dataset includes sample speech semantic vectors, sample text semantic vectors, and sample facial expression semantic vectors; Obtain the sample demand features and sample sentiment features corresponding to the sample interaction dataset. A modal feature extraction model is constructed, wherein the modal feature extraction model includes three extraction branches, namely a speech feature extraction branch, a text feature extraction branch, and a micro-expression feature extraction branch, and each extraction branch includes two output branches: demand feature and emotion feature; The modality feature extraction model is trained using the sample interaction dataset, the sample demand features, and the sample sentiment features until convergence, thus obtaining the trained modality feature extraction model.
[0027] First, a sample interaction dataset is obtained, which includes sample speech semantic vectors, sample text semantic vectors, and sample facial expression semantic vectors. The sample interaction dataset refers to a standardized dataset used to train the modal feature extraction model. It consists of multimodal sample vectors that have undergone semantic extraction and serves as the input basis for supervised training of the model.
[0028] Specifically, the sample interaction dataset integrates publicly available general datasets and self-built scenario datasets. The publicly available dataset uses the CMU-MOSEI multimodal emotion interaction dataset, while the self-built dataset consists of labeled data collected from historical interactions of emotion interaction agents. The dataset contains three types of standardized 5-dimensional semantic vectors: sample speech semantic vectors, sample text semantic vectors, and sample facial expression semantic vectors, covering all scenarios such as work, life, and emotional expression. For example, 10,000 sets of sample data are selected, including samples of work frustration and emotional expression needs consistent with the current scenario. Each set of samples corresponds to complete speech, text, and micro-expression semantic vectors.
[0029] Secondly, the sample demand features and sample sentiment features corresponding to the sample interaction dataset are obtained. The sample demand features and sample sentiment features together correspond to a preset 7-dimensional standardized semantic vector. The two are integrated to form a 7-dimensional supervised label vector, with values ranging from 0 to 1; a larger value indicates a higher feature strength.
[0030] The specific correspondences and meanings of the seven dimensions are as follows: The sample sentiment feature is a manually labeled 3-dimensional quantized vector, corresponding to dimensions 1-3 of the 7-dimensional standardized semantic vector. Specifically, dimension 1 represents positive sentiment intensity, dimension 2 represents neutral sentiment intensity, and dimension 3 represents negative sentiment intensity, used to accurately represent the user's emotional state. The sample demand feature is a manually labeled 4-dimensional quantized vector, corresponding to dimensions 4-7 of the 7-dimensional standardized semantic vector. Specifically, dimension 4 represents the intensity of the need to confide, dimension 5 represents the intensity of the need for comfort, dimension 6 represents the intensity of the need for advice, and dimension 7 represents the intensity of the need for information retrieval, used to accurately represent the user's diverse needs. The sample demand feature and the sample sentiment feature together constitute a 7-dimensional supervision label vector, serving as supervision labels for training the modal feature extraction model. This guides iterative optimization of model parameters, ensuring a high degree of match between the model's output 7-dimensional standardized semantic vector and the sample's true semantics, achieving a precise mapping from the high-dimensional original semantic vector to the standardized 7-dimensional feature vector.
[0031] Specifically, a professional annotation team was employed to perform double-blind annotation on the sample interaction dataset. The annotation process strictly followed the dimensional definition of the pre-defined 7-dimensional standardized semantic vector. For each set of sample semantic vectors, the corresponding 4-dimensional sample demand features and 3-dimensional sample sentiment features were annotated one by one. After the annotation was completed, consistency verification was performed to remove abnormal samples with excessive annotation deviations. Finally, qualified 7-dimensional annotation results were determined as standard supervision labels for model training.
[0032] For example, for a set of sample semantic vectors corresponding to "work frustration, need for confiding and comfort", the sample demand characteristics are labeled as: confiding demand intensity = 0.95, comfort demand intensity = 0.90, advice demand intensity = 0.80, information query demand intensity = 0.10; the sample emotional characteristics are: positive emotional intensity = 0.05, neutral emotional intensity = 0.05, negative emotional intensity = 0.90.
[0033] Next, a modal feature extraction model is constructed, comprising three extraction branches: a speech feature extraction branch, a text feature extraction branch, and a micro-expression feature extraction branch. Each extraction branch includes two output branches: demand features and emotion features. The modal feature extraction model is a trimodal parallel model built upon a fully connected neural network, where each modality extracts features independently without parameter sharing, ensuring the independence of modal feature extraction.
[0034] Specifically, the modal feature extraction model adopts a trimodal parallel structure. The network structures of the three feature extraction branches—speech, text, and micro-expression—are completely identical, all based on fully connected neural networks: the speech feature extraction branch has 128 neurons in its input layer to receive 128-dimensional speech audio features; the text feature extraction branch has 768 neurons in its input layer to receive 768-dimensional text word vector features; and the micro-expression feature extraction branch has 256 neurons in its input layer to receive 256-dimensional facial image features. The first hidden layer of each branch contains 64 neurons, using ReLU as the activation function to complete the initial feature extraction; the second hidden layer contains 32 neurons, also using the ReLU activation function, to achieve higher-order feature purification and redundant information filtering; the output layer of each branch uniformly has 7 neurons, using the Sigmoid activation function to constrain the output value to the range of 0-1, corresponding to a preset 7-dimensional standardized semantic vector, where dimensions 1-3 are positive emotional intensity, neutral emotional intensity, and negative emotional intensity, respectively, and dimensions 4-7 are the intensity of the need to confide, the intensity of the need for comfort, the intensity of the need for advice, and the intensity of the need for information retrieval, respectively.
[0035] Finally, the modal feature extraction model is trained using the sample interaction dataset, the sample demand features, and the sample sentiment features until convergence, yielding the trained modal feature extraction model. Supervised learning is employed, with the Adam adaptive moment estimation optimizer and a learning rate of 0.001. The mean squared error loss function (MSE) is used to adapt to the feature strength regression task. The sample interaction dataset, sample demand features, and sample sentiment features are input into the model for batch-by-batch iterative training. The convergence condition is that the training loss function decreases by less than 1 × 10^6 times within 10 consecutive epochs. -4 The model is determined to be convergent.
[0036] For example, when the modal feature extraction model was trained to the 86th epoch, the loss decreased steadily by 8 × 10 for 10 consecutive epochs. -5 Once the convergence condition is met, training stops, and the trained modal feature extraction model is obtained.
[0037] Then, the interaction dataset is input into the modality feature extraction model to obtain demand features and sentiment features. Demand features are the quantitative features of user interaction demands output by the modality feature extraction model. Sentiment features are the quantitative features of user emotional states output by the modality feature extraction model. The interaction dataset generated by S100—voice semantic vector V1, text semantic vector V2, and micro-expression semantic vector V3—is input into the corresponding branches of the trained modality feature extraction model; each branch independently infers and calculates, simultaneously outputting the demand features and sentiment features of that modality.
[0038] For example, the modal feature extraction model outputs the following results: the speech feature extraction branch outputs a 7-dimensional vector of [0.10, 0.05, 0.85, 0.90, 0.85, 0.20, 0.10], corresponding to the following sentiment features: positive 0.10, neutral 0.05, negative 0.85; and the following need features: confide 0.90, offer comfort 0.85, provide advice 0.20, and request information 0.10. The text feature extraction branch outputs a 7-dimensional vector of [0.08, 0.04, 0.88, 0.92, 0.86, 0.22, 0.1]. [1], corresponding to the following emotional features: positive 0.08, neutral 0.04, negative 0.88; demand features: confide 0.92, comfort 0.86, suggestion 0.22, information query 0.11; the micro-expression feature extraction branch outputs a 7-dimensional vector of [0.07, 0.03, 0.90, 0.89, 0.84, 0.19, 0.09], corresponding to the following emotional features: positive 0.07, neutral 0.03, negative 0.90; demand features: confide 0.89, comfort 0.84, suggestion 0.19, information query 0.09.
[0039] Secondly, correlation analysis is performed on the aforementioned demand features and emotional features to obtain the voice-text correlation, voice-expression correlation, and text-expression correlation.
[0040] Specifically, correlation analysis is performed on the demand features and the emotional features to obtain voice-text correlation, voice-expression correlation, and text-expression correlation, including: The feature intensity values of speech-corresponding emotion features, text-corresponding emotion features, and micro-expression-corresponding emotion features were extracted respectively. Calculate the similarity of the feature intensity values between the emotional features corresponding to speech and the emotional features corresponding to text to obtain the speech-text correlation. Calculate the similarity of the feature intensity values between the emotional features corresponding to speech and the emotional features corresponding to micro-expressions to obtain the speech-expression correlation degree; Calculate the similarity between the feature intensity values of the sentiment features corresponding to the text and the sentiment features corresponding to the micro-expressions to obtain the text-expression correlation.
[0041] First, feature intensity values are extracted for the emotional features corresponding to speech, text, and micro-expressions. Feature intensity values refer to the specific numerical values corresponding to positive, neutral, and negative emotions within the emotional features, reflecting the strength of the corresponding emotions. From the 7-dimensional standardized semantic vectors output by speech, text, and micro-expressions respectively, the first three dimensions corresponding to the emotional features are extracted, which are the feature intensity values for each modality's emotional features.
[0042] For example, the intensity values of voice emotion features are: positive 0.10, neutral 0.05, and negative 0.85; the intensity values of text emotion features are: positive 0.08, neutral 0.04, and negative 0.88; and the intensity values of micro-expression emotion features are: positive 0.07, neutral 0.03, and negative 0.90.
[0043] Secondly, the similarity between the feature intensity values of the corresponding emotional features in speech and the corresponding emotional features in text is calculated to obtain the speech-text correlation. The speech-text correlation measures the consistency between the speech modality and the text modality in emotional expression, and is calculated from the similarity of their emotional feature intensity values; a higher value indicates higher consistency. Cosine similarity is used to calculate the similarity between the speech emotional feature intensity values and the text emotional feature intensity values; the result is the speech-text correlation. For example, cosine similarity calculation for speech emotional feature intensity values [0.10, 0.05, 0.85] and text emotional feature intensity values [0.08, 0.04, 0.88] yields a speech-text correlation of 0.98.
[0044] Next, the similarity between the feature intensity values of the emotional features corresponding to speech and the emotional features corresponding to micro-expressions is calculated to obtain the speech-expression correlation degree. The speech-expression correlation degree measures the consistency between the speech modality and the micro-expression modality in emotional expression. Cosine similarity is used to calculate the similarity between the speech emotional feature intensity values and the micro-expression emotional feature intensity values; the result is the speech-expression correlation degree. For example, cosine similarity calculation is performed on the speech emotional feature intensity values [0.10, 0.05, 0.85] and the micro-expression emotional feature intensity values [0.07, 0.03, 0.90], resulting in a speech-expression correlation degree of 0.97.
[0045] Then, the similarity between the feature intensity values of the corresponding sentiment features in the text and the corresponding sentiment features in micro-expressions is calculated to obtain the text-expression correlation. The text-expression correlation measures the degree of consistency between the text modality and the micro-expression modality in emotional expression. Cosine similarity is used to calculate the similarity between the text sentiment feature intensity values and the micro-expression sentiment feature intensity values; the result is the text-expression correlation. For example, cosine similarity calculation is performed on the text sentiment feature intensity values [0.08, 0.04, 0.88] and the micro-expression sentiment feature intensity values [0.07, 0.03, 0.90], resulting in a text-expression correlation of 0.99.
[0046] Finally, based on the voice-text correlation, voice-expression correlation, and text-expression correlation, modal consistency analysis is performed, and the emotional feature corresponding to the modality with the highest correlation is determined as the comprehensive emotional feature. Modal consistency refers to the degree of agreement between emotional expressions across different modalities, directly reflected by the correlation level. The comprehensive emotional feature is the single-modal emotional feature selected from multimodal emotional features that best represents the user's true emotions, serving as the final emotional basis for subsequent state recognition.
[0047] Specifically, the average correlation score for each modality is calculated based on the pairwise correlation between modalities: the average of the speech-text correlation score and the speech-expression correlation score for the speech modality is used as the overall consistency score for the speech modality; the average of the speech-text correlation score and the text-expression correlation score for the text modality is used as the overall consistency score for the text modality; and the average of the speech-expression correlation score and the text-expression correlation score for the microexpression modality is used as the overall consistency score for the microexpression modality. The overall consistency scores of the three modalities are then compared, and the emotional feature corresponding to the modality with the highest overall consistency score is determined as the overall emotional feature.
[0048] For example, the speech-text correlation, speech-expression correlation, and text-expression correlation are 0.98, 0.97, and 0.99, respectively. Therefore: the overall consistency score for the speech modality is (0.98 + 0.97) / 2 = 0.975; the overall consistency score for the text modality is (0.98 + 0.99) / 2 = 0.985; and the overall consistency score for the micro-expression modality is (0.97 + 0.99) / 2 = 0.98. The text modality has the highest overall consistency score; therefore, text sentiment features are selected as the overall sentiment features.
[0049] In this embodiment of the invention, a three-modal parallel dual-output feature extraction model is constructed to achieve the decoupled and independent extraction of demand features and emotional features, avoiding mutual interference between the two types of features. Modal correlation is obtained by calculating the pairwise similarity of the intensity of each modal emotional feature, thus completing the quantitative analysis of multimodal emotional consistency. The optimal modal emotion is selected as the comprehensive emotional feature based on the highest correlation, which can effectively filter out single-modal noise and abnormal bias, improve the accuracy and stability of the comprehensive emotional feature, and provide a reliable emotional feature basis for subsequent user status recognition.
[0050] S300: Extract and analyze features from the scene context data to obtain scene features.
[0051] In this embodiment of the invention, feature extraction and analysis are performed on the scene context data to obtain scene features. Scene context data contains a large amount of historical interaction semantics, including redundant information unrelated to the current interaction. Direct use of this information would cause feature interference. Furthermore, existing technologies lack a quantitative calculation method for the correlation between historical semantics and the current interaction, making it impossible to distinguish the importance of historical scene information. This results in extracted scene features that are redundant, messy, and lack specificity, making it difficult to accurately support subsequent user state recognition. Therefore, it is necessary to use an attention mechanism to filter effective context, quantify the semantic fit, and sort them by importance to obtain concise, accurate scene requirement features and scene sentiment features that are highly relevant to the current interaction.
[0052] Step S300 in the method provided in this embodiment of the invention includes: The scene dataset is input into the context attention extraction model to filter out effective context semantic information in the scene dataset that has a sequential relationship with the current interaction semantics. Calculate the semantic fit parameter between the effective context semantic information and the current interaction content, wherein the semantic fit parameter is obtained by calculating the overlap ratio of the semantic features before and after the interaction. Based on the semantic fit parameters, the importance of the contextual semantic information is ranked, and the top N contextual semantic information are determined as scene features, wherein the scene features include scene demand features and scene emotional features.
[0053] First, the scene dataset is input into the context attention extraction model to filter out valid contextual semantic information in the scene dataset that has a sequential relationship with the current interaction semantics. The context attention extraction model is used to filter out valid information related to the current interaction semantics from the scene context data and remove irrelevant and redundant data.
[0054] Specifically, the scene dataset constructed in S100 is input into the context attention extraction model. The context attention extraction model calculates the attention weight of each piece of context data with the current user interaction semantics through the attention mechanism, and retains the context data with attention weights higher than a preset threshold as valid context semantic information.
[0055] For example, the scenario dataset D2={S1,S2,……} constructed in S100 is input into the context attention extraction model. The user's current interaction semantics are: being criticized, the solution not being approved, feeling bad, needing to confide and be comforted. The context attention extraction model filters the results: S1, S2, S3, and S4 are all strongly related to the current interaction and are effective context semantic information; the rest are irrelevant to daily life and are removed. Finally, the effective context semantic set D2'={S1,S2,S3,S4} is obtained.
[0056] Secondly, the semantic fit parameter between the effective contextual semantic information and the current interaction content is calculated. This semantic fit parameter is obtained by calculating the overlap ratio of preceding and following semantic features. The semantic fit parameter quantifies the degree of matching between the historical effective contextual semantics and the current interaction semantics, with a value ranging from 0 to 1. A higher value indicates a stronger connection between the historical semantics and the current interaction. The preceding semantic features refer to the core emotional and demand features contained in the effective contextual semantic information; the following semantic features refer to the emotional and demand features contained in the current interaction content. The overlap ratio calculation involves counting the number of overlapping features between the preceding and following semantic features, dividing this number by the total number of current interaction semantic features, and obtaining the ratio, which is the semantic fit parameter. The calculation formula is: Semantic Fit Parameter = Number of Overlapping Features / Total Number of Current Interaction Semantic Features.
[0057] Specifically, the semantic features of the current interactive content are extracted to determine the total number of features; the semantic features of each valid context are extracted one by one; the number of overlaps between each valid context and the core semantic features of the current interaction is counted; and the semantic fit parameter corresponding to each valid context is calculated using the overlap ratio calculation formula.
[0058] For example, the semantic features of the current interaction content are extracted, and the total number of features is determined to be 3, namely: negative emotion features, need for expression features, and need for comfort features. Semantic features of the effective context semantic set D2'={S1,S2,S3,S4} are extracted one by one: S1 features are negative emotion features and need for expression features, a total of 2; S2 features are negative emotion features, a total of 1; S3 features are negative emotion features, a total of 1; S4 features are negative emotion features and need for comfort features, a total of 2. The overlap ratio is calculated as follows: semantic fit parameter of S1 = 2 / 3 ≈ 0.67; semantic fit parameter of S2 = 1 / 3 ≈ 0.33; semantic fit parameter of S3 = 1 / 3 ≈ 0.33; semantic fit parameter of S4 = 2 / 3 ≈ 0.67.
[0059] Finally, based on the semantic fit parameters, the contextual semantic information is ranked by importance, and the top N contextual semantic information are determined as scene features. These scene features include scene demand features and scene sentiment features. Scene features are a set of features extracted from highly important historical contexts, containing scene demand features and scene sentiment features, representing user needs and emotional tendencies in historical interactions, respectively. Importance ranking refers to sorting the effective contextual semantic information according to the semantic fit parameters from largest to smallest. A higher semantic fit parameter indicates a stronger correlation between the context and the current interaction, and thus a higher scene importance. If the parameters are the same, they are ranked according to the temporal correlation between the context and the current interaction; that is, the closer the time, the higher the priority. N is a preset number to retain; in this embodiment, N=2, but it can be adjusted according to the actual scenario.
[0060] Specifically, the effective contextual semantic information is sorted from high to low according to the semantic fit parameter, and if the parameters are the same, they are sorted by time. A preset value of N is used to filter the top N effective contexts. Emotional features and demand features are extracted from the N contexts, and after merging and deduplication, they are determined as the final scene features.
[0061] For example, sorted in descending order by semantic fit parameters: S1 (0.67) > S4 (0.67) > S2 (0.33) = S3 (0.33); the first N=2 valid contexts are selected: S1 and S4; features are extracted and merged for deduplication: from S1: negative emotions, need to confide; from S4: negative emotions, need for comfort. After merging and deduplication, the final scene features are determined: scene emotional features: high intensity of negative emotions, weak neutral and positive emotions; scene need features: strong need to confide, strong need for comfort, weak need for suggestions and information query.
[0062] In this embodiment of the invention, a contextual attention mechanism is used to eliminate redundant semantics in the scene, ensuring the purity of effective information; semantic alignment parameters are calculated by quantifying the proportion of semantic overlap, enabling an objective distinction of the importance of historical scene information; based on the semantic alignment parameters, the top N key features are selected to obtain scene requirements and scene emotional features that are highly related to user semantic features, providing accurate and efficient scene basis for subsequent fusion of emotional features and scene features, and improving the scene adaptation accuracy of user state recognition.
[0063] S400: Integrate the comprehensive emotional features and the scene features to perform user state recognition and obtain user state data.
[0064] In this embodiment of the invention, the comprehensive sentiment features and the scene features are fused to perform user state recognition and obtain user state data. The comprehensive sentiment features only represent the user's current real-time sentiment, while the scene sentiment features only represent the user's historical scene sentiment. Simply overlaying or using either alone cannot reflect the difference between the credibility of the current modality and the importance of historical scenes, easily leading to distortion in user state recognition. Furthermore, different sentiment dimensions need to be weighted independently to ensure a reasonable sentiment distribution. Therefore, it is necessary to dynamically allocate fusion weights based on modality correlation and semantic fit parameters, and to calculate the weights separately for the same sentiment dimension, ultimately obtaining a fused sentiment feature that accurately reflects the user's overall state.
[0065] Step S400 in the method provided in this embodiment of the invention includes: Based on the speech-text correlation, the speech-expression correlation, the text-expression correlation, and the semantic fit parameters, the modality fusion weight and the scene fusion weight are obtained; Based on the modal fusion weights and the scene fusion weights, the feature intensity values of the comprehensive emotional features and the scene emotional features are weighted and calculated to perform user state recognition and obtain user state data, wherein the user state data includes fused emotional features.
[0066] First, based on the speech-text correlation, speech-expression correlation, text-expression correlation, and semantic fit parameters, modality fusion weights and scene fusion weights are obtained. The modality fusion weights characterize the reliability of the current comprehensive emotional features and are obtained by averaging the correlations of the speech-text, speech-expression, and text-expression sets, reflecting the level of consistency in multimodal emotions. The scene fusion weights characterize the importance of historical scene emotional features and are determined by the semantic fit parameters, reflecting the close connection between the scene and the current interaction.
[0067] Specifically, the arithmetic mean of the three sets of modal relevance is taken as the initial value of the original modal weight; the average value of the effective context semantic fitting parameters is calculated as the initial value of the scene fusion weight; the two initial values are normalized to obtain the final weight: modal fusion weight = modal average relevance / (modal average relevance + semantic fitting parameters), scene fusion weight = semantic fitting parameters / (modal average relevance + semantic fitting parameters).
[0068] For example, the voice-text correlation is 0.98, the voice-expression correlation is 0.97, and the text-expression correlation is 0.99; the average modal correlation is (0.98+0.97+0.99) / 3=0.98; the average semantic fit parameter is (0.67+0.67) / 2=0.67; normalized calculation: modal fusion weight = 0.98 / (0.98+0.67)≈0.59, scene fusion weight = 0.67 / (0.98+0.67)≈0.41.
[0069] Secondly, based on the modality fusion weights and the scene fusion weights, the feature intensity values of the comprehensive sentiment feature and the scene sentiment feature are weighted and calculated to perform user state recognition and obtain user state data. The user state data includes fused sentiment features. Fusion sentiment features refer to the final sentiment feature obtained by independently weighting and summing the comprehensive sentiment feature and the scene sentiment feature along the same dimension.
[0070] Specifically, the dimensions of the comprehensive emotional features and the scene emotional features are kept completely consistent, both including three dimensions: positive emotion, neutral emotion, and negative emotion. Each dimension is weighted independently: the value of a dimension after fusion = the value of the corresponding dimension of comprehensive emotion × modal fusion weight + the value of the corresponding dimension of scene emotion × scene fusion weight. The fused features of each dimension are combined to form the fused emotional features, which serve as user status data.
[0071] For example, given the comprehensive sentiment features: positive = 0.08, neutral = 0.04, negative = 0.88; and the scene sentiment features: positive = 0.05, neutral = 0.05, negative = 0.90; the weighted calculation of the fused sentiment features is as follows: fused positive sentiment = 0.08 × 0.59 + 0.05 × 0.41 ≈ 0.068; fused neutral sentiment = 0.04 × 0.59 + 0.05 × 0.41 ≈ 0.044; fused negative sentiment = 0.88 × 0.59 + 0.90 × 0.41 ≈ 0.888; the final fused sentiment features are [0.068, 0.044, 0.888], which constitute the user state data in this embodiment.
[0072] In this embodiment of the invention, the fusion weights are dynamically allocated based on modal correlation and semantic fit parameters, which can objectively reflect the reliability of the current emotion and the importance of the historical scene. By independently weighting and fusing the same emotion dimension, dimension confusion and emotion distortion are avoided, and finally, accurate, stable, and realistic fused emotion features are obtained, which effectively improves the accuracy of user state recognition and the rationality of the scene, and provides a reliable basis for generating response schemes that fit the user state.
[0073] S500: Based on the user status data and demand data, map and obtain the agent response scheme, and perform agent interaction.
[0074] In this embodiment of the invention, a response scheme for the intelligent agent is obtained by mapping user state data and demand data, and then the intelligent agent interacts. User state data only represents the fused emotional state, while the real-time demand characteristics of users and the demand characteristics of the scenario are relatively scattered, with interference from secondary demands. If a response is generated directly based on all demands, it will cause the response to deviate from the user's core needs. Therefore, it is necessary to first screen high-frequency demands, establish a mapping relationship with emotional states, and finally generate a response scheme that fits the user's emotions and core needs, ensuring accurate and appropriate interaction.
[0075] Step S500 in the method provided in this embodiment of the invention includes: Based on the scenario requirements and those requirements, obtain the requirements data; Construct a semantic mapping relationship between demand characteristics and user status data, the semantic mapping relationship being obtained based on historical demand data and historical associated responses; Synchronously match the current demand characteristics with user status data into the semantic mapping relationship, and retrieve the relevant response basis; Based on the associated response, generate an agent response plan and conduct agent interaction.
[0076] First, based on the scenario requirements and the aforementioned requirements, requirement data is obtained.
[0077] Specifically, based on the scenario requirements and those requirements, requirement data is obtained, including: Calculate the frequency of occurrence of demand elements in the aforementioned demand characteristics and the aforementioned scenario demand characteristics; The demand elements whose frequency of occurrence is greater than or equal to the frequency threshold are integrated into the demand data.
[0078] The frequency threshold is obtained based on semantic fit parameters.
[0079] First, calculate the frequency of occurrence of demand elements in the stated demand features and the stated scenario demand features. Demand elements are the smallest units constituting user needs, including four categories: need for expression, need for comfort, need for advice, and need for information retrieval. Demand features refer to the real-time demand features output by the multimodal system in S200. Scenario demand features refer to the scenario demand features extracted from the historical context in S300. Frequency of occurrence refers to the proportion of times the same demand element appears in both real-time and scenario demands relative to the total number of demand elements.
[0080] Specifically, summarize all elements of real-time and scenario-based needs, and calculate the frequency of occurrence = the number of times a specific need element appears / the total number of times all need elements appear. For example, S200 real-time need elements: venting, comforting, venting (3 in total); S300 scenario-based need element: venting (1 in total); total number of all need elements = 3 + 1 = 4. The frequency of occurrence for each need element is: venting need frequency = 3 / 4 = 0.75, comfort need frequency = 1 / 4 = 0.25, suggestion need frequency = 0, information query need frequency = 0.
[0081] Secondly, demand elements with an occurrence frequency greater than or equal to the occurrence frequency threshold are integrated into the demand data. The occurrence frequency threshold is obtained based on semantic fit parameters. The occurrence frequency threshold is a critical value used to distinguish between primary and secondary demands, directly determined by the average semantic fit parameters in S300. The demand data is a set of high-frequency demand elements, representing the user's true needs. The average semantic fit parameters of the effective context in S300 are taken as the occurrence frequency threshold; demand elements with an occurrence frequency greater than or equal to this threshold are selected and integrated into the demand data.
[0082] For example, the average semantic fit parameter of S300 is (0.67 + 0.67) / 2 = 0.67, which means the frequency threshold is 0.67. The frequency of the need to confide is 0.75 ≥ 0.67, which is a primary need, while the frequency of the need for comfort is 0.25 < 0.67, which is a secondary need and is therefore eliminated. The result is: Need data = {Confiding needs}.
[0083] Furthermore, a semantic mapping relationship between demand characteristics and user status data is constructed. This semantic mapping relationship is obtained based on historical demand data and historical associated response criteria. The semantic mapping relationship refers to the pre-established correspondence rules between user status, demand data, and response criteria. Historical demand data refers to high-frequency demands selected from historical interactions. Historical associated response criteria refer to the reasonable response rules corresponding to the corresponding status and demand in history.
[0084] Specifically, based on the three-dimensional fused emotional feature vector generated in historical interactions, historical demand data filtered by frequency, and verified effective historical interaction response content, the continuous fused emotional feature vector is first divided into discrete emotional state intervals such as high negative, medium negative, and low negative according to the distribution of positive, neutral, and negative emotional intensity. Then, the emotional state interval + historical demand data is used as the mapping key, and the corresponding historical related response is used as the mapping value to form a fixed and searchable one-to-one correspondence mapping rule.
[0085] For example, multiple sets of historical interaction data are extracted, and the historical fusion emotional features [0.05, 0.06, 0.90], [0.07, 0.05, 0.87] and other vectors are uniformly classified as high negative emotional states. The historical demand data {the need to confide, the need for comfort} are taken as the associated demand, and the effective comforting and listening response rules for such user states and needs in history are taken as the corresponding outputs. These are combined one by one to form multiple mapping entries, and finally the construction of a complete semantic mapping relationship is completed, providing a standard basis for the matching and retrieval of current user data.
[0086] Then, the current demand features and user status data are synchronously matched into a semantic mapping relationship to retrieve the relevant response criteria. The relevant response criteria refer to response rules that perfectly match the current user status and demand, used to generate a reply. The emotional features and demand data are integrated and input into the mapping relationship to obtain the corresponding response criteria. For example, user status data: integrated emotional features [0.068, 0.044, 0.888], demand data: {need to confide}, matching yields: Relevant response criteria: Empathize with the user's negative emotions, guide the user to confide, and listen patiently throughout.
[0087] Finally, an agent response plan is generated based on the associated response criteria to facilitate agent interaction. The agent response plan refers to a natural language response generated based on the response criteria. Agent interaction refers to using the response plan to converse with the user. The response criteria are transformed into natural, gentle, and emotionally appropriate dialogue statements. For example, a generated response plan might be: "I completely understand that you are feeling very upset right now. Please feel free to tell me what's bothering you, and I will listen attentively." In this embodiment of the invention, secondary needs are eliminated by filtering based on demand frequency, while retaining the user's primary needs; thresholds are adaptively set in conjunction with semantic fit parameters to ensure the rationality of demand filtering; and response rules are precisely matched through state-demand mapping relationships, so that the agent's output simultaneously fits the user's true emotional state and primary needs, making the agent's interaction more targeted, empathetic, and adaptable.
[0088] Through the specific implementation methods described above, the embodiments of the present invention achieve the following technical effects: This invention provides an emotional interaction intelligent agent interaction method based on a scene and modality dual attention mechanism. By synchronously collecting real-time user interaction and historical context data in a multimodal manner, it completes multimodal emotional feature extraction and consistency association analysis to determine reliable comprehensive emotional features. Based on context relevance, it filters and extracts accurate scene features, adaptively weights and fuses multimodal emotions and scene emotions to obtain real user state data. Through demand frequency filtering and semantic mapping, it generates response schemes that fit the user's emotions and core needs. The entire process effectively eliminates single-modal bias and historical context redundancy interference, achieving accurate identification of user emotional state and demand intentions, improving the empathy, pertinence, and scene adaptability of intelligent agent interaction, and ensuring that the interaction output fully matches the user's real emotions and core needs.
[0089] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An interaction method for emotion-based interactive intelligent agents based on a dual attention mechanism of scene and modality, characterized in that, include: Acquire interaction data and scene context data during the interaction between the emotional interaction intelligent agent and the user, and obtain the interaction dataset and scene dataset; Multimodal feature extraction is performed on the interactive dataset to obtain demand features and sentiment features, and correlation analysis is performed on the sentiment features to obtain comprehensive sentiment features; The scene context data is subjected to feature extraction and analysis to obtain scene features; By integrating the comprehensive emotional features and the scene features, user status recognition is performed to obtain user status data; Based on the user status data and demand data, a response plan for the intelligent agent is obtained and the intelligent agent interacts.
2. The emotional interaction agent interaction method based on scene and modality dual attention mechanism according to claim 1, characterized in that, Acquire interaction data and scene context data during the interaction between the emotional interaction agent and the user, and acquire interaction datasets and scene datasets, including: Collect multimodal raw interaction information during the interaction process between the user and the emotional interaction intelligent agent, wherein the multimodal raw interaction information includes voice interaction information, text interaction information and micro-expression imaging information; Semantic extraction is performed on the original multimodal interaction information, and the original interaction information is extracted into semantic vectors and integrated to form an interaction dataset; Extract the semantics of previous interactions before this interaction as the semantic context information of the interaction, collect and organize them into scene context data and form a scene dataset.
3. The emotional interaction agent interaction method based on scene and modality dual attention mechanism according to claim 1, characterized in that, Multimodal feature extraction is performed on the interactive dataset to obtain demand features and sentiment features. Correlation analysis is then performed on the sentiment features to obtain comprehensive sentiment features, including: Construct a modal feature extraction model, input the interaction dataset into the modal feature extraction model, and obtain demand features and sentiment features; The correlation analysis is performed on the aforementioned demand features and the aforementioned emotional features to obtain the voice-text correlation degree, voice-expression correlation degree, and text-expression correlation degree; Based on the speech-text correlation, speech-expression correlation, and text-expression correlation, modality consistency analysis is performed, and the sentiment feature corresponding to the modality with the highest correlation is determined as the comprehensive sentiment feature.
4. The emotional interaction agent interaction method based on scene and modality dual attention mechanism according to claim 3, characterized in that, Constructing a modal feature extraction model includes: Obtain a sample interaction dataset, wherein the sample interaction dataset includes sample speech semantic vectors, sample text semantic vectors, and sample facial expression semantic vectors; Obtain the sample demand features and sample sentiment features corresponding to the sample interaction dataset. A modal feature extraction model is constructed, wherein the modal feature extraction model includes three extraction branches, namely a speech feature extraction branch, a text feature extraction branch, and a micro-expression feature extraction branch, and each extraction branch includes two output branches: demand feature and emotion feature; The modality feature extraction model is trained using the sample interaction dataset, the sample demand features, and the sample sentiment features until convergence, thus obtaining the trained modality feature extraction model.
5. The emotional interaction agent interaction method based on scene and modality dual attention mechanism according to claim 3, characterized in that, The correlation analysis is performed on the aforementioned demand features and the aforementioned emotional features to obtain voice-text correlation, voice-expression correlation, and text-expression correlation, including: The feature intensity values of speech-corresponding emotion features, text-corresponding emotion features, and micro-expression-corresponding emotion features were extracted respectively. Calculate the similarity of the feature intensity values between the emotional features corresponding to speech and the emotional features corresponding to text to obtain the speech-text correlation. Calculate the similarity of the feature intensity values between the emotional features corresponding to speech and the emotional features corresponding to micro-expressions to obtain the speech-expression correlation degree; Calculate the similarity between the feature intensity values of the sentiment features corresponding to the text and the sentiment features corresponding to the micro-expressions to obtain the text-expression correlation.
6. The emotional interaction agent interaction method based on scene and modality dual attention mechanism according to claim 1, characterized in that, The scene context data is subjected to feature extraction and analysis to obtain scene features, including: The scene dataset is input into the context attention extraction model to filter out effective context semantic information in the scene dataset that has a sequential relationship with the current interaction semantics. Calculate the semantic fit parameter between the effective context semantic information and the current interaction content, wherein the semantic fit parameter is obtained by calculating the overlap ratio of the semantic features before and after the interaction. Based on the semantic fit parameters, the importance of the contextual semantic information is ranked, and the top N contextual semantic information are determined as scene features, wherein the scene features include scene demand features and scene emotional features.
7. The emotional interaction agent interaction method based on scene and modality dual attention mechanism according to claim 5, characterized in that, By integrating the comprehensive emotional features and the scene features, user state recognition is performed to obtain user state data, including: Based on the speech-text correlation, the speech-expression correlation, the text-expression correlation, and the semantic fit parameters, the modality fusion weight and the scene fusion weight are obtained; Based on the modal fusion weights and the scene fusion weights, the feature intensity values of the comprehensive emotional features and the scene emotional features are weighted and calculated to perform user state recognition and obtain user state data, wherein the user state data includes fused emotional features.
8. The emotional interaction agent interaction method based on scene and modality dual attention mechanism according to claim 1, characterized in that, Based on the user state data and demand data, a response scheme for the intelligent agent is obtained through mapping, and intelligent agent interaction is performed, including: Based on the scenario requirements and those requirements, obtain the requirements data; Construct a semantic mapping relationship between demand characteristics and user status data, the semantic mapping relationship being obtained based on historical demand data and historical associated responses; Synchronously match the current demand characteristics with user status data into the semantic mapping relationship, and retrieve the relevant response basis; Based on the associated response, generate an agent response plan and conduct agent interaction.
9. The emotional interaction agent interaction method based on scene and modality dual attention mechanism according to claim 8, characterized in that, Based on the scenario requirements and those requirements, requirement data is obtained, including: Calculate the frequency of occurrence of demand elements in the aforementioned demand characteristics and the aforementioned scenario demand characteristics; The demand elements whose frequency of occurrence is greater than or equal to the frequency threshold are integrated into the demand data.
10. The interaction method for emotion-based interactive intelligent agents based on scene and modality dual attention mechanism according to claim 9, characterized in that, The frequency threshold is obtained based on semantic fit parameters.