Emotional interaction method and device based on multimodal data fusion
By integrating a sensor array to collect and process multimodal data, and combining a cross-modal causal reasoning engine and an emotion behavior causal mapping library, the problem of emotion recognition misjudgment of embodied smart devices in complex scenarios is solved, achieving more accurate user emotion understanding and personalized services.
Patent Information
- Application Number
- CN202511029629.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-25
AI Technical Summary
The accuracy of single-modal emotion recognition of existing embodied intelligent devices decreases in complex scenarios, and the lack of causal relationship analysis between multimodal signals leads to a high misjudgment rate. It is impossible to accurately understand the real driving factors of user emotions, thereby reducing decision-making accuracy.
By integrating the sensor array to collect multimodal data in real time, preprocessing and time alignment are performed, and the data is input into the cross-modal causal reasoning engine for emotional state label analysis and event attribution analysis, and the optimal interaction strategy is generated by combining the preset emotional behavior causal mapping library.
It improves the accuracy of emotion recognition, reduces the misjudgment rate, and enhances the decision-making accuracy of embodied smart devices, enabling more accurate understanding of user emotional needs and providing personalized services.
Smart Images

Figure CN120524447B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent decision-making technology for embodied intelligent devices, and in particular to an emotional interaction method and device based on multimodal data fusion. Background Art
[0002] In today's society, embodied intelligent devices (such as humanoid robots and service robots) are widely used in various fields, including home, healthcare, education, and commerce. By interacting naturally with humans, these devices can provide more personalized and attentive services, thereby improving user experience and work efficiency. For example, in the home environment, humanoid robots provide daily care and emotional support to the elderly.
[0003] In related technologies, existing emotion interaction systems mainly rely on single-modal data for emotion recognition and interaction strategy generation; some systems attempt to combine voice and facial expression modalities for emotion recognition through association analysis to improve recognition accuracy.
[0004] However, on the one hand, the accuracy of emotion recognition of a single modality decreases significantly in complex scenarios (such as noise interference); on the other hand, emotion classification is achieved through correlation analysis without considering the causal relationship between multimodal signals. This analysis method lacking causality leads to a high misjudgment rate in emotion recognition and is unable to accurately understand the true driving factors of user emotions, thereby reducing the decision-making accuracy of embodied smart devices. Summary of the Invention
[0005] The embodiments of this application provide a method and apparatus for emotional interaction based on multimodal data fusion. To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is provided below. This summary is not intended to be a comprehensive review, identify key / important elements, or delineate the scope of protection for these embodiments. Its sole purpose is to present some concepts in a simplified form, serving as a prelude to the detailed description that follows.
[0006] In a first aspect, an embodiment of the present application provides an emotional interaction method based on multimodal data fusion, which is applied to an embodied intelligent device. The method includes:
[0007] By integrating the sensor array in the embodied smart device into real-time acquisition and preprocessing of multimodal data, time-aligned multimodal features are obtained.
[0008] Input the multimodal features into a preset cross-modal causal reasoning engine, output the user's current emotional state label and the current causal factors that lead to the current emotional state label, and the preset cross-modal causal reasoning engine is used to perform current emotional state label analysis and event attribution analysis;
[0009] Based on the current emotional state label and current causal factors, the preset emotional behavior causal mapping library is queried to obtain the current optimal emotional interaction strategy instruction set. The preset emotional behavior causal mapping library is a triple consisting of historical emotional state labels, historical causal factors, and historical interaction strategies that have been verified to be effective.
[0010] In a second aspect, an embodiment of the present application provides an emotional interaction device based on multimodal data fusion, the device comprising:
[0011] The modal data preprocessing module is used to collect and preprocess multimodal data in real time through the sensor array integrated in the embodied smart device to obtain time-aligned multimodal features;
[0012] An engine reasoning module, configured to input multimodal features into a preset cross-modal causal reasoning engine, and output the user's current emotional state label and the current causal factors that led to the current emotional state label. The preset cross-modal causal reasoning engine is used to perform current emotional state label analysis and event attribution analysis;
[0013] The instruction set determination module is used to query the preset emotional behavior causal mapping library based on the current emotional state label and current causal factors to obtain the current optimal emotional interaction strategy instruction set. The preset emotional behavior causal mapping library is a triple consisting of historical emotional state labels, historical causal factors, and historical interaction strategies that have been verified to be effective.
[0014] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0015] In the embodiment of the present application, on the one hand, the embodied intelligent device collects and pre-processes multimodal data in real time through an integrated sensor array to obtain time-aligned multimodal features. These features contain rich information, providing a comprehensive data basis for emotion recognition, reducing misjudgments caused by the limitations of single-modal data, and thus improving the accuracy of emotion recognition. On the other hand, the multimodal features are input into a preset cross-modal causal reasoning engine, which can perform current emotional state label analysis and event attribution analysis, output the user's current emotional state label and the causal factors that lead to the emotional state. Through cross-modal causal reasoning, the causal relationship between multimodal signals is fully considered, so that the user's emotional state can be more accurately identified. At the same time, combined with the preset emotional behavior causal mapping library, the misjudgment rate is effectively reduced, and the decision-making accuracy of the embodied intelligent device is improved.
[0016] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0018] Figure 1 This is a flowchart of a method for emotional interaction based on multimodal data fusion provided in an embodiment of the present application;
[0019] Figure 2 This is a schematic diagram of sensors deployed on an embodied smart device provided in an embodiment of the present application;
[0020] Figure 3 This is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0021] Figure 4 Schematic diagram of a preset cross-modal causal reasoning engine provided in an embodiment of the present application;
[0022] Figure 5 This is a flowchart of a method for training an intent recognition model provided in an embodiment of the present application;
[0023] Figure 6 This is a schematic structural diagram of an emotional interaction device based on multimodal data fusion provided in an embodiment of the present application;
[0024] Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The following description and the drawings sufficiently illustrate specific embodiments of the application to enable those skilled in the art to practice them.
[0026] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0027] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0028] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances. In addition, in the description of this application, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0029] At present, existing emotion interaction systems mainly rely on single-modal data for emotion recognition and interaction strategy generation; some systems also attempt to combine voice and facial expression modalities for emotion recognition through association analysis to improve recognition accuracy.
[0030] The applicant of this application realized that, on the one hand, the accuracy of emotion recognition of a single modality decreases significantly in complex scenarios (such as noise interference); on the other hand, emotion classification is achieved through correlation analysis without considering the causal relationship between multimodal signals. This analysis method lacking causal relationship leads to a high misjudgment rate of emotion recognition and is unable to accurately understand the real driving factors of user emotions, thereby reducing the decision-making accuracy of embodied smart devices.
[0031] In order to solve the above problems, the present application provides an emotional interaction method and device based on multimodal data fusion to solve the problems existing in the above-mentioned related technical problems. In an embodiment of the present application, on the one hand, the embodied intelligent device collects and preprocesses multimodal data in real time through an integrated sensor array to obtain time-aligned multimodal features. These features contain rich information, provide a comprehensive data basis for emotion recognition, reduce misjudgments caused by the limitations of single modal data, and thus improve the accuracy of emotion recognition. On the other hand, the multimodal features are input into a preset cross-modal causal reasoning engine, which can perform current emotional state label analysis and event attribution analysis, output the user's current emotional state label and the causal factors that lead to the emotional state. Through cross-modal causal reasoning, the causal relationship between multimodal signals is fully considered, so that the user's emotional state can be more accurately identified. At the same time, combined with the preset emotional behavior causal mapping library, the misjudgment rate is effectively reduced, and the decision-making accuracy of the embodied intelligent device is improved. The following uses an exemplary embodiment for detailed description.
[0032] The following will be combined with the Figure 1 -Attached Figure 5This article details the emotional interaction method based on multimodal data fusion provided in an embodiment of the present application. This method can be implemented using a computer program and run on an emotional interaction device based on multimodal data fusion, which is based on the von Neumann architecture. This computer program can be integrated into an application or run as a standalone tool application.
[0033] See Figure 1 , provides a flow chart of an emotional interaction method based on multimodal data fusion for an embodiment of the present application, which is applied to embodied intelligent devices. Figure 1 As shown, the method of the embodiment of the present application includes the following steps:
[0034] S101, the embodied smart device collects and preprocesses multimodal data in real time through a sensor array integrated in the embodied smart device to obtain time-aligned multimodal features;
[0035] Among them, embodied intelligent devices are intelligent devices with physical forms and perceptual interaction capabilities, such as humanoid robots and service robots. These devices can interact with the environment through their own sensors and actuators to achieve various tasks and services. Sensor arrays are a group of sensors integrated into embodied intelligent devices to collect different types of environmental and user data. For example Figure 2 As shown in Figure 1, these sensors include cameras, microphones, and other types of sensors. Multimodal data is sensor data from different sensors or different types, such as visual data (images, videos), auditory data (speech, sounds), tactile data (pressure, temperature), and motion data (acceleration, posture).
[0036] In some embodiments of the present application, multiple sensors are integrated into the embodied smart device. For example, a camera is used to collect visual data, such as the user's facial expressions and body movements. A microphone is used to collect auditory data, such as the user's voice and ambient sounds. During operation of the embodied smart device, the sensor array is activated to collect data in real time at a set sampling frequency. For example, the camera collects image data at a frequency of 30 frames per second, and the microphone collects audio data at a frequency of 44.1kHz per second. Data preprocessing and time alignment are performed on the collected multimodal data to obtain time-aligned multimodal features.
[0037] Specifically, when preprocessing multimodal data, remove invalid or abnormal data points. For example, remove blurred images captured by a camera or noisy signals captured by a microphone. Use filtering algorithms to remove noise from the data, such as using a low-pass filter to remove high-frequency noise from audio data. Convert data to a format suitable for subsequent processing, such as converting image data from RGB to grayscale to reduce the data volume. Extract useful features from the raw data, such as extracting facial key features from images and speech features from audio.
[0038] Specifically, when time-aligning multimodal data, data of different modalities are aligned to the same timeline based on the timestamps carried by the multimodal data itself. For example, the image data captured by the camera is aligned with the audio data captured by the microphone to ensure that the image and audio are synchronized in time.
[0039] In an embodiment of the present application, by preprocessing and time-aligning multimodal data, time-aligned multimodal features can be obtained, providing a high-quality data foundation for subsequent emotion recognition and interaction strategy generation.
[0040] S102, the embodied intelligent device inputs the multimodal features into a preset cross-modal causal reasoning engine, outputs the user's current emotional state label and the current causal factors that lead to the current emotional state label, and the preset cross-modal causal reasoning engine is used to perform current emotional state label analysis and event attribution analysis.
[0041] Among them, a preset cross-modal causal reasoning engine is used to process multimodal data and analyze the user's emotional state and behavior through causal reasoning. The engine is able to process data from different modalities and find causal relationships between multimodal data. The current emotional state label refers to the user's current emotional state derived by the system based on multimodal feature analysis. The emotional state label can be a specific emotional category, such as "happy", "sad", "angry", "surprised", etc. The current causal factor refers to the cause or triggering event that leads to the user's current emotional state. These factors can be environmental factors, user behavior, device interaction, etc.
[0042] For example Figure 3 As shown, the embodied intelligent device collects and preprocesses multimodal data, inputs the preprocessed features into a preset cross-modal causal reasoning engine to perform current emotional state label analysis and event attribution analysis, and outputs the current emotional state label and current causal factors.
[0043] Among them, for example Figure 4 As shown, the preset cross-modal causal reasoning engine includes an emotional state label analysis layer, a trigger event recognition layer, an evidence feature extraction layer, and an event attribution analysis layer.
[0044] In some embodiments of the present application, the multimodal features are input into a preset cross-modal causal reasoning engine, and the specific process of outputting the user's current emotional state label and the current causal factors that lead to the current emotional state label includes: the emotional state label analysis layer analyzes the user's current emotional state label based on the multimodal features; the trigger event recognition layer analyzes the coordinated changes of the multimodal features on the timeline to identify the triggering events that cause emotional changes, and obtains each emotional triggering event with a timestamp; the evidence feature extraction layer extracts the emotional state features and environmental features of the current interactive environment at each timestamp from the multimodal features to obtain a subset of multimodal evidence features related to each emotional triggering event; the event attribution analysis layer performs event attribution analysis on each emotional triggering event based on the multimodal evidence feature subset to obtain the current causal factors that lead to the current emotional state label.
[0045] Among them, coordinated changes on the timeline refer to the synchronous changes of multimodal features in a time series. These changes may indicate the occurrence of a specific event. By analyzing these coordinated changes, the triggering events that lead to changes in emotional state can be identified, providing a basis for causal analysis. Emotional state features are features extracted from multimodal features that can reflect the user's emotional state, such as facial expression features and voice intonation features. These features directly reflect the user's emotional state and are an important basis for emotional state analysis. Environmental features are features extracted from multimodal features that can reflect the current interaction environment, such as ambient sound, lighting conditions, temperature, etc. These features provide the environmental context for changes in emotional state and help to more comprehensively understand the causes of emotional changes.
[0046] In an embodiment of the present application, the embodied intelligent device can not only identify the user's current emotional state, but also analyze the specific reasons that lead to the emotional state, so that the device can more accurately understand the user's emotional needs and thus provide more targeted interaction strategies. For example, if the device recognizes that the user is currently in an "angry" state, and through causal analysis finds that it is caused by an improper operation of the device, the device can immediately adjust the operation, apologize to the user and provide a solution, thereby effectively alleviating the user's emotions and improving the user experience. This emotional interaction method based on causal reasoning not only improves the accuracy of emotion recognition, but also enhances the decision-making intelligence of the device, enabling the device to better adapt to complex and changing interaction scenarios and provide more personalized and considerate services.
[0047] Among them, the emotional state label analysis layer is deployed with a pre-trained emotional state analysis model, and the pre-trained emotional state analysis model is a mathematical model for analyzing the user's emotional state.
[0048] In some embodiments of the present application, the user's current emotional state label is analyzed based on multimodal features, including: inputting the multimodal features into a pre-trained emotional state analysis model to predict the current emotional state label corresponding to the multimodal data; and outputting the user's current emotional state label.
[0049] For example, multimodal features are input into the emotional state label analysis layer, and a pre-trained emotional state analysis model (such as a deep neural network) is used to analyze the multimodal features and output the user's current emotional state label, such as "happy", "sad", etc.
[0050] In some embodiments of the present application, a pre-trained emotional state analysis model is generated according to the following steps, including: obtaining historical multimodal data collected by the sensor array of the embodied smart device; labeling the historical multimodal data with emotional state labels to obtain model training samples; using a neural network to create an emotional state analysis model; inputting the model training samples into the emotional state analysis model to perform machine learning on the emotional state analysis model and output a model loss value; when the model loss value reaches a minimum, generating a pre-trained emotional state analysis model; or when the model loss value does not reach a minimum, backpropagating the model loss value to update the parameters of the model, and continuing to execute the step of inputting the model training samples into the emotional state analysis model until the model loss value reaches a minimum.
[0051] In some embodiments of the present application, event attribution analysis is performed on each emotion triggering event based on a multimodal evidence feature subset to obtain the current causal factors that lead to the current emotion state label. The specific process includes: generating a current causal factor candidate set for each emotion triggering event based on the multimodal evidence feature subset; establishing a causal graph using the elements in the current causal factor candidate set for each emotion triggering event as nodes; analyzing the sequence of node event occurrence and the node correlation matrix between each node in the causal graph; generating the current causal factors that lead to the current emotion state label based on the sequence of node event occurrence and the node correlation matrix between each node.
[0052] Specifically, each element in the current set of candidate causal factors is first defined as a node in a causal graph. Based on causal relationships, directed edges are defined between nodes. Directed edges indicate the direction of the causal relationship, that is, one node (cause) leads to another node (effect). All nodes and edges are combined into a causal graph. The event timestamps of each node (causal factor) in the causal graph are analyzed to determine the order in which the events occurred. The correlation matrix between the nodes in the causal graph is calculated using statistical methods (such as the Pearson correlation coefficient) or data-based similarity calculation methods.
[0053] In some embodiments of the present application, the specific process of generating the current causal factor candidate set of each emotion triggering event based on the multimodal evidence feature subset includes: obtaining a pre-set evidence feature knowledge base, which stores the mapping relationship between the original evidence feature sequence and the original causal factor set; calculating the feature similarity between the multimodal evidence feature subset and each original evidence feature sequence; determining the original evidence feature sequence with the greatest feature similarity; and according to the original evidence feature sequence with the greatest feature similarity, obtaining the corresponding original causal factor set from the mapping relationship as the current causal factor candidate set of each emotion triggering event.
[0054] Specifically, the evidence feature knowledge base is a predefined data structure that stores the mapping relationship between the original evidence feature sequence and the original causal factor set. For example, the knowledge base can be a table, where each row represents an original evidence feature sequence and each column represents a causal factor.
[0055] For example, suppose the multimodal evidence feature subset is a vector , the original evidence feature sequence is a vector ,in Indicates the The feature similarity can be calculated by cosine similarity, as shown in the following formula:
[0056]
[0057] in, is a vector and vector The dot product operation, and is a vector and vector The Euclidean norm of .
[0058] In some embodiments of the present application, based on the order of occurrence of node events between nodes and the node correlation matrix, the specific process of generating the current causal factors that lead to the appearance of the current emotional state label includes: calculating the correlation between each node according to the node correlation matrix between each node; deleting nodes with correlation less than or equal to 0 in the causal graph to obtain an alternative node sequence; establishing directed edges for the alternative node sequence according to the order of occurrence of node events between each node to obtain a node causal chain; and using the causal factors in the node causal chain as the current causal factors that lead to the appearance of the current emotional state label.
[0059] Specifically, assume a correlation matrix is ,in Representation node and nodes The correlation between The calculation formula is:
[0060]
[0061] According to the calculated correlation matrix , delete the nodes whose correlation is less than or equal to 0. By traversing the matrix And check each element to achieve it.
[0062] For example, for:
[0063]
[0064] Combined with the relevant formula, the correlation matrix is calculated for:
[0065]
[0066] Traversing the matrix C, we find that there are no nodes with a correlation less than or equal to 0, so the candidate node sequence is all nodes. Assuming that the order of node events is [C1, C2, C3], we re-establish the directed edge C1 → C2 → C3. The nodes C1, C2, C3 in the node causal chain are the current causal factors.
[0067] In the embodiment of the present application, the correlation is converted into a correlation degree to more intuitively represent the degree of correlation between nodes. It is ensured that only nodes with a correlation degree greater than 0 are considered, that is, it is considered that there is a positive correlation between these nodes. According to the order in which the node events occur, directed edges are established to form a node causal chain. The nodes in the node causal chain are the current causal factors that lead to the current emotional state label.
[0068] S103, the embodied intelligent device queries the preset emotion behavior causal mapping library based on the current emotion state label and current causal factors to obtain the current optimal emotion interaction strategy instruction set. The preset emotion behavior causal mapping library is a triple consisting of historical emotion state labels, historical causal factors, and historical interaction strategies that have been verified to be effective.
[0069] In some embodiments of the present application, based on the current emotional state label and the current causal factors, the specific process of querying the preset emotional behavior causal mapping library to obtain the current optimal emotional interaction strategy instruction set includes: querying the preset emotional behavior causal mapping library for the triples to which the historical emotional state label is the same as the current emotional state label as multiple alternative triples; calculating the similarity between the historical causal factors and the current causal factors in each alternative triple; in the case where there is a historical causal factor with the highest similarity, using the historical interaction strategy in the triple to which the historical causal factor with the highest similarity belongs as the current optimal emotional interaction strategy instruction set; or, in the case where the similarity is less than or equal to 0, triggering the manual response strategy to send the current emotional state label and causal factors to the client for manual labeling, and receiving the current optimal emotional interaction strategy instruction set manually input. The preset emotional behavior causal mapping library is shown in Table 1.
[0070] Table 1
[0071]
[0072] Specifically, assuming the current emotional state is labeled "sad" and the current causal factor is {C1, C2}, we search the mapping library for all triplets with the emotional state label "sad": (sad, {C1, C2}, strategy A) and (sad, {C2, C3}, strategy B). We then calculate the similarity between the current causal factor {C1, C2} and each historical causal factor: the similarity with {C1, C2} is 1.0, and the similarity with {C2, C3} is 0.5. The historical causal factor with the highest similarity is {C1, C2}, and its corresponding triplet is (sad, {C1, C2}, strategy A). The current optimal emotional interaction strategy instruction set is "strategy A."
[0073] "Strategy A" could be, for example:
[0074] Voice: Immediately stop the current instruction and apologize clearly: "I'm sorry for confusing you." Repeat the core steps using a slower pace and simpler vocabulary.
[0075] Expression: The avatar shows a concerned expression (slightly furrowed eyebrows, focused eyes); the robot's head tilts slightly to indicate listening.
[0076] Action: The robot may take a half step back (to reduce the sense of pressure) or point a finger at a key operating part (to aid understanding).
[0077] Task adjustment: temporarily reduce the complexity of the task and split the current steps.
[0078] Specifically, if the similarity of the most similar historical causal factors is less than or equal to 0, the manual response strategy is triggered. The current emotional state label and causal factors are sent to the human annotator client, who then inputs the optimal emotional interaction strategy instruction set. This process ensures that the system can select the optimal emotional interaction strategy based on historical data, and can also provide accurate responses through manual intervention when there is insufficiently similar historical data.
[0079] In some embodiments of the present application, a preset emotion behavior causal mapping library is generated according to the following steps: historical multimodal data collected by the sensor array of the embodied smart device and at least one set of historical emotion interaction strategy instruction sets are obtained; historical multimodal data are preprocessed to obtain time-aligned historical multimodal features; historical multimodal features are input into a preset cross-modal causal reasoning engine to output the user's historical emotion state label and the historical causal factors that lead to the current emotion state label; from multiple sets of historical emotion interaction strategy instruction sets, a historical emotion interaction strategy instruction set used to characterize user satisfaction is determined as a historical interaction strategy that has been historically verified to be valid; a triple relationship between historical emotion state labels, historical causal factors, and historical interaction strategies that have been historically verified to be valid is stored to obtain a preset emotion behavior causal mapping library.
[0080] In this embodiment of the application, by continuously accumulating and analyzing historical multimodal data and emotional interaction strategies of embodied intelligent devices in interactive scenarios, the system can accurately identify emotional states and their causal factors, filter out effective interaction strategies, and build a pre-set emotional behavior causal mapping library containing emotional states, causal factors, and effective strategies. This process enables intelligent devices to quickly select the optimal strategy based on rich empirical data when faced with new emotional interactions, thereby more accurately responding to user emotional needs and improving user experience and system performance.
[0081] Furthermore, after obtaining the current optimal emotional interaction strategy instruction set, the current optimal emotional interaction strategy instruction set is executed. After the execution is completed, multimodal data is collected for analysis again, and the strategy is optimized in real time based on user emotional feedback (such as changes in voice tone and adjustments in body movements) to form a dynamic closed loop of emotional interaction.
[0082] In the embodiment of the present application, on the one hand, the embodied intelligent device collects and pre-processes multimodal data in real time through an integrated sensor array to obtain time-aligned multimodal features. These features contain rich information, providing a comprehensive data basis for emotion recognition, reducing misjudgments caused by the limitations of single-modal data, and thus improving the accuracy of emotion recognition. On the other hand, the multimodal features are input into a preset cross-modal causal reasoning engine, which can perform current emotional state label analysis and event attribution analysis, output the user's current emotional state label and the causal factors that lead to the emotional state. Through cross-modal causal reasoning, the causal relationship between multimodal signals is fully considered, so that the user's emotional state can be more accurately identified. At the same time, combined with the preset emotional behavior causal mapping library, the misjudgment rate is effectively reduced, and the decision-making accuracy of the embodied intelligent device is improved.
[0083] See Figure 5 , which is a flow chart of a method for training an intent recognition model according to an embodiment of the present application. Figure 5 As shown, the method of the embodiment of the present application may include the following steps:
[0084] S201, obtaining historical multimodal data collected by a sensor array of an embodied smart device;
[0085] S202, labeling emotional state labels for historical multimodal data to obtain model training samples;
[0086] S203, creating an emotional state analysis model using a neural network;
[0087] S204, inputting the model training sample into the emotional state analysis model to perform machine learning on the emotional state analysis model and outputting a model loss value;
[0088] S205, when the model loss value reaches the minimum, generate a pre-trained emotional state analysis model; or when the model loss value does not reach the minimum, backpropagate the model loss value to update the model parameters, and continue to execute the step of inputting the model training sample into the emotional state analysis model until the model loss value reaches the minimum.
[0089] In the embodiments of this application, by acquiring and annotating historical multimodal data, a neural network is used to create an emotional state analysis model, and machine learning is used to continuously optimize the model parameters until the model loss value is minimized, thereby generating a high-precision emotional state analysis model. This process not only improves the accuracy of emotion recognition, but also enhances the generalization ability of the model, enabling embodied smart devices to more accurately understand user emotions and provide personalized services, thereby improving user experience and system performance.
[0090] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0091] See Figure 6 , which shows a schematic diagram of the structure of an emotional interaction device based on multimodal data fusion provided by an exemplary embodiment of the present application. The emotional interaction device based on multimodal data fusion can be implemented as all or part of an electronic device through software, hardware, or a combination of both. The device 1 includes a modal data preprocessing module 10, an engine reasoning module 20, and an instruction set determination module 30.
[0092] The modal data preprocessing module 10 is used to collect and preprocess multimodal data in real time through the sensor array integrated in the embodied smart device to obtain time-aligned multimodal features;
[0093] An engine reasoning module 20 is configured to input multimodal features into a preset cross-modal causal reasoning engine and output the user's current emotional state label and the current causal factors that lead to the current emotional state label. The preset cross-modal causal reasoning engine is configured to perform current emotional state label analysis and event attribution analysis.
[0094] The instruction set determination module 30 is used to query the preset emotional behavior causal mapping library based on the current emotional state label and the current causal factors to obtain the current optimal emotional interaction strategy instruction set. The preset emotional behavior causal mapping library is a triple consisting of historical emotional state labels, historical causal factors, and historical interaction strategies that have been verified to be effective.
[0095] It should be noted that the emotional interaction device based on multimodal data fusion provided in the above embodiment only uses the division of the above functional modules as an example when executing the emotional interaction method based on multimodal data fusion. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the emotional interaction device based on multimodal data fusion provided in the above embodiment and the emotional interaction method based on multimodal data fusion embodiment belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
[0096] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0097] In the embodiment of the present application, on the one hand, the embodied intelligent device collects and pre-processes multimodal data in real time through an integrated sensor array to obtain time-aligned multimodal features. These features contain rich information, providing a comprehensive data basis for emotion recognition, reducing misjudgments caused by the limitations of single-modal data, and thus improving the accuracy of emotion recognition. On the other hand, the multimodal features are input into a preset cross-modal causal reasoning engine, which can perform current emotional state label analysis and event attribution analysis, output the user's current emotional state label and the causal factors that lead to the emotional state. Through cross-modal causal reasoning, the causal relationship between multimodal signals is fully considered, so that the user's emotional state can be more accurately identified. At the same time, combined with the preset emotional behavior causal mapping library, the misjudgment rate is effectively reduced, and the decision-making accuracy of the embodied intelligent device is improved.
[0098] The present application also provides a computer-readable medium having program instructions stored thereon, which, when executed by a processor, implement the emotional interaction method based on multimodal data fusion provided by the above-mentioned various method embodiments.
[0099] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the emotion interaction method based on multimodal data fusion of the above-mentioned various method embodiments.
[0100] See Figure 7 , is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 7 As shown, the electronic device 1000 may include: at least one processor 1001 , at least one network interface 1004 , a user interface 1003 , a memory 1005 , and at least one communication bus 1002 .
[0101] The communication bus 1002 is used to implement the connection and communication between these components.
[0102] The user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface.
[0103] The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0104] The processor 1001 may include one or more processing cores. The processor 1001 utilizes various interfaces and circuits to connect various components within the electronic device 1000. It executes instructions, programs, code sets, or instruction sets stored in the memory 1005, and accesses data stored in the memory 1005 to perform various functions and process data within the electronic device 1000. Optionally, the processor 1001 may be implemented in hardware using at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 1001 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the processor 1001 and implemented on a separate chip.
[0105] Among them, the memory 1005 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 1005 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 1005 may also be optionally at least one storage system located away from the aforementioned processor 1001. As Figure 7 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an emotion interaction application based on multimodal data fusion.
[0106] exist Figure 7In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an input interface for the user and obtain user input data; and the processor 1001 can be used to call the emotion interaction application based on multimodal data fusion stored in the memory 1005 and specifically perform the following operations:
[0107] By integrating the sensor array in the embodied smart device into real-time acquisition and preprocessing of multimodal data, time-aligned multimodal features are obtained.
[0108] Input the multimodal features into a preset cross-modal causal reasoning engine, output the user's current emotional state label and the current causal factors that lead to the current emotional state label, and the preset cross-modal causal reasoning engine is used to perform current emotional state label analysis and event attribution analysis;
[0109] Based on the current emotional state label and current causal factors, the preset emotional behavior causal mapping library is queried to obtain the current optimal emotional interaction strategy instruction set. The preset emotional behavior causal mapping library is a triple consisting of historical emotional state labels, historical causal factors, and historical interaction strategies that have been verified to be effective.
[0110] In one embodiment, when the processor 1001 inputs the multimodal features into a preset cross-modal causal reasoning engine and outputs the user's current emotional state label and the current causal factors that lead to the current emotional state label, the processor 1001 specifically performs the following operations:
[0111] The emotional state label analysis layer analyzes the user's current emotional state label based on multimodal features;
[0112] The trigger event recognition layer analyzes the coordinated changes of multimodal features on the time axis to identify the trigger events that cause emotional changes and obtain each emotional trigger event with a timestamp;
[0113] The evidence feature extraction layer extracts the emotional state features reflecting the user's emotional state at each timestamp and the environmental features of the current interaction environment from the multimodal features, obtaining a subset of multimodal evidence features related to each emotion triggering event;
[0114] The event attribution analysis layer performs event attribution analysis on each emotion triggering event based on the multimodal evidence feature subset to obtain the current causal factors that lead to the current emotion state label.
[0115] In one embodiment, when performing event attribution analysis on each emotion triggering event based on the multimodal evidence feature subset to obtain the current causal factor leading to the current emotion state label, the processor 1001 specifically performs the following operations:
[0116] Generate the current causal factor candidate set of each emotion triggering event based on the multimodal evidence feature subset;
[0117] Using the elements in the current causal factor candidate set of each emotion triggering event as nodes, a causal graph is established;
[0118] Analyze the sequence of node events between nodes in the causal graph and the correlation matrix between nodes;
[0119] Based on the order of occurrence of node events between nodes and the correlation matrix between nodes, the current causal factors that lead to the current emotional state label are generated.
[0120] In one embodiment, when the processor 1001 generates a current candidate set of causal factors for each emotion triggering event based on the multimodal evidence feature subset, the processor 1001 specifically performs the following operations:
[0121] Obtaining a pre-set evidence feature knowledge base, the evidence feature knowledge base storing a mapping relationship between an original evidence feature sequence and an original causal factor set;
[0122] Calculate the feature similarity between the multimodal evidence feature subset and each original evidence feature sequence;
[0123] Determine the original evidence feature sequence with the greatest feature similarity;
[0124] According to the original evidence feature sequence with the largest feature similarity, the corresponding original causal factor set is obtained from the mapping relationship as the current causal factor candidate set of each emotion triggering event.
[0125] In one embodiment, when the processor 1001 generates the current causal factor leading to the current emotional state label based on the order of occurrence of node events between nodes and the node correlation matrix, the processor 1001 specifically performs the following operations:
[0126] Calculate the correlation between each node according to the node correlation matrix between each node;
[0127] Delete the nodes with correlation less than or equal to 0 in the causal graph to obtain a sequence of candidate nodes;
[0128] According to the order of occurrence of node events between nodes, a directed edge is established for the candidate node sequence to obtain a node causal chain;
[0129] The causal factors in the node causal chain are taken as the current causal factors that lead to the current emotional state label.
[0130] In one embodiment, when analyzing the user's current emotional state label based on the multimodal features, the processor 1001 specifically performs the following operations:
[0131] Input the multimodal features into the pre-trained emotional state analysis model to predict the current emotional state label corresponding to the multimodal data;
[0132] Output the user's current emotional state label.
[0133] In one embodiment, when generating a pre-trained emotional state analysis model, the processor 1001 specifically performs the following operations:
[0134] Obtain historical multimodal data collected by the sensor array of embodied smart devices;
[0135] Label the emotional state of historical multimodal data to obtain model training samples;
[0136] Use neural networks to create an emotional state analysis model;
[0137] Input the model training samples into the emotional state analysis model to perform machine learning on the emotional state analysis model and output the model loss value;
[0138] When the model loss value reaches the minimum, a pre-trained emotional state analysis model is generated; or when the model loss value does not reach the minimum, the model loss value is back-propagated to update the model parameters, and the step of inputting the model training samples into the emotional state analysis model is continued until the model loss value reaches the minimum.
[0139] In one embodiment, when the processor 1001 queries the preset emotion behavior causal mapping library based on the current emotion state label and the current causal factors to obtain the current optimal emotion interaction strategy instruction set, the processor 1001 specifically performs the following operations:
[0140] From the preset emotion behavior causal mapping library, query the triples belonging to the historical emotion state label that is the same as the current emotion state label as multiple candidate triples;
[0141] Calculate the similarity between the historical causal factors and the current causal factors in each alternative triple;
[0142] In the case of a historical causal factor with the highest similarity, the historical interaction strategy in the triplet to which the historical causal factor with the highest similarity belongs is used as the current optimal emotional interaction strategy instruction set; or,
[0143] When the similarity is less than or equal to 0, the manual response strategy is triggered to send the current emotional state label and causal factors to the client for manual labeling, and receive the current optimal emotional interaction strategy instruction set input by manual input.
[0144] In one embodiment, the processor 1001 further performs the following operations:
[0145] Obtain historical multimodal data collected by the sensor array of the embodied smart device and at least one set of historical emotional interaction strategy instruction sets;
[0146] Preprocess historical multimodal data to obtain time-aligned historical multimodal features;
[0147] Input historical multimodal features into a preset cross-modal causal reasoning engine to output the user's historical emotional state label and the historical causal factors that led to the current emotional state label;
[0148] From multiple sets of historical emotion interaction strategy instruction sets, determine a historical emotion interaction strategy instruction set used to characterize user satisfaction as a historical interaction strategy that has been historically verified to be effective;
[0149] The triple relationship between historical emotional state labels, historical causal factors, and historically verified effective historical interaction strategies is stored to obtain a preset emotional behavior causal mapping library.
[0150] In the embodiment of the present application, on the one hand, the embodied intelligent device collects and pre-processes multimodal data in real time through an integrated sensor array to obtain time-aligned multimodal features. These features contain rich information, providing a comprehensive data basis for emotion recognition, reducing misjudgments caused by the limitations of single-modal data, and thus improving the accuracy of emotion recognition. On the other hand, the multimodal features are input into a preset cross-modal causal reasoning engine, which can perform current emotional state label analysis and event attribution analysis, output the user's current emotional state label and the causal factors that lead to the emotional state. Through cross-modal causal reasoning, the causal relationship between multimodal signals is fully considered, so that the user's emotional state can be more accurately identified. At the same time, combined with the preset emotional behavior causal mapping library, the misjudgment rate is effectively reduced, and the decision-making accuracy of the embodied intelligent device is improved.
[0151] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiment methods can be implemented by instructing the relevant hardware through a computer program. The program for emotional interaction based on multimodal data fusion can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-described methods. The storage medium of the program for emotional interaction based on multimodal data fusion can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0152] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. An emotional interaction method based on multimodal data fusion, characterized in that: Applied to an embodied intelligent device, the method includes: The sensor array integrated in the embodied smart device collects and preprocesses multimodal data in real time to obtain time-aligned multimodal features; Inputting the multimodal features into a preset cross-modal causal reasoning engine, outputting the user's current emotional state label and the current causal factors that lead to the current emotional state label, wherein the preset cross-modal causal reasoning engine is used to perform current emotional state label analysis and event attribution analysis; The preset cross-modal causal reasoning engine includes an emotional state label analysis layer, a trigger event recognition layer, an evidence feature extraction layer, and an event attribution analysis layer; Inputting the multimodal features into a preset cross-modal causal reasoning engine to output the user's current emotional state label and the current causal factors leading to the current emotional state label includes: The emotional state label analysis layer analyzes the user's current emotional state label based on the multimodal features; The trigger event recognition layer analyzes the coordinated changes of the multimodal features on the time axis to identify the trigger events that cause emotional changes and obtain each emotional trigger event with a time stamp; The evidence feature extraction layer extracts the emotional state features reflecting the user's emotional state and the environmental features of the current interaction environment at each time stamp from the multimodal features, and obtains a subset of multimodal evidence features related to each emotion triggering event; The event attribution analysis layer performs event attribution analysis on each emotion triggering event based on the multimodal evidence feature subset to obtain the current causal factors that lead to the appearance of the current emotion state label; Based on the current emotional state label and the current causal factors, the preset emotional behavior causal mapping library is queried to obtain the current optimal emotional interaction strategy instruction set. The preset emotional behavior causal mapping library is a triple consisting of historical emotional state labels, historical causal factors, and historical interaction strategies that have been historically verified to be effective.
2. The method according to claim 1, characterized in that The performing of event attribution analysis on each emotion triggering event based on the multimodal evidence feature subset to obtain the current causal factors leading to the appearance of the current emotion state label includes: generating a current causal factor candidate set for each emotion triggering event based on the multimodal evidence feature subset; Establishing a causal graph using elements in the current causal factor candidate set of each emotion triggering event as nodes; Analyze the occurrence sequence of node events between nodes in the causal graph and the correlation matrix between nodes; Based on the occurrence sequence of node events between the nodes and the node correlation matrix, a current causal factor leading to the appearance of the current emotional state label is generated.
3. The method according to claim 2, characterized in that Generating a current causal factor candidate set for each emotion triggering event based on the multimodal evidence feature subset includes: Obtaining a pre-set evidence feature knowledge base, wherein the evidence feature knowledge base stores a mapping relationship between an original evidence feature sequence and an original causal factor set; Calculating feature similarity between the multimodal evidence feature subset and each original evidence feature sequence; Determine the original evidence feature sequence with the greatest feature similarity; According to the original evidence feature sequence with the greatest feature similarity, the corresponding original causal factor set is obtained from the mapping relationship as the current causal factor candidate set of each emotion triggering event.
4. The method according to claim 2, characterized in that The generating of the current causal factors leading to the appearance of the current emotional state label based on the occurrence sequence of node events between the nodes and the node correlation matrix includes: Calculating the correlation between the nodes according to the node correlation matrix between the nodes; Deleting nodes in the causal graph whose correlation is less than or equal to 0 to obtain a candidate node sequence; According to the order in which node events between the nodes occur, a directed edge is established for the candidate node sequence to obtain a node causal chain; The causal factors in the node causal chain are used as current causal factors that lead to the appearance of the current emotional state label.
5. The method according to claim 1, characterized in that The emotional state tag analysis layer is equipped with a pre-trained emotional state analysis model, which is a mathematical model for analyzing the user's emotional state; Analyzing the user's current emotional state label based on the multimodal features includes: Inputting the multimodal features into a pre-trained emotional state analysis model to predict a current emotional state label corresponding to the multimodal data; Output the user's current emotional state label.
6. The method according to claim 5, characterized in that Follow these steps to generate a pre-trained emotional state analysis model, including: Acquiring historical multimodal data collected by a sensor array of the embodied smart device; Labeling the historical multimodal data with emotional state labels to obtain model training samples; Use neural networks to create an emotional state analysis model; Inputting the model training sample into the emotional state analysis model to perform machine learning on the emotional state analysis model and outputting a model loss value; When the model loss value reaches the minimum, a pre-trained emotional state analysis model is generated; or when the model loss value does not reach the minimum, the model loss value is back-propagated to update the model parameters, and the step of inputting the model training sample into the emotional state analysis model is continued until the model loss value reaches the minimum.
7. The method according to claim 1, characterized in that The method of querying a preset emotional behavior causal mapping library based on the current emotional state label and the current causal factors to obtain the current optimal emotional interaction strategy instruction set includes: Querying the preset emotion behavior causal mapping library for triples belonging to historical emotion state labels identical to the current emotion state label as multiple candidate triples; Calculating the similarity between the historical causal factors and the current causal factors in each candidate triple; In the case of a historical causal factor with the highest similarity, the historical interaction strategy in the triplet to which the historical causal factor with the highest similarity belongs is used as the current optimal emotional interaction strategy instruction set; or, When the similarity is less than or equal to 0, a manual response strategy is triggered to send the current emotional state label and the causal factors to the client for manual labeling, and receive the current optimal emotional interaction strategy instruction set manually input.
8. The method according to claim 1, characterized in that To generate a preset emotion-behavior causal mapping library, follow these steps: Acquiring historical multimodal data collected by the sensor array of the embodied intelligent device and at least one set of historical emotion interaction strategy instruction sets; Preprocessing the historical multimodal data to obtain time-aligned historical multimodal features; Inputting the historical multimodal features into the preset cross-modal causal reasoning engine, outputting the user's historical emotional state label and the historical causal factors that led to the current emotional state label; Determining a historical emotion interaction strategy instruction set for representing user satisfaction from the at least one set of historical emotion interaction strategy instruction sets as a historical interaction strategy that has been verified to be effective; The triple relationship between the historical emotional state labels, historical causal factors, and historically verified effective historical interaction strategies is stored to obtain a preset emotional behavior causal mapping library.
9. An emotional interaction device based on multimodal data fusion implemented using the method according to any one of claims 1 to 8, characterized in that: The device comprises: The modal data preprocessing module is used to collect and preprocess multimodal data in real time through the sensor array integrated in the embodied smart device to obtain time-aligned multimodal features; An engine reasoning module, configured to input the multimodal features into a preset cross-modal causal reasoning engine, and output the user's current emotional state label and the current causal factors that lead to the current emotional state label. The preset cross-modal causal reasoning engine is configured to perform current emotional state label analysis and event attribution analysis. The instruction set determination module is used to query the preset emotional behavior causal mapping library based on the current emotional state label and the current causal factors to obtain the current optimal emotional interaction strategy instruction set. The preset emotional behavior causal mapping library is a triple consisting of historical emotional state labels, historical causal factors and historical interaction strategies that have been verified to be effective.