Multi-modal sentiment analysis method, intelligent question and answer method and intelligent question and answer device

Through the multimodal emotion analysis method combined with speech and facial image information, and the multimodal emotion recognition model is used to solve the problem of inaccurate single mode emotion recognition, and the service quality and efficiency of human-computer interaction are improved.

CN120541570APending Publication Date: 2025-08-26INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510635484.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, it is not accurate enough to obtain the user's emotional state information through voice information, which affects the service quality and efficiency of human-computer interaction.

Method used

The multimodal emotion analysis method is adopted, combining the user's voice information and facial image information, and the training multimodal emotion recognition model is used to obtain more accurate emotional state information.

Benefits of technology

The service quality and efficiency in human-computer interaction are improved, and more accurate emotional state information is obtained through multimodal information fusion analysis, and the Q&A strategy is adjusted to improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541570A_ABST
    Figure CN120541570A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a multi-modal sentiment analysis method, an intelligent question answering method and an intelligent question answering device. The multi-modal sentiment analysis method comprises the steps that firstly, multi-modal information of a user is acquired, and the multi-modal information comprises first voice information and first face image information of the user; inputting the multi-modal information of the user into the trained multi-modal emotion recognition model to obtain emotion state information corresponding to the multi-modal information; the multi-modal emotion recognition model is obtained by training based on a multi-modal information sample set, and the multi-modal information sample set comprises a voice information sample and a facial image information sample; according to the method, the user voice information and the face information of the user are fused and analyzed through the multi-mode emotion recognition model, so that the obtained emotion state information of the user is more accurate, and the service quality and the interaction efficiency are improved when man-machine interaction is carried out according to the emotion state information and the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a multimodal sentiment analysis method, an intelligent question-answering method, and an intelligent question-answering device. Background Art

[0002] With the development of artificial intelligence and deep learning, human-computer interaction technology has become more mature and intelligent, and is gradually being applied to various industries. For example, current intelligent question-answering technology is widely used in technical fields such as intelligent customer service and intelligent guidance to improve user service efficiency while reducing labor costs.

[0003] In related technologies, during intelligent human-computer interaction, user voice information is typically acquired and then features extracted using deep learning models to obtain the user's question-and-answer information and emotional state information. This allows the Q&A strategy to be adjusted based on the user's emotional state, improving service quality. However, user emotional information is often not simply expressed through voice information, resulting in inaccurate extracted emotional state information, which impacts both service quality and the efficiency of human-computer interaction. Summary of the Invention

[0004] The present application provides a multimodal sentiment analysis method, an intelligent question-answering method, and an intelligent question-answering device, which can solve the technical problem in related technologies that obtaining user's emotional state information only through voice information is not accurate enough.

[0005] In a first aspect, an embodiment of the present application provides a multimodal sentiment analysis method, the method comprising:

[0006] Acquiring multimodal information of a user, the multimodal information including first voice information and first facial image information of the user;

[0007] Inputting the user's multimodal information into a trained multimodal emotion recognition model to obtain emotional state information corresponding to the multimodal information;

[0008] In which, the multimodal emotion recognition model is trained based on a multimodal information sample set, wherein the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and the facial image information samples; the emotional state information labels are used to reflect the user's emotional state.

[0009] In a second aspect, the present application further provides an intelligent question-answering method based on emotion recognition, the method comprising:

[0010] During the question-and-answer interaction with the user, obtaining multimodal information of the user, the multimodal information including first voice information and first facial image information of the user;

[0011] Inputting the user's multimodal information into a trained multimodal emotion recognition model to obtain emotional state information corresponding to the multimodal information;

[0012] Adjusting the question-answering strategy for the user based on the user's emotional state information; the question-answering strategy includes the tone and wording of the question-answering;

[0013] In which, the multimodal emotion recognition model is trained based on a multimodal information sample set, wherein the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and the facial image information samples; the emotional state information labels are used to reflect the user's emotional state.

[0014] In a third aspect, the present application further provides an intelligent question-answering device, comprising:

[0015] An information collection module, configured to obtain multimodal information of the user during a question-and-answer interaction with the user, the multimodal information including the user's first voice information and first facial image information;

[0016] An emotional state recognition module, configured to input the user's multimodal information into a trained multimodal emotion recognition model to obtain emotional state information corresponding to the multimodal information;

[0017] An intelligent question-answering module, configured to adjust a question-answering strategy for a user based on the user's emotional state information; the question-answering strategy includes the tone and wording of the question and answer;

[0018] In which, the multimodal emotion recognition model is trained based on a multimodal information sample set, wherein the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and the facial image information samples; the emotional state information labels are used to reflect the user's emotional state.

[0019] In a fourth aspect, the present application further provides an electronic device, comprising:

[0020] one or more processors;

[0021] a memory for storing one or more programs,

[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement the multimodal sentiment analysis method as described in any one of the above items; or implement the intelligent question-answering method based on emotion recognition as described in any one of the above items.

[0023] In a fifth aspect, the present application also provides a storage medium on which a computer program is stored, which, when executed by a processor, implements a multimodal sentiment analysis method as described in any one of the above items; or implements an intelligent question-answering method based on emotion recognition as described in any one of the above items.

[0024] In a sixth aspect, the present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements a multimodal sentiment analysis method as described in any one of the above items; or implements an intelligent question-answering method based on sentiment recognition as described in any one of the above items.

[0025] The embodiment of the present application proposes a multimodal emotion analysis method, which first obtains the multimodal information of the user, wherein the multimodal information includes the user's first voice information and first facial image information; then the user's multimodal information is input into a trained multimodal emotion recognition model to obtain the emotional state information corresponding to the multimodal information; wherein the multimodal emotion recognition model is trained based on a multimodal information sample set, wherein the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and facial image information samples; the emotional state information labels are used to reflect the user's emotional state. It can be seen that the embodiment of the present application uses the multimodal emotion recognition model to perform a fusion analysis of the user's voice information and the user's facial information, making the obtained user's emotional state information more accurate, thereby improving the service quality and interaction efficiency when conducting human-computer interaction with the user based on the emotional state information. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 One of the flow charts of the multimodal sentiment analysis method provided in the embodiment of the present application;

[0027] Figure 2 A flowchart of a multimodal emotion recognition model training method provided in an embodiment of the present application;

[0028] Figure 3 One of the schematic diagrams of the multimodal emotion recognition model structure provided in the embodiment of this application;

[0029] Figure 4 The second flowchart of the multimodal sentiment analysis method provided in the embodiment of the present application;

[0030] Figure 5The third flowchart of the multimodal sentiment analysis method provided in the embodiment of the present application;

[0031] Figure 6 A flowchart of an intelligent question-answering method based on emotion recognition provided in an embodiment of the present application;

[0032] Figure 7 A schematic diagram of the structure of the intelligent question-answering device provided in an embodiment of the present application;

[0033] Figure 8 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present application and are not intended to limit the present application. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions of the present application, not all of the structures.

[0035] With the rapid development of artificial intelligence (AI), emotion recognition has become a key research area in the field of human-computer interaction. Related technologies primarily rely on single-modal information, such as speech or facial expressions. However, these methods often suffer from insufficient recognition accuracy and poor robustness in complex and diverse real-world scenarios. In point-of-sale service scenarios, customers' emotional states are often expressed through a variety of channels, including speech and facial expressions. Therefore, single-modal emotion recognition methods struggle to fully capture users' true emotions.

[0036] In order to overcome the deficiencies in the above-mentioned related technologies, the present application provides a multimodal emotion analysis method, which first obtains the multimodal information of the user, wherein the multimodal information includes the user's first voice information and first facial image information; then the user's multimodal information is input into a trained multimodal emotion recognition model to obtain the emotional state information corresponding to the multimodal information; wherein the multimodal emotion recognition model is trained based on a multimodal information sample set, wherein the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and facial image information samples; the emotional state information labels are used to reflect the user's emotional state. It can be seen that in the embodiment of the present application, the multimodal emotion recognition model is used to fuse and analyze the user's voice information and the user's facial information, so that the obtained user's emotional state information is more accurate, thereby improving the service quality and interaction efficiency when conducting human-computer interaction with the user based on the emotional state information.

[0037] Figure 1This is one of the flow charts of the multimodal sentiment analysis method provided in the embodiments of the present application. This method can be performed by a multimodal sentiment analysis device or electronic device. The device or electronic device can be implemented by software and / or hardware. The device or electronic device can be integrated into any smart device with network communication function. Figure 1 As shown, the multimodal sentiment analysis method provided in this embodiment may include the following steps:

[0038] S101. Acquire multimodal information of a user, where the multimodal information includes first voice information and first facial image information of the user.

[0039] In one application scenario, the user's first voice information can be obtained through a voice acquisition device, and the user's first facial image information can be obtained through a video or image acquisition device, and then the collected user's first voice information and first facial image information are sent to a multimodal emotion analysis device.

[0040] For example, in one application scenario, when a user is conducting intelligent question-and-answer sessions with intelligent customer service through a smart terminal (such as a mobile phone, notebook, tablet, or smart wearable device), the user's current first voice information can be collected through the microphone on the smart terminal, and the user's first facial image information can be collected through the camera on the smart terminal.

[0041] S102. Input the user's multimodal information into the trained multimodal emotion recognition model to obtain the emotion state information corresponding to the multimodal information; wherein the multimodal emotion recognition model is trained based on a multimodal information sample set, and the multimodal information sample set includes voice information samples and facial image information samples, as well as emotion state information labels corresponding to the voice information samples and facial image information samples; the emotion state information labels are used to reflect the user's emotion state.

[0042] Among them, for the training of the multimodal emotion recognition model, a labeled training method can be adopted in this embodiment. Figure 2 For a flow chart of a multimodal emotion recognition model training method provided in this application embodiment, please refer to Figure 2 As shown, the multimodal emotion recognition model training method includes:

[0043] S201. Construct a multimodal information sample set, which includes multiple voice information samples and emotional state information labels corresponding to each voice information sample; and also includes multiple facial image information samples and emotional state information labels corresponding to each facial image information sample.

[0044] The emotional state information tags are keywords that can represent the user's emotions. For example, the emotional state information tags may include words such as tranquility, joy, impatience, anger, and indifference.

[0045] S202: Use the multimodal information sample set to train the initial multimodal emotion recognition model, and end the training when it is determined that the recognition accuracy of the multimodal emotion recognition model meets the preset requirements.

[0046] Specifically, each labeled speech information sample and facial image information sample is input into the initial multimodal emotion recognition model, and the model parameters in the initial multimodal emotion recognition model are trained through iteration so that the model converges and the training is terminated when it is determined that the recognition accuracy of the multimodal emotion recognition model meets the preset requirements, thereby obtaining a trained multimodal emotion recognition model.

[0047] It can be understood that the trained multimodal emotion recognition model in this embodiment is an overall model, which can include multiple sub-models with different functions. Each sub-model can be trained separately, so that each sub-model has different functionality to form a multimodal emotion recognition model, thereby realizing the function of emotional state recognition of the input multimodal information.

[0048] Figure 3 One of the multimodal emotion recognition model structure diagrams provided in this application embodiment is shown in Figure 3 As shown, in one embodiment, the multimodal emotion recognition model 300 includes a first feature extraction model 303, a second feature extraction model 304, a language emotion recognition model 305 and a micro-expression recognition model 306; for each specific model, the corresponding training set is used in advance to train each model, and after the training is completed, the corresponding function can be realized.

[0049] The multimodal emotion recognition model structure provided in this embodiment mainly includes a language emotion recognition model 305 and a micro-expression recognition model 306. The language emotion recognition model 305 and the micro-expression recognition model 306 can respectively perform emotion recognition on the first voice information and the first facial image information to obtain corresponding emotion state information.

[0050] Based on the multimodal emotion recognition model structure provided in this embodiment, the initial language emotion recognition model can be trained using the above-mentioned labeled voice information samples to obtain a trained language emotion recognition model 305, and the initial micro-expression recognition model can be trained using the above-mentioned labeled facial image information samples to obtain a trained micro-expression recognition model 306. The training method in this embodiment can refer to the relevant labeled model training method and will not be repeated here.

[0051] In one embodiment, a convolutional neural network (CNN) can be trained using labeled speech information samples to obtain a trained language emotion recognition model 305; and a recurrent neural network (RNN) can be trained using labeled facial image information samples to obtain a trained micro-expression recognition model 306. The trained language emotion recognition model 305 and micro-expression recognition model 306 are then combined to construct a multimodal emotion recognition model to accurately identify the customer's emotional state.

[0052] The convolutional neural network (CNN) is trained using labeled speech information samples to obtain the trained language emotion recognition model 305, which can be expressed by the following expression:

[0053] Model micro-expression =Train(D micro-expression ) (1)

[0055] Among them, Model micro-expression Denotes the language emotion recognition model obtained after training, D micro-expression Represents a convolutional neural network, and Train represents a training function.

[0056] Among them, the recurrent neural network (RNN) is trained using labeled facial image information samples to obtain a trained micro-expression recognition model which can be expressed by the following relationship:

[0057] Model voice =Train(D voice ) (2)

[0059] Among them, D voice Represents a recurrent neural network, Model voice Represents the trained micro-expression recognition model, and Train represents the training function.

[0060] The following relationship can be used to construct a multimodal emotion recognition model using the trained language emotion recognition model 305 and micro-expression recognition model 306:

[0061] Model multimodal =Combine(Model micro-expression ,Model voice ) (3)

[0063] Among them, Model micro-expressionRepresents the language emotion recognition model obtained after training, Model voice Represents the trained micro-expression recognition model, Model multimodal It represents the combined multimodal emotion recognition model, and Combine represents the combination function.

[0064] Figure 4 The second flowchart of the multimodal sentiment analysis method provided in the embodiment of the present application is as follows: Figure 5 For the third flow chart of the multimodal sentiment analysis method provided in the embodiment of this application, please refer to Figure 4 and Figure 5 As shown, the user's multimodal information is input into the trained multimodal emotion recognition model to obtain the emotional state information corresponding to the multimodal information, specifically including:

[0065] S401: Inputting first speech information into a first feature extraction model to perform feature extraction on the first speech information to obtain emotion representation vocabulary included in the first speech information.

[0066] Among them, feature extraction is performed on the user's voice information, and the user's emotional representation vocabulary that can be obtained includes intonation, tone, volume, speaking speed, anxiety, impatience and other words that express emotions.

[0067] S402: Input the first facial image information into a second feature extraction model, perform feature extraction on the first facial image information, and obtain micro-expression information included in the first facial image information.

[0068] The micro-expression information obtained by extracting features from the first facial image information mainly includes feature point information that can represent the user's micro-expression. The user's micro-expression type is determined based on this feature point information. The micro-expression type can include micro-expression features corresponding to the following various emotion tags:

[0069] 1. Happy: The corners of the mouth turn up, the cheeks rise, the eyelids shrink, and "crow's feet" form at the tail of the eyes;

[0070] 2. Sadness: squinting eyes, tightening eyebrows, lowering the corners of the mouth, and raising or tightening the chin;

[0071] 3. Anger: eyebrows droop, forehead wrinkled, eyelids and lips tense;

[0072] 4. Disgust: sniffing, raising the upper lip, drooping eyebrows, squinting, sometimes including rolling the eyes, drooping the corners of the eyes, twitching the muscles in the philtrum of the upper lip, and looking down;

[0073] 5. Surprise: jaw dropped, lips and mouth relaxed, eyes wide open, eyelids and eyebrows slightly raised.

[0074] S403: Input the emotion representation vocabulary into the trained language emotion recognition model to obtain first emotion state information corresponding to the emotion representation vocabulary.

[0075] For example, if the emotion representation vocabulary input into the trained language emotion recognition model 305 is rapid speech with a raised tone, the output first emotion state information may be impatience or irritability.

[0076] S404: Input the micro-expression information into the trained micro-expression recognition model to obtain second emotional state information corresponding to the micro-expression information.

[0077] For example, if the micro-expression information input into the trained micro-expression recognition model 306 is drooping eyebrows, wrinkled forehead, and tense eyelids and lips, the output second emotional state information may be anger.

[0078] For further information, please see Figure 3 As shown, the multimodal emotion recognition model provided in this embodiment also includes an information fusion model 307; the output first emotion state information and the second emotion state information are input into the information fusion model 307, and the information fusion model 307 is used to fuse the first emotion state information and the second emotion state information to obtain the user's emotion state information.

[0079] For example, the first emotional state information output by the language emotion recognition model 305 is impatience, and the second emotional state information output by the micro-expression recognition model 306 is anger. The information fusion model 307 then fuses the first emotional state information and the second emotional state information to obtain the user's emotional state information as angry and impatient.

[0080] In one embodiment, see Figure 3 As shown, the multimodal emotion recognition model 300 also includes a first preprocessing module 301 and a second preprocessing module 302.

[0081] The first speech information is input into the first preprocessing module 301, which is used to perform a first preprocessing on the first speech information and input the first speech information after the first preprocessing into the first feature extraction model 303; the first preprocessing method may include at least one of silence removal, denoising processing, standardization processing and speech enhancement processing.

[0082] The first facial image information is input into a second preprocessing module 302, which is used to perform a second preprocessing on the first facial image information and input the first facial image information after the second preprocessing into a second feature extraction model; the second preprocessing method includes at least one of image normalization, image enhancement, posture correction, illumination standardization, face detection, face cropping and background removal.

[0083] By preprocessing the first voice information and the first facial image information as described above, the proposed first emotional state information and the second emotional state information are made more accurate.

[0084] In one embodiment, the denoising process is performed on the first voice information, which can be expressed by the following expression:

[0085] s clean =s original -s noise (4)

[0087] Among them, s original represents the first speech information before denoising, s noise represents the doped noise, s clean Indicates the first speech information after denoising.

[0088] In one embodiment, the speech enhancement process is performed on the first speech information after denoising, which can be expressed by the following expression:

[0089]

[0090] Among them, s background Represents background voice information, s clean Represents the first speech information after denoising, s enhanced Indicates the enhanced first voice information.

[0091] In one embodiment, feature extraction is performed on the first speech information after speech enhancement processing, which can be expressed by the following expression:

[0092] f MFCC =MFCC(s enhanced ) (6)

[0094] Among them, f MFCC It represents feature information extracted from the first speech information, and the MFCC (Mel Frequency Cepstral Coefficients) function represents feature extraction.

[0095] In one embodiment, performing face detection on the first facial image information can be expressed using the following expression:

[0096] Face = DetectFace(I) (7)

[0098] Wherein, Face represents the facial feature information obtained by detection, I represents the first facial image information, and the DetectFace function represents face detection.

[0099] In one embodiment, the following expression can be used to represent the location of feature points of the detected facial feature information:

[0100] P = LocateLandmarks(Face) (8)

[0102] Wherein, P represents the facial feature point information located in the first facial image information in this embodiment, and LocateLandmarks represents the feature point location function.

[0103] In one embodiment, performing image enhancement processing on the first facial image information can be expressed using the following expression:

[0104] I enhanced =Enhance(I) (9)

[0106] Among them, I enhanced represents the first facial image information after enhancement processing, and Enhance represents the image enhancement function.

[0107] It can be seen that in the embodiment of the present application, the user's voice information and the user's facial information are fused and analyzed through a multimodal emotion recognition model, so that the obtained user's emotional state information is more accurate, thereby improving the service quality and interaction efficiency when conducting human-computer interaction with the user based on the emotional state information.

[0108] The following provides an intelligent question-and-answer method based on emotion recognition. This method can be applied to a variety of human-computer interaction scenarios, especially in point-of-sale service scenarios, such as self-service machines, self-service business processing machines, or intelligent customer service based on intelligent agents. Some technical features of the intelligent question-and-answer method based on emotion recognition provided below can refer to the multimodal emotion analysis methods provided in the above embodiments and will not be repeated below.

[0109] Figure 6 For a flow chart of the intelligent question-answering method based on emotion recognition provided in the embodiment of this application, please refer to Figure 6 As shown, the intelligent question-answering method based on emotion recognition specifically includes:

[0110] S601. During a question-and-answer interaction with a user, obtain multimodal information of the user, where the multimodal information includes first voice information and first facial image information of the user.

[0111] S602. Input the user's multimodal information into the trained multimodal emotion recognition model to obtain the emotion state information corresponding to the multimodal information; wherein the multimodal emotion recognition model is trained based on a multimodal information sample set, and the multimodal information sample set includes voice information samples and facial image information samples, as well as emotion state information labels corresponding to the voice information samples and facial image information samples; the emotion state information labels are used to reflect the user's emotion state.

[0112] S603. Adjust the question-answering strategy for the user according to the user's emotional state information; the question-answering strategy includes the tone and wording of the question-answering.

[0113] For example, if the system identifies a customer expressing dissatisfaction or anxiety, it can automatically adjust the tone and wording of its response, or transfer the call to a human service representative for comfort and explanation. The system can also analyze the customer's emotional preferences and service needs based on their historical data, providing targeted service recommendations to enhance the customer's service experience.

[0114] The question-answering strategy for the user is adjusted based on the identified user's emotional state information, which can be specifically expressed by the following expression:

[0115] S=AdjustStrategy(E) (10)

[0117] Among them, E represents the question-answering strategy before adjustment, S represents the question-answering strategy before adjustment, and Adjuststrategy represents the question-answering strategy adjustment function.

[0118] In one embodiment, during the operation of the above-mentioned intelligent question-answering method based on emotion recognition, user feedback data can be continuously collected and analyzed, and the multimodal emotion recognition model 300 can be regularly trained and updated based on the collected and analyzed user feedback data, so that the emotion information recognition accuracy of the multimodal emotion recognition model 300 is higher.

[0119] In one embodiment, during the operation of the above-mentioned intelligent question-answering method based on emotion recognition, a new algorithm model or model training method can also be introduced to optimize the working efficiency and emotion information recognition accuracy of the multimodal emotion recognition model.

[0120] According to the intelligent question-answering method based on emotion recognition provided in the embodiment of the present application, the user's voice information and the user's facial information are fused and analyzed through a multimodal emotion recognition model, so that the obtained user's emotional state information is more accurate. Therefore, when conducting human-computer interaction with the user based on the emotional state information, the question-answering strategy for the user can be adjusted to improve service quality and interaction efficiency.

[0121] The following provides an intelligent question-and-answer device that can be applied in a variety of human-computer interaction scenarios, particularly in point-of-sale service scenarios, such as self-service machines, self-service business processing machines, or intelligent customer service based on intelligent agents. Some technical features of the intelligent question-and-answer method based on emotion recognition provided below can be referenced to the intelligent question-and-answer method based on emotion recognition provided in the above embodiments and will not be further described below.

[0122] Figure 7 For a schematic diagram of the structure of the intelligent question-answering device provided in the embodiment of this application, please refer to Figure 7 As shown, the intelligent question-answering device includes:

[0123] The information collection module 701 is used to obtain the user's multimodal information during the question-and-answer interaction with the user, where the multimodal information includes the user's first voice information and first facial image information.

[0124] The emotional state recognition module 702 is used to input the user's multimodal information into the trained multimodal emotion recognition model to obtain the emotional state information corresponding to the multimodal information; wherein the multimodal emotion recognition model is trained based on a multimodal information sample set, and the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and the facial image information samples; the emotional state information labels are used to reflect the user's emotional state.

[0125] The intelligent question-answering module 703 is used to adjust the question-answering strategy for the user based on the user's emotional state information; the question-answering strategy includes the tone and wording of the question and answer.

[0126] In one embodiment, the multimodal emotion recognition model includes a first feature extraction model, a second feature extraction model, a language emotion recognition model, and a micro-expression recognition model.

[0127] Among them, the user's multimodal information is input into the trained multimodal emotion recognition model to obtain the emotion state information corresponding to the multimodal information, specifically including: inputting the first voice information into the first feature extraction model to perform feature extraction on the first voice information to obtain the emotion representation vocabulary included in the first voice information; inputting the first facial image information into the second feature extraction model to perform feature extraction on the first facial image information to obtain the micro-expression information included in the first facial image information; inputting the emotion representation vocabulary into the trained language emotion recognition model to obtain the first emotion state information corresponding to the emotion representation vocabulary; inputting the micro-expression information into the trained micro-expression recognition model to obtain the second emotion state information corresponding to the micro-expression information; inputting the first emotion state information and the second emotion state information into the information fusion model, the information fusion model is used to perform information fusion on the first emotion state information and the second emotion state information to obtain the user's emotion state information.

[0128] According to the intelligent question-and-answer device based on emotion recognition provided in the embodiment of the present application, the user's voice information and the user's facial information are fused and analyzed through a multimodal emotion recognition model, so that the obtained user's emotional state information is more accurate. Therefore, when conducting human-computer interaction with the user based on the emotional state information, the question-and-answer strategy for the user can be adjusted to improve service quality and interaction efficiency.

[0129] An embodiment of the present invention also provides a computer program product, which includes a computer program, which, when executed by a processor, implements the multimodal sentiment analysis method provided in the above embodiments; or, when executed by a processor, implements the intelligent question-answering method based on emotion recognition provided in the above embodiments.

[0130] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer program products, which can include one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] Figure 8 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, with reference to Figure 8 , Figure 8 The electronic device 12 shown is only an example and should not limit the functions and scope of use of the embodiments of the present application. Figure 8 As shown, electronic device 12 is implemented as a general-purpose computing device. Components of electronic device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).

[0132] Bus 18 represents at least one of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of such architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0133] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0134] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 8 Not shown, often called a "hard drive"). Although Figure 8 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present application.

[0135] A program / utility 40 having a set (at least one) of program modules 46 may be stored, for example, in memory 28. Such program modules 46 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 46 generally implement the functions and / or methods of the embodiments described herein.

[0136] The electronic device 12 may also communicate with one or more external devices 14 (e.g., a keyboard, a pointing device, a display 24, etc.), one or more devices that enable a user to interact with the electronic device 12, and / or any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed through an input / output (I / O) interface 22. Furthermore, the electronic device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the electronic device 12 via the bus 18. It should be understood that although Figure 8 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0137] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the multimodal sentiment analysis method or the intelligent question-answering method based on sentiment recognition provided in the embodiment of the present invention. The multimodal sentiment analysis method includes: obtaining multimodal information of the user, the multimodal information including the user's first voice information and first facial image information; inputting the user's multimodal information into a trained multimodal sentiment recognition model to obtain the sentiment state information corresponding to the multimodal information; wherein the multimodal sentiment recognition model is trained based on a multimodal information sample set, the multimodal information sample set including voice information samples and facial image information samples, and sentiment state information labels corresponding to the voice information samples and facial image information samples; the sentiment state information labels are used to reflect the user's sentiment state. Among them, the intelligent question and answer method based on emotion recognition includes: obtaining the user's multimodal information during the question and answer interaction with the user, the multimodal information including the user's first voice information and first facial image information; inputting the user's multimodal information into the trained multimodal emotion recognition model to obtain the emotional state information corresponding to the multimodal information; adjusting the question and answer strategy for the user according to the user's emotional state information; the question and answer strategy includes the tone of the question and answer and the wording of the question and answer; wherein, the multimodal emotion recognition model is trained based on a multimodal information sample set, the multimodal information sample set including voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and facial image information samples; the emotional state information labels are used to reflect the user's emotional state.

[0138] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a multimodal sentiment analysis method or an intelligent question-answering method based on emotion recognition as provided in all the embodiments of the present invention: wherein the multimodal sentiment analysis method includes: obtaining multimodal information of a user, the multimodal information including the user's first voice information and first facial image information; inputting the user's multimodal information into a trained multimodal emotion recognition model to obtain emotion state information corresponding to the multimodal information; wherein the multimodal emotion recognition model is trained based on a multimodal information sample set, the multimodal information sample set including voice information samples and facial image information samples, and emotion state information labels corresponding to the voice information samples and facial image information samples; the emotion state information labels are used to reflect the user's emotional state. The intelligent question-answering method based on emotion recognition includes: obtaining multimodal information from the user during a question-answering interaction with the user, the multimodal information including the user's first voice information and first facial image information; inputting the user's multimodal information into a trained multimodal emotion recognition model to obtain emotional state information corresponding to the multimodal information; and adjusting the question-answering strategy for the user based on the user's emotional state information; the question-answering strategy includes the tone and wording of the question and answer; wherein the multimodal emotion recognition model is trained based on a multimodal information sample set, the multimodal information sample set including voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and facial image information samples; the emotional state information labels are used to reflect the user's emotional state. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic device, apparatus, or component that is electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction-executing electronic device, apparatus, or device.

[0139] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction-executing electronic device, apparatus, or device.

[0140] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0141] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0142] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will appreciate that the present invention is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions are possible for those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present invention. The scope of the present invention is determined by the scope of the appended claims.

Claims

1. A multimodal sentiment analysis method, characterized in that: The method comprises: Acquiring multimodal information of a user, the multimodal information including first voice information and first facial image information of the user; Inputting the user's multimodal information into a trained multimodal emotion recognition model to obtain emotional state information corresponding to the multimodal information; In which, the multimodal emotion recognition model is trained based on a multimodal information sample set, wherein the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and the facial image information samples; the emotional state information labels are used to reflect the user's emotional state.

2. The multimodal sentiment analysis method according to claim 1, characterized in that The multimodal emotion recognition model includes a first feature extraction model, a second feature extraction model, a language emotion recognition model and a micro-expression recognition model; The step of inputting the multimodal information of the user into the trained multimodal emotion recognition model to obtain the emotion state information corresponding to the multimodal information includes: Inputting the first speech information into the first feature extraction model to perform feature extraction on the first speech information to obtain emotion representation vocabulary included in the first speech information; Inputting the first facial image information into the second feature extraction model, performing feature extraction on the first facial image information, and obtaining micro-expression information included in the first facial image information; Inputting the emotion representation vocabulary into the trained language emotion recognition model to obtain first emotion state information corresponding to the emotion representation vocabulary; Inputting the micro-expression information into a trained micro-expression recognition model to obtain second emotional state information corresponding to the micro-expression information; The user's emotional state information includes the first emotional state information and the second emotional state information.

3. The multimodal sentiment analysis method according to claim 2, characterized in that: The multimodal emotion recognition model also includes an information fusion model; The step of inputting the multimodal information of the user into a trained multimodal emotion recognition model to obtain the emotional state information corresponding to the multimodal information further includes: The first emotional state information and the second emotional state information are input into the information fusion model, and the information fusion model is used to fuse the first emotional state information and the second emotional state information to obtain the emotional state information of the user.

4. The multimodal sentiment analysis method according to claim 2, characterized in that: The multimodal emotion recognition model further includes a first preprocessing module and a second preprocessing module; The step of inputting the multimodal information of the user into the trained multimodal emotion recognition model to obtain the emotion state information corresponding to the multimodal information further includes: Inputting the first speech information into the first preprocessing module, the first preprocessing module being configured to perform a first preprocessing on the first speech information and inputting the first speech information after the first preprocessing into the first feature extraction model; the first preprocessing comprising at least one of silence removal, denoising, normalization, and speech enhancement; The first facial image information is input into the second preprocessing module, and the second preprocessing module is used to perform a second preprocessing on the first facial image information, and input the first facial image information after the second preprocessing into the second feature extraction model; the second preprocessing includes at least one of image normalization, image enhancement, posture correction, illumination standardization, face detection, face cropping and background removal.

5. An intelligent question-answering method based on emotion recognition, characterized in that: The method comprises: During the question-and-answer interaction with the user, obtaining multimodal information of the user, the multimodal information including first voice information and first facial image information of the user; Inputting the user's multimodal information into a trained multimodal emotion recognition model to obtain emotional state information corresponding to the multimodal information; Adjusting the question-answering strategy for the user based on the user's emotional state information; the question-answering strategy includes the tone and wording of the question-answering; In which, the multimodal emotion recognition model is trained based on a multimodal information sample set, wherein the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and the facial image information samples; the emotional state information labels are used to reflect the user's emotional state.

6. The intelligent question-answering method based on emotion recognition according to claim 5, characterized in that: The multimodal emotion recognition model includes a first feature extraction model, a second feature extraction model, a language emotion recognition model and a micro-expression recognition model; The step of inputting the multimodal information of the user into the trained multimodal emotion recognition model to obtain the emotion state information corresponding to the multimodal information includes: Inputting the first speech information into the first feature extraction model to perform feature extraction on the first speech information to obtain emotion representation vocabulary included in the first speech information; Inputting the first facial image information into the second feature extraction model, performing feature extraction on the first facial image information, and obtaining micro-expression information included in the first facial image information; Inputting the emotion representation vocabulary into the trained language emotion recognition model to obtain first emotion state information corresponding to the emotion representation vocabulary; The micro-expression information is input into a trained micro-expression recognition model to obtain second emotional state information corresponding to the micro-expression information.

7. An intelligent question-answering device, characterized in that: The device comprises: An information collection module, configured to obtain multimodal information of the user during a question-and-answer interaction with the user, the multimodal information including the user's first voice information and first facial image information; An emotional state recognition module, configured to input the user's multimodal information into a trained multimodal emotion recognition model to obtain emotional state information corresponding to the multimodal information; An intelligent question-answering module, configured to adjust a question-answering strategy for a user based on the user's emotional state information; the question-answering strategy includes the tone and wording of the question and answer; In which, the multimodal emotion recognition model is trained based on a multimodal information sample set, wherein the multimodal information sample set includes voice information samples and facial image information samples, as well as emotional state information labels corresponding to the voice information samples and the facial image information samples; the emotional state information labels are used to reflect the user's emotional state.

8. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the multimodal sentiment analysis method as described in any one of claims 1 to 4; or implement the intelligent question-answering method based on emotion recognition as described in claim 5 or 6.

9. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the multimodal sentiment analysis method as described in any one of claims 1 to 4; or implements the intelligent question-answering method based on sentiment recognition as described in claim 5 or 6.

10. A computer program product comprising a computer program, characterized in that When executed by a processor, the computer program implements the multimodal sentiment analysis method as described in any one of claims 1 to 4; or implements the intelligent question-answering method based on sentiment recognition as described in claim 5 or 6.