Emotion recognition method and emotion analysis system

Through an image acquisition device integrating infrared and color cameras, a multimodal emotion recognition method combining facial and hand images and voice data, a lightweight neural network processing is used to solve the problem of low accuracy in emotional recognition among traffic participants, and improve traffic safety and human-vehicle interaction experience.

CN120472435AInactive Publication Date: 2025-08-12ZHEJIANG LEAPMOTOR TECH CO LTD +1

Patent Information

Application Number
CN202510987501.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the accuracy of emotion recognition of traffic participants is low, especially in dynamic environments, and the accuracy of facial expression recognition is low, and the multimodal fusion is insufficient, resulting in an increase in traffic safety threat.

Method used

An image acquisition device integrating infrared cameras and color cameras is adopted, combining facial and hand image data and voice data, multi-modal emotion recognition is performed through neural network models, and a lightweight neural network after knowledge distillation is used for real-time processing, and multiple emotion recognition results are fused to improve accuracy.

Benefits of technology

It significantly improves the accuracy of emotional recognition, reduces misjudgment in dynamic environments, and enhances traffic safety and human-vehicle interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472435A_ABST
    Figure CN120472435A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an emotion recognition method and an emotion analysis system, and relates to the technical field of emotion analysis. The method comprises the following steps: acquiring image data and voice data of a to-be-detected target, wherein the image data comprises a first image; a part image of the to-be-measured target is extracted from the first image, the part image comprises a first part, and the first part is at least part of the exposed part of the to-be-measured target; performing emotion recognition based on the part image by using a first model to obtain a first recognition result; performing emotion recognition based on the voice data by using a second model to obtain a second recognition result; and determining the target emotion of the to-be-detected target based on the first recognition result and the second recognition result. Therefore, the emotion recognition accuracy can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of sentiment analysis technology, and in particular to an emotion recognition method and a sentiment analysis system. Background Art

[0002] Traffic safety is gaining increasing attention, and the emotional state of traffic participants is a key factor influencing traffic safety. Emotions that impact traffic safety primarily refer to negative emotions caused by driving situations like traffic congestion. These negative emotions can lead to traffic accidents, posing a threat to both traffic safety and social security. Therefore, accurately detecting the emotional state of traffic participants is crucial for improving traffic safety. Summary of the Invention

[0003] The embodiments of the present application provide an emotion recognition method and a sentiment analysis system to improve the accuracy of emotion recognition.

[0004] In a first aspect, an embodiment of the present application provides an emotion recognition method, comprising: Acquiring image data and voice data of a target to be measured, where the image data includes a first image; Extracting a part image of the object to be measured from the first image, the part image including a first part, the first part being at least a portion of an exposed part of the object to be measured; Using the first model, emotion recognition is performed based on the part image to obtain a first recognition result; Using the second model, emotion recognition is performed based on the voice data to obtain a second recognition result; Based on the first recognition result and the second recognition result, a target emotion of the target to be measured is determined.

[0005] In one embodiment, the image data is collected within a preset time period and further includes a plurality of second images collected before the first image; The above-mentioned emotion recognition method also includes: Determining an image sequence based on the plurality of second images and the first image, wherein the images in the image sequence are arranged in order of acquisition time from earliest to latest; Using the third model, emotion recognition is performed based on the image sequence to obtain a third recognition result; Determining a target emotion of the target to be measured based on the first recognition result and the second recognition result includes: Based on the first recognition result, the second recognition result, and the third recognition result, a target emotion of the target to be measured is determined.

[0006] In one embodiment, the above-mentioned emotion recognition method is applied to a sentiment analysis system, which includes an image acquisition device, which integrates an infrared camera and a color camera; The first image includes a first infrared image captured by the infrared camera and a first color image captured by the color camera, and the first infrared image and the first color image are captured at the same time; Each second image includes a second color image captured by a color camera; Determining an image sequence based on the plurality of second images and the first image includes: Based on the respective second color images and the first color image, an image sequence is determined.

[0007] In one embodiment, the first recognition result, the second recognition result, and the third recognition result all include recognized emotion labels and corresponding first probability values, and the first model, the second model, and the third model are all configured with weight values; Determining a target emotion of the target to be measured based on the first recognition result, the second recognition result, and the third recognition result, including: In response to at least some of the emotion labels being different, normalizing the first probability values based on the weight values to generate second probability values corresponding to the emotion labels; The emotion indicated by the emotion label corresponding to the maximum value among the second probability values is determined as the target emotion of the target to be measured.

[0008] In one embodiment, normalizing each first probability value based on each weight value to generate a second probability value corresponding to each emotion tag includes: Performing weighted summation on each first probability value based on each weight value to obtain a target sum value; For each emotion tag, the product of the first probability value corresponding to the emotion tag and the associated weight value is determined, and the ratio of the product to the target sum value is determined as the second probability value corresponding to the emotion tag.

[0009] In one embodiment, the above-mentioned emotion recognition method is applied to a sentiment analysis system, and the sentiment analysis system includes a neural network processing unit; The first model is a neural network after knowledge distillation and is deployed in a neural network processing unit.

[0010] In one embodiment, the emotion recognition method further includes: Extract features from speech data to generate speech features; Using the second model, emotion recognition is performed based on speech data, including: The second model is used to perform emotion recognition based on speech features.

[0011] In one embodiment, the emotion recognition method further includes: Searching for a target interaction strategy including a target emotion among a plurality of preset interaction strategies; wherein the plurality of interaction strategies include interaction system information and different emotion information; In response to finding the target interaction strategy, an interaction instruction is generated based on the target interaction strategy, and the interaction instruction is sent to the interaction system indicated by the interaction system information in the target interaction strategy.

[0012] In one embodiment, the part image further includes a second part, which is at least a portion of an exposed part of the object to be measured; The above-mentioned emotion recognition method also includes: Extracting key points of the first part from the image of the first part, and generating first part features including the extracted corresponding key points; Extracting second part key points from the image of the second part, and generating second part features including the extracted corresponding key points; Using the first model, emotion recognition is performed based on body part images, including: The first model is used to perform emotion recognition based on the fusion result of the first part feature and the second part feature.

[0013] In a second aspect, an embodiment of the present application provides a sentiment analysis system, comprising: an acquiring unit configured to acquire image data and voice data of a target to be measured, wherein the image data includes a first image; an extraction unit configured to extract a part image from the first image, where the part image includes a first part, and the first part is at least a portion of an exposed part of the object to be measured; a first recognition unit configured to perform emotion recognition based on the part image using a first model to obtain a first recognition result; a second recognition unit configured to perform emotion recognition based on the voice data using a second model to obtain a second recognition result; The determination unit is configured to determine the target emotion of the target to be measured based on the first recognition result and the second recognition result.

[0014] In the solution provided in the embodiment of the present application, image data and voice data of the target to be measured can be obtained, and then a part image can be extracted from the first image included in the image data, the part image including a first part, the first part being at least part of the exposed part of the target to be measured. Then, using a first model, emotion recognition is performed based on the part image to obtain a first recognition result, and using a second model, emotion recognition is performed based on the voice data to obtain a second recognition result. Then, based on the first recognition result and the second recognition result, the target emotion of the target to be measured is determined. In this way, emotion recognition can be performed based on multimodal data of the target to be measured (such as the first part and voice, etc.), and the target emotion of the target to be measured can be determined by combining multiple emotion recognition results, thereby effectively improving the accuracy of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings will make the technical solutions and other beneficial effects of the present application apparent.

[0016] Figure 1 This is a flow chart of the emotion recognition method in the embodiment of the present application; Figure 2 is another flow chart of the emotion recognition method in an embodiment of the present application; Figure 3 is a schematic diagram of an application scenario of the emotion recognition method in an embodiment of the present application; Figure 4 Schematic diagram of the structure of the sentiment analysis system in the embodiment of the present application; Figure 5 This is a schematic diagram of the causal chain derived from the accuracy improvement; Figure 6 It is a schematic diagram of the causal chain derived from real-time improvement; Figure 7 This is a schematic diagram of the causal chain of environmental robustness derived from adaptability improvement.

[0017] Reference numerals: 401 - acquisition unit, 402 - extraction unit, 403 - first recognition unit, 404 - second recognition unit, 405 - determination unit, 406 - third recognition unit, 407 - image acquisition device, 408 - microphone array. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0019] In the description of this application, it should be noted that, unless otherwise specified or limited, the term "and / or" herein is merely a description of an association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the character " / " herein, unless otherwise specified, generally indicates that the associated objects are in an "or" relationship.

[0020] As mentioned above, accurately detecting the emotional state of traffic participants is crucial for improving traffic safety. Related technologies typically rely solely on facial expressions for emotion recognition, resulting in low accuracy. Furthermore, these technologies suffer from insufficient multimodal fusion. For example, they rely solely on facial expressions, ignoring the collaborative emotional signals from body language (such as gestures) and in-vehicle voices. They also process facial, voice, and gesture signals independently, lacking spatiotemporal alignment (e.g., in scenarios where angry voices but calm expressions are present, collaborative judgment is impossible). The embodiments of the present application provide an emotion recognition method and a sentiment analysis system, which can perform emotion recognition based on multimodal data of the target to be measured, and determine the target emotion of the target to be measured by combining multiple emotion recognition results, so as to effectively improve the accuracy of emotion recognition.

[0021] In some embodiments, the present application is applied to an in-vehicle emotion analysis system. In this case, traffic participants are mainly vehicle drivers, vehicle auxiliary control personnel, vehicle coordinators and other people in the vehicle.

[0022] In some embodiments, the present application is applied to aircraft or other transportation equipment.

[0023] In some embodiments, the target to be detected is a traffic participant, which may include a driver, a traffic assistant, a traffic security personnel, etc.

[0024] In some embodiments, the first part includes the face of the traffic participant. Generally, the face is at least part of the exposed part. In this application, the face can be used as the exposed part as long as the part image can be collected.

[0025] In some embodiments, the second part includes the hands of the traffic participant. Generally, the hands are at least part of the exposed parts. In this application, the hands can be used as the exposed parts as long as the part image can be collected.

[0026] Figure 1 This is a flow chart of the emotion recognition method in the embodiment of the present application. The method can be executed by the sentiment analysis system. The method includes the following steps: S101: Acquire image data and voice data of a target to be measured, where the image data includes a first image; S103: extracting a part image of the object to be measured from the first image, where the part image includes a first part, and the first part is at least a portion of an exposed part of the object to be measured; S105: Using the first model, performing emotion recognition based on the body part image to obtain a first recognition result; S107: Using the second model, performing emotion recognition based on the voice data to obtain a second recognition result; S109: Determine the target emotion of the target to be measured based on the first recognition result and the second recognition result.

[0027] exist Figure 1 In the solution provided by the corresponding embodiment, image data and voice data of the target to be measured can be obtained, and then a part image of the target to be measured is extracted from the first image included in the image data, the part image including a first part, which is at least part of the exposed part of the target to be measured. Then, using a first model, emotion recognition is performed based on the part image to obtain a first recognition result, and using a second model, emotion recognition is performed based on the voice data to obtain a second recognition result. Then, based on the first recognition result and the second recognition result, the target emotion of the target to be measured is determined. In this way, emotion recognition can be performed based on multimodal data of the target to be measured (such as the first part and voice, etc.), and the target emotion of the target to be measured can be determined by combining multiple emotion recognition results, thereby effectively improving the accuracy of emotion recognition.

[0028] The explanation of the target to be measured, the first part, and the emotion analysis system can be referred to the relevant description in the previous text, which will not be repeated here. Below, taking the target to be measured as a driver and the first part as the face as an example, the above steps S101 to S109 are explained.

[0029] In step S101, image data and voice data of the driver are acquired, where the image data includes a first image. The emotion analysis system includes an image acquisition device and a voice acquisition device. The image data is acquired using the image acquisition device, and the voice data is acquired using the voice acquisition device. Furthermore, to avoid interference and thereby improve the accuracy of driver emotion recognition, both the image acquisition device and the voice acquisition device are directional acquisition devices. That is, the image acquisition device only acquires images of the main passenger area, and the voice acquisition device only acquires voice from the main passenger area. The voice acquisition device includes, but is not limited to, a microphone array.

[0030] Currently, related technologies suffer from the following limitations: Lighting changes and driver posture interference can lead to low facial expression recognition accuracy (e.g., squinting in bright light can be misinterpreted as anger). Extreme lighting (e.g., sudden changes in brightness when entering or exiting a tunnel) and occlusion (e.g., from masks or sunglasses) can cause facial features to fail, limiting recognition accuracy. To minimize this limitation, in one embodiment, the image acquisition device in the embodiments of the present application can be an integrated infrared camera and a color camera. An infrared camera is a visual device that uses infrared radiation for imaging, independent of visible light, and can capture scenes in complete darkness or low light. The color camera can include an RGB camera. An RGB camera is an optical electronic device that captures color images by capturing light signals from the three primary colors of red, green, and blue. When an image acquisition device integrates both an infrared camera and an RGB camera, the device can be referred to as an infrared-RGB dual-mode camera.

[0031] In addition, when the image acquisition device integrates an infrared camera and a color camera, the first image includes a first infrared image acquired by the infrared camera and a first color image acquired by the color camera, and the first infrared image and the first color image are acquired at the same time. Specifically, when the color camera is an RGB camera, the first color image is an RGB image.

[0032] In step S103, a driver's part image is extracted from the first image, where the part image includes the face. It should be understood that the part image specifically includes the face image. Furthermore, when the first image includes a first infrared image and a first color image, the driver's part image can be extracted from each of the first infrared image and the first color image for use in emotion recognition. This allows for enhanced robustness by fusing the infrared image with the color image, compensating for facial occlusions in the color image using the infrared image, and avoiding misjudgments due to light sensitivity, thereby improving emotion recognition accuracy.

[0033] In one embodiment, the above-mentioned part image also includes a second part, and the second part is at least a part of the exposed part of the driver. The second part may include a hand, for example, and the above-mentioned part image specifically includes a facial image and a hand image. Below, taking the second part as the hand as an example, the solution provided by the embodiment of the present application will be further introduced. It should be noted that by extracting the driver's facial image and hand image from the first image, and using the driver's facial image, hand image and voice data for emotion recognition, it can avoid missing key emotions and significantly improve the accuracy of emotion recognition.

[0034] Taking the example of a first image comprising a first infrared image and a first color image, in one embodiment, the driver's facial image and hand image can be extracted from the first infrared image and the first color image, respectively, for use in emotion recognition. This allows the fusion of infrared and color images to enhance robustness and avoid misjudgments caused by illumination sensitivity. Furthermore, by combining multimodal driver data (such as face, gestures, and voice) for emotion recognition, key emotions can be avoided from being missed, significantly improving emotion recognition accuracy.

[0035] In step S105, emotion recognition is performed based on the part image using the first model to obtain a first recognition result. The first recognition result includes an identified emotion label and a first probability value corresponding to the emotion label. In practice, the first model can correspond to multiple emotion labels. The first model can be used to predict the probability value of the driver under each of the multiple emotion labels based on the part image, and the maximum probability value is used as the first probability value, and the emotion label corresponding to the first probability value is used as the identified emotion label. Among them, any one of the multiple emotion labels can indicate one of the following emotions: anger, irritation, fatigue, happiness, normal, boredom, etc. It should be understood that the multiple emotion labels can be set according to actual needs and are not specifically limited here.

[0036] The first model can include a graph neural network (GNN). In practice, a GNN is a deep learning model specifically designed to process graph-structured data. Unlike traditional neural networks, GNNs are designed to solve the problem of modeling data in non-Euclidean space.

[0037] In one embodiment, when the first image includes a first infrared image and a first color image, the first model may include a Generative Adversarial Network (GAN). The GAN can reconstruct the eye area obscured by sunglasses using the infrared depth map. Furthermore, the first model may include dynamic HDR (High Dynamic Range) fusion capabilities. HDR addresses the issue of overexposed highlights or dark shadows in high-contrast scenes, restoring the rich details and natural gradations seen by the human eye.

[0038] In one embodiment, the first model is used to model facial-gesture associations (e.g., frown + fist -> anger). For example, when using the first model to perform emotion recognition based on facial and hand images, if the facial image shows a frown and the hand image shows a fist, anger can be recognized.

[0039] In one embodiment, the sentiment analysis system includes a neural network processing unit (NPU), and the first model is a neural network after knowledge distillation (such as MobileNetV3), and is deployed in the neural network processing unit. Among them, the NPU is a hardware core dedicated to AI (Artificial Intelligence) acceleration. The parameter volume of the neural network after knowledge distillation is, for example, less than 1MB (Megabyte). It should be pointed out that the relevant technology usually adopts the traditional CNN (Convolutional Neural Network) model. The traditional CNN model has a large amount of calculation, and it is difficult to achieve low-latency response on the on-board chip, and there is a problem of insufficient real-time performance. Compared with the relevant technology, the solution provided in the embodiment of the present application can realize NPU parallel execution and lightweight reasoning, ensure the real-time performance of the model calculation, and achieve low-latency response.

[0040] MobileNetV3 is an efficient neural network architecture for mobile devices and the third generation of the MobileNet series. MobileNetV3 is designed specifically for the computational limitations of mobile devices and embedded systems, aiming to achieve the best balance between computational cost and model performance by designing an efficient, lightweight network structure. When the first model is a MobileNetV3 after knowledge distillation, MobileNetV3 can be called a student model, and the teacher model corresponding to the student model can be, for example, ResNet101. Among them, ResNet101 is a classic model in the deep residual network (Residual Network, ResNet) series and belongs to the deep learning architecture in the field of computer vision. The core innovation of ResNet101 is that it solves the gradient vanishing problem in deep neural network training through "residual connection", enabling the model to build an extremely deep network structure (such as 101 layers) while maintaining good performance.

[0041] It's important to note that when performing model compression based on distillation loss (for example, using ResNet101 as the teacher model and MobileNetV3 as the student model), dynamic weight quantization (FP32 → INT8) is performed to convert the high-precision floating-point model to a low-precision integer model while preserving model performance as much as possible. FP32 represents 32-bit floating-point numbers, and INT8 represents 8-bit integer numbers.

[0042] When the aforementioned part image includes a facial image, in step S105, emotion recognition can be performed based on the facial image using the first model to obtain a first recognition result. Furthermore, facial key points can be extracted from the facial image to generate facial features including the extracted facial key points. Emotion recognition can then be performed based on the facial features using the first model. Furthermore, to ensure the accuracy of feature extraction, the facial image can be preprocessed before facial key point extraction is performed on the preprocessed facial image. This preprocessing includes, but is not limited to, alignment and normalization. In one embodiment, the facial image can be preprocessed using an ISP (Image Signal Processor). In practice, an ISP typically involves a series of operations applied to the raw data captured by an image sensor in a digital camera or other device. These operations may include preprocessing, demosaicing, color correction, exposure correction, sharpening, and compression.

[0043] The extracted facial landmarks are typically distributed across key facial regions, including the eyebrows, eyes, nose, mouth, and jawline. Furthermore, facial features also include head posture. As an implementation, a 3D Dense Face Alignment (3DDFA) model can be used to extract facial landmarks from facial images. The 3DDFA model can be located within the first model, where its input consists of a facial image; alternatively, the 3DDFA model can be located outside the first model, where its input consists of facial features.

[0044] When the part image as mentioned above includes a facial image and a hand image, in step S105, the first model can be used to perform emotion recognition based on the facial image and the hand image to obtain a first recognition result. Furthermore, facial key points can be extracted from the facial image to generate facial features including the extracted facial key points, and gesture skeleton points can be extracted from the hand image to generate gesture features including the extracted gesture skeleton points. Then, the first model can be used to perform emotion recognition based on the fusion result of the facial features and the gesture features to obtain a first recognition result. Furthermore, in order to ensure the effectiveness of feature extraction, the facial image and the hand image can be preprocessed separately, such as using ISP to preprocess the facial image and the hand image separately, and then facial key points can be extracted from the preprocessed facial image, and gesture skeleton points can be extracted from the preprocessed hand image. The preprocessing here includes but is not limited to alignment and normalization.

[0045] Among them, for the content related to facial features, please refer to the relevant explanations in the previous article and will not be repeated here. When extracting gesture skeleton points from hand images, for example, the MediaPipe model can be used to track gesture skeleton points in real time, thereby extracting gesture skeleton points from hand images. MediaPipe is an open source cross-platform machine learning framework that focuses on solving real-time multimedia understanding tasks on mobile terminals and edge devices. Through modular design and pre-trained models, MediaPipe supports the rapid development of various visual, audio, and sensor data processing applications. The MediaPipe model can be located inside or outside the first model, and is not specifically limited here.

[0046] For example, when the aforementioned 3DDFA model and MediaPipe model are located inside the first model, the input of the first model includes facial images and hand images. When the 3DDFA model and MediaPipe model are located outside the first model, the input of the first model includes facial features and gesture features.

[0047] In step S107, emotion recognition is performed based on the voice data using the second model to obtain a second recognition result. The second recognition result includes the recognized emotion label and the first probability value corresponding to the emotion label. In practice, the second model can correspond to multiple emotion labels. The second model can be used to predict the probability value of the driver under each of the multiple emotion labels based on the voice data, and the maximum probability value is used as the first probability value, and the emotion label corresponding to the first probability value is used as the recognized emotion label. The emotion labels corresponding to the second model may be the same as or different from the emotion labels corresponding to the first model, and no specific limitation is made here.

[0048] In one embodiment, to improve the accuracy of the second recognition result, the speech data may be preprocessed, and then emotion recognition may be performed based on the preprocessed speech data using the second model. The preprocessing herein includes but is not limited to speech noise reduction processing.

[0049] When using the second model to perform emotion recognition based on speech data (such as preprocessed speech data), as an implementation method, the speech data can be converted into corresponding text, such as using Automatic Speech Recognition (ASR) technology to convert the speech data into corresponding text, and then text emotion features are extracted from the text. Then, the second model is used to perform emotion recognition based on the text emotion features. In this implementation method, the input of the second model can include the text emotion features. In addition, the second model can include any one of the models such as Logistic Regression, Support Vector Machine (SVM), and Random Forest.

[0050] As another implementation, feature extraction can be performed on speech data (e.g., preprocessed speech data) to generate speech features, and then a second model can be used to perform emotion recognition based on the speech features. The speech features may include, but are not limited to, voiceprint features. In this implementation, the input of the second model may include the speech features. Furthermore, the second model may include any of a convolutional neural network, a recurrent neural network (RNN), a transformer, a support vector machine, and the like.

[0051] In step S109 , the target emotion of the driver is determined based on the first recognition result and the second recognition result.

[0052] Specifically, when the emotion labels included in the first recognition result and the second recognition result are the same, the emotion indicated by the emotion label included in either the first recognition result or the second recognition result may be determined as the target emotion of the driver.

[0053] When the first and second recognition results include different emotion labels, each first probability value may be normalized based on the weight values of the first and second models to generate second probability values corresponding to the emotion labels. The emotion indicated by the emotion label corresponding to the maximum value among the second probability values is determined as the driver's target emotion. When emotion recognition is performed using only the first and second models, the sum of the weight values of the first and second models may be equal to 1.

[0054] Furthermore, a weighted sum of the first probability values can be performed based on the weight values to obtain a target sum value. Then, for each emotion label, the product of the first probability value corresponding to the emotion label and the associated weight value is determined, and the ratio of the product to the target sum value is determined as the second probability value corresponding to the emotion label.

[0055] Furthermore, the second probability value corresponding to each emotion label can be determined by the following formula (1): (1) Wherein, j can represent the model for calculating the second probability value corresponding to the currently recognized emotion tag. represents the first probability value corresponding to the emotion label identified by model j, represents the second probability value corresponding to the emotion label identified by model j, Represents the weight value of model j. For any one of the first model and the second model i , Representation Model i The first probability value corresponding to the identified emotion label, Representation Model i The weight value of . i is a natural number within [1, N], where N represents the number of models used for emotion recognition. It should be understood that when only the first model and the second model are used for emotion recognition, N=2.

[0056] In one embodiment, in order to achieve passenger zone detection and thus effectively improve the recognition accuracy of the driver's emotions, the emotion analysis system can determine whether there is a passenger in the driver's seat based on the pressure change of the driver's seat in the vehicle in which it is located, and if it is determined that there is a passenger in the driver's seat, the passenger is determined to be the driver, and then the following steps are performed: Figure 1 The emotion recognition method shown in the figure can include a seat pressure sensor installed on the driver's seat. Based on the information collected by the seat pressure sensor, the pressure change of the driver's seat can be determined. This allows the seat pressure sensor to suppress non-driver signals.

[0057] In one embodiment, the sentiment analysis system may be pre-configured with multiple interaction strategies, each of which includes interaction system information and different emotion information. If the sentiment analysis system is an in-vehicle sentiment analysis system, the interaction system indicated by the interaction system information is located in the same vehicle as the sentiment analysis system. This interaction system may, for example, be an in-vehicle infotainment system, a voice interaction system, an air conditioning system, or a seating system. It should be understood that any two interaction strategies may include different emotion information, but the interaction system information included may be the same or different. Furthermore, the interaction strategies may also include instruction information.

[0058] For example, for any of the aforementioned interaction strategies R, if the emotional information in interaction strategy R includes happiness, the interaction system information in interaction strategy R may indicate an in-vehicle infotainment system. The instruction information in interaction strategy R may, for example, indicate entertainment recommendations, audiobook playback, or music playback. If the emotional information in interaction strategy R includes anger or rage, the interaction system information in interaction strategy R may indicate a voice interaction system. The instruction information in interaction strategy R may, for example, indicate a danger warning. Furthermore, the interaction system may also include the in-vehicle infotainment system. The instruction information may also indicate the playback of soothing music. If the emotional information in interaction strategy R includes a relatively low emotion, such as boredom, the interaction system information in interaction strategy R may indicate a voice interaction system. The instruction information in interaction strategy R may, for example, indicate voice interaction with the driver. The voice interaction system may include an emotion macromodel, and the instruction information may indicate the use of this emotion macromodel for voice interaction. It should be understood that, based on this emotion macromodel, the voice interaction system is able to understand and perceive user emotions and provide corresponding emotional support and services.

[0059] After determining the driver's target emotion, the system searches for a target interaction strategy that includes the target emotion among the aforementioned interaction strategies. If a target interaction strategy is found, interaction instructions are generated based on the target interaction strategy and sent to the interaction system indicated by the interaction system information in the target interaction strategy. This helps soothe the driver's emotions, ensures driving safety, and enhances the driver-vehicle interaction experience.

[0060] Figure 2 This is another flow chart of the emotion recognition method in the embodiment of the present application. The method can be performed by the sentiment analysis system. The method includes the following steps: S201: Acquire image data and voice data of a target to be measured, where the image data is collected within a preset time period and includes a first image and a plurality of second images collected before the first image; S203: extracting a part image of the object to be measured from the first image, where the part image includes a first part, and the first part is at least a portion of an exposed part of the object to be measured; S205: Using the first model, performing emotion recognition based on the body part image to obtain a first recognition result; S207: Using the second model, performing emotion recognition based on the voice data to obtain a second recognition result; S209: Determine an image sequence based on the second images and the first image, and arrange the images in the image sequence in order of acquisition time from earliest to latest; S211: Using the third model, performing emotion recognition based on the image sequence to obtain a third recognition result; S213: Determine the target emotion of the target to be measured based on the first recognition result, the second recognition result, and the third recognition result.

[0061] exist Figure 2 In the corresponding embodiments, by using the first model to perform emotion recognition based on the part image of the target to be measured, using the second model to perform emotion recognition based on the voice data of the target to be measured, and using the third model to perform emotion recognition based on the image sequence of the target to be measured, emotion recognition can be performed based on the multimodal data of the target to be measured, and the target emotion of the target to be measured can be determined by combining multiple emotion recognition results, thereby significantly improving the accuracy of emotion recognition.

[0062] In step S201, the preset time period can be the last 3 seconds, 5 seconds, 10 seconds, or 20 seconds, etc., and can be set according to actual needs and is not specifically limited here. When the sentiment analysis system includes an image acquisition device that integrates an infrared camera and a color camera, the first image includes a first infrared image acquired by the infrared camera and a first color image acquired by the color camera, and the first infrared image and the first color image are acquired at the same time. Each second image includes a second color image acquired by the color camera.

[0063] For a more detailed explanation of step S201 and explanations of steps S203 to S207, please refer to the relevant descriptions in the previous text, which will not be repeated here.

[0064] In step S209, an image sequence may be determined based on each second image and the first image, with each image in the image sequence arranged in order of earliest acquisition time. In one embodiment, each second image may be arranged with the first image in order of earliest acquisition time to form an image sequence. In another embodiment, when the first image includes a first infrared image and a first color image, and each second image includes a second color image, an image sequence may be determined based on each second color image and the first color image, such as by arranging each second color image with the first color image in order of earliest acquisition time to form an image sequence.

[0065] In step S211, emotion recognition is performed based on the image sequence using the third model to obtain a third recognition result. The third recognition result includes the recognized emotion label and the first probability value corresponding to the emotion label. In practice, the third model can correspond to multiple emotion labels. The third model can be used to predict the probability value of the target under each of the multiple emotion labels based on the image sequence, and the maximum probability value is used as the first probability value, and the emotion label corresponding to the first probability value is used as the recognized emotion label. The emotion labels corresponding to the first model, the second model, and the third model can be the same or different, and are not specifically limited here.

[0066] The third model may include, but is not limited to, an LSTM (Long Short-Term Memory) network. LSTM networks process temporal changes. For example, if the subject yawns continuously during image data acquisition, after generating an image sequence based on the image data, the LSTM network can be used to identify an emotion tag indicating fatigue based on the image sequence.

[0067] As an implementation method, feature extraction can be performed on the image sequence to generate an image feature sequence. A third model can then be used to perform emotion recognition based on the image feature sequence to obtain a third recognition result. The image features in the image feature sequence include first part features. Furthermore, the image features may also include second part features. It should be understood that when the first part is the face, the first part features are facial features; and when the second part is the hand, the second part features are gesture features.

[0068] In step S213 , the target emotion of the target to be measured is determined based on the first recognition result, the second recognition result, and the third recognition result.

[0069] Specifically, when the emotion labels included in the first recognition result, the second recognition result, and the third recognition result are the same, the emotion indicated by the emotion label included in any one of the first recognition result, the second recognition result, and the third recognition result can be determined as the target emotion of the target to be measured.

[0070] When at least some of the emotion labels included in the first, second, and third recognition results are different, each first probability value may be normalized based on the weight values of the first, second, and third models to generate second probability values corresponding to each emotion label. The emotion indicated by the emotion label corresponding to the maximum value among the second probability values is determined as the target emotion of the target. The sum of the weight values of the first, second, and third models may be equal to 1.

[0071] Furthermore, a weighted sum of the first probability values can be performed based on the weight values to obtain a target sum value. Then, for each emotion label, the product of the first probability value corresponding to the emotion label and the associated weight value is determined, and the ratio of the product to the target sum value is determined as the second probability value corresponding to the emotion label. Furthermore, the second probability value corresponding to each emotion label can be determined using the aforementioned formula (1).

[0072] Next, we take the emotion analysis system including infrared-RGB dual-mode camera and microphone array, the target to be measured is the driver, and the part images as mentioned above include facial images and hand images as an example. Figure 3 ,right Figure 2 The corresponding embodiments provide examples for illustration. Figure 3 It is a schematic diagram of an application scenario of the emotion recognition method in an embodiment of the present application.

[0073] like Figure 3 As shown, the solution provided by the embodiment of the present application involves multimodal input. Specifically, an infrared-RGB dual-mode camera can be used to collect image data of the driver within a preset time period, and a microphone array can be used to collect voice data of the driver, such as voice data within the preset time period, and the image data and voice data can be preprocessed separately. Specifically, the image data includes a first infrared image and a first RGB image with the same acquisition time, and also includes multiple second RGB images acquired before the first RGB image. When preprocessing the image data, the driver's facial image and hand image can be extracted from the first infrared image and the first RGB image respectively, and the extracted facial image and hand image can be aligned and normalized respectively. In addition, the second RGB images and the first RGB images are formed into an image sequence in the order of acquisition time from early to late. Optionally, each image in the image sequence can also be aligned and normalized. When preprocessing the voice data, the voice data can be subjected to voice noise reduction processing.

[0074] Next, feature extraction can be performed on the preprocessed facial image, hand image, and voice data. For example, a 3DDFA model can be used to extract facial key points from the facial image, generating facial features including the extracted facial key points. The MediaPipe model can be used to track gesture skeleton points in real time, thereby extracting gesture skeleton points from the hand image and generating gesture features including the extracted gesture skeleton points. Alternatively, feature extraction can be performed on an image sequence to generate an image feature sequence, where the image features in the image feature sequence include facial features. Furthermore, the image features can also include gesture features.

[0075] Then, spatiotemporal fusion can be performed based on the various generated features. For example, after obtaining facial features and gesture features based on the preprocessed facial and hand images, emotion recognition can be performed based on the fusion result of the facial and gesture features using the first model described above to obtain a first recognition result. Emotion recognition can be performed based on speech features using the second model to obtain a second recognition result. Emotion recognition can be performed based on the image feature sequence using the third model to obtain a third recognition result. The first, second, and third recognition results each include an identified emotion label and a corresponding first probability value, and the first, second, and third models are each configured with a weight value. Subsequently, in response to at least some of the emotion labels being different, a multimodal weighted score can be calculated. For example, a weighted sum of the first probability values can be performed based on the weight values to obtain a target sum value. Then, for each emotion label, the product of the first probability value corresponding to the emotion label and the associated weight value is determined, and the ratio of this product to the target sum value is determined as the second probability value corresponding to the emotion label.

[0076] Then, an emotional decision can be made based on the multimodal weighted scores, that is, the second probability values. For example, the emotion indicated by the emotion label corresponding to the maximum value among the second probability values can be determined as the driver's target emotion.

[0077] Next, to help soothe the driver's emotions, ensure driving safety, and enhance the driver-vehicle interaction experience, the interaction system can be triggered to respond emotionally. For example, a target interaction strategy that includes the target emotion can be searched among the multiple interaction strategies described above. Upon finding the target interaction strategy, an interaction instruction is generated based on the target interaction strategy and sent to the interaction system indicated by the interaction system information in the target interaction strategy.

[0078] Figure 4 This is a schematic diagram of the structure of the sentiment analysis system in the embodiment of this application. Figure 4 As shown, the sentiment analysis system includes: An acquisition unit 401 is configured to acquire image data and voice data of a target to be measured, where the image data includes a first image; An extraction unit 402 is configured to extract a part image of the target to be measured from the first image, where the part image includes a first part, and the first part is at least a portion of an exposed part of the target to be measured; The first recognition unit 403 is configured to use the first model to perform emotion recognition based on the part image to obtain a first recognition result; The second recognition unit 404 is configured to use the second model to perform emotion recognition based on the voice data to obtain a second recognition result; The determination unit 405 is configured to determine the target emotion of the target to be measured based on the first recognition result and the second recognition result.

[0079] In one embodiment, the image data is collected within a preset time period and further includes a plurality of second images collected before the first image; The above apparatus further includes a third identification unit 406, which is configured to: Determining an image sequence based on the plurality of second images and the first image, wherein the images in the image sequence are arranged in order of acquisition time from earliest to latest; Using the third model, emotion recognition is performed based on the image sequence to obtain a third recognition result; The determination unit 405 is configured to determine the target emotion of the target to be measured based on the first recognition result and the second recognition result, including: The determination unit 405 is configured to determine the target emotion of the target to be measured based on the first recognition result, the second recognition result, and the third recognition result.

[0080] In one embodiment, the sentiment analysis system includes an image acquisition device 407 , which is integrated with an infrared camera and a color camera; The first image includes a first infrared image captured by the infrared camera and a first color image captured by the color camera, and the first infrared image and the first color image are captured at the same time; Each second image includes a second color image captured by a color camera; The third recognition unit 406 is configured to determine an image sequence based on the plurality of second images and the first image, including: The third recognition unit 406 is configured to determine an image sequence based on the second color image and the first color image.

[0081] In one embodiment, the sentiment analysis system further includes a voice acquisition device (e.g. Figure 4 The microphone array 408 shown in FIG), the above-mentioned voice data is collected by using a voice collection device.

[0082] In one embodiment, the first recognition result, the second recognition result, and the third recognition result all include recognized emotion labels and corresponding first probability values, and the first model, the second model, and the third model are all configured with weight values; The determination unit 405 is configured to determine the target emotion of the target to be measured based on the first recognition result, the second recognition result, and the third recognition result, including: In response to at least some of the emotion labels being different, normalizing the first probability values based on the weight values to generate second probability values corresponding to the emotion labels; The emotion indicated by the emotion label corresponding to the maximum value among the second probability values is determined as the target emotion of the target to be measured.

[0083] In one embodiment, the determining unit 405 is configured to perform normalization processing on each first probability value based on each weight value to generate a second probability value corresponding to each emotion tag, including: Performing weighted summation on each first probability value based on each weight value to obtain a target sum value; For each emotion tag, the product of the first probability value corresponding to the emotion tag and the associated weight value is determined, and the ratio of the product to the target sum value is determined as the second probability value corresponding to the emotion tag.

[0084] In one embodiment, the sentiment analysis system includes a neural network processing unit; The first model is a neural network after knowledge distillation and is deployed in a neural network processing unit.

[0085] In one embodiment, the second identification unit 404 is further configured to: Extract features from speech data to generate speech features; The second model is used to perform emotion recognition based on speech features.

[0086] In one embodiment, the determining unit 405 is further configured to: Searching for a target interaction strategy including a target emotion among a plurality of preset interaction strategies; wherein the plurality of interaction strategies include interaction system information and different emotion information; In response to finding the target interaction strategy, an interaction instruction is generated based on the target interaction strategy, and the interaction instruction is sent to the interaction system indicated by the interaction system information in the target interaction strategy.

[0087] In one embodiment, the part image further includes a second part, which is at least a portion of an exposed part of the object to be measured; The first identification unit 403 is further configured to: Extracting key points of the first part from the image of the first part, and generating first part features including the extracted corresponding key points; Extracting second part key points from the image of the second part, and generating second part features including the extracted corresponding key points; The first model is used to perform emotion recognition based on the fusion result of the first part feature and the second part feature.

[0088] It should be noted that other aspects and implementation details of the sentiment analysis system provided in the embodiment of the present application are the same or similar to the emotion recognition method described above and will not be repeated here.

[0089] The embodiment of the present application further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following is achieved: Figure 1 or Figure 2 Describe the emotion recognition method.

[0090] The embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the following is achieved: Figure 1 or Figure 2 Describe the emotion recognition method.

[0091] The present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the following Figure 1 or Figure 2 Describe the emotion recognition method.

[0092] As described above, in the solution provided in the embodiments of this application, the emotion analysis system may include an infrared-RGB dual-mode camera, a microphone array, and an NPU acceleration module. It also uses a GNN+LSTM to fuse the spatiotemporal features of facial, gesture, and voice. The GNN can be deployed in a lightweight manner, such as using distillation loss-based model compression (teacher model: ResNet101; student model: MobileNetV3) and dynamic weight quantization (FP32→INT8). Furthermore, infrared images can be used to compensate for facial occlusions in RGB images and perform passenger segmentation detection, enabling interference-resistant emotion recognition. This approach can address issues such as poor adaptability to dynamic environments, insufficient multimodal fusion, and insufficient real-time performance, significantly improving emotion recognition accuracy.

[0093] The advantages of the embodiments of the present application are deduced below by reasoning.

[0094] 1. Derivation of Accuracy Improvement Figure 5 This is a schematic diagram of the causal chain derived from the accuracy improvement. Figure 5 , and provide quantitative evidence.

[0095] Single-mode misjudgment case: Analyzing only the face: Smiling expressions (actually sneers) → easily misinterpreted as positive emotions; Added gesture signals: Detecting a clenched fist + white knuckles → corrected to anger, effectively improving recognition accuracy.

[0096] Mathematical expression: Comprehensive accuracy = 1-∏(1-P_i); where P_i is the independent accuracy of each modality; Assume that face P_1=0.8, speech P_2=0.7, and gesture P_3=0.75 → multimodal accuracy = 1-(0.2×0.3×0.25)=98.5%.

[0097] 2. Derivation of real-time performance improvement Figure 6 This is a causal chain diagram for real-time improvement. Figure 6 and Table 1 below for key data comparison.

[0098] Table 1 Key data comparison example

[0099] 3. Derivation of Adaptability Improvement Figure 7 This is a schematic diagram of the environmental robustness causal chain derived from adaptability improvement. Figure 7 The diagram introduces the experimental verification related content.

[0100] Test scenario: Underground garage (illuminance <5lux): Traditional RGB camera: Facial key point detection failure rate 68%; Infrared-RGB fusion: key point completeness rate 92% (increased by 40%).

[0101] Physical mechanism: Infrared wavelengths penetrate the reflective layer of the windshield to avoid glare interference; The GAN network reconstructs the eye area blocked by sunglasses through the infrared depth map.

[0102] The comprehensive advantages are summarized in Table 2 below: Table 2 Example of comprehensive advantage summary

[0103] It should be noted that a delay of <50ms meets the threshold for human perception of no lag (research has confirmed that within 100ms is an "instantaneous response").

[0104] Through the above deduction, the technical advantages of the embodiments of the present application form a closed-loop logic chain, which has both theoretical rigor and experimental data support.

[0105] In addition, in a test on a certain car model, the solution provided in the embodiment of the present application has a composite emotion recognition accuracy of about 92% (the related technology is about 78%).

[0106] The above description is only a partial implementation of the embodiments of the present application and does not constitute any form of limitation to the application. The protection scope of the embodiments of the present application is not limited thereto. Any simple modifications, equivalent changes and modifications that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed in the embodiments of the present application should be covered within the protection scope of the embodiments of the present application.

Claims

1. An emotion recognition method, characterized in that: The emotion recognition method comprises: Acquiring image data and voice data of a target to be measured, wherein the image data includes a first image; Extracting a part image of the object to be measured from the first image, the part image including a first part, the first part being at least a portion of an exposed part of the object to be measured; Using the first model, emotion recognition is performed based on the body part image to obtain a first recognition result; Using the second model, performing emotion recognition based on the voice data to obtain a second recognition result; The target emotion of the target to be measured is determined based on the first recognition result and the second recognition result.

2. The emotion recognition method according to claim 1, characterized in that The image data is collected within a preset time period and also includes a plurality of second images collected before the first image; The emotion recognition method further includes: Determine an image sequence based on the plurality of second images and the first image, wherein the images in the image sequence are arranged in order of acquisition time from earliest to latest; Using a third model, performing emotion recognition based on the image sequence to obtain a third recognition result; Determining a target emotion of the target to be measured based on the first recognition result and the second recognition result includes: The target emotion is determined based on the first recognition result, the second recognition result, and the third recognition result.

3. The emotion recognition method according to claim 2, characterized in that The emotion recognition method is applied to a sentiment analysis system, wherein the sentiment analysis system includes an image acquisition device, wherein the image acquisition device integrates an infrared camera and a color camera; The first image includes a first infrared image captured by the infrared camera and a first color image captured by the color camera, and the first infrared image and the first color image are captured at the same time; Each of the second images includes a second color image captured by the color camera; Determining an image sequence based on the plurality of second images and the first image, comprising: The image sequence is determined based on each of the second color image and the first color image.

4. The emotion recognition method according to claim 2, characterized in that The first recognition result, the second recognition result, and the third recognition result all include recognized emotion labels and corresponding first probability values, and the first model, the second model, and the third model are all configured with weight values; Determining the target emotion based on the first recognition result, the second recognition result, and the third recognition result includes: In response to at least some of the emotion labels being different, normalizing the first probability values based on the weight values to generate second probability values corresponding to the emotion labels; The emotion indicated by the emotion label corresponding to the maximum value among the second probability values is determined as the target emotion.

5. The emotion recognition method according to claim 4, characterized in that Normalizing each of the first probability values based on each of the weight values to generate a second probability value corresponding to each of the emotion tags includes: Performing a weighted summation on each of the first probability values based on each of the weight values to obtain a target sum value; For each emotion tag, a product of the first probability value corresponding to the emotion tag and the associated weight value is determined, and a ratio of the product to the target sum value is determined as a second probability value corresponding to the emotion tag.

6. The emotion recognition method according to claim 1, characterized in that The emotion recognition method is applied to a sentiment analysis system, wherein the sentiment analysis system includes a neural network processing unit; The first model is a neural network after knowledge distillation and is deployed in the neural network processing unit.

7. The emotion recognition method according to claim 1, characterized in that Also includes: Performing feature extraction on the voice data to generate voice features; The utilizing the second model to perform emotion recognition based on the speech data includes: The second model is used to perform emotion recognition based on the speech features.

8. The emotion recognition method according to claim 1, characterized in that: Also includes: Searching for a target interaction strategy including the target emotion among a plurality of preset interaction strategies; wherein the plurality of interaction strategies include interaction system information and different emotion information; In response to finding the target interaction strategy, an interaction instruction is generated based on the target interaction strategy, and the interaction instruction is sent to the interaction system indicated by the interaction system information in the target interaction strategy.

9. The emotion recognition method according to any one of claims 1 to 8, characterized in that: The part image further includes a second part, which is at least a portion of the exposed part of the object to be measured; The emotion recognition method further includes: Extracting first part key points from the image of the first part, and generating first part features including the extracted corresponding key points; Extracting second part key points from the image of the second part, and generating second part features including the extracted corresponding key points; The using the first model to perform emotion recognition based on the body part image includes: The first model is used to perform emotion recognition based on a fusion result of the first part features and the second part features.

10. A sentiment analysis system, characterized in that: The sentiment analysis system includes: an acquiring unit configured to acquire image data and voice data of a target to be measured, wherein the image data includes a first image; an extraction unit configured to extract a part image of the target to be measured from the first image, wherein the part image includes a first part, and the first part is at least a part of an exposed part of the target to be measured; a first recognition unit configured to perform emotion recognition based on the part image using a first model to obtain a first recognition result; a second recognition unit configured to perform emotion recognition based on the voice data using a second model to obtain a second recognition result; A determination unit is configured to determine the target emotion of the target to be measured based on the first recognition result and the second recognition result.

Citation Information

Patent Citations

  • Multi-source fusion visual perception cabin system

    CN114155592A

  • Method, device and equipment for identifying driver emotion, medium and vehicle

    CN114973209A

  • Emotion recognition method, device and equipment based on audio and video

    CN115376559A

  • Face attribute discrimination method and device

    CN116453194A

  • Vehicle control method, training method and device, vehicle, server and medium

    CN116645959A

Cited By

  • AI-based visual emotion analysis method and system

    CN121415438A