A method and device for human emotion recognition
By combining audio and image emotion recognition, this multimodal fusion emotion recognition method solves the technical problems of single-modal emotion recognition in existing technologies, realizes the technical means of emotion recognition, and solves the misjudgment problems caused by noise interference and facial occlusion in traditional single-modal emotion recognition technology, achieving highly accurate and objective emotion recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIVIEW TECH CO LTD
- Filing Date
- 2025-08-04
- Publication Date
- 2026-07-17
AI Technical Summary
Traditional single-modal emotion recognition technology is susceptible to environmental noise interference. Pure audio recognition cannot establish a direct correlation between sound signals and people in the picture, while pure visual recognition relies on complete and clear facial images, leading to misjudgment of emotions and a decline in recognition performance.
A multimodal fusion method is adopted to perform emotion recognition by combining audio and images. Audio recognition is used to determine the target emotion type and confidence level. By combining the collision of facial and emotion key points in the monitoring screen, high-quality images are captured for image emotion scoring, and the final emotion recognition result is determined.
It improves the accuracy and objectivity of emotion recognition by using multimodal data fusion to detect and capture images reflecting human emotions in real time, thus enhancing the reliability of recognition.
Smart Images

Figure CN121483308B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular to a method and device for recognizing human emotions. Background Technology
[0002] In fields such as intelligent security and smart cities, surveillance equipment plays a crucial role in recognizing the emotions of moving targets. Modern surveillance equipment utilizes computer vision and deep learning technologies to perform real-time emotion recognition of moving targets. The system captures facial images, extracts facial features, and classifies basic emotions such as anger, joy, and sadness using pre-trained neural network models (such as CNN and LSTM). Surveillance equipment can also collect audio from targets. Audio emotion recognition technology analyzes the acoustic characteristics of speech signals (such as pitch, speech rate, energy, and spectrum) and combines this with deep learning models (such as LSTM and Transformer) to identify the speaker's emotional state.
[0003] Traditional single-modal emotion recognition technology has significant drawbacks: pure audio recognition is easily interfered with by environmental noise (such as vehicle noise and background human voices) and cannot establish a direct association between sound signals and people in the picture, leading to misjudgment of the emotion subject; pure visual recognition relies on complete and clear facial images, and the recognition effect drops sharply when the person is not fully in the picture, or when the side view or face is obscured. Summary of the Invention
[0004] This application provides a method and device for human emotion recognition, which uses multimodal fusion of high-quality audio and high-quality captured images of the target person to improve the accuracy of emotion recognition.
[0005] According to one aspect of this application, a method for recognizing human emotions is provided. The method includes: performing emotion recognition on the collected audio of a target person, determining the identified target emotion type and its corresponding first confidence level; during real-time monitoring of the target person using a monitoring device, performing keypoint collision analysis on facial keypoints and emotional keypoints detected in the monitoring image; wherein the target emotion type is reflected in specific expressions on the emotional keypoints of the person's face; if the collision result indicates that the monitoring image is suitable for recognizing the target emotion type, then capturing an image of the target person and determining an image emotion score based on the captured image; and determining the human emotion recognition result of the target person based on the first confidence level corresponding to the target emotion type and the image emotion score.
[0006] According to one aspect of this application, a person emotion recognition device is provided, the device comprising: an audio recognition module, used to perform emotion recognition on the collected audio of a target person, and determine the identified target emotion type and the corresponding first confidence level; a key point collision module, used to perform key point collision between detected facial key points and emotion key points of the target person in the monitoring screen during real-time monitoring of the target person through a monitoring device; wherein the target emotion type is reflected in the specific expression of the emotion key points on the person's face; an image emotion scoring determination module, used to capture the target person if the collision result indicates that the monitoring screen is suitable for identifying the target emotion type, and determine the image emotion score based on the captured image; and a person emotion recognition result determination module, used to determine the person emotion recognition result of the target person based on the first confidence level corresponding to the target emotion type and the image emotion score.
[0007] According to another aspect of this application, an electronic device is provided, the electronic device comprising:
[0008] At least one processor; and
[0009] A memory that is communicatively connected to at least one processor; wherein,
[0010] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor can perform the human emotion recognition method of any embodiment of this application.
[0011] According to another aspect of this application, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the human emotion recognition method of any embodiment of this application.
[0012] According to another aspect of this application, a computer program product is provided, which includes a computer program that, when executed by a processor, implements the human emotion recognition method of any embodiment of this application.
[0013] The technical solution of this application embodiment performs emotion recognition on the collected audio of a target person, determines the identified target emotion type and the corresponding first confidence level; during real-time monitoring of the target person through a monitoring device, key point collision is performed between the target person's facial key points and emotional key points detected in the monitoring screen; wherein, the target emotion type is reflected in the specific expression of the emotional key points on the person's face; if the collision result indicates that the monitoring screen is suitable for identifying the target person's target emotion type, the target person is captured, and an image emotion score is determined based on the captured image; based on the first confidence level corresponding to the target emotion type and the image emotion score, the emotion recognition result of the target person is determined. The above solution can detect in real time whether the monitoring screen presents an image that can intuitively reflect the person's emotion by colliding the target person's facial key points and emotional key points, and capture images when the monitoring screen is suitable for identifying the target emotion type, obtaining captured images that are more referential for emotion recognition. Based on the joint recognition of audio and captured images, multimodal data fusion recognition is achieved, improving the objectivity and accuracy of emotion recognition.
[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A flowchart illustrating a method for recognizing human emotions provided in an embodiment of this application; Figure 2 A flowchart illustrating a method for recognizing human emotions provided in another embodiment of this application; Figure 3 A flowchart illustrating a method for recognizing human emotions, as provided in another embodiment of this application; Figure 4 A flowchart illustrating a method for recognizing human emotions provided in another embodiment of this application; Figure 5 This is a schematic diagram of the structure of a human emotion recognition device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0017] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0018] It should be noted that the terms "first," "second," "third," "fourth," "actual," "preset," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0019] It should be noted that the emotion recognition of the person in this application embodiment is performed with the authorization and permission of the target person, and will not be disclosed without the permission of the target person, will not infringe on the target person's portrait rights, will not be used for illegal purposes or purposes that harm the interests of the target person, will not be used for personalized analysis of the target person or product promotion, and will not affect the normal life of the target person.
[0020] Figure 1 This is a flowchart illustrating a method for recognizing human emotions according to an embodiment of this application. This embodiment is applicable to situations requiring the recognition of the emotions of a target person in a scene. The method can be executed by a human emotion recognition device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0021] S110. Perform emotion recognition on the collected audio of the target person, and determine the identified target emotion type and the corresponding first confidence level.
[0022] Microphones can be deployed within the scene where the target person is located. These microphones can be integrated into the monitoring equipment or deployed independently. The microphone's audio capture range must cover the monitoring equipment's field of view, meaning that all sounds emitted within the monitoring equipment's field of view can be captured by the deployed microphone. The type of microphone is not limited and can be selected based on the actual application scenario. For more comprehensive sound capture, omnidirectional microphones can be used when the scene area is larger than a preset threshold and the number of people in the scene is less than a threshold, such as on a one-way road in a school. When the density of people is greater than a preset density, and it is necessary to pinpoint the exact location of people and accurately separate different sounds, such as in school classrooms or reading rooms, microphone arrays can be used.
[0023] In this embodiment, in real-world scenarios, a "sound before image" phenomenon commonly occurs when a target person approaches a monitoring device—that is, the sound signal emitted by the person (such as dialogue or footsteps) is captured before the visual image. Emotion recognition can be performed on the first-captured audio of the target person to identify the target's target emotion type and its corresponding first confidence level. Target emotion types include, for example, typical emotion categories: happiness, anger, sadness, fear, disgust, etc. The first confidence level reflects the probability that the target person's emotion is actually the target emotion type.
[0024] Specifically, the process of identifying a target person based on their audio may include, for the acquired audio frame sequence, dividing the audio frame sequence into frames using preset parameters to generate an audio frame sequence {y}. n (t)}. The preset parameters can be set to a 25ms Hamming window and a 10ms frame shift. Mel frequency cepstral coefficients are extracted to generate feature vectors. Specifically, for each frame after framing, the spectrum of each frame is calculated. Where n is the nth frame, T = 25ms is the frame length, and k is the index of the discrete frequency point, corresponding to the frequency component number after Fourier transform. The linear frequency is converted to the Mel frequency, and the logarithmic energy log(∑|x) is calculated. n (k)| 2 ·H m (k)), where H m (k) represents the m-th Mel filter response. Extract the first 12 dimensions of the MFCC features to generate a feature vector. That is, for {y n For each frame in (t)}, generate a feature vector X. soundA bidirectional LSTM neural network is used, with the input being 10 consecutive frames of stacked MFCC features (12×10 dimensions), and the output being a probability distribution of 6 emotion categories S={s1,s2,…,s6}, corresponding to happiness, anger, sadness, surprise, fear, and calmness; the loss function is cross-entropy. Where t r For the actual emotion labels, the optimizer uses Adam. r represents the r-th category in the classification task. For example, in an emotion classification task, if the emotion labels include "happy", "sad", "angry", "neutral", etc., then r corresponds to the index of these specific categories (r = 1, 2, ..., R).
[0025] The emotion type identified from the audio may include at least one. One of them can be selected as the target emotion type for subsequent evaluation, or the emotion type with the highest confidence can be selected as the target emotion type, or each of the identified emotion types can be used as the target emotion type.
[0026] S120. During the real-time monitoring of the target person through the monitoring equipment, the key points of the target person's face and the key points of emotion will be detected in the monitoring screen and the key point collision will be performed; among them, the target's emotional type is reflected in the specific expression of the emotional key points of the person's face.
[0027] The monitoring equipment is used to monitor the target person. The field of view of the monitoring equipment corresponds to the sound collection range of the microphone. In other words, the audio collected in the above process is emitted by the target person appearing in the monitoring screen of the monitoring equipment.
[0028] In this embodiment, the monitoring device performs real-time monitoring of the target person without triggering image capture in real time. Instead, during real-time monitoring, it performs key point collision analysis on the facial key points and emotional key points detected in the monitoring screen to determine whether any images in the monitoring screen are helpful for identifying the target person's emotions. The principle for determining the emotional key points is that when specific expressions are displayed on the emotional key points of a person's face, it reflects that the person's emotion is the target emotion type; conversely, the target emotion type is reflected in specific expressions on the emotional key points of a person's face. For example, for the target emotion type "happiness," when the corners of the mouth are turned up at the left and right corners of the mouth and the midpoint of the upper and lower lips, when the orbicularis oculi muscle is contracted at the outer corner of the eye and the key point of the orbicularis oculi muscle, and when the apple cheek area is lifted upward to form a full shape at the key point of the apple cheek area, it reflects that the person's emotion is happiness. In the above case, the emotional key points reflecting the target emotion type "happiness" include the left and right corners of the mouth, the midpoint of the upper and lower lips, the outer corner of the eye, the key point of the orbicularis oculi muscle, and the key point of the apple cheek area. For the target emotion type "sadness," the following points indicate sadness: eyebrows that appear converging and downward-pressed at the inner and outer corners of the eyebrows; eyelids that appear tightened and corners of the mouth that appear pulled down at the midpoints of the upper and lower eyelids and inner corners of the eyes; and corners of the mouth that appear pulled down at the midpoints of the lower lip. In these cases, the key points reflecting the emotion type "sadness" include the inner and outer corners of the eyebrows, the inner corners of the eyebrows, the corners of the mouth that appear pulled down, and the midpoint of the lower lip.
[0029] When emotional key points can reflect a person's emotions to a certain extent, the facial key points of the target person detected in the real-time image can be compared with the emotional key points to determine whether there are facial key points in the current monitoring image that are helpful for the emotional recognition of the target person, and whether the current monitoring image is highly relevant for the emotional recognition of the target person.
[0030] In this application embodiment, the process of collating the facial key points and emotional key points of the target person in the monitoring screen may include positional matching based on the facial key points and emotional key points, or it may include quantity matching of the facial key points and emotional key points.
[0031] S130. If the collision result indicates that the monitoring screen is suitable for identifying the target person's emotional type, then the target person is captured, and an image emotion score is determined based on the captured image.
[0032] For example, by collating the facial key points of the target person with the emotional key points, a preliminary assessment of the surveillance footage can be made to determine whether the surveillance footage is suitable for identifying the target person's emotional type. In other words, whether the surveillance footage contains facial key points that are helpful in identifying the target emotional type, so as to avoid capturing images and using them for emotion recognition when the surveillance footage contains images that are unsuitable for recognition, such as side profiles, wearing masks, or wearing glasses.
[0033] If the collision result indicates that the monitoring footage is suitable for identifying the target person's emotion type, then the target person is captured, resulting in a captured image. The emotion score is then determined based on this captured image. The captured image can be at least one frame, and the image used to determine the emotion score can be one of those frames, such as the first frame, or any selected frame, or the frame with the highest quality score. Multiple frames can also be used; for example, all captured images can be used to determine the emotion score, or images with quality scores exceeding a preset threshold can be selected.
[0034] Determining image emotion scores based on captured images involves identifying the type of emotion and its confidence level reflected in the content of the captured image. This can be done by recognizing whether the captured image matches the target emotion type and its confidence level, or by identifying multiple possible emotion types and their corresponding confidence levels, thus determining the image emotion score. Alternatively, it can be done by quantifying the actions at facial key points corresponding to emotional key points in the captured image to determine the image emotion score.
[0035] S140. Determine the character emotion recognition result of the target character based on the first confidence level corresponding to the target emotion type and the image emotion score.
[0036] In this embodiment, multimodal data is fused, combining audio and captured images to comprehensively identify the emotions of a target person, thereby improving the accuracy of emotion recognition. Specifically, the recognition result obtained from audio-based emotion recognition is the target emotion type and its corresponding first confidence level. The recognition result obtained from captured image-based emotion recognition is represented by an image emotion score. The emotion recognition result of the target person can be determined based on the first confidence level corresponding to the target emotion type and the image emotion score. When the image emotion score is positively correlated with the confidence level of the target emotion type, determining the emotion recognition result of the target person based on the first confidence level and the image emotion score can be achieved by performing a gain operation on the first confidence level and the image emotion score, such as weighted summation or product. Conversely, when the confidence level of the target emotion type reflected by the image emotion score is negatively correlated, determining the emotion recognition result of the target person based on the first confidence level and the image emotion score can be achieved by performing a loss operation on the first confidence level and the image emotion score, such as weighted subtraction or quotient.
[0037] The technical solution of this application embodiment performs emotion recognition on the collected audio of a target person, determines the identified target emotion type and the corresponding first confidence level; during real-time monitoring of the target person through a monitoring device, key point collision is performed between the target person's facial key points and emotional key points detected in the monitoring screen; wherein, the target emotion type is reflected in the specific expression of the emotional key points on the person's face; if the collision result indicates that the monitoring screen is suitable for identifying the target person's target emotion type, the target person is captured, and an image emotion score is determined based on the captured image; based on the first confidence level corresponding to the target emotion type and the image emotion score, the emotion recognition result of the target person is determined. The above solution can detect in real time whether the monitoring screen presents an image that can intuitively reflect the person's emotion by colliding the target person's facial key points and emotional key points, and capture images when the monitoring screen is suitable for identifying the target emotion type. Based on the joint recognition of audio and captured images, multimodal data fusion recognition is achieved, improving the objectivity and accuracy of emotion recognition.
[0038] Figure 2 This is a flowchart illustrating a method for recognizing human emotions, provided as another embodiment of this application. This embodiment is an optimization based on the above embodiment; solutions not described in detail in this embodiment are found in the above embodiment. Figure 2 As shown, the method in this embodiment of the application specifically includes the following steps:
[0039] S210. Perform emotion recognition on the collected audio of the target person, and determine the identified target emotion type and the corresponding first confidence level.
[0040] S220. During the real-time monitoring of the target person through the monitoring equipment, the number of facial key points that match the emotional key points among the facial key points detected in the monitoring screen is counted; wherein, the method for determining whether the facial key points match the emotional key points is that the relative coordinates of the facial key points in the target person's facial image and the relative coordinates of the emotional key points in the preset facial image are within a preset error range.
[0041] For example, in a surveillance video, facial key points of a target person are detected, identifying at least one facial key point and its corresponding confidence score. Specifically, an algorithm (such as MTCNN) is used to detect the target person's face in the surveillance video, outputting bounding boxes (Bbox = (x, y, w, h). Overlapping boxes are removed using non-maximum suppression (NMS), retaining facial regions with a confidence score ≥ a confidence threshold (e.g., 0.8). Based on an OpenFace or Dlib model, M (e.g., 68) facial key point coordinates and confidence scores are extracted from the detected facial regions. Where c j ∈[0,1] represents the confidence level corresponding to the j-th key point.
[0042] By matching facial key points with emotional key points, the number of facial key points that can be matched with emotional key points is determined, thereby reflecting the completeness and representativeness of the facial features of the target person contained in the surveillance footage.
[0043] One method for matching facial keypoints with emotional keypoints involves iterating through each facial keypoint. During this process, the relative coordinates of each facial keypoint in the target person's facial image are compared with the relative coordinates of each emotional keypoint in a preset facial image. If the error between the relative coordinates of an emotional keypoint in the preset facial image and the relative coordinates of the facial keypoint relative to the target person's facial image is within a preset error range, then the facial keypoint is considered to be matched with that emotional keypoint. The target person's facial image can be a minimum bounding box image containing the target person's face, and the preset facial image can be a pre-set minimum bounding box image containing a preset person's face. The preset error range can be determined based on actual conditions, for example, set to 0-10 pixels, so that an error within the preset error range reflects that the relative coordinates of the facial keypoint and the emotional keypoint in the facial image are close.
[0044] In this embodiment of the application, counting the number of facial key points that match the emotional key points among the facial key points of the target person detected in the monitoring screen includes:
[0045] For the facial key points of the target person detected in the monitoring screen, among the facial key points that match the emotional key points, the number of facial key points whose corresponding second confidence level is greater than a preset confidence threshold is counted; wherein, the second confidence level is the confidence level corresponding to the facial key points of the target person detected in the monitoring screen.
[0046] In this embodiment, to further constrain the collision principle between facial key points and emotional key points, the number of facial key points with a second confidence level greater than a preset confidence threshold can be counted among those matching emotional key points. This allows for the selection of facial key points with higher reliability and greater reference value for emotion recognition. The second confidence level is the output confidence level corresponding to the identified facial key points when detecting facial key points of a target person in a monitoring screen. Specifically, it is expressed by the formula: Where i is the i-th facial keypoint, N is the total number of facial keypoints, and P detect P is a set of facial key points. target c is a set of key emotional points. i τ represents the second confidence level of the i-th facial key point. conf For the preset confidence threshold, ∏(·) is an indicator function that reflects the number of facial key points that satisfy the conditions in parentheses.
[0047] S230. Determine the collision result based on the stated quantity.
[0048] For example, the collision result can be determined based on the number of facial key points that match emotional key points, thus determining whether the surveillance footage is suitable for identifying the target person's emotional type. Generally, the more facial key points that match emotional key points, the more facial key points in the target person's facial image that can reflect the target emotional type, and the more suitable it is for identifying the target person's emotional type.
[0049] In this embodiment of the application, determining the collision result based on the quantity includes:
[0050] If the number is greater than a preset threshold, the collision result is determined to be that the monitoring screen is suitable for identifying the target person's emotional type.
[0051] Otherwise, the collision result is determined to be that the monitoring footage is not suitable for identifying the target person's emotional type.
[0052] The preset quantity threshold can be determined according to the actual situation. For example, it can be set as the product of the preset proportion of the number of emotional key points corresponding to the target emotion type. The preset proportion can be, for example, 80%. For example, if the number of facial key points and emotional key points determined by collision in any of the above embodiments is greater than the preset quantity threshold, the collision result is determined to be that the monitoring screen is suitable for identifying the target emotion type of the target person; otherwise, the collision result is determined to be that the monitoring screen is not suitable for identifying the target emotion type of the target person.
[0053] S240. If the collision result indicates that the monitoring screen is suitable for identifying the target person's emotional type, then the target person is captured, and an image emotion score is determined based on the captured image.
[0054] S250. Based on the first confidence level corresponding to the target emotion type and the image emotion score, determine the emotion recognition result of the target person.
[0055] This application provides a method for recognizing human emotions. It counts the number of facial key points that match emotional key points among those detected in a surveillance video. The matching method is based on the relative coordinates of the facial key points in the target person's facial image, ensuring the error between these coordinates and the relative coordinates of the emotional key points in a preset facial image is within a preset error range. A collision result is determined based on the number of matches. By matching facial key points with emotional key points, the number of matches determines whether the surveillance video is suitable for recognizing the target person's emotional type. This method detects surveillance footage with high reference value for emotional recognition of the target person in real-time monitoring, solving the problem of incomplete or missing key features for emotion recognition due to random snapshots, thus affecting the accuracy of emotion recognition.
[0056] Figure 3 This is a flowchart illustrating a method for recognizing human emotions, provided as another embodiment of this application. This embodiment is an optimization based on the above embodiments; solutions not described in detail in this embodiment are found in the above embodiments. Figure 3 As shown, the method in this embodiment of the application specifically includes the following steps:
[0057] S310. Perform emotion recognition on the collected audio of the target person, and determine the identified target emotion type and the corresponding first confidence level.
[0058] S320. During the real-time monitoring of a target person through monitoring equipment, the key points of the target person's face and the key points of their emotions will be detected in the monitoring screen and collide with each other; among them, the target's emotional type is reflected in the specific expression of the emotional key points of the person's face.
[0059] S330. If the collision result indicates that the monitoring image is suitable for identifying the target person's emotional type, then the target person is captured on camera.
[0060] S340. For the facial key points in the captured image that match the emotional key points, determine the key point weights corresponding to the facial key points based on a mapping table; wherein, the mapping table contains the correspondence between the emotional key points and the key point weights corresponding to the target emotional type, and the key point weights reflect the importance of the specific expression on the emotional key points to the determination of the target emotional type.
[0061] The mapping table is a pre-established table containing the correspondence between emotional key points and key point weights. Each emotion type corresponds to one mapping table, which includes at least the emotional key points that affect the identification of that emotion type, as well as the key point weights corresponding to the emotional key points. The key point weights reflect the importance of a specific expression at an emotional key point to the determination of the emotion type. The mapping tables corresponding to different emotion types are shown in Tables 1 and 2. Table 1 represents the mapping table corresponding to the emotion type "happiness," and Table 2 represents the mapping table corresponding to the emotion type "sadness." Mapping tables for other emotion types are not listed one by one. In this embodiment, since the facial key points used for emotion recognition in the captured image are facial key points that match the emotional key points corresponding to the target emotion type, the key point weights corresponding to the facial key points are determined according to the mapping table corresponding to the target emotion type. The mapping table corresponding to the target emotion type contains the key point weights corresponding to the emotional key points, that is, the key point weights corresponding to the facial key points that match the emotional key points.
[0062] Table 1
[0063]
[0064] Table 2
[0065]
[0066]
[0067] S350. Determine the image emotion score based on the second confidence level of the facial key points and the key point weights corresponding to the facial key points.
[0068] For example, an image emotion score can be determined based on the second confidence score of facial key points and the corresponding key point weights. This means evaluating the image emotion score corresponding to the target emotion type in a facial image by using the second confidence score corresponding to facial key point recognition and the importance of facial key points in emotion recognition. In the specific calculation process, the second confidence score and the key point weight must correspond to the same facial key point for the calculation to be performed.
[0069] In this embodiment of the application, the image emotion score is determined based on the second confidence level of the facial key points and the key point weights corresponding to the facial key points, including: for the same facial key point, using the key point weight as the confidence weight corresponding to the second confidence level; for the second confidence levels corresponding to each facial key point, performing a weighted summation of each second confidence level based on the corresponding confidence weight to obtain the image emotion score.
[0070] Specifically, in calculating the image emotion score based on the second confidence level of facial key points and their corresponding key point weights, for the same facial key point, the key point weight corresponding to that key point can be used as the confidence weight of the second confidence level of that facial key point. Each facial key point's second confidence level corresponds to a confidence weight. The second confidence levels of all facial key points are then weighted and summed to obtain the image emotion score. Among them, V l Give the image emotion score for the captured image in frame l. w represents the second confidence level of the i-th facial key point in the l-th frame of the captured image. i Let be the keypoint weight corresponding to the i-th facial keypoint.
[0071] In this embodiment of the application, determining the image emotion score based on the captured image includes: if there are at least two captured images, then determining the image emotion score based on each captured image; and taking the maximum image emotion score among all image emotion scores.
[0072] For example, when a snapshot is triggered, multiple snapshots may be executed consecutively to obtain at least two snapshot images. Each snapshot image, after processing as described in the above embodiments, will obtain its own corresponding image emotion score. The highest image emotion score can be taken for subsequent steps, i.e., V. max =max{V1,V2,…,V L This allows for the selection of information from more realistic and referential captured images for subsequent processes, thereby improving the accuracy of emotion recognition.
[0073] S360. Based on the first confidence level corresponding to the target emotion type and the image emotion score, determine the emotion recognition result of the target person.
[0074] This application provides a method for recognizing facial emotions. For facial key points in a captured image that match emotional key points, a mapping table is used to determine the key point weights corresponding to the facial key points. The mapping table contains the correspondence between emotional key points and key point weights corresponding to a target emotion type. The key point weights reflect the importance of specific expressions at the emotional key points in determining the target emotion type. An image emotion score is determined based on a second confidence level of the facial key points and their corresponding key point weights. By using the key point weights in the mapping table corresponding to the target emotion type to reflect the influence of facial key points in target emotion type recognition, and applying this to the second confidence level, the image emotion score corresponding to the captured image is calculated. This accurately determines the probability that the target person identified from the captured image matches the target emotion type, achieving fine-grained recognition of subtle facial key points in the target person's facial image, and thus realizing objective and accurate recognition of the target emotion type.
[0075] Figure 4 This is a flowchart illustrating a method for recognizing human emotions, provided as another embodiment of this application. This embodiment is an optimization based on the above embodiments; solutions not described in detail in this embodiment are found in the above embodiments. Figure 4 As shown, the method in this embodiment of the application specifically includes the following steps:
[0076] S410. Perform emotion recognition on the collected audio of the target person, and determine the identified target emotion type and the corresponding first confidence level.
[0077] S410. During the real-time monitoring of a target person through monitoring equipment, the key points of the target person's face and the key points of emotion will be detected in the monitoring screen and the key point collision will be performed; among them, the target's emotional type is reflected in the specific expression of the emotional key points of the person's face.
[0078] S430. If the collision result indicates that the monitoring screen is suitable for identifying the target person's emotional type, then the target person is captured, and an image emotion score is determined based on the captured image.
[0079] S440. Determine the first weight of the audio participating in emotion recognition and the second weight of the captured image participating in emotion recognition.
[0080] In this embodiment, multimodal data fusion recognition is performed on audio and captured images to improve the accuracy of emotion recognition. During emotion recognition, the importance of audio and captured images may differ. Therefore, a first weight for audio and a second weight for captured images in emotion recognition can be adaptively determined. For example, the first weight for audio can be determined based on the audio volume. Generally, the louder the audio, the higher the confidence level of the target person's emotion reflected in the audio, and the larger the corresponding first weight. Conversely, the softer the audio, the lower the confidence level of the target person's emotion reflected in the audio, and the smaller the corresponding first weight. The second weight for captured images in emotion recognition can be determined based on the image quality score of the captured image. Generally, a higher image quality score reflects better performance in terms of image clarity, resolution, pose angle, lighting, and occlusion, which is more conducive to image recognition. Therefore, a larger second weight can be set. Conversely, a lower image quality score requires a smaller second weight. The sum of the first and second weights should equal 1.
[0081] In this embodiment of the application, determining the first weight of the audio participating in emotion recognition and the second weight of the captured image participating in emotion recognition includes: determining the ratio between the size information of the image region of the target person in the captured image and the size information of the captured image; if the ratio is less than a first preset ratio, then the first weight is determined to be greater than the second weight; if the ratio is greater than or equal to the first preset ratio and less than the second preset ratio, then the first weight is determined to be equal to the second weight; if the ratio is greater than or equal to the second preset ratio, then the first weight is determined to be less than the second weight.
[0082] For example, during real-time monitoring, the closer the target person is to the monitoring device, the larger the target person's image area is reflected in the captured image, and the higher the confidence level of the information identified from the captured image. Conversely, the farther the target person is from the monitoring device, the smaller the target person's image area is reflected in the captured image, and the lower the confidence level of the information identified from the captured image. Therefore, a second weight for the captured image in emotion recognition can be determined based on the size information of the target person's image area in the captured image. Specifically, the proportional relationship between the size information of the target person's image area in the captured image and the size information of the captured image can be determined. Size information includes, but is not limited to, length, width, area, perimeter, etc. If the proportional relationship is less than a first preset ratio, it reflects that the target person's image area accounts for a small proportion in the captured image, that is, the target person is far from the monitoring device. In this case, the second weight is set to be less than the first weight, that is, the importance of the captured image in emotion recognition is less than the importance of audio in emotion recognition. If the ratio is greater than or equal to the first preset ratio but less than the second preset ratio, it indicates that the target person's image area occupies a moderate proportion in the captured image, meaning the target person is at a moderate distance from the monitoring device. In this case, the first weight is set to equal the second weight, meaning the importance of the captured image in emotion recognition is equal to the importance of the audio. If the ratio is greater than the second preset ratio, it indicates that the target person's image area occupies a large proportion in the captured image, meaning the target person is close to the monitoring device. In this case, the second weight is set to greater than the first weight, meaning the importance of the captured image in emotion recognition is greater than the importance of the audio.
[0083] In this embodiment, if there are at least two captured images, the image emotion score is determined based on each captured image. The image with the highest emotion score is selected, and the captured image corresponding to the highest emotion score is applied to the weight determination process. Furthermore, if multiple captured images (a total of K images, K <= L) all have the highest emotion score, the captured image with the best quality score is selected based on the image quality evaluation (clarity, size, angle, lighting, occlusion, etc.). Capture image quality evaluation function:
[0084] QualityScore = α1·S s +α2·S q +α3·S p +α4·S l +α5·S o
[0085] Among them, S s Sharpness assessment value, α1 sharpness assessment value and its corresponding weight. S q Resolution evaluation value, weights corresponding to the α2 resolution evaluation value. pAttitude angle evaluation value, and the weight corresponding to the α3 attitude angle evaluation value. l Illumination assessment value, and the weight corresponding to the α4 illumination assessment value. S o Occlusion assessment value, and the weight corresponding to the α5 occlusion assessment value.
[0086] S450. Based on the first weight and the second weight, the first confidence score and the image emotion score are weighted and summed, and the emotion recognition result is determined according to the weighted summation result.
[0087] For example, the first weight is the weight corresponding to the audio, that is, the weight corresponding to the first confidence level of the target emotion type identified by the audio. The second weight is the weight corresponding to the captured image, that is, the weight corresponding to the image emotion score. The first confidence level and the image emotion score can be weighted and summed to determine the emotion recognition result of the target person. For example, a preset result threshold can be set in advance. If the weighted sum threshold is greater than the preset result threshold, the emotion recognition result of the target person is determined to be the target emotion type. Otherwise, the emotion recognition result of the target person is determined not to be the target emotion type, and recognition continues from S410 based on other emotion types identified by the audio.
[0088] In this embodiment of the application, the process of determining the target emotion type includes: for at least one emotion type identified based on audio, where the first confidence level exceeds a preset first confidence threshold, traversing each emotion type and taking the currently traversed emotion type as the target emotion type; correspondingly, determining the person's emotion recognition result based on the weighted summation result includes:
[0089] The target emotion type corresponding to the maximum value of the weighted summation result is taken as the emotion recognition result of the person.
[0090] For example, in the process of emotion recognition in audio, multiple emotion types may be identified. For each identified emotion type, the emotion type with a first confidence level exceeding a preset first confidence threshold is selected. The selected emotion types are iterated through, and the currently iterated emotion type is taken as the target emotion type. Subsequent processes are then executed, including the processes starting from S110, S210, S310, and S410 in the above embodiments. Correspondingly, after executing the above processes, a weighted summation result is calculated for each target emotion type. The target emotion type corresponding to the maximum value of the weighted summation result is taken as the emotion recognition result of the target person. The first confidence level S corresponding to the target emotion type identified in the audio is fused with the maximum image emotion score V. max The final weighted sum is:
[0091] C = β·S + (1-β)·V max ;
[0092] The target person's emotion recognition result was determined to be an emotion. final =argmax2(C). Where β is the first weight and (1-β) is the second weight. argmax(C) represents the target sentiment type corresponding to the largest C.
[0093] This application provides a method for recognizing human emotions. It determines a first weight for the audio input and a second weight for the captured image input. Based on the first and second weights, a weighted sum is performed on the first confidence level and the image emotion score. The human emotion recognition result is determined based on the weighted summation result. This method can adaptively set different weights according to the contribution of different modalities to the credibility of emotion recognition, achieving weighted fusion of multimodal data. It more comprehensively identifies the emotions of the target person based on their multimodal data, thereby improving the accuracy and credibility of emotion recognition.
[0094] Figure 5 This is a schematic diagram of a human emotion recognition device provided in an embodiment of this application. This device can execute the human emotion recognition method provided in any embodiment of this application, and possesses the corresponding functional modules and beneficial effects of the method. For example... Figure 5 As shown, the device includes:
[0095] The audio recognition module 510 is used to perform emotion recognition on the collected audio of the target person, determine the identified target emotion type and the corresponding first confidence level; the key point collision module 520 is used to perform key point collision between the facial key points and emotional key points detected in the monitoring screen during real-time monitoring of the target person through the monitoring device; wherein, the target emotion type is reflected in the specific expression of the emotional key points on the person's face; the image emotion scoring determination module 530 is used to capture the target person if the collision result shows that the monitoring screen is suitable for identifying the target emotion type, and determine the image emotion score based on the captured image; the person emotion recognition result determination module 540 is used to determine the person emotion recognition result of the target person based on the first confidence level corresponding to the target emotion type and the image emotion score.
[0096] In this embodiment, the keypoint collision module 520 performs keypoint collision on the facial keypoints and emotional keypoints of the target person detected in the monitoring screen. This includes: counting the number of facial keypoints that match the emotional keypoints among the facial keypoints of the target person detected in the monitoring screen; wherein, the method for determining whether the facial keypoint matches the emotional keypoint is that the relative coordinates of the facial keypoint in the target person's facial image and the relative coordinates of the emotional keypoint in a preset facial image have an error within a preset error range; and determining the collision result based on the number of matches.
[0097] In this embodiment of the application, the key point collision module 520 counts the number of facial key points that match the emotion key points among the facial key points of the target person detected in the monitoring screen, including: for the facial key points of the target person detected in the monitoring screen, among the facial key points that match the emotion key points, count the number of facial key points whose corresponding second confidence is greater than a preset confidence threshold; wherein, the second confidence is the confidence corresponding to the facial key points of the target person detected in the monitoring screen.
[0098] In this embodiment of the application, the key point collision module 520 determines the collision result based on the number, including: if the number is greater than a preset number threshold, then the collision result is determined to be that the monitoring screen is suitable for identifying the target emotion type of the target person; otherwise, the collision result is determined to be that the monitoring screen is not suitable for identifying the target emotion type of the target person.
[0099] In this embodiment, the image emotion scoring determination module 530 determines an image emotion score based on a captured image, including: determining the key point weights corresponding to the facial key points that match the emotion key points in the captured image based on a mapping table; wherein the mapping table contains the correspondence between the emotion key points and the key point weights corresponding to the target emotion type, and the key point weights reflect the importance of specific expressions on the emotion key points to the determination of the target emotion type; and determining the image emotion score based on the second confidence level of the facial key points and the key point weights corresponding to the facial key points.
[0100] In this embodiment, the image emotion scoring determination module 530 determines the image emotion score based on the second confidence level of the facial key points and the key point weights corresponding to the facial key points, including: for the same facial key point, using the key point weight as the confidence weight corresponding to the second confidence level; for the second confidence levels corresponding to each facial key point, performing a weighted summation of each second confidence level based on the corresponding confidence weight to obtain the image emotion score.
[0101] In this embodiment of the application, the image emotion score determination module 530 determines the image emotion score based on the captured image, including: if there are at least two frames in the captured image, then determine the image emotion score based on each captured image; and take the maximum image emotion score among all the image emotion scores.
[0102] In this embodiment of the application, the character emotion recognition result determination module 540 determines the character emotion recognition result of the target character based on the first confidence level corresponding to the target emotion type and the image emotion score, including: determining the first weight of the audio participating in emotion recognition and the second weight of the captured image participating in emotion recognition; performing a weighted summation on the first confidence level and the image emotion score based on the first weight and the second weight, and determining the character emotion recognition result based on the weighted summation result.
[0103] In this embodiment, the emotion recognition result determination module 540 determines a first weight for the audio to participate in emotion recognition and a second weight for the captured image to participate in emotion recognition, including: determining the ratio between the size information of the image region of the target person in the captured image and the size information of the captured image; if the ratio is less than a first preset ratio, then the first weight is determined to be greater than the second weight; if the ratio is greater than or equal to the first preset ratio and less than the second preset ratio, then the first weight is determined to be equal to the second weight; if the ratio is greater than or equal to the second preset ratio, then the first weight is determined to be less than the second weight.
[0104] In this embodiment of the application, the audio recognition module 510 is specifically used to: for an emotion type whose first confidence level exceeds a preset first confidence level threshold among at least one emotion type identified based on audio, traverse each emotion type and take the currently traversed emotion type as the target emotion type; correspondingly, the character emotion recognition result determination module 540 determines the character emotion recognition result based on the weighted summation result, including: taking the target emotion type corresponding to the maximum value of the weighted summation result as the character emotion recognition result.
[0105] The human emotion recognition device provided in this application embodiment can execute a human emotion recognition method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of executing the method.
[0106] Figure 6A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0107] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0108] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless human emotion recognition transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0109] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as human emotion recognition methods.
[0110] In some embodiments, the human emotion recognition method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the human emotion recognition method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the human emotion recognition method by any other suitable means (e.g., by means of firmware).
[0111] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0112] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable human emotion recognition device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0113] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input). The systems and techniques described herein can be implemented in computing systems that include back-end components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or a web browser through which the user interacts with embodiments of the systems and techniques described herein), or computing systems that include any combination of such back-end components, middleware components, or front-end components. System components can be interconnected via digital data communication (e.g., communication networks) in any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), blockchain networks, and the Internet. Computing systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service ecosystem to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0114] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the human emotion recognition method as provided in any embodiment of this application. In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0115] It should be understood that the various processes shown above can be used, with steps rearranged, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired information of the technical solution of this application can be achieved, and this is not limited herein. The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for recognizing human emotions, characterized in that, The method includes: Emotion recognition is performed on the collected audio of the target person to determine the type of the identified emotion and the corresponding first confidence level. During real-time monitoring of a target person using surveillance equipment, key point collisions will be performed on the target person's facial key points and emotional key points detected in the surveillance footage. The target emotion type is reflected in the specific expression of the emotional key points on the person's face. The emotional key points refer to facial key points related to the target emotion type. Detecting key point collisions on the target person's facial key points and emotional key points in the surveillance footage includes matching the positions of facial key points and emotional key points, or matching the number of facial key points and emotional key points. If the collision result indicates that the monitoring screen is suitable for identifying the target person's emotional type, then the target person is captured, and an image emotion score is determined based on the captured image. Based on the first confidence level corresponding to the target emotion type and the image emotion score, the emotion recognition result of the target person is determined.
2. The method according to claim 1, characterized in that, The system will detect facial key points and emotional key points of the target person in the monitoring footage and perform key point collision analysis, including: The number of facial key points that match the emotional key points among the facial key points of the target person detected in the monitoring screen is counted; wherein, the method for judging the match between the facial key points and the emotional key points is that the relative coordinates of the facial key points in the target person's facial image are within a preset error range. The collision result is determined based on the stated quantity.
3. The method according to claim 2, characterized in that, The number of facial key points that match the emotional key points among the facial key points detected in the surveillance footage of the target person is counted, including: For the facial key points of the target person detected in the monitoring screen, among the facial key points that match the emotional key points, the number of facial key points whose corresponding second confidence level is greater than a preset confidence threshold is counted; wherein, the second confidence level is the confidence level corresponding to the facial key points of the target person detected in the monitoring screen.
4. The method according to claim 2, characterized in that, Determining the collision result based on the stated quantity includes: If the number is greater than a preset threshold, the collision result is determined to be that the monitoring screen is suitable for identifying the target person's emotional type. Otherwise, the collision result is determined to be that the monitoring footage is not suitable for identifying the target person's emotional type.
5. The method according to claim 1, characterized in that, Determine the image emotion score based on the captured image, including: For facial key points in the captured image that match the emotional key points, the key point weights corresponding to the facial key points are determined based on a mapping table; wherein, the mapping table contains the correspondence between emotional key points and key point weights corresponding to the target emotional type, and the key point weights reflect the importance of specific expressions on the emotional key points to the determination of the target emotional type. The image emotion score is determined based on the second confidence level of the facial key points and the key point weights corresponding to the facial key points.
6. The method according to claim 5, characterized in that, The image emotion score is determined based on the second confidence level of the facial key points and the key point weights corresponding to the facial key points, including: For the same facial key point, the weight of the key point is used as the confidence weight corresponding to the second confidence level; For each facial key point, the second confidence level is weighted and summed based on the corresponding confidence level weight to obtain the image emotion score.
7. The method according to claim 1, characterized in that, Based on the first confidence level corresponding to the target emotion type and the image emotion score, the emotion recognition result of the target person is determined, including: A first weight for the audio to participate in emotion recognition and a second weight for the captured image to participate in emotion recognition are determined respectively. Based on the first weight and the second weight, the first confidence score and the image emotion score are weighted and summed, and the human emotion recognition result is determined according to the weighted summation result.
8. The method according to claim 7, characterized in that, Determine the first weight of the audio in emotion recognition and the second weight of the captured image in emotion recognition, including: Determine the proportional relationship between the size information of the image region of the target person in the captured image and the size information of the captured image; If the ratio is less than the first preset ratio, then the first weight is determined to be greater than the second weight. If the ratio is greater than or equal to the first preset ratio and less than the second preset ratio, then the first weight is determined to be equal to the second weight. If the ratio is greater than or equal to the second preset ratio, then the first weight is determined to be less than the second weight.
9. The method according to claim 7, characterized in that, The process of determining the target emotion type includes: For an emotion type whose first confidence score exceeds a preset first confidence threshold among at least one emotion type identified based on audio, iterate through each emotion type and take the currently traversed emotion type as the target emotion type; Accordingly, the result of human emotion recognition is determined based on the weighted summation result, including: The target emotion type corresponding to the maximum value of the weighted summation result is taken as the emotion recognition result of the person.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the human emotion recognition method according to any one of claims 1-9.