AU-based happiness degree detection method and device, equipment and medium

Through a multi-dimensional factor comprehensive detection method based on face action units, expression types and happiness degree scores, the target loss function is constructed and the model is trained, which solves the problem of low detection accuracy of happiness emotion in the existing technology, and achieves a more accurate and stable happiness degree detection effect.

CN120148083APending Publication Date: 2025-06-13PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196952.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, it is difficult for emotion detection models to accurately grasp the difference in strength and weakness of happy emotions, resulting in inaccurate output scoring, lack of continuity and poor robustness, especially when facing complex or mixed emotions samples.

Method used

By obtaining face pictures and labeling based on face action units, expression types and happiness degree score scores, we construct expression classification loss functions, face action units loss functions and happiness degree score loss functions, and combine these loss functions as target loss functions to train the backbone network to obtain the happiness degree detection model.

Benefits of technology

It significantly improves the accuracy and stability of the model detection results, making the degree of happiness detection more accurate and reliable, and can effectively compensate for the problem of insufficient data labeling in the comprehensive detection of multi-dimensional factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148083A_ABST
    Figure CN120148083A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and medical health, and provides an AU-based happiness degree detection method, device, equipment and medium, which can label face pictures based on face action units, expression types and happiness degree scores, so that a trained happiness degree detection model can perform detection by integrating multi-dimensional factors, and the detection efficiency is improved. The defects of insufficient model continuity and poor robustness caused by the fact that data lacks degree labeling and only depends on simple labels are effectively overcome; in the training process, the constructed expression classification loss function, the face action unit loss function and the happiness degree score loss function are used as target loss, the accuracy of a model detection result and the model stability are remarkably improved, the user happiness degree can be detected in a more detailed mode in the field of medical health, and the user experience is improved. Therefore, the method has high application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and medical health, and particularly to a method, device, equipment and medium for detecting the degree of happiness based on AU. Background Art

[0002] In the field of sentiment analysis, the degree of happiness refers to a score that can quantitatively represent the degree of happiness output when a face image is used as input. Nowadays, such algorithms have received high attention and application requirements in multiple application scenarios, such as in social media sentiment monitoring, intelligent customer service evaluation of customer satisfaction, mental health auxiliary analysis, etc., and have high application value.

[0003] However, in the prior art, the current industry lacks degree annotation for data, and the open-source data only has binary labels of whether it is happy or not, lacking intermediate transition state annotation. This makes it difficult for the model to accurately grasp the intensity difference of happy emotions. If such data is used to train the model, it will result in inaccurate output scores, lack of continuity, and poor robustness, and it is also prone to misjudgment when facing complex or mixed emotion samples. Summary of the Invention

[0004] In view of the above, it is necessary to provide a method, device, equipment and medium for detecting the degree of happiness based on AU, aiming to solve the problems of low detection accuracy and inaccurate detection results during emotion detection.

[0005] A method for detecting the degree of happiness based on AU, the method for detecting the degree of happiness based on AU includes:

[0006] Obtain the collected face pictures, and perform label processing on the face pictures based on facial action units, expression types and happiness degree scores to obtain training samples;

[0007] Construct an expression classification loss function, a facial action unit loss function and a happiness degree score loss function;

[0008] Combine the expression classification loss function, the facial action unit loss function and the happiness degree score loss function to obtain a target loss function;

[0009] Based on the target loss function, and use the training samples to train the backbone network to obtain an initial model;

[0010] Adjust the output data of the initial model to obtain a happiness degree detection model;

[0011] In response to a happiness degree detection instruction for the target face in the picture to be processed, input the picture to be processed into the happiness degree detection model for processing to obtain the target expression and target happiness degree score of the target face.

[0012] An AU-based happiness detection device, the AU-based happiness detection device includes:

[0013] A processing unit, configured to obtain the collected face images, and perform tagging processing on the face images based on facial action units, expression types, and happiness degree scores to obtain training samples;

[0014] A construction unit, configured to construct an expression classification loss function, a facial action unit loss function, and a happiness degree score loss function;

[0015] A combination unit, configured to combine the expression classification loss function, the facial action unit loss function, and the happiness degree score loss function to obtain a target loss function;

[0016] A training unit, configured to train a backbone network based on the target loss function and using the training samples to obtain an initial model;

[0017] An adjustment unit, configured to adjust the output data of the initial model to obtain a happiness degree detection model;

[0018] The processing unit is further configured to, in response to a happiness degree detection instruction for a target face in a to-be-processed image, input the to-be-processed image into the happiness degree detection model for processing to obtain the target expression and the target happiness degree score of the target face.

[0019] A computer device, the computer device includes:

[0020] A memory, storing at least one instruction; and

[0021] A processor, executing the instructions stored in the memory to implement the AU-based happiness detection method.

[0022] A computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in a computer device to implement the AU-based happiness detection method.

[0023] As can be seen from the above technical solutions, the present invention can perform tagging processing on face images based on facial action units, expression types, and happiness degree scores, enabling the trained happiness degree detection model to comprehensively detect multi-dimensional factors, effectively making up for the deficiencies of insufficient model continuity and poor robustness caused by the lack of degree annotation in data and relying only on simple labels. During the training process, the constructed expression classification loss function, facial action unit loss function, and happiness degree score loss function are used as the target losses, significantly improving the accuracy of the model detection results and the model stability. In the field of medical and health, it can assist in more refined detection of the user's happiness degree, so it has high application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flowchart of a preferred embodiment of the AU-based happiness degree detection method of the present invention.

[0025] Figure 2 It is a functional module diagram of a preferred embodiment of the AU-based happiness degree detection device of the present invention.

[0026] Figure 3 It is a schematic structural diagram of a computer device of a preferred embodiment for implementing the AU-based happiness degree detection method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] As Figure 1 shown, it is a flowchart of a preferred embodiment of the AU-based happiness degree detection method of the present invention. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.

[0029] The AU-based happiness degree detection method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0030] The computer device can be any electronic product that can interact with users, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.

[0031] The computer device may also include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0032] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0033] Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0034] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0035] The network where the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.

[0036] Specifically, the method for detecting the degree of happiness based on AU (Action Unit) includes:

[0037] S10, obtaining the collected face image, and performing tagging processing on the face image based on the action unit of the face, the type of expression, and the score of the degree of happiness to obtain a training sample.

[0038] In this embodiment, an open source data set may be obtained as the face image, and the face image may also be collected using a web crawler technology.

[0039] In this embodiment, the facial action unit refers to the basic components of facial expressions, which express different emotions and expressions through the movement of facial muscles. These action units can be combined to express all possible facial expressions, such as frowning, pursing lips, etc.

[0040] In this embodiment, the face image is labeled based on the face action unit, expression type and happiness score to obtain training samples including:

[0041] Detect the facial action unit to which the face belongs in each face image;

[0042] According to the facial action unit to which the face belongs in each face image, select AU6 (Cheek Raiser) or AU12 (Lip Corner Puller) as the first label to mark the corresponding face image;

[0043] Detecting the facial expression type of each face image; wherein the facial expression type includes a happy type and an unhappy type;

[0044] Using the expression type of the face in each face image as the second label to mark the corresponding face image;

[0045] Obtaining a face picture whose second label is a happy type, and calculating a happiness degree score of the face in the obtained face picture;

[0046] Using the happiness level score of the face in the acquired face picture as a third label to mark the corresponding face picture;

[0047] All the marked face images are integrated to obtain the training samples.

[0048] Among them, AU6 is used to describe the degree of upward movement of the cheekbones, which is usually related to smiling or happy expressions; AU12 is used to describe pulling up the corners of the mouth, which is also usually related to smiling or expressing happiness.

[0049] Therefore, this embodiment selects AU6 and AU12 to mark the face picture, and subsequently trains the model based on the markings, so that the model can more accurately detect whether the face in the face picture is happy.

[0050] Among them, the expression types include happy type and unhappy type, of course, not limited to the literal "happy" or "unhappy", "happy" and "unhappy", "joyful" and "unhappy" and other words that represent emotions of happiness or not can be used as labels, or, pairs of symbols (such as a represents happy type, b represents unhappy type) or numbers (such as 1 represents happy type, 0 represents unhappy type) can be directly used as labels. The present invention does not limit the actual identifier used for the label.

[0051] The happiness score of the face in the face image obtained by calculation includes:

[0052] For each of the acquired face images, obtain an AU6 intensity value and an AU12 intensity value of each face image;

[0053] The larger intensity value is obtained from the AU6 intensity value and the AU12 intensity value of each face image as the happiness degree score of the corresponding face image.

[0054] It can be seen that only the happy face pictures have labels representing the happiness scores.

[0055] Furthermore, the AU6 intensity value and the AU12 intensity value may be calculated in the following manner, but not limited to:

[0056] 1) Calculate using the coordinates of facial key points.

[0057] For example, the strength of AU12 (lip corner uplift) can be measured by calculating the degree of lip stretching, and the change ratio of the distance between the lip corner point and the mouth center can be defined. The greater the distance change, the higher the strength. For AU6 (cheek lift), it can be measured by the change in the relative position of the key point under the eye socket and the specific point of the cheek. For example, the greater the distance between the two points is reduced, the higher the strength of AU6.

[0058] 2) Extract pixel values ​​from facial images.

[0059] For example, when analyzing AU6, the texture and brightness changes in the cheek area can be used as indicators. Through image filtering and statistical analysis of the changes in pixel values ​​in this area, the more obvious the change, the higher the intensity.

[0060] Of course, in other embodiments, the AU6 intensity value and the AU12 intensity value may also be predicted by deep learning, machine learning, etc., which will not be elaborated here.

[0061] In this embodiment, it is also possible to obtain open-source data with only binary labels of whether one is happy to construct training samples. Since the open-source data already has labels of whether one is happy, it can effectively save the time cost, computing power cost, and labor cost in the subsequent labeling process. Moreover, since the open-source data has been widely used, while effectively improving the labeling efficiency, it can also ensure the accuracy of the labels.

[0062] S11. Construct an expression classification loss function, a facial action unit loss function, and a happiness degree score loss function.

[0063] In this embodiment, the construction of the expression classification loss function, the facial action unit loss function, and the happiness degree score loss function includes:

[0064] The expression classification loss function is constructed using the following formula:

[0065]

[0066] where L em represents the expression classification loss; y e represents the label value of the expression type; represents the predicted value of the expression type detected by the model;

[0067] The facial action unit loss function is constructed using the following formula:

[0068]

[0069] where L AU represents the facial action unit loss; y au represents the label value of the facial action unit; represents the predicted value of the facial action unit detected by the model; γ represents a hyperparameter, which is a coefficient used to balance the uneven distribution of the training samples;

[0070] The happiness degree score loss function is constructed using the following formula:

[0071]

[0072] where L score represents the happiness degree score loss; y score represents the label value of the happiness degree score; represents the predicted value of the happiness degree score detected by the model.

[0073] Through the above embodiments, an expression classification loss function, a facial action unit loss function, and a happiness degree score loss function are respectively established, so that the multi-dimensional loss situation can be fully considered during the model training process, making the model more accurate and more stable.

[0074] S12. Combine the expression classification loss function, the facial action unit loss function, and the happiness score loss function to obtain a target loss function.

[0075] In this embodiment, the step of combining the expression classification loss function, the facial action unit loss function, and the happiness score loss function to obtain a target loss function includes:

[0076] Configure weight values for the expression classification loss function, the facial action unit loss function, and the happiness score loss function according to the influence degrees of these functions, to obtain a first weight corresponding to the expression classification loss function, a second weight corresponding to the facial action unit loss function, and a third weight corresponding to the happiness score loss function;

[0077] Calculate the weighted sum of the expression classification loss function, the facial action unit loss function, and the happiness score loss function according to the first weight, the second weight, and the third weight, to obtain the target loss function.

[0078] For example: The target loss function can be expressed as follows:

[0079] Loss = λ 1 L em + λ 2 L AU + λ 3 L score ;

[0080] where Loss represents the target loss function, λ 1 represents the first weight, λ 2 represents the second weight, and λ 3 represents the third weight.

[0081] In the above embodiment, by configuring the weight values of each loss function according to the influence degree of each loss on the result and calculating the weighted sum of each loss function, the rationality and effectiveness of the constructed loss function can be further improved, so that the trained model has a better detection effect.

[0082] S13. Based on the target loss function, use the training samples to train the backbone network to obtain an initial model.

[0083] In this embodiment, the step of using the training samples to train the backbone network based on the target loss function to obtain an initial model includes:

[0084] Using the second label and the third label as supervision signals to train the backbone network until the loss value of the target loss function converges, and then stop training to obtain the initial model.

[0085] Specifically, using the label representing the expression type and the label representing the score of the degree of happiness as supervision signals to train the model can further improve the training effect of the model.

[0086] S14. Adjust the output data of the initial model to obtain a happiness degree detection model.

[0087] It can be understood that AU can be used as intermediate supervision information during training to help the model better learn the association between facial features and expressions. For example, certain action units such as AU6 and AU12 have a certain connection with expressions such as happiness. By learning these action units, the model can more accurately understand the mapping relationship between facial muscle movement patterns and expressions, thereby improving the accuracy of the model in expression classification and happiness degree scoring.

[0088] However, the focus of the model inference stage is to quickly and accurately give the expression classification and happiness degree score. Since the output of AU requires additional computing resources and time, and it is not directly necessary for the final inference result (expression classification and happiness degree scoring), removing its output can simplify the inference process, improve the inference efficiency, and enable the model to respond more quickly in practical applications.

[0089] Therefore, in this embodiment, the output data of the initial model can also be adjusted to remove the output of the facial action unit, that is, the AU output.

[0090] Specifically, the adjusting the output data of the initial model to obtain a happiness degree detection model includes:

[0091] When defining the forward propagation function of the initial model, control the initial model to only return the detected expression classification and the score of the degree of happiness, and not return the facial action unit.

[0092] S15. In response to a happiness degree detection instruction for the target face in the to-be-processed picture, input the to-be-processed picture into the happiness degree detection model for processing to obtain the target expression and the target score of the degree of happiness of the target face.

[0093] In this embodiment, the to-be-processed picture can be a patient image collected through a specified medical platform in the medical field, or a real-time image collected during the customer service process in the financial field, or a customer image, etc.

[0094] For example, when the output result is: happy, 9, it indicates that the emotion corresponding to the target face is happy and the degree is relatively high; when the output result is: not happy, it indicates that the emotion corresponding to the target face is not happy, and at this time, the happy degree score is not output.

[0095] Through the above embodiments, it is not only possible to accurately identify whether the face in the face picture is happy, but also to give the happy degree to assist in more refined analysis and processing in the field of medical and health. At the same time, by comprehensively detecting the happy degree in combination with multi-dimensional factors such as facial action units, expression types, and happy degree scores, the accuracy and stability of the happy degree scoring are significantly improved, and good application potential and innovation value are shown in related technical fields of sentiment analysis (such as the detection of the emotions of customer service and users in the financial field) and medical and health fields.

[0096] For example: in the field of medical and health, when conducting psychological counseling for users, the happy degree detection method of this embodiment can accurately evaluate the real-time happy degree of users to assist medical staff such as psychologists to better understand the user's state and adjust the psychological counseling strategy in a timely manner according to the user's real-time emotional state, thereby improving the treatment effect.

[0097] It can be seen from the above technical solutions that the present invention can perform tagging processing on face pictures based on facial action units, expression types, and happy degree scores, so that the trained happy degree detection model can comprehensively detect multi-dimensional factors, effectively making up for the deficiencies of insufficient model continuity and poor robustness caused by the lack of degree annotation in data and relying only on simple labels; in the training process, using the constructed expression classification loss function, facial action unit loss function, and happy degree score loss function as the target loss significantly improves the accuracy of the model detection result and the model stability, and can assist in more refined detection of the happy degree of users in the field of medical and health, so it has high application value.

[0098] As Figure 2 shown, it is a functional module diagram of a preferred embodiment of the AU-based happy degree detection device of the present invention. The AU-based happy degree detection device 11 includes a processing unit 110, a construction unit 111, a combination unit 112, a training unit 113, and an adjustment unit 114. The modules / units referred to in the present invention refer to a series of computer program segments that can be executed by a processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in the subsequent embodiments. Specifically:

[0099] The processing unit 110 is used to acquire the collected face pictures and perform tagging processing on the face pictures based on facial action units, expression types, and happy degree scores to obtain training samples.

[0100] In this embodiment, an open-source dataset can be obtained as the face images, and the face images can also be collected by using web crawler technology.

[0101] In this embodiment, the facial action unit refers to the basic component of facial expressions, and different emotions and expressions are represented by the movement of facial muscles. These action units can be combined to represent all possible expressions of facial expressions, such as frowning, pursing the lips, etc.

[0102] In this embodiment, the processing unit 110 performs tagging processing on the face images based on the facial action units, expression types, and happiness degree scores, and the obtained training samples include:

[0103] Detect the facial action unit to which the face in each face image belongs;

[0104] Select AU6 (Cheek Raiser) or AU12 (Lip Corner Puller) as the first label according to the facial action unit to which the face in each face image belongs, and mark the corresponding face image;

[0105] Detect the expression type to which the face in each face image belongs; wherein, the expression types include happy type and unhappy type;

[0106] Use the expression type to which the face in each face image belongs as the second label to mark the corresponding face image;

[0107] Obtain the face images with the second label being the happy type, and calculate the happiness degree score of the face in the obtained face images;

[0108] Use the happiness degree score of the face in the obtained face images as the third label to mark the corresponding face images;

[0109] Integrate all the marked face images to obtain the training samples.

[0110] Among them, AU6 is used to describe the degree of upward movement of the cheekbones, which is usually related to a smiling or happy expression; AU12 is used to describe pulling up the corners of the mouth, which is usually also related to smiling or expressing happiness.

[0111] Therefore, in this embodiment, AU6 and AU12 are selected to mark the face images, and the model can be trained according to this mark, enabling the model to more accurately detect whether the face in the face image is happy.

[0112] Among them, the expression types include happy types and unhappy types. Of course, it is not limited to the literal "happy" or "unhappy". Words expressing emotions such as "happy" and "unhappy", "joyful" and "joyless" can be used as labels. Or, paired symbols (such as a representing the happy type and b representing the unhappy type) or numbers (such as 1 representing the happy type and 0 representing the unhappy type) can also be directly used as labels. The present invention places no restrictions on the actual identifiers used for the labels.

[0113] Among them, the calculated happiness degree scores of the human faces in the obtained human face pictures include:

[0114] For each human face picture in the obtained human face pictures, obtain the AU6 intensity value and the AU12 intensity value of each human face picture;

[0115] Obtain the larger intensity value from the AU6 intensity value and the AU12 intensity value of each human face picture as the happiness degree score of the corresponding human face picture.

[0116] It can be seen that only the human face pictures of the happy type have labels representing the happiness degree scores.

[0117] Moreover, the AU6 intensity value and the AU12 intensity value can be calculated in, but not limited to, the following ways:

[0118] 1) Calculate using the facial key point coordinates.

[0119] For example: Measure the intensity of AU12 (lip corner pulling up) by calculating the degree of lip stretching. The change ratio of the distance between the lip corner point and the center of the mouth can be defined. The greater the distance change, the higher the intensity. For AU6 (cheek raising), it can be measured by the change in the relative position of the key point below the eye socket and a specific point on the cheek. For example, the greater the degree of reduction in the distance between the two points, the higher the AU6 intensity.

[0120] 2) Extract from the facial image pixel values.

[0121] For example: When analyzing AU6, the texture and brightness changes in the cheek area can be used as indicators. By image filtering and statistical analysis of the pixel value changes in this area, the more obvious the changes, the higher the intensity.

[0122] Of course, in other embodiments, the AU6 intensity value and the AU12 intensity value can also be predicted by means of deep learning, machine learning, etc., which will not be elaborated here.

[0123] In this embodiment, open-source data with only binary labels indicating whether a person is happy can also be obtained to construct training samples. Since the open-source data already has labels indicating whether a person is happy, it can effectively save the time cost, computing power cost, and labor cost in the subsequent labeling process. Moreover, since the open-source data has been widely used, while effectively improving the labeling efficiency, it can also ensure the accuracy of the labels.

[0124] The construction unit 111 is configured to construct an expression classification loss function, a facial action unit loss function, and a happiness degree score loss function.

[0125] In this embodiment, the construction of the expression classification loss function, the facial action unit loss function, and the happiness degree score loss function by the construction unit 111 includes:

[0126] The expression classification loss function is constructed using the following formula:

[0127]

[0128] where L em represents the expression classification loss; y e represents the label value of the expression type; represents the predicted value of the expression type detected by the model;

[0129] The facial action unit loss function is constructed using the following formula:

[0130]

[0131] where L AU represents the facial action unit loss; y au represents the label value of the facial action unit; represents the predicted value of the facial action unit detected by the model; γ represents a hyperparameter, which is a coefficient used to balance the uneven distribution of the training samples;

[0132] The happiness degree score loss function is constructed using the following formula:

[0133]

[0134] where L score represents the happiness degree score loss; y score represents the label value of the happiness degree score; represents the predicted value of the happiness degree score detected by the model.

[0135] Through the above embodiments, an expression classification loss function, a facial action unit loss function, and a happiness degree score loss function are respectively established, so that in the model training process, multi-dimensional loss situations can be fully considered, making the model more accurate and more stable.

[0136] The combination unit 112 is configured to combine the expression classification loss function, the facial action unit loss function, and the happiness score loss function to obtain a target loss function.

[0137] In this embodiment, the combination unit 112 combines the expression classification loss function, the facial action unit loss function, and the happiness score loss function to obtain a target loss function, which includes:

[0138] Configuring weight values for the expression classification loss function, the facial action unit loss function, and the happiness score loss function according to the influence degrees of the expression classification loss function, the facial action unit loss function, and the happiness score loss function, to obtain a first weight corresponding to the expression classification loss function, a second weight corresponding to the facial action unit loss function, and a third weight corresponding to the happiness score loss function;

[0139] Calculating the weighted sum of the expression classification loss function, the facial action unit loss function, and the happiness score loss function according to the first weight, the second weight, and the third weight to obtain the target loss function.

[0140] For example: the target loss function can be expressed as follows:

[0141] Loss = λ 1 L em + λ 2 L AU + λ 3 L score ;

[0142] where Loss represents the target loss function, λ 1 represents the first weight, λ 2 represents the second weight, and λ 3 represents the third weight.

[0143] In the above embodiment, configuring the weight values of each loss function according to the influence degree of each loss on the result and calculating the weighted sum of each loss function can further improve the rationality and effectiveness of the constructed loss function, so that the trained model has a better detection effect.

[0144] The training unit 113 is configured to train the backbone network based on the target loss function and using the training samples to obtain an initial model.

[0145] In this embodiment, the training unit 113 trains the backbone network based on the target loss function and using the training samples to obtain an initial model, which includes:

[0146] Use the second label and the third label as supervision signals to train the backbone network, and stop training until the loss value of the target loss function converges, obtaining the initial model.

[0147] Specifically, using the label representing the expression type and the label representing the happiness degree score as supervision signals to train the model can further improve the training effect of the model.

[0148] The adjustment unit 114 is used to adjust the output data of the initial model to obtain a happiness degree detection model.

[0149] It can be understood that AU can be used as intermediate supervision information during the training process to help the model better learn the association between facial features and expressions. For example, there is a certain connection between specific action units such as AU6 and AU12 and expressions such as happiness. By learning these action units, the model can more accurately understand the mapping relationship between facial muscle movement patterns and expressions, thereby improving the accuracy of the model in expression classification and happiness degree scoring.

[0150] However, the focus of the model inference stage is to quickly and accurately give the expression classification and happiness degree score. Since the output of AU requires additional computing resources and time, and it is not directly necessary for the final inference result (expression classification and happiness degree scoring), removing its output can simplify the inference process, improve the inference efficiency, and enable the model to respond more quickly in practical applications.

[0151] Therefore, this embodiment can also adjust the output data of the initial model to remove the output of the facial action unit, that is, the AU output.

[0152] Specifically, the adjustment unit 114 adjusts the output data of the initial model to obtain a happiness degree detection model, including:

[0153] When defining the forward propagation function of the initial model, control the initial model to only return the detected expression classification and happiness degree score, and not return the facial action unit.

[0154] The processing unit 110 is further configured to, in response to a happiness degree detection instruction for the target face in the to-be-processed picture, input the to-be-processed picture into the happiness degree detection model for processing, obtaining the target expression and target happiness degree score of the target face.

[0155] In this embodiment, the to-be-processed picture can be a patient image collected through a specified medical platform in the medical field, or a real-time image collected during the customer service process in the financial field, or a customer image, etc.

[0156] For example, when the output result is: happy, 9, it indicates that the emotion corresponding to the target face is happy and the degree is relatively high; when the output result is: not happy, it indicates that the emotion corresponding to the target face is not happy, and at this time, the happy degree score is not output.

[0157] Through the above embodiments, not only can it accurately identify whether the face in the face picture is happy, but also it can give the degree of happiness to assist in more refined analysis and processing in the field of medical and health. At the same time, by comprehensively detecting the degree of happiness in combination with multi-dimensional factors such as facial action units, expression types, and happiness degree scores, the accuracy and stability of the happiness degree scoring are significantly improved, and good application potential and innovation value are demonstrated in related technical fields of sentiment analysis (such as the detection of the emotions of customer service and users in the financial field) and medical and health fields.

[0158] For example: in the field of medical and health, when conducting psychological counseling for users, using the happiness degree detection method of this embodiment can accurately evaluate the real-time happiness degree of users to assist medical staff such as psychologists to better understand the user's state, and timely adjust the psychological counseling strategy according to the user's real-time emotional state, thereby improving the treatment effect.

[0159] It can be seen from the above technical solutions that the present invention can perform tagging processing on face pictures based on facial action units, expression types, and happiness degree scores, enabling the trained happiness degree detection model to comprehensively detect multi-dimensional factors, effectively making up for the deficiencies of insufficient model continuity and poor robustness caused by the lack of degree annotation in data and relying only on simple labels; during the training process, using the constructed expression classification loss function, facial action unit loss function, and happiness degree score loss function as the target loss significantly improves the accuracy of the model detection result and the model stability, and can assist in more refined detection of the happiness degree of users in the field of medical and health, so it has high application value.

[0160] As Figure 3 shown, it is a schematic structural diagram of a computer device of a preferred embodiment for implementing the AU-based happiness degree detection method of the present invention.

[0161] The computer device 1 may include a memory 12, a processor 13, and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as an AU-based happiness degree detection program.

[0162] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1, and does not constitute a limitation on the computer device 1. The computer device 1 can be either a bus structure or a star structure. The computer device 1 can also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the computer device 1 can also include input / output devices, network access devices, etc.

[0163] It should be noted that the computer device 1 is only an example. Other existing or future possible electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are hereby incorporated by reference.

[0164] Among them, the memory 12 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 12 can be an internal storage unit of the computer device 1 in some embodiments. For example, the mobile hard disk of the computer device 1. The memory 12 can also be an external storage device of the computer device 1 in other embodiments. For example, a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the computer device 1. The memory 12 can be used not only to store application software installed on the computer device 1 and various types of data, such as the code of the happiness degree detection program based on AU, etc., but also to temporarily store the data that has been output or will be output.

[0165] The processor 13 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the computer device 1, connecting various components of the entire computer device 1 through various interfaces and lines, and by running or executing the programs or modules stored in the memory 12 (such as executing the happiness degree detection program based on AU, etc.), and calling the data stored in the memory 12, to execute various functions of the computer device 1 and process data.

[0166] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-described embodiments of various AU-based happiness degree detection methods, such as Figure 1 the steps shown.

[0167] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a processing unit 110, a construction unit 111, a combination unit 112, a training unit 113, and an adjustment unit 114.

[0168] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the AU-based happiness degree detection methods described in various embodiments of the present invention.

[0169] If the modules / units integrated in the computer device 1 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-described method embodiments may be implemented.

[0170] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory, etc.

[0171] Further, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of blockchain nodes, etc.

[0172] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information on a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0173] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, in Figure 3 it is only represented by a single straight line, but it does not mean that there is only one bus or one type of bus. The bus is arranged to realize the connection and communication between the memory 12 and at least one processor 13, etc.

[0174] Although not shown, the computer device 1 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 13 through a power management device, so as to realize functions such as charging management, discharging management, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, a power status indicator, etc. The computer device 1 may further include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0175] Furthermore, the computer device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.

[0176] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device 1 and to display a visual user interface.

[0177] It should be understood that the above embodiments are only for illustrative purposes and are not limited by this structure in the scope of the patent application.

[0178] Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the computer device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0179] In combination with Figure 1 , the memory 12 in the computer device 1 stores a plurality of instructions to implement a method for detecting the happiness level based on AU, and the processor 13 can execute the plurality of instructions to implement:

[0180] Obtain the collected face images, and perform tagging processing on the face images based on facial action units, expression types, and happiness level scores to obtain training samples;

[0181] Construct an expression classification loss function, a facial action unit loss function, and a happiness level score loss function;

[0182] Combine the expression classification loss function, the facial action unit loss function, and the happiness level score loss function to obtain a target loss function;

[0183] Based on the target loss function and using the training samples to train the backbone network to obtain an initial model;

[0184] Adjust the output data of the initial model to obtain a happiness level detection model;

[0185] In response to a happiness level detection instruction for the target face in the to-be-processed image, input the to-be-processed image into the happiness level detection model for processing to obtain the target expression and target happiness level score of the target face.

[0186] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1Descriptions of relevant steps in corresponding embodiments are not elaborated here.

[0187] It should be noted that all data involved in this case are legally obtained. The non-company software tools or components appearing in the embodiments of this application are only for illustrative introduction and do not represent actual use.

[0188] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0189] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0190] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0191] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional module.

[0192] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0193] Therefore, in any case, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

[0194] In addition, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. A plurality of elements or devices described in the present invention can also be implemented by one element or device through software or hardware. The terms first, second, etc. are used to denote names and do not denote any particular order.

[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for detecting happiness based on AU, characterized in that: The AU-based happiness detection method includes: Acquire collected face pictures, and label the face pictures based on face action units, expression types, and happiness scores to obtain training samples; Construct expression classification loss function, facial action unit loss function and happiness score loss function; Combining the expression classification loss function, the facial action unit loss function and the happiness score loss function to obtain a target loss function; Based on the target loss function, a backbone network is trained using the training samples to obtain an initial model; Adjusting the output data of the initial model to obtain a happiness degree detection model; In response to a happiness detection instruction for a target face in a to-be-processed image, the to-be-processed image is input into the happiness detection model for processing to obtain a target expression and a target happiness score for the target face.

2. The AU-based happiness detection method according to claim 1, characterized in that: The face images are labeled based on the face action units, expression types and happiness scores to obtain training samples including: Detect the facial action unit to which the face belongs in each face image; According to the facial action unit to which the face in each face image belongs, select AU6 or AU12 as the first label to mark the corresponding face image; Detecting the expression type of the face in each face image; wherein the expression type includes a happy type and an unhappy type; Using the expression type of the face in each face image as the second label to mark the corresponding face image; Obtaining a face picture whose second label is a happy type, and calculating a happiness level score of the face in the obtained face picture; Using the happiness level score of the face in the acquired face picture as a third label to mark the corresponding face picture; All the labeled face images are integrated to obtain the training samples.

3. The AU-based happiness detection method according to claim 2, characterized in that: The happiness score of the face in the face picture obtained by calculation includes: For each of the acquired face images, obtain an AU6 intensity value and an AU12 intensity value of each face image; The larger intensity value is obtained from the AU6 intensity value and the AU12 intensity value of each face image as the happiness degree score of the corresponding face image.

4. The AU-based happiness detection method according to claim 1, characterized in that: The construction of expression classification loss function, face action unit loss function and happiness score loss function includes: The expression classification loss function is constructed using the following formula: Among them, L em represents the expression classification loss; y e The tag value indicating the type of expression; The predicted value representing the type of expression detected by the model; The face action unit loss function is constructed using the following formula: Among them, L AU represents the face action unit loss; y au Represents the label value of the face action unit; represents the predicted value of the facial action unit detected by the model; γ represents a hyperparameter, which is a coefficient used to balance the uneven distribution of the training samples; The happiness score loss function is constructed using the following formula: Among them, L score Indicates the loss of happiness score; y score The label value representing the happiness score; Represents the predicted value of the happiness score detected by the model.

5. The AU-based happiness detection method according to claim 1, characterized in that: The target loss function obtained by combining the expression classification loss function, the face action unit loss function and the happiness score loss function comprises: According to the influence degree of the expression classification loss function, the facial action unit loss function and the happiness score loss function, weight values ​​are configured for the expression classification loss function, the facial action unit loss function and the happiness score loss function to obtain a first weight corresponding to the expression classification loss function, a second weight corresponding to the facial action unit loss function, and a third weight corresponding to the happiness score loss function; The weighted sum of the expression classification loss function, the facial action unit loss function and the happiness score loss function is calculated according to the first weight, the second weight and the third weight to obtain the target loss function.

6. The AU-based happiness detection method according to claim 2, characterized in that: The step of training the backbone network based on the target loss function and using the training samples to obtain the initial model includes: The backbone network is trained using the second label and the third label as supervisory signals, and the training is stopped when the loss value of the target loss function reaches convergence, thereby obtaining the initial model.

7. The AU-based happiness detection method according to claim 1, characterized in that: The adjusting the output data of the initial model to obtain the happiness degree detection model comprises: When defining the forward propagation function of the initial model, the initial model is controlled to return only the detected expression classification and happiness score, and not the facial action unit.

8. A happiness detection device based on AU, characterized in that: The AU-based happiness detection device comprises: A processing unit, used to obtain collected face pictures, and label the face pictures based on face action units, expression types and happiness scores to obtain training samples; A construction unit is used to construct an expression classification loss function, a face action unit loss function, and a happiness score loss function; A combining unit, used for combining the expression classification loss function, the face action unit loss function and the happiness score loss function to obtain a target loss function; A training unit, used to train a backbone network based on the target loss function and using the training samples to obtain an initial model; An adjustment unit, used for adjusting the output data of the initial model to obtain a happiness degree detection model; The processing unit is also used to respond to a happiness detection instruction for a target face in a to-be-processed image, input the to-be-processed image into the happiness detection model for processing, and obtain a target expression and a target happiness score for the target face.

9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes the instructions stored in the memory to implement the AU-based happiness detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the AU-based happiness detection method according to any one of claims 1 to 7.