Facial expression recognition method and medical care robot

By adopting feature extraction methods of active appearance model and perspective information adjustment in facial expression recognition, combined with the dual training technology of the target random forest model, the problem of low recognition accuracy and reliability caused by diversification of perspectives in traditional technologies is solved, and higher facial expression recognition accuracy and accurate adjustment of medical and nursing robot working modes are achieved.

CN120220211AActive Publication Date: 2025-06-27河北博健科技有限公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510313966.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

Traditional facial expression recognition technology has low accuracy and reliability in the absence of side data sets and diversified perspectives.

Method used

The feature extraction method based on the active appearance model is adopted, and the expression characteristics are adjusted in combination with perspective information, and the target random forest model is used to double training using real data and data generated by the generative adversarial network to improve the accuracy of expression recognition.

Benefits of technology

By considering the perspective information and the dual training model, the recognition error caused by the perspective is reduced, the accuracy and reliability of facial expression recognition are improved, so as to more accurately judge the emotional state of the target person and adjust the working mode of the medical and nursing robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220211A_ABST
    Figure CN120220211A_ABST
Patent Text Reader

Abstract

The invention provides a facial expression recognition method and a medical care robot, and belongs to the technical field of image recognition, and the method comprises the steps: carrying out the feature extraction of a first image based on an active appearance model, and obtaining a first expression feature; wherein the first image is a face image of the target person; determining a target expression feature based on the view angle of the first image and the first expression feature, and inputting the target expression feature into a target random forest model to obtain an expression recognition result of the target person; wherein the target random forest model is a model obtained through training of a first data set and a second data set, data in the first data set are real expression feature data and corresponding expression recognition results, and data in the second data set are expression feature data generated based on a generative adversarial network model and corresponding expression recognition results; and determining the working mode of the medical care robot based on the expression recognition result of the target person. The accuracy and reliability of facial expression recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of image recognition, and more particularly, relates to a facial expression recognition method and a medical and elderly care robot. Background Art

[0002] In the current context of the accelerating development of an aging society, the combination of medical care and elderly care is becoming increasingly important. As a key auxiliary tool in this model, medical and elderly care robots can provide all-round care and services for the elderly.

[0003] Traditional facial expression recognition relies on traditional machine learning algorithms. However, due to the lack of side datasets and the diverse perspectives obtained when recognizing facial expressions, the accuracy and reliability of facial expression recognition are relatively low.

[0004] Therefore, an accurate and reliable facial expression recognition method is needed. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a facial expression recognition method and a medical and elderly care robot to improve the accuracy and reliability of facial expression recognition.

[0006] In the first aspect of the embodiments of the present disclosure, a facial expression recognition method is provided, which is applied to a medical and elderly care robot and includes: Performing feature extraction on a first image based on an active appearance model to obtain first expression features; wherein, the first image is a facial image of a target person; Determining target expression features based on the perspective of the first image and the first expression features, and inputting the target expression features into a target random forest model to obtain an expression recognition result of the target person; wherein, the target random forest model is a model trained by a first dataset and a second dataset, the data in the first dataset are real expression feature data and corresponding expression recognition results, and the data in the second dataset are expression feature data generated based on a generative adversarial network model and corresponding expression recognition results; Determining the working mode of the medical and elderly care robot based on the expression recognition result of the target person.

[0007] In the second aspect of the embodiments of the present disclosure, a facial expression recognition device is provided, which is applied to a medical and elderly care robot and includes: A feature extraction module, configured to perform feature extraction on a first image based on an active appearance model to obtain first expression features; wherein, the first image is a facial image of a target person; An expression recognition module, configured to determine target expression features based on the perspective of the first image and the first expression features, and input the target expression features into a target random forest model to obtain an expression recognition result of the target person; wherein, the target random forest model is a model trained by a first data set and a second data set, the data in the first data set are real expression feature data and corresponding expression recognition results, and the data in the second data set are expression feature data generated based on a generative adversarial network model and corresponding expression recognition results; A control module, configured to determine the working mode of the medical care robot based on the expression recognition result of the target person.

[0008] In a third aspect of the embodiments of the present disclosure, a medical care robot is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned facial expression recognition method are implemented.

[0009] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned facial expression recognition method are implemented.

[0010] The beneficial effects of the facial expression recognition method and the medical care robot provided by the embodiments of the present disclosure are as follows: By considering the perspective of the first image and adjusting the expression features accordingly, the present disclosure can reduce the recognition error caused by the perspective and enhance the robustness of the present disclosure. The target random forest model is trained by both real data and data generated by a generative adversarial network, which enhances its generalization ability and improves the accuracy and reliability of facial expression recognition. When the accuracy and reliability of facial expression recognition are very high, the medical care robot can more accurately judge the emotional state of the target person, thereby adjusting its working mode and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is a schematic flowchart of a facial expression recognition method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of a facial expression recognition device provided by an embodiment of the present disclosure; Figure 3Schematic block diagram of a medical and elderly care robot provided by an embodiment of the present disclosure. Detailed implementation manners

[0013] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0014] To make the purpose, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.

[0015] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a facial expression recognition method provided by an embodiment of the present disclosure. This method is applied to a medical and elderly care robot and includes: S101: Extract features from the first image based on the active appearance model to obtain the first expression feature; wherein, the first image is a facial image of a target person.

[0016] In this embodiment, the active appearance model is established by statistical analysis of a large number of face images to establish a shape model and a texture model of the face. The shape model describes the distribution and variation rules of key feature points of the face (such as the corners of the eyes, the corners of the mouth, the tip of the nose, etc.); the texture model describes the gray scale or color information on the surface of the face. When applied, the active appearance model will find the shape and texture that best match the model in the input face image, so as to extract features that can reflect information such as facial expressions and postures of the face. The first image refers to the facial image of the target person, which can be obtained based on the camera configured on the medical and elderly care robot. The target person refers to the person to be recognized, such as the elderly in a nursing home or patients undergoing rehabilitation in a hospital, etc. The first expression feature refers to the feature data that can reflect the current facial expression of the target person, such as the four corner points of the eyes, the end points of the eyebrows, the points of the corners of the mouth, etc.

[0017] S102: Determine the target expression feature based on the viewing angle of the first image and the first expression feature, and input the target expression feature into the target random forest model to obtain the expression recognition result of the target person; wherein, the target random forest model is a model trained with a first data set and a second data set. The data in the first data set are real expression feature data and corresponding expression recognition results, and the data in the second data set are expression feature data generated based on the generative adversarial network model and corresponding expression recognition results.

[0018] In this embodiment, the viewing angle of the first image can be determined by certain features in the image, such as simple features like the number of eyes or the number of ears. The viewing angle of the first image (such as front view, side view, etc.) affects the integrity and accuracy of the expression features. In different viewing angles, some facial expression features may be blocked or distorted, resulting in information loss. Therefore, it is necessary to process the first expression feature in combination with the viewing angle of the first image to obtain more comprehensive and accurate expression features.

[0019] Specifically, it is explained that determining the target expression feature based on the viewing angle of the first image and the first expression feature includes: In response to the viewing angle of the first image being a front view angle, determining the first expression feature as the target expression feature; In response to the viewing angle of the first image being a side view angle, generating a second expression feature based on the first expression feature, and determining the first expression feature and the second expression feature as the target expression feature; wherein, the second expression feature corresponds to the front view angle of the first image.

[0020] In this embodiment, considering that if the first image is a front view angle, since the front view angle can relatively completely display the facial expression information, the first expression feature can usually better reflect the expression of the target person. At this time, the first expression feature can be directly determined as the target expression feature. In a non-front view angle, the first expression feature is incomplete and needs to be processed. For example, based on the first expression feature, a second expression feature can be generated through a mapping or generation method (such as based on a pre-trained mapping relationship or generation model), and then the first expression feature and the second expression feature are combined to form the target expression feature to supplement the missing expression information. The second expression feature is the expression feature corresponding to the front view angle of the first expression feature.

[0021] After the foregoing description, the target expression feature can be determined. Then, inputting the target expression feature into the target random forest model, the expression recognition result of the target person can be obtained. Among them, the random forest is an ensemble learning method composed of multiple decision trees. Each decision tree learns based on different feature subsets and sample subsets during the training process, and finally makes a final classification decision by comprehensively (such as voting) the output results of multiple decision trees. In this disclosure, the target random forest model will judge the expression category of the target person according to the input target expression feature.

[0022] The training data of the target random forest model are the first data set and the second data set. The first data set contains real expression feature data and corresponding expression recognition results. The data in the first data set are collected in the actual scenario and have relatively high authenticity and reliability. The second data set is the expression feature data and corresponding expression recognition results generated based on the generative adversarial network model.

[0023] It should be noted that, due to the scarcity of side-face expression data in reality, the proportion of expression feature data from the frontal view in the first dataset is greater than that from the side view. The purpose of the generative adversarial network is to generate and augment the expression feature data from the side view. Therefore, the proportion of expression feature data from the side view in the second dataset is greater than that from the frontal view, that is, the proportion of expression feature data from the side view in the first dataset is less than that in the second dataset. The data in the second dataset may or may not include the expression feature data from the frontal view. That is, whether the generative adversarial network generates the expression feature data from the frontal view can be determined according to the quantity of the expression feature data from the frontal view in the first dataset. If the quantity is greater than the preset quantity, generation may not be performed. If it is less than the preset quantity, generation can be carried out until the preset quantity or a certain value higher than the preset quantity is reached. However, it is not advisable to generate too much, so as not to let the random forest model learn too much generated data during training, resulting in a reduction in the training effect.

[0024] S103: Determine the working mode of the medical care robot based on the expression recognition result of the target person.

[0025] In this embodiment, considering that different expressions often reflect different physical and mental states and needs of the target person, the medical care robot needs to adjust its working mode according to these different situations. For example: In response to the expression recognition result of the target person being a positive emotion, set the working mode of the medical care robot to the companionship and entertainment mode; In response to the expression recognition result of the target person being a negative emotion, determine the working mode of the medical care robot as the soothing and caring mode; In response to the expression recognition result of the target person being pain, determine the working mode of the medical care robot as the emergency handling mode.

[0026] In this embodiment, the companionship and entertainment mode can be playing relaxing and pleasant music, having a pleasant conversation with the target person, telling jokes or recommending some interesting videos. The soothing and caring mode can be actively communicating with the target person, giving comforting and encouraging words, asking if help is needed, or providing some relaxation guidance, such as deep breathing exercises, meditation guidance, etc., to help the target person relieve negative emotions. The emergency handling mode can be promptly notifying medical staff or relevant caregivers and informing them that the target person may have physical discomfort.

[0027] As can be seen from the above, by considering the perspective of the first image and adjusting the expression features accordingly, the present disclosure can reduce the recognition error caused by the perspective and enhance the robustness of the present disclosure. The target random forest model is trained twice with real data and data generated by the generative adversarial network, which enhances its generalization ability and improves the accuracy and reliability of facial expression recognition. When the accuracy and reliability of facial expression recognition are very high, the medical care robot can more accurately judge the emotional state of the target person, thereby adjusting its working mode and improving the user experience.

[0028] As can be seen from the foregoing, the second expression feature is generated based on the first expression feature. Specifically, in one embodiment of the present disclosure, generating the second expression feature based on the first expression feature includes: Generating the second expression feature based on the first expression feature and the first mapping relationship, where the first mapping relationship is determined based on the third data set, and the data in the third data set includes multiple groups of expression feature data with the same expression but different perspectives.

[0029] In this embodiment, the data in the third data set are multiple groups of expression feature data with the same expression but different perspectives. The geometric transformation parameters between different feature points can be statistically obtained by analyzing a large number of side and front face image data. Specifically, for the perspective conversion from the side to the front, parameters such as the translation amount, rotation angle, and scaling ratio of the feature points in the x, y, and z axis directions need to be determined. Through the least squares method for fitting and optimization, the feature point mapping based on these parameters can best conform to the actual face morphology, and the first mapping relationship is obtained.

[0030] Alternatively, the second expression feature can also be generated based on a neural network model. Specifically, generating the second expression feature based on the first expression feature includes: Generating the second expression feature based on the first neural network model and the first expression feature, where the first neural network model is trained with the expression feature data in the first number of side and front face images.

[0031] As can be seen from the above, by introducing the first mapping relationship, the present disclosure can generate the second expression feature based on the first expression feature. The first mapping relationship is obtained by analyzing multiple groups of expression feature data with the same expression but different perspectives in the third data set, ensuring the accuracy and reliability of the present disclosure in generating the expression feature data of the front perspective based on the side perspective, thereby improving the accuracy and reliability of the target random forest model, and thus achieving the improvement of the accuracy and reliability of facial expression recognition.

[0032] In one embodiment of the present disclosure, inputting the target expression feature into the target random forest model to obtain the expression recognition result of the target person includes: In response to the perspective of the first image being a side perspective, inputting the first expression feature into the target random forest model to obtain a first voting result, and inputting the second expression feature into the target random forest model to obtain a second voting result; The first voting result and the second voting result are weighted to determine the expression recognition result of the target person.

[0033] In this embodiment, considering that the perspective of the first image is a side perspective, and the expression feature data about the side perspective in the training data of the target random forest is generated based on the generative adversarial network, not only the first expression feature can be input into the target random forest model, but also the second expression feature can be input into the target random forest model to achieve better recognition effect.

[0034] The principle is: as mentioned above, the target random forest model is obtained by training the first data set and the second data set. The first data set contains more expression feature data from the frontal perspective, and the second data set contains more expression feature data from the side perspective. Therefore, the target random forest model can recognize both the expression feature data from the frontal perspective and the expression feature data from the side perspective.

[0035] The random forest model is composed of multiple decision trees. Each decision tree will classify and judge the first expression feature according to the rules it has learned, and give a voting result (that is, a "vote" to determine whether the expression belongs to a certain expression category). The voting results of all decision trees are combined to get the first voting result, which represents the model's judgment on the target person's expression based on the side expression features.

[0036] Similarly, the second expression feature is input into the target random forest model, and each decision tree in the model analyzes and votes on it, and finally summarizes the second voting result, which reflects the model's judgment on the target person's expression based on the frontal perspective expression feature.

[0037] Since the first voting result is based on the facial features of the side view, and some facial features of the side view may be blocked or deformed, the accuracy and reliability of its judgment are limited; the second voting result is based on the generated facial features of the front view, but there may be certain errors in the generation process. Therefore, we cannot simply use one of the voting results as the final facial recognition result, and we need to consider both voting results comprehensively.

[0038] More deeply, considering that the perspective of the first image is a side perspective, and most of the training data under the side perspective is generated data, and the expression feature data of the front perspective is essentially generated, the distribution of weights during weighted calculation needs to be set more accurately.

[0039] Therefore, in an embodiment of the present disclosure, before calculating the weighted values of the first voting result and the second voting result, it further includes: Determining a first facial expression recognition result based on the first voting result; Determining a generated weight based on the first facial expression recognition result, where the sum of the generated weight and the mapped weight is 1, the generated weight is the weight corresponding to the first voting result, and the mapped weight is the weight corresponding to the second voting result.

[0040] In this embodiment, the weights of the first voting result and the second voting result should be related to their reliability. On this basis, considering that the second voting result is obtained based on the mapping relationship or the aforementioned first neural network model, for different facial expressions, the stability of the second voting result is relatively strong, that is, for different facial expressions, the generation effects have little difference.

[0041] For the generative adversarial network model, when generating different facial expressions, there are differences in its difficulty and effect. Simple facial expressions such as smiling and frowning have relatively simple and clear facial muscle movement patterns, and the features are relatively intuitive. After the generator learns a large amount of relevant data, it is relatively easy to capture the key features of these facial expressions, thereby generating relatively realistic facial expression images, and the generation effect is usually good. However, complex facial expressions such as contempt and disdain involve subtle coordinated changes of multiple facial muscle groups, and the expression methods of different individuals may vary significantly. The generator needs to learn more complex feature combinations and change rules, which increases the difficulty of generation and may lead to deficiencies in details, naturalness, etc. of the generated facial expression images, and the generation effect is relatively poor.

[0042] Therefore, the corresponding facial expression can be determined according to the first facial expression recognition result, and then the weight in the weighted calculation can be determined based on the generation effect of this facial expression.

[0043] It can be concluded from the above that the present disclosure analyzes the generation effects of different facial expressions and determines the generated weight and the mapped weight accordingly, enabling the present disclosure to perform dynamic adjustment according to the complexity and generation difficulty of different facial expressions, thereby improving the accuracy and reliability of the weighted calculation result, and further enhancing the accuracy and reliability of facial expression recognition.

[0044] As can be seen from the foregoing, the weights corresponding to the first voting result and the second voting result should be determined according to the generation effect of the facial expression corresponding to the first facial expression recognition result. Specifically, determining the generated weight based on the first facial expression recognition result includes: In response to the generation effect evaluation value being greater than the first evaluation value, increasing the reference value of the generated weight by the first step length to obtain the generated weight; In response to the generation effect evaluation value being less than the second evaluation value, increasing the reference value of the mapped weight by the second step length to obtain the mapped weight; Among them, the generation effect evaluation value is the evaluation value of the generation effect of the expression corresponding to the first expression recognition result, and the first evaluation value is greater than the second evaluation value.

[0045] The first evaluation value and the second evaluation value can be set according to experience.

[0046] In this embodiment, the evaluation value of the generation effect of the generative adversarial network model for different expressions is calculated, and then the weight distribution is determined according to this evaluation value. That is, after the generative adversarial network model is generated, the following steps should be executed: Calculate the mutual information between the expression feature data of various types generated by the generative adversarial network model and the corresponding real expression feature data respectively to obtain the first mutual information; Take the first mutual information as the evaluation value of the generation effect of this expression.

[0047] Mutual information can measure the correlation between two variables. Calculate the mutual information between the generated image and the real image. The higher the mutual information, the stronger the correlation between the generated image and the real image, and the closer the generated expression is to the real expression.

[0048] When the generation effect evaluation value is greater than the first evaluation value, it indicates that the effect of the side expression features generated by the generative adversarial network model is better and more credible in expression recognition. At this time, increase the reference value of the generation weight according to the first step length to obtain the final generation weight; at the same time, in order to make the sum of the weights equal to 1, reduce the reference value of the mapping weight according to the same first step length to obtain the final mapping weight, so as to give the first voting result a greater weight and make its influence on the final expression recognition result greater.

[0049] When the generation effect evaluation value is less than the second evaluation value, it indicates that the effect of the side expression features generated by the generative adversarial network model is poor. At this time, increase the reference value of the mapping weight according to the second step length to obtain the final mapping weight; in order to make the sum of the weights equal to 1, reduce the reference value of the generation weight according to the same second step length to obtain the final generation weight, so that the frontal view expression features (the second voting result) generated based on the mapping relationship have a greater influence in the final expression recognition. The first step length, the second step length, the reference value of the mapping weight, and the reference value of the generation weight can all be determined according to the parameters for solving similar problems or the data in the experimental process.

[0050] It should be noted that if the generation effect evaluation value is between the first evaluation value and the second evaluation value, no adjustment may be made, that is, in response to the generation effect evaluation value being greater than or equal to the second evaluation value and less than or equal to the first evaluation value, determine the reference value of the generation weight as the generation weight and determine the reference value of the mapping weight as the mapping weight.

[0051] The weighted calculation of the first voting result and the second voting result can be used to obtain which expression has the highest number of votes, and the expression with the highest number of votes is output as the target expression recognition result. The process of determining the first expression recognition result based on the first voting result is the same, and the expression with the highest number of votes is output as the result.

[0052] In an embodiment of the present disclosure, in addition to the above step size adjustment method, the generation weight can also be determined based on the first formula, and the first formula is , where is the generation weight, is the minimum value of the generation weight, is the maximum value of the generation weight, is the generation effect evaluation value, is the parameter for adjusting the exponential decay, The larger is when is small, the faster decreases, and when is large, , and can be set according to experience.

[0053] The logic of the first formula is that represents the adjustable range of the generation weight, is the exponential function when is small, approaches 1, approaches 0, and at this time the value of the adjustment part is small, and the generation weight is close to , indicating that the generation weight is low when the generation effect is poor; when increases, approaches 0, approaches 1, and the value of the adjustment part approaches , and the generation weight is close to , meaning that the generation weight is high when the generation effect is good. Normalize the entire adjustment part to ensure that when , the adjustment part can reach , so that the generation weight can reach the maximum value .

[0054] As can be seen from the above, the present disclosure calculates the mutual information between the expression feature data generated by the generative adversarial network model and the real expression feature data as the evaluation value of the generation effect, which can objectively measure the quality of the generated expression. According to the generation effect evaluation value, the generation weight and the mapping weight are dynamically adjusted, so that when the quality of the generated expression is high, a greater weight is given to the first voting result, thereby improving the accuracy and credibility of expression recognition. On the contrary, when the quality of the generated expression is low, the weight of the mapping weight is increased to reduce the recognition error caused by poor generation effect and improve the accuracy and reliability of facial expression recognition.

[0055] A facial expression recognition method corresponding to the above embodiment Figure 2 is a structural block diagram of a facial expression recognition device provided by an embodiment of the present disclosure. For the sake of illustration, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The facial expression recognition device 20 is applied to a medical and nursing robot and includes: a feature extraction module 21, an expression recognition module 22, and a control module 23.

[0056] Among them, the feature extraction module 21 is used to extract features from the first image based on the active appearance model to obtain the first expression feature; wherein, the first image is a facial image of a target person; The expression recognition module 22 is used to determine the target expression feature based on the perspective of the first image and the first expression feature, and input the target expression feature into the target random forest model to obtain the expression recognition result of the target person; wherein, the target random forest model is a model trained by the first data set and the second data set, and the data in the first data set are real expression feature data and corresponding expression recognition results, and the data in the second data set are expression feature data generated based on the generative adversarial network model and corresponding expression recognition results; The control module 23 is used to determine the working mode of the medical and nursing robot based on the expression recognition result of the target person.

[0057] In an embodiment of the present disclosure, the expression recognition module 22 is specifically configured to, in response to the perspective of the first image being a front perspective, determine the first expression feature as the target expression feature; In response to the perspective of the first image being a side perspective, generate a second expression feature based on the first expression feature, and determine the first expression feature and the second expression feature as the target expression feature; wherein, the second expression feature corresponds to the front perspective of the first image.

[0058] In an embodiment of the present disclosure, the expression recognition module 22 is specifically further configured to, in response to the perspective of the first image being a side perspective, input the first expression feature into the target random forest model to obtain a first voting result, and input the second expression feature into the target random forest model to obtain a second voting result; Perform weighted calculation on the first voting result and the second voting result to determine the facial expression recognition result of the target person.

[0059] In an embodiment of the present disclosure, a facial expression recognition device 20 further includes: a weight determination module, configured to determine a first facial expression recognition result based on the first voting result; Determine a generated weight based on the first facial expression recognition result, and the sum of the generated weight and the mapped weight is one. The generated weight is the weight corresponding to the first voting result, and the mapped weight is the weight corresponding to the second voting result.

[0060] In an embodiment of the present disclosure, the weight determination module is specifically configured to, in response to the generated effect evaluation value being greater than the first evaluation value, increase the reference value of the generated weight by a first step length to obtain the generated weight; In response to the generated effect evaluation value being less than the second evaluation value, increase the reference value of the mapped weight by a second step length to obtain the mapped weight; Wherein, the generated effect evaluation value is the evaluation value of the generation effect of the expression corresponding to the first facial expression recognition result, and the first evaluation value is greater than the second evaluation value.

[0061] In an embodiment of the present disclosure, the facial expression recognition module 22 is specifically further configured to generate a second facial expression feature based on the first facial expression feature and the first mapping relationship. The first mapping relationship is determined based on a third data set, and the data in the third data set includes multiple groups of facial expression feature data with the same expression but different perspectives.

[0062] In an embodiment of the present disclosure, the control module 23 is specifically configured to, in response to the facial expression recognition result of the target person being a positive emotion, set the working mode of the medical care robot to the companion entertainment mode; In response to the facial expression recognition result of the target person being a negative emotion, determine the working mode of the medical care robot as the comfort and care mode; In response to the facial expression recognition result of the target person being pain, determine the working mode of the medical care robot as the emergency handling mode.

[0063] See Figure 3 , Figure 3 is a schematic block diagram of a medical care robot provided in an embodiment of the present disclosure. As Figure 3The medical and elderly care robot 300 in the present embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 complete mutual communication through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, for example Figure 2 the functions of the modules 21 to 23 shown.

[0064] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0065] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0066] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0067] In specific implementation, the processors 301, input devices 302, and output devices 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first and second embodiments of a facial expression recognition method provided by the embodiments of the present disclosure, and may also execute the implementation manner of the medical and elderly care robot described in the embodiments of the present disclosure, which will not be elaborated here.

[0068] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0069] The computer-readable storage medium can be the internal storage unit of the medical and elderly care robot in any of the foregoing embodiments, such as the hard disk or memory of the medical and elderly care robot. The computer-readable storage medium can also be an external storage device of the medical and elderly care robot, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the medical and elderly care robot. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the medical and elderly care robot. The computer-readable storage medium is used to store the computer program and other programs and data required by the medical and elderly care robot. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0070] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0071] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described medical and elderly care robot and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0072] In several embodiments provided by this application, it should be understood that the disclosed medical and elderly care robot and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, or can also be in the form of electrical, mechanical or other connections.

[0073] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can also be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.

[0074] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0075] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or replacements, and these modifications or replacements should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A facial expression recognition method, characterized in that: Applied to medical robots, including: Extracting features from a first image based on an active appearance model to obtain a first expression feature; wherein the first image is a facial image of a target person; Determine a target expression feature based on the viewing angle of the first image and the first expression feature, and input the target expression feature into a target random forest model to obtain an expression recognition result of the target person; The target random forest model is a model trained by a first data set and a second data set, the data in the first data set are real expression feature data and corresponding expression recognition results, and the data in the second data set are expression feature data and corresponding expression recognition results generated based on a generative adversarial network model; The working mode of the medical robot is determined based on the expression recognition result of the target person.

2. A facial expression recognition method as claimed in claim 1, characterized in that: The determining of a target expression feature based on the viewing angle of the first image and the first expression feature comprises: In response to the perspective of the first image being a frontal perspective, determining the first expression feature as a target expression feature; In response to the perspective of the first image being a side perspective, a second expression feature is generated based on the first expression feature, and the first expression feature and the second expression feature are determined as target expression features; wherein the second expression feature corresponds to the frontal perspective of the first image.

3. A facial expression recognition method as claimed in claim 2, characterized in that, The step of inputting the target expression feature into a target random forest model to obtain an expression recognition result of the target person includes: In response to the perspective of the first image being a side perspective, inputting the first expression feature into a target random forest model to obtain a first voting result, and inputting the second expression feature into the target random forest model to obtain a second voting result; The first voting result and the second voting result are weightedly calculated to determine the expression recognition result of the target person.

4. A facial expression recognition method as claimed in claim 3, characterized in that, Before performing weighted calculation on the first voting result and the second voting result, the method further includes: Determine a first expression recognition result based on the first voting result; A generation weight is determined based on the first expression recognition result, the sum of the generation weight and the mapping weight is one, the generation weight is the weight corresponding to the first voting result, and the mapping weight is the weight corresponding to the second voting result.

5. A facial expression recognition method as claimed in claim 4, characterized in that, The determining of generating weights based on the first expression recognition result includes: In response to the generation effect evaluation value being greater than the first evaluation value, increasing the reference value of the generation weight according to the first step to obtain the generation weight; In response to the generation effect evaluation value being less than the second evaluation value, increasing the reference value of the mapping weight according to a second step length to obtain a mapping weight; Among them, the generation effect evaluation value is an evaluation value of the generation effect of the expression corresponding to the first expression recognition result, and the first evaluation value is greater than the second evaluation value.

6. A facial expression recognition method as claimed in claim 2, characterized in that, The generating a second expression feature based on the first expression feature comprises: A second expression feature is generated based on the first expression feature and a first mapping relationship, wherein the first mapping relationship is determined based on a third data set, and the data in the third data set are multiple groups of expression feature data with the same expression but different viewing angles.

7. A facial expression recognition method as claimed in claim 1, characterized in that: The determining of the working mode of the medical robot based on the expression recognition result of the target person includes: In response to the expression recognition result of the target person being a positive emotion, setting the working mode of the medical robot to a companionship and entertainment mode; In response to the expression recognition result of the target person being a negative emotion, determining the working mode of the medical robot to be a comforting and caring mode; In response to the expression recognition result of the target person being in pain, the working mode of the medical robot is determined to be an emergency processing mode.

8. A facial expression recognition device, characterized in that: Applied to medical robots, including: A feature extraction module, used to extract features from a first image based on an active appearance model to obtain a first expression feature; wherein the first image is a facial image of a target person; An expression recognition module, used to determine a target expression feature based on the viewing angle of the first image and the first expression feature, and input the target expression feature into a target random forest model to obtain an expression recognition result of the target person; wherein the target random forest model is a model trained by a first data set and a second data set, the data in the first data set are real expression feature data and corresponding expression recognition results, and the data in the second data set are expression feature data generated based on a generative adversarial network model and corresponding expression recognition results; The control module is used to determine the working mode of the medical robot based on the expression recognition result of the target person.

9. A medical robot, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Facial expression recognition method

    CN105469080A

  • Facial expression recognition method, device and equipment based on deep learning

    CN112149651A

  • Fusion-based expression recognition method and device and electronic equipment

    CN112836654A

  • Facial expression classification method and system

    CN119150095A

  • Non-Contact Biometric Facial Recognition applied Management System of each Individual Livestock

    KR102558350B1