Emotion recognition method and system based on facial expression semi-supervised learning
Through the semi-supervised learning method of facial expressions, combined with OpenPose and expert group review, a semi-supervised learning infant expression recognition network was built, which solved the problem of medium and high cost labeling of infant expression recognition and achieved efficient and accurate infant emotion recognition.
Patent Information
- Application Number
- CN202410063903.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-29
AI Technical Summary
In the existing infant expression recognition algorithm, the collection and labeling process of action unit label data sets is expensive and the quality is poor, resulting in inaccurate recognition results and difficult to meet the needs of low-cost and high-quality training.
Using a semi-supervised learning method based on facial expressions, OpenPose is used to mark facial key points, and manually mark and review it in combination with the relationship between facial action units and key points, to build a semi-supervised learning expression recognition network to reduce the number of samples and the difficulty of labeling.
The low-cost training of infant expression recognition models has been achieved, which improves the efficiency and accuracy of expression recognition, especially the recognition accuracy of infant emotions has reached more than 94%.
Smart Images

Figure CN120388400A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of facial expression recognition, and particularly relates to an emotion recognition method and system based on semi-supervised learning of facial expressions. Background Art
[0002] Facial expression recognition refers to extracting the facial expressions of the object to be recognized from a set of static images or dynamic video sequences and dividing them into action unit parts. Combining the pediatric clinical and infant psychology theories and experiences of expert groups for manual annotation and review of partial data, and through deep learning, the calculator can automatically judge the emotion of the object to be recognized. Facial expression recognition based on deep learning is a research hotspot in the current fields of computer vision and affective computing.
[0003] Facial expressions can naturally and directly express people's emotions without relying on language. Research shows that in daily communication, emotions convey more than 50% of the information, and these emotions are often expressed through subtle facial expressions, such as slightly opening the lips and raising the cheeks when happy, and lowering the eyebrows and raising the chin when sad. Therefore, the internationally renowned psychologist Ekman and his research partners proposed a correspondence between different expressions and different facial muscle movements - the Facial Expression Coding System. Based on the anatomical features of the face, the human face is segmented into multiple independent and interrelated action units (ActionUnits, AU), and the main areas controlled by each action unit, the movement characteristics of the action units, and their corresponding expressions are studied.
[0004] Currently, the vast majority of research is based on adults. For infants and young children who lack the ability to express themselves verbally, they express their emotions and needs more through facial expressions than adults. And infancy is a sensitive period for the regulation and development of human language and behavioral abilities. Good care experiences in the early life can create a good foundation for their lifelong social behavior and communication abilities. If not understood by the caregiver, the physiological or pathophysiological needs of infants such as pain, cold, heat, and hunger cannot be met in a timely manner, and the hidden risks are great. Long-term accumulation may cause abnormal programming of the infants' social behavior and communication abilities themselves, leading to disorders in their behavior and language. Epidemiological investigations have found that the incidence of behavioral and language disorders before the age of 6 is 10% - 21%. Therefore, it is very important to understand how to read the facial expressions of infants and young children, which is the cornerstone for infants to grow up healthily. Recognizing the emotions of infants and young children is not only a technology but also a bridge for psychological communication with infants.
[0005] For the recognition of facial expressions and emotions of infants and young children, there are the following difficulties:
[0006] (1) Incomplete development of facial muscles: The facial muscles of infants are under development, and the facial expression movements are not as developed and coordinated as those of adults. Therefore, adult expression training cannot be directly used.
[0007] (2) Differences in facial expressions: During development, there are certain changes in the development of emotions, and there are also certain changes in facial expressions. Infants need to be distinguished from toddlers and children as well.
[0008] (3) Incomplete hair development: The feature extraction of the eyebrow area is relatively difficult, and it is necessary to focus on extracting and analyzing the changes in facial muscles and action units.
[0009] (4) The annotators lack professional qualifications, and the credibility of the annotation results is low, resulting in low quality of most expression datasets and inaccurate recognition results.
[0010] In summary, in the training process of existing algorithms for infant and toddler emotion recognition, the collection and annotation of the action unit label datasets required are costly and of poor quality, unable to meet the low-cost and high-quality training requirements of infant and toddler emotion recognition algorithms, and it is difficult to ensure their popularization in families and clinics. Currently, there are very few people who master infant expression and analysis techniques. After returning from studying at the Harvard University Infant Behavior Research Institute, the applicant, based on clinical practice of contacting tens of thousands of infants, pioneered the infant expression database and the mother-infant interaction behavior observation and evaluation system, and is one of the earliest teams in China to conduct pioneering research on the neural activity synchronization of mother-infant interaction using fNIRS, with rich pediatric clinical and infant and toddler psychology experience. Through extracting facial features, combining expert experience, and adopting the method of AI extraction + manual annotation and review, the present invention provides theoretical and technical guidance for the difficulties in analyzing, evaluating, and precisely intervening in infant emotions. Summary of the Invention
[0011] The technical problem to be solved by the present invention is to provide a method and system for emotion recognition based on semi-supervised learning of facial expressions to solve the technical problems of high cost in collecting and annotating infant and toddler expression samples based on action units and low reliability of annotation results at the present stage, which can reduce the craving of infant and toddler emotion recognition models for manually annotated data and realize the low-cost training of infant and toddler emotion recognition networks.
[0012] The present invention adopts the following technical solutions:
[0013] A method for emotion recognition based on semi-supervised learning of facial expressions, comprising the following steps:
[0014] S1. Convert video data into a video frame sequence and preprocess the video frame sequence;
[0015] S2. Use OpenPose to perform facial key point annotation on the video frame sequence obtained in step S1 to construct an expression dataset with key point annotation;
[0016] S3. Select some key frames from the expression dataset with key point annotations obtained in step S2, and perform manual annotation and review of action units by combining facial action units and the relationship between facial key points and action units to establish an expression and emotion dataset;
[0017] S4. Based on the expression dataset with key point annotations obtained in step S2 and the expression and emotion dataset obtained in step S3, construct an expression recognition network for semi-supervised learning based on manual annotation and review by a panel of experts on partial facial action units;
[0018] S5. Input the newly collected video data into the semi-supervised learning expression recognition network obtained in step S4 to obtain the emotion recognition result.
[0019] Preferably, step S1 is specifically as follows:
[0020] Convert the video into a video frame sequence, preprocess the converted video frame sequence, obtain the effective time period, then use Opencv to frame the effective time period in the video, and after reading the video, save each frame of the video within the effective time period.
[0021] More preferably, the preprocessing of the converted video frame sequence is specifically as follows:
[0022] Write the time period with a complete face into a file and perform rotation and / or cropping. Crop the background area by judging the position of the infant's face to increase the proportion of the infant in the image, and at the same time rotate to make the infant's face located above the image.
[0023] Preferably, step S2 is specifically as follows:
[0024] Extract features from the given face image, find the key points of the face, and use OpenPose to detect the facial key points; then extract the internal key points except the facial contour; the key point labels of all images are saved in the same file, and each line in the file corresponds to the position of the key points (x1, y1, x2, y2...) in one image.
[0025] More preferably, the key points include: eyebrows, eyes, nose, and mouth. Rotate the entire image according to the key point coordinates of the inner corners of the two eyes to make the positions of the two eyes on the same horizontal line and correct the infant's face.
[0026] Preferably, step S3 is specifically as follows:
[0027] Divide the partial muscle movements of the face according to the anatomical characteristics of the human face, represent each movement as one or more action units, and define an action unit set;
[0028] The Facial Action Coding System defines action units based on facial anatomical features, and there is a priori positional relationship between the center of each action unit and facial key points;
[0029] According to the definition of action units and the relationship between action units and key points, action unit labels are annotated for some images, and the action unit labels of all images are saved in a file. Each line corresponds to whether six types of action units appear in an image. If it appears, it is marked as 1, otherwise it is marked as 0.
[0030] More preferably, each action unit is a basic facial action related to one or more local facial muscle actions; in a facial expression, only one action unit appears or multiple action units appear simultaneously.
[0031] Preferably, step S4 is specifically as follows:
[0032] S401. Decompose the features obtained from the image into features independent of key points and features related to key points;
[0033] S402. Regenerate the feature maps of the dataset obtained in step S3 and the feature maps of the dataset obtained in step S2 into a newly generated image containing the transferred labeled image action unit labels and the appearance of the unlabeled image;
[0034] S403. Obtain the newly generated image by combining the information related to action units in the labeled image with the information unrelated to action units in the unlabeled image and maximizing the detection performance of action units in the unlabeled image.
[0035] Preferably, in step S5, the expression data for which action units are not recognized is added as new data to the infant expression dataset with key point annotations constructed in step S2.
[0036] In a second aspect, an embodiment of the present invention provides an emotion recognition system based on semi-supervised learning of facial expressions, including:
[0037] A data module that converts video data into a video frame sequence and preprocesses the video frame sequence;
[0038] A labeling module that uses OpenPose to perform facial key point labeling on the video frame sequence obtained by the data module and constructs an expression dataset with key point annotations;
[0039] A selection module that selects some key frames from the expression dataset with key point annotations obtained by the labeling module, combines facial action units and the relationship between facial key points and action units to perform manual annotation and review of action units, and establishes an expression and emotion dataset;
[0040] A network module constructs an expression recognition network for semi-supervised learning based on the expression dataset with key point annotations obtained by the annotation module and the expression and emotion dataset obtained by the selection module, and is manually annotated and reviewed by an expert group of partial facial action units.
[0041] A recognition module inputs the newly collected video data into the semi-supervised learning expression recognition network obtained by the network module to obtain an emotion recognition result.
[0042] Compared with the prior art, the present invention has at least the following beneficial effects:
[0043] An emotion recognition method based on semi-supervised learning of facial expressions uses OpenPose to perform facial key point annotation on the video frame sequence obtained in step S1, selects some key frames, combines facial action units and the relationship between facial key points and action units for manual annotation and review of action units, constructs an expression recognition network for semi-supervised learning based on manual annotation and review by an expert group of partial facial action units, and obtains an emotion recognition result. The method of the present invention reduces the demand for samples and the annotation difficulty, realizes more efficient training of the emotion recognition method, and thus improves the deployment efficiency.
[0044] Further, the facial area of the infant is cropped out and enlarged, and areas such as the background are cropped off, increasing the proportion of the infant in the image and improving the recognition efficiency.
[0045] Further, OpenPose is used for facial key point detection, the key point coordinates of the inner corners of the two eyes are used, and the entire image is rotated so that the positions of the two eyes are on the same horizontal line to correct the face of the infant.
[0046] Further, an expression dataset with annotations by an expert group of action units is constructed. The cost of manually marking AUs is huge, and often only some samples in the dataset have complete AU labels, while the remaining samples do not have or only have partial AU labels; and the currently existing and publicly available datasets are mainly for adults and lack infant datasets. Therefore, the present invention uses a camera to record the video of the infant's medical treatment, divides it into a video frame sequence, performs key point annotation on it using OpenPose, and constitutes an infant dataset with key point annotation and manual annotation by an expert group of action units by extracting key frames and performing AU annotation, which reduces the number of samples to be annotated, and the intermediate frame data can be interpolated through the key frames, improving the annotation efficiency.
[0047] Furthermore, an infant expression recognition network based on semi-supervised learning with partial action unit annotation is constructed. When using deep learning technology to solve problems, the biggest difficulty is that there are many parameters to be trained in the model. Therefore, a large amount of training data is required. Although there is a large amount of data, this data often has no labels, which makes it impossible to train machine learning models. Aiming at the lack of infant datasets with AU labels, this project adopts a semi-supervised learning method to reduce the workload of manually labeling action units for infant expression data, and uses the dataset that has been marked and reviewed by the existing action unit expert group to achieve semi-supervised learning of action units on the infant expression dataset.
[0048] It can be understood that the beneficial effects of the second aspect can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0049] In summary, the present invention reduces the difficulty of annotation, comprehensively utilizes the labeled data and unlabeled data, and constructs a semi-supervised learning method to improve the accuracy of infant facial emotion recognition.
[0050] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0051] Figure 1 Schematic diagram of the method of the present invention;
[0052] Figure 2 Schematic diagram of the definition of facial key points;
[0053] Figure 3 Schematic diagram of the detection effect of facial key points;
[0054] Figure 4 Schematic diagram of the comparison before and after facial correction, where (a) is before correction and (b) is after correction;
[0055] Figure 5 Schematic diagram of the definition of facial action units;
[0056] Figure 6 Schematic diagram of using action unit combinations to define expressions;
[0057] Figure 7 Schematic diagram of the relationship between AU and key points;
[0058] Figure 8 Schematic diagram of the basic structure of the network;
[0059] Figure 9 Schematic diagram of a computer device provided by an embodiment of the present invention;
[0060] Figure 10A block diagram of a chip provided according to an embodiment of the present invention. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] In the description of the present invention, it should be understood that the terms "including" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0063] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0064] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the preceding and following related objects.
[0065] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0066] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0067] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the accompanying drawings. These figures are not drawn to scale, where certain details are enlarged for the purpose of clear expression, and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art can additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0068] The present invention provides an emotion recognition method based on semi-supervised learning of facial expressions, which constructs an infant emotion recognition network based on semi-supervised learning with manual annotation and review by an expert group; the expert interprets the network recognition results; it solves the problems of the lack of infant datasets in the current infant emotion recognition task, especially the lack of data for infants aged 0 to 6 months, as well as the large workload of facial action unit annotation, and the annotators do not have pediatric clinical professional experience and the qualifications of infant psychologists, resulting in low credibility of the annotation results and poor recognition effects. It provides an idea for the training of infant emotion recognition under partial samples, realizes the low-cost training of the infant facial expression and emotion recognition network, and provides theoretical and technical guidance for interpreting infant emotions.
[0069] Please refer to Figure 1 , an emotion recognition method based on semi-supervised learning of facial expressions according to the present invention, includes the following steps:
[0070] S1. Collect a large number of infant videos, convert them into a video frame sequence, automatically locate the infant's face position, and then perform preprocessing such as rotation and cropping;
[0071] Using a camera to collect infant videos requires a large amount of high-quality infant expression data as the basis for network training.
[0072] First, convert the infant video into a frame sequence, write the time period with the complete face of the infant into a file, after obtaining the effective time period, use Opencv to split the frames of the effective time period in the video. After reading the video, judge whether each frame is within the effective time period, and if so, save it.
[0073] Since the research object of the present invention is the expressions of infants, the picture data used contains many non-expression features, and the proportion of infants in the whole picture is too small, which will greatly reduce the accuracy of expression recognition. Therefore, the image is cropped by locating the infant's face; and due to the influence of the camera position during shooting, the infant's face is located at the bottom of the image, so the image is rotated.
[0074] S2. For the pre - processed images of the video frame sequence converted in step S1, use OpenPose to perform facial key point annotation on all the images, and construct an infant expression dataset with key point annotation;
[0075] Facial key points refer to the main feature points of each part of the face, generally contour points and corner points. Please refer to Figure 2 for the definition of internal key points other than the facial contour.
[0076] Facial key point detection refers to finding the key points of the face from a given face image through feature extraction. Using OpenPose to detect facial key points is the core in the process of expression recognition. If the key points of the infant's face cannot be extracted, or are interfered by factors such as the loss of key point features, it will seriously affect the accuracy of infant expression recognition. Use OpenPose to detect key points in the image, and the detection results are as Figure 3 shown;
[0077] Subsequently, extract the internal key points other than the facial contour; the key point labels of all images are saved in the same file, and each line in the file corresponds to the positions of the key points (x1, y1, x2, y2...) in one image.
[0078] The key points include: eyebrows, eyes, nose, mouth, etc.
[0079] According to the coordinates of the key points at the inner corners of the two eyes, rotate the entire image so that the positions of the two eyes are on the same horizontal line; then locate the infant's facial area according to the key point positions, and crop the background, body and other areas. As Figure 4 shown in the comparison diagram before and after the transformation, the transformed image is used as the input of the network; finally, perform corresponding position transformation on the key point coordinates as well.
[0080] S3. For the infant expression dataset with key point annotation constructed in step S2, select some key frames, and have a panel of experts perform manual annotation and review of action units based on rich experience, in combination with the definition of facial action units and the relationship between facial key points and action units, to establish an infant expression and emotion dataset manually marked and reviewed by a trained and qualified panel of experts;
[0081] Divide some of the facial muscle movements according to the anatomical characteristics of the human face, represent each movement as one or more action units (AUs), and define an action unit set, as Figure 5 shown.
[0082] Among them, each AU is a basic facial action related to one or more local facial muscle movements. In a facial expression, only one AU appears or multiple AUs appear simultaneously.Figure 6 The AUs that may appear in each expression are listed in the table. These AUs appear simultaneously or partially simultaneously in a certain expression, and these AUs can be combined into a variety of different facial expressions. For example, a happy expression is represented by the combination of AU6, AU12, and AU25, and an angry expression is represented by the combination of AU4, AU5, AU7, and AU23.
[0083] The Facial Action Coding System defines AUs based on facial anatomical features, and there is a priori positional relationship between the center of each AU and the facial key points. As Figure 7 shown, the position of AU1 is defined at 1 / 2 scale above the inner side of the eyebrows, and the position of AU4 is defined at 1 / 3 scale below the center of the eyebrows. The position of the center of each AU can be accurately defined using facial key points, and local features related to the AU can be extracted therefrom.
[0084] According to the definition of the AU and the relationship between the AU and the key points, AU labels are annotated for some images, and the AU labels of all images are saved in a file. Each line corresponds to whether six types of AUs appear in an image. If they appear, they are marked as 1, otherwise they are marked as 0.
[0085] S4. For the dataset of infant expressions with key point annotations constructed in step S2, combined with the dataset of infant expressions and emotions established in step S3 with manual marking and review by a trained and qualified expert group, construct an infant expression recognition network based on semi-supervised learning with manual annotation and review by an expert group of partial facial action units;
[0086] The present invention proposes a key point adversarial loss algorithm as follows:
[0087] S401. Decompose the features obtained from the image into features irrelevant to the key points and features relevant to the key points;
[0088] S402. Regenerate the decomposed labeled feature images (feature maps of the dataset in step S3) and unlabeled feature images (feature maps of the dataset in step S2) into a newly generated image containing the transferred AU labels of the labeled image and the appearance of the unlabeled image;
[0089] Among them, the labeled image has accurate AU and key point labels, while the unlabeled image only has accurate key point labels.
[0090] S403. The newly generated image in step S402 is obtained by combining the AU-related information in the labeled image with the AU-irrelevant information in the unlabeled image and maximizing the AU detection performance of the unlabeled image.
[0091] Please refer to Figure 8, after extracting the basic features of a pair of labeled data and unlabeled data, it is further decomposed into features related to key points and features unrelated to key points. After exchanging the features related to key points, a new image is regenerated. At this time, the target domain image contains the features related to key points of the labeled data and the features unrelated to key points of the unlabeled data. Since the action units are related to the key points, the recombination of the information related to the action units in the labeled data and the information unrelated to the action units in the unlabeled data can learn the action units of the automatically generated labeled data on the newly generated image by maximizing the action unit detection performance of the unlabeled data.
[0092] S5. For the infant expression recognition network based on semi-supervised learning with partial action unit annotation constructed in step S4, after identifying the action units of the newly collected infant expression data, an expert interprets the results, and the infant expression data with poor recognition effect is added as new data to the training of the network.
[0093] Using the infant expression recognition network based on semi-supervised learning with partial action unit annotation constructed in step S4, identify the action units of the newly collected infant expression data. The expert interprets the network recognition confidence results of each action unit, and the infant expression data with poor recognition effect is added as new data to the infant expression dataset with key point annotation constructed in step S2).
[0094] In another embodiment of the present invention, an emotion recognition system based on semi-supervised learning of facial expressions is provided. This system can be used to implement the above-mentioned emotion recognition method based on semi-supervised learning of facial expressions. Specifically, the emotion recognition system based on semi-supervised learning of facial expressions includes a data module, an annotation module, a selection module, a network module, and an identification module.
[0095] Among them, the data module converts video data into a video frame sequence, locates the facial position and then performs preprocessing;
[0096] The annotation module uses OpenPose to perform facial key point annotation on the video frame sequence obtained by the data module, and constructs an expression dataset with key point annotation;
[0097] The selection module selects some key frames from the expression dataset with key point annotation obtained by the annotation module, combines facial action units and the relationship between facial key points and action units to perform manual annotation and review of action units, and establishes an expression and emotion dataset;
[0098] The network module constructs an expression recognition network based on semi-supervised learning with manual annotation and review by a partial facial action unit expert group based on the expression dataset with key point annotation obtained by the annotation module and the expression and emotion dataset obtained by the selection module;
[0099] The recognition module inputs the newly collected video data into the semi-supervised learning facial expression recognition network obtained by the network module to obtain the emotion recognition result.
[0100] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention can be used for the operations of the emotion recognition method based on semi-supervised learning of facial expressions, including:
[0101] Convert the video data into a video frame sequence and preprocess the video frame sequence; use OpenPose to perform facial key point annotation on the video frame sequence to construct an expression data set with key point annotation; select some key frames from the expression data set with key point annotation, and combine the facial action units and the relationship between the facial key points and the action units to perform manual annotation and review of the action units, and establish an expression and emotion data set; based on the expression data set with key point annotation and the expression and emotion data set, construct a semi-supervised learning facial expression recognition network based on manual annotation and review by a partial facial action unit expert group; input the newly collected video data into the semi-supervised learning facial expression recognition network to obtain the emotion recognition result.
[0102] Please refer to Figure 9, the terminal device is a computer device. The computer device 60 in this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, it implements the method for calculating the fluid composition in the reservoir stimulation wellbore in the embodiment. To avoid repetition, it will not be elaborated here one by one. Alternatively, when the computer program 63 is executed by the processor 61, it implements the functions of each model / unit in the emotion recognition system based on semi-supervised learning of facial expressions in the embodiment. To avoid repetition, it will not be elaborated here one by one.
[0103] The computer device 60 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 8 merely examples of the computer device 60, which do not constitute a limitation to the computer device 60, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.
[0104] The so-called processor 61 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0105] The memory 62 may be an internal storage unit of the computer device 60, such as the hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk equipped on the computer device 60, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0106] Furthermore, the memory 62 may also include both the internal storage unit and the external storage device of the computer device 60. The memory 62 is used to store the computer program and other programs and data required by the computer device. The memory 62 may also be used to temporarily store the data that has been output or will be output.
[0107] Please refer to Figure 10 , the terminal device is a chip. The chip 600 of this embodiment includes a processor 622, the number of which can be one or more, and a memory 632 for storing computer programs executable by the processor 622. The computer programs stored in the memory 632 may include one or more modules each corresponding to a set of instructions. In addition, the processor 622 may be configured to execute the computer program to perform the above-mentioned emotion recognition method based on semi-supervised learning of facial expressions.
[0108] In addition, the chip 600 may further include a power supply component 626 and a communication component 650. The power supply component 626 may be configured to perform power management of the chip 600, and the communication component 650 may be configured to implement communication of the chip 600, for example, wired or wireless communication. In addition, the chip 600 may further include an input / output (I / O) interface 658. The chip 600 may operate based on an operating system stored in the memory 632.
[0109] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the terminal device for storing programs and data. It can be understood that the computer-readable storage medium here may include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space that stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions may be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here may be a high-speed RAM memory or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory.
[0110] One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the emotion recognition method based on semi-supervised learning of facial expressions in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor as follows:
[0111] Convert the video data into a sequence of video frames and preprocess the sequence of video frames; use OpenPose to perform facial key point annotation on the sequence of video frames to construct an expression dataset with key point annotation; select some key frames from the expression dataset with key point annotation, and combine the facial action units and the relationship between the facial key points and the action units to perform manual annotation and review of the action units, and establish an expression and emotion dataset; based on the expression dataset with key point annotation and the expression and emotion dataset, construct an expression recognition network for semi-supervised learning based on manual annotation and review by a partial facial action unit expert group; input the newly collected video data into the semi-supervised learning expression recognition network to obtain the emotion recognition result.
[0112] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0113] The laboratory version of the present invention has been field-tested in a hospital outpatient scenario, can run successfully and correctly judge the emotions of infants, with an accuracy rate of over 94%, and can assist medical staff and parents in keenly perceiving the emotional changes and needs of infants.
[0114] To ensure a comprehensive evaluation of the performance of the present invention, 100 representative samples, including infants of different months of age, genders and emotional states, were selected. Field tests were conducted in a hospital outpatient environment to better evaluate the actual application effect and verify the robustness of the system. Professional pediatric clinicians and psychologists conducted on-site evaluations and judgments of the expressions and emotions of infants to provide the standard answers for the recognition results. By comparing the system recognition results, the accuracy rate was calculated based on the ratio of the number of samples correctly judged by the system to the total number of samples. Through this objective calculation method, the performance of the system of the present invention was accurately evaluated. The final result was that 94 cases were consistent with the expert judgment results, that is, the accuracy rate reached 94%.
[0115] In summary, the emotion recognition method and system based on semi-supervised learning of facial expressions according to the present invention, based on the semi-supervised learning algorithm, readjusts the emotion classification criteria according to the unique emotion characteristics of infants, and uses a mixed infant facial expression dataset of unlabeled data and partially labeled data for training, which can effectively improve the accuracy and stability of emotion recognition, and provides a new solution for emotion analysis in the field of artificial intelligence, especially for the problem of difficult emotion analysis of infants.
[0116] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be described in detail here.
[0117] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0118] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0119] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0120] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0121] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0122] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0123] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or Figure 1 one or more blocks.
[0124] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or Figure 1 one or more blocks.
[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or Figure 1 one or more blocks.
[0126] The above content is only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.
Claims
1. A method for emotion recognition based on semi-supervised learning of facial expressions, characterized in that It includes the following steps: S1. Convert the video data into a video frame sequence and preprocess the video frame sequence; S2. Use OpenPose to perform facial key point annotation on the video frame sequence obtained in step S1, and construct an expression dataset with key point annotation; S3. Select some key frames from the expression dataset with key point annotation obtained in step S2, and perform manual annotation and review of action units in combination with facial action units and the relationship between facial key points and action units to establish an expression and emotion dataset; S4. Based on the expression dataset with key point annotation obtained in step S2 and the expression and emotion dataset obtained in step S3, construct an expression recognition network for semi-supervised learning based on manual annotation and review by a partial facial action unit expert group; S5. Input the newly collected video data into the semi-supervised learning expression recognition network obtained in step S4 to obtain an emotion recognition result.
2. The emotion recognition method based on semi-supervised learning of facial expressions according to claim 1, wherein Step S1 is specifically as follows: Convert the video into a video frame sequence, preprocess the converted video frame sequence, and after obtaining the effective time period, use Opencv to frame the effective time period in the video. After reading the video, save each frame of the video within the effective time period.
3. The emotion recognition method based on semi-supervised learning of facial expressions according to claim 2, characterized in that, The preprocessing of the converted video frame sequence is specifically as follows: Write the time period with a complete face into a file and perform rotation and / or cropping. By judging the position of the infant's face, crop the background area to increase the proportion of the infant in the image. At the same time, rotate to make the infant's face located above the image.
4. The emotion recognition method based on semi-supervised learning of facial expressions according to claim 1, characterized in that Step S2 is specifically as follows: Extract features from the given face image, find the key points of the face, and use OpenPose to detect the facial key points; then extract the internal key points except the facial contour; the key point labels of all images are saved in the same file, and each line in the file corresponds to the position of the key points (x1, y1, x2, y2...) in one image.
5. The emotion recognition method based on semi-supervised learning of facial expressions according to claim 4, characterized in that the key points It includes: Eyebrows, eyes, nose, and mouth. Rotate the entire image according to the key point coordinates of the inner corners of the two eyes to make the positions of the two eyes on the same horizontal line, and correct the infant's face.
6. The emotion recognition method based on semi-supervised learning of facial expressions according to claim 1, characterized in that Step S3 is specifically as follows: Divide the partial muscle movements of the face according to the anatomical characteristics of the human face, represent each movement as one or more action units, and define an action unit set; The facial action coding system defines action units based on facial anatomical characteristics, and there is a prior positional relationship between the center of each action unit and the facial key points; According to the definition of action units and the relationship between action units and key points, label the action unit labels for some images, and save the action unit labels of all images in a file. Each line corresponds to whether six types of action units appear in one image. If they appear, they are marked as 1, otherwise they are marked as 0.
7. The emotion recognition method based on semi-supervised learning of facial expressions according to claim 6, wherein Each action unit is a basic facial action related to one or more local facial muscle movements; in a facial expression, only one action unit appears or multiple action units appear simultaneously.
8. The emotion recognition method based on semi-supervised learning of facial expressions according to claim 1, characterized in that, Step S4 is specifically as follows: S401. Decompose the features obtained from the image into features irrelevant to the key points and features relevant to the key points; S402. Regenerate the feature maps of the dataset obtained in step S3 and the feature maps of the dataset obtained in step S2 into a newly generated image that includes the transferred labeled image action unit labels and the appearance of the unlabeled image; S403. Obtain the newly generated image by combining the information related to the action units in the labeled image with the information unrelated to the action units in the unlabeled image and maximizing the action unit detection performance of the unlabeled image.
9. The emotion recognition method based on semi-supervised learning of facial expressions according to claim 1, characterized in that In step S5, the expression data for which the action units are not recognized is added as new data to the infant expression dataset with key point annotations constructed in step S2.
10. An emotion recognition system based on semi-supervised learning of facial expressions, characterized in that, It includes: A data module that converts video data into a sequence of video frames, preprocesses the data after locating the facial positions; An annotation module that uses OpenPose to perform facial key point annotation on the sequence of video frames obtained by the data module, and constructs an expression dataset with key point annotations; A selection module that selects some key frames from the expression dataset with key point annotations obtained by the annotation module, performs manual annotation and review of the action units by combining the facial action units and the relationship between the facial key points and the action units, and establishes an expression and emotion dataset; A network module that constructs an expression recognition network for semi-supervised learning based on the manual annotation and review of some facial action unit expert groups, based on the expression dataset with key point annotations obtained by the annotation module and the expression and emotion dataset obtained by the selection module; A recognition module that inputs the newly collected video data into the semi-supervised learning expression recognition network obtained by the network module to obtain the emotion recognition result.
Citation Information
Cited By
Expression recognition model training method and system based on head posture and facial key points
CN120913009A