Facial expression recognition methods, devices, equipment and readable storage media
By combining an attention mechanism model of facial and environmental feature images in facial expression recognition, the problem of low accuracy in facial expression recognition is solved, and higher accuracy in expression recognition is achieved.
Patent Information
- Application Number
- CN202211198439.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-09-29
AI Technical Summary
The accuracy of facial expression recognition in existing technologies is low, mainly due to the differences between each individual and the differences in the expression of different individuals, which makes recognition difficult.
A pre-defined attention mechanism model is used to extract facial feature images and environmental feature images of the target person from the image to be identified. Combined with local and global attention mechanism models, facial expression recognition is performed through facial feature images and environmental feature images, reducing the negative impact of facial feature images and increasing the objective factors of environmental feature images as the basis for expression recognition.
It improves the accuracy of facial expression recognition by incorporating environmental feature images as the basis for expression recognition, reducing the influence of facial feature images, increasing relevant feature information for recognition, and thus improving the accuracy of recognition.
Smart Images

Figure CN115482573B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of facial recognition technology, and in particular to a facial expression recognition method, apparatus, device, and readable storage medium. Background Technology
[0002] Facial expressions typically reflect a person's inner emotions, and understanding these expressions can facilitate communication between people. Currently, methods exist for recognizing facial expressions by analyzing facial features. However, due to individual differences, different individuals may express the same expression differently, and even within the same individual, different expressions may not differ significantly. This presents challenges for facial recognition technology and results in its relatively low accuracy.
[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, device, and readable storage medium for facial expression recognition, aiming to solve the technical problem of low accuracy in current facial expression recognition technologies.
[0005] To achieve the above objectives, the present invention provides a facial expression recognition method, which includes the following steps:
[0006] Collect images of the target person to be identified;
[0007] Based on a preset attention mechanism model, the facial feature image of the target person and the environmental feature image of the environment in which the target person is located are extracted from the image to be identified;
[0008] Based on the facial feature image and the environmental feature image, facial expression recognition is performed on the target person to obtain the target person's current expression.
[0009] Furthermore, the preset attention mechanism model includes a local attention mechanism model and a global attention mechanism model. The step of extracting the facial feature image of the target person and the environmental feature image of the environment in which the target person is located from the image to be identified based on the preset attention mechanism model includes:
[0010] The local attention mechanism model is used to identify the facial region of the target person in the image to be identified;
[0011] The facial feature image is extracted from the facial region using the local attention mechanism model.
[0012] The facial region is removed from the image to be identified to obtain the environmental region;
[0013] The environmental feature image is extracted from the environmental region using the global attention mechanism model, wherein the environmental feature image includes at least one environmental feature unit.
[0014] Furthermore, the facial feature image is configured with facial weights, and the environmental feature image is configured with an environmental weight set, the environmental weight set including at least one environmental weight corresponding to an environmental feature unit. After the step of extracting the environmental feature image from the environmental region through the global attention mechanism model, the method further includes:
[0015] The environmental weight adjustment parameter is obtained by comparing the maximum environmental weight in each of the environmental feature units with the facial weight.
[0016] The environmental weight adjustment parameter is multiplied by each of the environmental weights to generate new environmental weights;
[0017] Ignore environmental feature units in the environmental feature image whose environmental weight is less than a preset ratio of the facial weight, and generate a new environmental feature image.
[0018] Furthermore, the step of performing facial expression recognition on the target person based on the facial feature image and the environmental feature image includes:
[0019] Synchronize the facial weights and each of the environmental weights to make the facial weights and each of the environmental weights identical;
[0020] The combined feature image is obtained by combining the facial feature image after weight synchronization and the environmental feature image after weight synchronization.
[0021] The combined feature image is input into a preset expression classification model to identify the current expression of the target person, wherein the preset expression classification model includes facial features and environmental features.
[0022] Furthermore, the step of acquiring the image of the target person to be identified includes:
[0023] The original image of the target person in the corresponding field of view of the AR glasses is acquired through AR glasses;
[0024] The original image is preprocessed using a super-resolution reconstruction network to generate an image to be identified.
[0025] Furthermore, after the step of performing facial expression recognition on the target person based on the facial feature image and the environmental feature image, the method further includes:
[0026] Generate the relative position of the facial feature image in the image to be identified;
[0027] Based on the relative position, the target person's current expression is displayed and output through AR glasses.
[0028] Furthermore, the step of displaying and outputting the target person's current expression through AR glasses based on the relative position includes:
[0029] Based on the current facial expression, a preset communication suggestion message is matched;
[0030] The current expression and the preset communication suggestion information are displayed at their relative positions within the display field of view of the AR glasses.
[0031] Furthermore, to achieve the above objectives, the present invention also provides a facial expression recognition device, the facial expression recognition device comprising:
[0032] The acquisition module is used to acquire images of the target person to be identified.
[0033] The extraction module is used to extract the facial feature image of the target person and the environmental feature image of the environment in which the target person is located from the image to be identified based on a preset attention mechanism model;
[0034] The recognition module is used to perform facial expression recognition on the target person based on the facial feature image and the environmental feature image, so as to obtain the current expression of the target person.
[0035] In addition, to achieve the above objectives, the present invention also provides a facial expression recognition device, the facial expression recognition device comprising: a memory, a processor, and a facial expression recognition program stored in the memory and executable on the processor, wherein the facial expression recognition program, when executed by the processor, implements the steps of the facial expression recognition method as described above.
[0036] In addition, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a facial expression recognition program, which, when executed by a processor, implements the steps of the facial expression recognition method described above.
[0037] This invention proposes a facial expression recognition method, apparatus, device, and readable storage medium. Compared to existing methods that only recognize expressions based on the face, this application acquires an image of a target person to be recognized; extracts facial feature images of the target person and environmental feature images of the target person's environment from the image based on a preset attention mechanism model; and performs facial expression recognition on the target person based on the facial feature images and the environmental feature images to obtain the target person's current expression. In reality, the environment can influence a target person's inner emotions, thereby affecting their facial expressions. This invention incorporates environmental feature images (objective factors) as one of the bases for expression recognition, reducing the influence of facial feature images on the results and mitigating the negative impact of relying solely on facial feature images (personal factors) on the recognition results, thus improving the accuracy of facial expression recognition. Furthermore, the inclusion of environmental feature images increases the relevant feature information for expression recognition, further enhancing the accuracy of expression recognition. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention;
[0039] Figure 2 This is a flowchart illustrating the first embodiment of the facial expression recognition method of the present invention;
[0040] Figure 3 This is a flowchart illustrating the second embodiment of the facial expression recognition method of the present invention;
[0041] Figure 4 This is a flowchart illustrating the third embodiment of the facial expression recognition method of the present invention;
[0042] Figure 5 This is a schematic diagram illustrating the effect of AR glasses in the facial expression recognition method of the present invention.
[0043] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0044] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0045] like Figure 1 As shown, Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention.
[0046] The device in this invention embodiment can be VR glasses, or it can be a smartphone, PC, tablet computer, e-book reader, MP4 (Moving Picture Experts Group Audio Layer IV) player, portable computer, or other portable terminal device with image acquisition and display functions.
[0047] like Figure 1 As shown, the device may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0048] Optionally, the device may also include a camera, RF (Radio Frequency) circuitry, sensors, audio circuitry, a WiFi module, and so on. Sensors may include light sensors, motion sensors, and other sensors. Specifically, light sensors may include ambient light sensors and proximity sensors. The ambient light sensor can adjust the display brightness according to the ambient light level, while the proximity sensor can turn off the display and / or backlight when the mobile terminal is moved to the ear. As a type of motion sensor, a gravity accelerometer can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the mobile terminal's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition functions (such as pedometers, taps), etc. Of course, the mobile terminal may also be equipped with other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, which will not be elaborated here.
[0049] Those skilled in the art will understand that Figure 1 The device structure shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0050] like Figure 1As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a facial expression recognition program.
[0051] exist Figure 1 In the device shown, network interface 1004 is mainly used to connect to the backend server and communicate data with it; user interface 1003 is mainly used to connect to the client (user terminal) and communicate data with it; while processor 1001 can be used to call the facial expression recognition program stored in memory 1005 and perform the following operations:
[0052] Collect images of the target person to be identified;
[0053] Based on a preset attention mechanism model, the facial feature image of the target person and the environmental feature image of the environment in which the target person is located are extracted from the image to be identified;
[0054] Based on the facial feature image and the environmental feature image, facial expression recognition is performed on the target person to obtain the target person's current expression.
[0055] Furthermore, the processor 1001 can call the facial expression recognition program stored in the memory 1005 and also perform the following operations:
[0056] The preset attention mechanism model includes a local attention mechanism model and a global attention mechanism model. The step of extracting the facial feature image of the target person and the environmental feature image of the environment in which the target person is located from the image to be identified based on the preset attention mechanism model includes:
[0057] The local attention mechanism model is used to identify the facial region of the target person in the image to be identified;
[0058] The facial feature image is extracted from the facial region using the local attention mechanism model.
[0059] The facial region is removed from the image to be identified to obtain the environmental region;
[0060] The environmental feature image is extracted from the environmental region using the global attention mechanism model, wherein the environmental feature image includes at least one environmental feature unit.
[0061] Furthermore, the processor 1001 can call the facial expression recognition program stored in the memory 1005 and also perform the following operations:
[0062] The facial feature image is configured with facial weights, and the environmental feature image is configured with an environmental weight set, wherein the environmental weight set includes at least one environmental weight corresponding to an environmental feature unit. After the step of extracting the environmental feature image from the environmental region using the global attention mechanism model, the method further includes:
[0063] The environmental weight adjustment parameter is obtained by comparing the maximum environmental weight in each of the environmental feature units with the facial weight.
[0064] The environmental weight adjustment parameter is multiplied by each of the environmental weights to generate new environmental weights;
[0065] Ignore environmental feature units in the environmental feature image whose environmental weight is less than a preset ratio of the facial weight, and generate a new environmental feature image.
[0066] Furthermore, the processor 1001 can call the facial expression recognition program stored in the memory 1005 and also perform the following operations:
[0067] The step of recognizing facial expressions of the target person based on the facial feature image and the environmental feature image includes:
[0068] Synchronize the facial weights and each of the environmental weights to make the facial weights and each of the environmental weights identical;
[0069] The combined feature image is obtained by combining the facial feature image after weight synchronization and the environmental feature image after weight synchronization.
[0070] The combined feature image is input into a preset expression classification model to identify the current expression of the target person, wherein the preset expression classification model includes facial features and environmental features.
[0071] Furthermore, the processor 1001 can call the facial expression recognition program stored in the memory 1005 and also perform the following operations:
[0072] The steps for acquiring the image of the target person to be identified include:
[0073] The original image of the target person in the corresponding field of view of the AR glasses is acquired through AR glasses;
[0074] The original image is preprocessed using a super-resolution reconstruction network to generate an image to be identified.
[0075] Furthermore, the processor 1001 can call the facial expression recognition program stored in the memory 1005 and also perform the following operations:
[0076] After the step of performing facial expression recognition on the target person based on the facial feature image and the environmental feature image, the method further includes:
[0077] Generate the relative position of the facial feature image in the image to be identified;
[0078] Based on the relative position, the target person's current expression is displayed and output through AR glasses.
[0079] Furthermore, the processor 1001 can call the facial expression recognition program stored in the memory 1005 and also perform the following operations:
[0080] The step of displaying and outputting the target person's current expression through AR glasses based on the relative position includes:
[0081] Based on the current facial expression, a preset communication suggestion message is matched;
[0082] The current expression and the preset communication suggestion information are displayed at their relative positions within the display field of view of the AR glasses.
[0083] Reference Figure 2 The first embodiment of the facial expression recognition method of the present invention includes:
[0084] Step S10: Acquire the image of the target person to be identified;
[0085] It should be noted that facial expressions can reflect a person's inner emotions to a certain extent. Therefore, facial recognition can be used to infer a person's emotions, thus enabling the monitoring of their behavior. For example, in medical monitoring scenarios, facial expression recognition technology can be used to identify a patient's expressions to infer their emotions, allowing for timely intervention for patients exhibiting abnormal emotions. In security detection scenarios, facial expression recognition technology can predict the emotions of a target person, preventing extreme situations. It can also be applied to AR (Augmented Reality) glasses, enabling users to communicate by recognizing facial expressions. However, current facial expression recognition methods typically rely solely on the target person's facial area. While this achieves the goal of facial expression recognition, it neglects information related to the expression, such as the environment, leading to inaccurate results. Therefore, this embodiment addresses this issue by incorporating environmental features into facial expression recognition to improve accuracy.
[0086] In this embodiment, a visual sensor installed in the aforementioned scenario can acquire real-time images of the target person to be identified. The target person can be a patient, a person being communicated with, or someone appearing in the visual sensor. Furthermore, the images to be identified can be real-time images of the target person during their daily life activities, such as talking to others or exercising. Therefore, the acquired images to be identified typically include the target person themselves and the environment in which they are located.
[0087] Step S20: Extract the facial feature image and environmental feature image of the target person from the image to be identified based on a preset attention mechanism model;
[0088] In this implementation, a preset attention mechanism model will be used to extract facial feature images of the target person's face region and environmental feature images of the target person's environment from the image to be recognized. Currently, in the field of image recognition technology, there is an attention mechanism technique. This technique can highlight key regions in the image to be recognized, such as assigning different weights to different regions. A higher weight indicates that the image in that region contains more or more important information. Correspondingly, during feature extraction, regions with relatively high weights will be extracted (after assigning weights to the image to be recognized through the attention mechanism, these regions will be encoded; the encoding process converts the original image into image data that can be recognized by the recognition or classification model, and then feature extraction will be performed on the encoded image data; this process will be referred to as feature extraction in the following sections). Specifically, a preset attention mechanism model generates weights for different regions in the image to be recognized, and based on these weights, the facial feature image and the environmental feature image are extracted from the image. After assigning different weights to different regions in the image to be recognized through the preset attention mechanism model, the maximum weight is determined, and the image regions corresponding to weights greater than a preset proportion of the maximum weight are extracted as features to obtain the feature image. It is understandable that when extracting features from the environment in which the target person is located, since the environment contains a large number of environmental feature units, if each environmental feature unit is used as the basis for expression recognition, it will increase the amount of computation and increase the computation time. Moreover, some environmental feature units with low weight in the environment will not be of effective help in recognizing the target person's expression. Therefore, only the regions with relatively high weight are extracted to obtain feature images, thereby improving the system's operating efficiency.
[0089] Furthermore, because the intersection of facial images and environmental images is complex, attention mechanisms typically assign high weight to them. Therefore, the target person's face is extracted as a feature into the feature image. In practical applications, a simple attention mechanism does not distinguish between facial and environmental feature images. Since current face recognition technologies only identify the facial features of the target person, we can first use face recognition technology to identify the target person's facial feature image, and then remove the aforementioned facial feature image from the feature images extracted by the attention mechanism model to obtain the environmental feature image. Since mature face recognition technologies already exist, the process of extracting facial feature images will not be elaborated upon here.
[0090] Step S30: Perform facial expression recognition on the target person based on the facial feature image and the environmental feature image to obtain the target person's current expression.
[0091] In this embodiment, the facial expressions of the target person are identified based on the extracted facial feature images and environmental feature images. The model for recognizing expressions can be a classification model, such as a neural network classification model or an SVM (Support Vector Machine) classification model. There is no limitation on the specific type of classification model; the classification model is trained by extracting facial features and environmental features from the training samples. It is understood that when recognizing the facial expressions of the target person in the image to be recognized, in addition to using facial features as the basis for expression recognition, the environmental features of the target person are also included. While facial expressions are a subjective expression of the target person's inner emotions, under normal circumstances, the environment in which the target person is located will affect their emotions. For example, being in a warm environment and a dangerous environment will have different effects on the target person's emotions, thus affecting the facial expressions they display. That is, the target person is more likely to show a happy expression in a warm environment and a fearful expression in a dangerous environment. Therefore, in this embodiment, objective factors are included as one of the bases for expression recognition, reducing the negative impact of personal factors on the expression recognition results, thereby improving the accuracy of facial expression recognition.
[0092] As an example, the preset attention mechanism model includes a local attention mechanism model and a global attention mechanism model. The step of extracting the facial feature image and environmental feature image of the target person from the image to be identified based on the preset attention mechanism model includes:
[0093] Step S210: Identify the facial region of the target person in the image to be identified using the local attention mechanism model;
[0094] Step S220: Extract the facial feature image from the facial region using the local attention mechanism model;
[0095] Step S230: Remove the facial region from the image to be identified to obtain the environment region;
[0096] Step S240: Extract the environmental feature image from the environmental region using the global attention mechanism model, wherein the environmental feature image includes at least one environmental feature unit.
[0097] In this embodiment, the preset attention mechanism model includes a local attention mechanism model and a global attention mechanism model. The local attention mechanism model extracts the facial region of the target person, and then extracts facial feature images from the facial region. It should be noted that the feature extraction of the target person's face in this embodiment can be implemented based on existing face recognition technology. It is understood that although the local attention mechanism model only extracts facial feature images, the extracted facial feature images also correspond to weights generated by the attention mechanism based on the entire image to be recognized. The generation principle can be referred to the above content and will not be repeated here. The aforementioned facial region in the image to be recognized is removed, for example, by setting the pixel values of the facial region to zero. After removing the facial region from the image to be recognized, the environment region is obtained, and then the global attention mechanism model extracts environmental feature images from the environment region. The environmental feature image includes at least one environmental feature unit, such as a characteristic object in the environment where the target person is located or a characteristic scene in which the target person is located. It is understood that in practical applications, when performing feature extraction on the entire image to be recognized based on the attention mechanism model, regions with relatively high weights will be extracted. The facial region of the target person is given a high weight. However, if the weight of the facial region is too high, the environmental feature image may contain too little information or too few environmental feature units. For example, the maximum weight of the facial region during extraction is 100, and the preset proportion is 30%. Therefore, only environmental feature units with a weight greater than 30 are retained in the environmental feature image. If the weights of most environmental feature units in the environmental feature image are distributed between 10 and 20, a large number of environmental feature units will be ignored. Therefore, in this embodiment, when extracting the environmental feature image, the facial region is removed from the image to be identified before extracting the environmental feature image, thereby avoiding the facial region's influence on the extraction of the environmental feature image and ensuring the integrity of the environmental feature image information.
[0098] As an example, the facial feature image further includes facial weights corresponding to the facial feature image, and the environmental feature image further includes a set of environmental weights corresponding to the environmental feature image, wherein the set of environmental weights includes at least one environmental weight corresponding to an environmental feature unit. After the step of extracting the environmental feature image from the environmental region through the global attention mechanism model, the method further includes:
[0099] Step S01: Compare the maximum environmental weight in each of the environmental feature units with the facial weight to obtain the environmental weight adjustment parameter;
[0100] Step S02: Multiply the environmental weight adjustment parameter by each of the environmental weights to generate new environmental weights;
[0101] Step S03: Ignore the environmental feature units in the environmental feature image whose new environmental weight is less than the preset ratio of the facial weight, and generate a new environmental feature image.
[0102] In this embodiment, the facial feature image corresponds to a facial weight, and each environmental feature unit in the environmental feature image corresponds to an environmental weight. To avoid an excessive number of environmental feature units increasing the computational load during expression recognition, the environmental feature units are filtered. The maximum environmental weight in the environmental feature unit is compared with the facial weight to obtain an environmental weight adjustment parameter. The environmental weight adjustment parameter can be the ratio of the facial weight to the maximum environmental weight. For example, if the maximum environmental weight is 80 and the facial weight is 100, then the ratio of the facial weight to the maximum environmental weight is 1.25. The environmental weight adjustment parameter is multiplied by each of the environmental weights to obtain new environmental weights. For example, multiplying the original maximum environmental weight by the environmental weight adjustment parameter results in a new maximum environmental weight of 100, which is the same as the facial weight. Then, environmental feature units in the environmental feature image whose new environmental weights are less than a preset proportion (e.g., 30%) of the facial weights are ignored to reduce the number of environmental feature units and speed up expression recognition.
[0103] This embodiment provides a facial expression recognition method. Compared to existing methods that only recognize expressions based on the face, this embodiment acquires an image of a target person to be recognized; extracts facial feature images and environmental feature images of the target person from the image to be recognized based on a preset attention mechanism model; and performs facial expression recognition on the target person based on the facial feature images and the environmental feature images to obtain the target person's current expression. In reality, the environment can influence the target person's inner emotions, thereby affecting their facial expressions. This invention incorporates environmental feature images (objective factors) as one of the bases for expression recognition, reducing the influence of facial feature images on the results and mitigating the negative impact of relying solely on facial feature images (personal factors) on the expression recognition results, thus improving the accuracy of facial expression recognition. Furthermore, incorporating environmental feature images inherently increases the relevant feature information for expression recognition, resulting in even higher accuracy.
[0104] Furthermore, refer to Figure 3 Based on the first embodiment of the facial expression recognition method of the present invention, a second embodiment of the facial expression recognition method of the present invention is proposed.
[0105] The step of recognizing facial expressions of the target person based on the facial feature image and the environmental feature image includes:
[0106] Step S310: Synchronize the facial weights and each of the environmental weights to make the facial weights and each of the environmental weights the same;
[0107] Step S320: Combine the facial feature image after weight synchronization and the environmental feature image after weight synchronization to obtain the combined feature image;
[0108] Step S330: Input the combined feature image into a preset expression classification model to identify the current expression of the target person, wherein the preset expression classification model includes facial features and environmental features.
[0109] In this embodiment, the obtained facial feature image and environmental feature image are combined to obtain a combined feature image, which facilitates the subsequent expression recognition classification model to perform recognition based on a single image. For example, the facial feature image and environmental feature image are input into a preset fusion neural network model to be combined into a single image. It is understood that there are already relatively mature image synthesis technologies, so they will not be elaborated here. Furthermore, the facial feature image and environmental feature image used to obtain the combined feature image correspond to facial weights and various environmental weights, respectively. To further increase the influence of environmental features on the expression recognition results in subsequent processes, the facial weights and various environmental weights are synchronized to make them identical. The facial feature image with synchronized weights and the environmental feature image with synchronized weights are then combined to obtain the combined feature image. This combined feature image will also serve as the input to the aforementioned preset expression classification model. It is understood that the preset expression classification model contains both facial features and environmental features. Multiple training samples used to train the preset expression classification model all contain both facial features and environmental features; these training samples can be images or videos that simultaneously contain both facial feature images and environmental feature images.
[0110] Furthermore, refer to Figure 4 Based on the first embodiment of the facial expression recognition method of the present invention, a third embodiment of the facial expression recognition method of the present invention is proposed.
[0111] Step S110: Acquire the original image of the target person in the field of view corresponding to the AR glasses through the AR glasses;
[0112] Step S120: Preprocess the original image based on the super-resolution reconstruction network to generate an image to be identified;
[0113] Step S200: Extract the facial feature image of the target person and the environmental feature image of the environment in which the target person is located from the image to be identified based on a preset attention mechanism model;
[0114] Step S300: Perform facial expression recognition on the target person based on the facial feature image and the environmental feature image to obtain the current expression of the target person;
[0115] Step S410: Generate the relative position of the facial feature image in the image to be identified;
[0116] Step S420: Display and output the current expression of the target person through AR glasses based on the relative position.
[0117] It should be noted that while understanding facial expressions is easy for the general population, it can be challenging for certain individuals, such as people with autism. Most people with autism already lack the ability to communicate with others, and the inability to understand facial expressions makes communication even more difficult.
[0118] This embodiment provides a facial expression recognition method applicable to AR glasses. After a user wears the AR glasses, the glasses can use their own visual camera to capture the view seen through the glasses in real time (i.e., the corresponding field of view of the AR glasses), obtaining the original image. When a person appears in the original image, they are identified as the target person, and facial expression recognition is performed on them. To further improve the accuracy of expression recognition, a super-resolution reconstruction network (such as the SRCNN algorithm) is used to preprocess the acquired original image to improve image clarity, obtaining the image to be recognized. Based on the processed image to be recognized, the target person's expression is identified to obtain the target person's current expression. The specific recognition process can be referred to in the first embodiment, and will not be repeated here. Then, based on the relative position of the facial feature image within the image to be recognized, the target person's image is displayed on the AR glasses. It can be understood that the image to be recognized is the image seen by the user through the AR glasses. The current expression is displayed in the aforementioned relative position within the AR glasses' field of view. For example, if the facial feature image is located in the center of the image to be recognized, the current expression is also displayed in the center of the AR glasses' field of view. This allows the user to see the target person's current expression (the current expression appears as text near the target person's face) while viewing the target person through the AR glasses. (See reference...) Figure 5 This includes AR glasses and the rendered images viewed on them. The target person's face is labeled with their current expression, helping users understand others' expressions and thus facilitating normal communication. It should be noted that the AR glasses in this embodiment can be used not only by patients with the condition but also to assist the general public in social interactions, such as between teachers and students, or between couples.
[0119] As an example, the step of displaying the target person's current expression through AR glasses based on the relative position includes:
[0120] Step S421: Match preset communication suggestion prompts based on the current facial expression;
[0121] Step S422: Display the current expression and the preset communication suggestion information at their relative positions within the display field of view of the AR glasses.
[0122] In this embodiment, in addition to displaying the target person's facial expression, it can also display explanations related to the expression, such as the target person's current emotional state and the emotional tendency under that state (i.e., the aforementioned preset communication suggestion prompts). A mapping relationship is established between different facial expressions and different communication suggestion prompts, and the corresponding communication suggestion prompts are obtained based on the current facial expression. After determining the current facial expression and the corresponding communication suggestion prompts, the current facial expression and the communication suggestion prompts are displayed synchronously at the aforementioned relative positions within the AR glasses' display field of view, allowing the user to communicate based on the facial expression or the communication suggestion prompts.
[0123] Furthermore, this invention also proposes a facial expression recognition device, which includes:
[0124] The acquisition module is used to acquire images of the target person to be identified.
[0125] The extraction module is used to extract the facial feature image of the target person and the environmental feature image of the environment in which the target person is located from the image to be identified based on a preset attention mechanism model;
[0126] The recognition module is used to perform facial expression recognition on the target person based on the facial feature image and the environmental feature image, so as to obtain the current expression of the target person.
[0127] Optionally, the preset attention mechanism model includes a local attention mechanism model and a global attention mechanism model, and the acquisition module is further used for:
[0128] The local attention mechanism model is used to identify the facial region of the target person in the image to be identified;
[0129] The facial feature image is extracted from the facial region using the local attention mechanism model.
[0130] The facial region is removed from the image to be identified to obtain the environmental region;
[0131] The environmental feature image is extracted from the environmental region using the global attention mechanism model, wherein the environmental feature image includes at least one environmental feature unit.
[0132] Optionally, the facial feature image is configured with facial weights, and the environmental feature image is configured with an environmental weight set, wherein the environmental weight set includes at least one environmental weight corresponding to an environmental feature unit, and the extraction module is further configured to:
[0133] The environmental weight adjustment parameter is obtained by comparing the maximum environmental weight in each of the environmental feature units with the facial weight.
[0134] The environmental weight adjustment parameter is multiplied by each of the environmental weights to generate new environmental weights;
[0135] Ignore environmental feature units in the environmental feature image whose environmental weight is less than a preset ratio of the facial weight, and generate a new environmental feature image.
[0136] Optionally, the recognition module is also used for:
[0137] Synchronize the facial weights and each of the environmental weights to make the facial weights and each of the environmental weights identical;
[0138] The combined feature image is obtained by combining the facial feature image after weight synchronization and the environmental feature image after weight synchronization.
[0139] The combined feature image is input into a preset expression classification model to identify the current expression of the target person, wherein the preset expression classification model includes facial features and environmental features.
[0140] Optionally, the acquisition module is further configured to:
[0141] The original image of the target person in the corresponding field of view of the AR glasses is acquired through AR glasses;
[0142] The original image is preprocessed using a super-resolution reconstruction network to generate an image to be identified.
[0143] Optionally, the identification module is further configured to:
[0144] Generate the relative position of the facial feature image in the image to be identified;
[0145] Based on the relative position, the target person's current expression is displayed and output through AR glasses.
[0146] Optionally, the identification module is further configured to:
[0147] Based on the current facial expression, a preset communication suggestion message is matched;
[0148] The current expression and the preset communication suggestion information are displayed at their relative positions within the display field of view of the AR glasses.
[0149] The facial expression recognition device provided by this invention employs the facial expression recognition method described in the above embodiments, aiming to solve the technical problem of low accuracy in current facial expression recognition technologies. Compared with the prior art, the beneficial effects of the facial expression recognition device provided by this invention are the same as those of the facial expression recognition method described in the above embodiments, and other technical features in this facial expression recognition device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0150] Furthermore, this invention also proposes a facial expression recognition device, which includes: a memory, a processor, and a facial expression recognition program stored in the memory and executable on the processor. When the facial expression recognition program is executed by the processor, it implements the steps of the facial expression recognition method described above.
[0151] The specific implementation of the facial expression recognition device of the present invention is basically the same as the embodiments of the facial expression recognition method described above, and will not be repeated here.
[0152] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a facial expression recognition program, wherein the facial expression recognition program, when executed by a processor, implements the steps of the facial expression recognition method described above.
[0153] The specific implementation of the readable storage medium of the present invention is basically the same as the embodiments of the facial expression recognition method described above, and will not be repeated here.
[0154] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0155] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0156] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, AR glasses, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0157] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A facial expression recognition method, characterized in that, The facial expression recognition method includes the following steps: Collect images of the target person to be identified; Based on a preset attention mechanism model, the facial feature image of the target person and the environmental feature image of the environment in which the target person is located are extracted from the image to be identified; Based on the facial feature image and the environmental feature image, facial expression recognition is performed on the target person to obtain the target person's current expression; The facial feature image is configured with facial weights, and the environmental feature image is configured with an environmental weight set. The environmental weight set includes at least one environmental weight corresponding to an environmental feature unit in the environmental feature image. The facial weights and the environmental weight set are used to filter environmental feature units in the environmental feature image to generate a new environmental feature image.
2. The facial expression recognition method as described in claim 1, characterized in that, The preset attention mechanism model includes a local attention mechanism model and a global attention mechanism model. The step of extracting the facial feature image of the target person and the environmental feature image of the environment in which the target person is located from the image to be identified based on the preset attention mechanism model includes: The local attention mechanism model is used to identify the facial region of the target person in the image to be identified; The facial feature image is extracted from the facial region using the local attention mechanism model. The facial region is removed from the image to be identified to obtain the environmental region; The environmental feature image is extracted from the environmental region using the global attention mechanism model, wherein the environmental feature image includes at least one environmental feature unit.
3. The facial expression recognition method as described in claim 2, characterized in that, After the step of extracting the environmental feature image from the environmental region using the global attention mechanism model, the method further includes: The environmental weight adjustment parameter is obtained by comparing the maximum environmental weight in each of the environmental feature units with the facial weight. The environmental weight adjustment parameter is multiplied by each of the environmental weights to generate new environmental weights; Ignore environmental feature units in the environmental feature image whose environmental weight is less than a preset ratio of the facial weight, and generate a new environmental feature image.
4. The facial expression recognition method as described in claim 3, characterized in that, The step of recognizing facial expressions of the target person based on the facial feature image and the environmental feature image includes: Synchronize the facial weights and each of the environmental weights to make the facial weights and each of the environmental weights identical; The facial feature image and the environmental feature image after weight synchronization are combined to obtain a combined feature image; The combined feature image is input into a preset expression classification model to identify the current expression of the target person, wherein the preset expression classification model includes facial features and environmental features.
5. The facial expression recognition method as described in claim 1, characterized in that, The steps for acquiring the image of the target person to be identified include: The original image of the target person in the corresponding field of view of the AR glasses is acquired through AR glasses; The original image is preprocessed using a super-resolution reconstruction network to generate an image to be identified.
6. The facial expression recognition method as described in claim 5, characterized in that, After the step of performing facial expression recognition on the target person based on the facial feature image and the environmental feature image, the method further includes: Generate the relative position of the facial feature image in the image to be identified; Based on the relative position, the target person's current expression is displayed and output through AR glasses.
7. The facial expression recognition method as described in claim 6, characterized in that, The step of displaying and outputting the target person's current expression through AR glasses based on the relative position includes: Based on the current facial expression, a preset communication suggestion message is matched; The current expression and the preset communication suggestion information are displayed at their relative positions within the display field of view of the AR glasses.
8. A facial expression recognition device, characterized in that, The facial expression recognition device includes: The acquisition module is used to acquire images of the target person to be identified. The extraction module is used to extract the facial feature image of the target person and the environmental feature image of the environment in which the target person is located from the image to be identified based on a preset attention mechanism model; The recognition module is used to perform facial expression recognition on the target person based on the facial feature image and the environmental feature image, so as to obtain the current expression of the target person; The facial feature image is configured with facial weights, and the environmental feature image is configured with an environmental weight set. The environmental weight set includes at least one environmental weight corresponding to an environmental feature unit in the environmental feature image. The facial weights and the environmental weight set are used to filter environmental feature units in the environmental feature image to generate a new environmental feature image.
9. A facial expression recognition device, characterized in that, The facial expression recognition device includes: a memory, a processor, and a facial expression recognition program stored in the memory and executable on the processor. When the facial expression recognition program is executed by the processor, it implements the steps of the facial expression recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a facial expression recognition program, which, when executed by a processor, implements the steps of the facial expression recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Information processing device, control method, and program
CN107405120A
Emotion recognition apparatus and method, head-mounted display device, and storage medium
CN109145861A