Facial local area-based expression recognition method and device, equipment and medium

By obtaining the target face area in facial expression recognition, performing key point positioning and local area cropping, and combining with the channel pruning model, the redundancy problem of feature extraction in facial expression recognition is solved, achieving more efficient and accurate expression recognition.

CN120375447APending Publication Date: 2025-07-25HUNAN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510508873.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, facial expression recognition methods have redundant features extraction problems, resulting in waste of computing resources and inefficient recognition.

Method used

By obtaining the target face area, positioning the key point, calculating the local area enclosure box and cropping the expression analysis image, and using the facial expression classification model processed by channel pruning for identification.

Benefits of technology

It effectively reduces the redundancy of feature extraction, improves the accuracy and efficiency of expression recognition, and makes the model lighter and more efficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375447A_ABST
    Figure CN120375447A_ABST
Patent Text Reader

Abstract

The invention discloses an expression recognition method and device based on a facial local area, equipment and a medium, and the method comprises the steps: obtaining a target human face area in image data, carrying out the key point positioning of the human face area, selecting a key point combination related to a specific facial expression, and carrying out the key point positioning of the target human face area; and calculating a surrounding frame of the local area of the face, cutting out a facial expression analysis image, and processing the facial expression analysis image by adopting a facial expression classification model subjected to channel pruning processing. Therefore, the redundancy of feature extraction is effectively reduced, the accuracy and efficiency of expression recognition are improved, and the model is lighter and more efficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision generation, and in particular, to an expression recognition method, device, equipment and medium based on facial local regions. Background Art

[0002] Currently, in the fields of artificial intelligence and computer vision, facial expression recognition technology infers human emotional states by analyzing facial features and has important application values in fields such as intelligent interaction and mental health assessment. Traditional technical solutions based on traditional image processing methods usually rely on global face feature analysis, use manually designed feature extraction algorithms combined with machine learning for classification, and at the same time, based on deep learning, the recognition accuracy has been improved through automatic feature learning. However, there is generally a problem of redundant feature extraction. That is, the feature calculation of non-expression-related regions in traditional methods will introduce interference noise, and deep learning methods need to process a large number of irrelevant pixels when the full face is input, resulting in waste of computing resources. There is an urgent need to develop a new type of expression recognition method that can reduce the computational redundancy of feature extraction while ensuring the recognition accuracy. Summary of the Invention

[0003] Embodiments of the present invention provide an expression recognition method, device, equipment and medium based on facial local regions, aiming to solve the problem of redundant feature extraction in traditional facial expression recognition methods under the existing technology.

[0004] In a first aspect, an embodiment of the present invention provides an expression recognition method based on facial local regions, including: obtaining image data including a human face; performing face detection on the image data to determine a target face region; performing key point positioning on the target face region and obtaining key point coordinates; based on a preset key point combination corresponding to a selected facial expression, calculating a facial local region bounding box according to the key point coordinates, and cropping to obtain a facial expression analysis image; inputting the facial expression analysis image into a facial expression classification model processed by channel pruning, and outputting a facial expression recognition result.

[0005] In a second aspect, an embodiment of the present invention further provides an expression recognition device based on facial local regions, which is used to execute the above-mentioned expression recognition method based on facial local regions.

[0006] In a third aspect, an embodiment of the present invention further provides a computer device, where the computer device includes a memory and a processor connected to the memory; the memory is used to store a computer program; the processor is used to run the computer program stored in the memory to execute the steps of the above-mentioned expression recognition method based on facial local regions.

[0007] Fourthly, an embodiment of the present invention further provides a computer-readable storage medium. The storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the steps of the above-mentioned facial expression recognition method based on a local facial area can be implemented.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0009] In the technical solution of the present invention, by obtaining a target face area in image data, performing key point positioning on the face area, selecting a combination of key points related to a specific facial expression, calculating the bounding box of the local facial area, cropping the facial expression analysis image, and processing the facial expression analysis image with a facial expression classification model that has undergone channel pruning. Thus, the redundancy of feature extraction is effectively reduced, the accuracy and efficiency of facial expression recognition are improved, and the model becomes more lightweight and efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 It is a flowchart of the facial expression recognition method based on a local facial area provided by the present invention;

[0012] Figure 2 It is the first sub-flowchart of the facial expression recognition method based on a local facial area provided by the present invention;

[0013] Figure 3 It is the second sub-flowchart of the facial expression recognition method based on a local facial area provided by the present invention;

[0014] Figure 4 It is the third sub-flowchart of the facial expression recognition method based on a local facial area provided by the present invention;

[0015] Figure 5 It is the fourth sub-flowchart of the facial expression recognition method based on a local facial area provided by the present invention;

[0016] Figure 6 It is the fifth sub-flowchart of the facial expression recognition method based on a local facial area provided by the present invention;

[0017] Figure 7 It is the fifth sub-flowchart of the facial expression recognition method based on a local facial area provided by the present invention;

[0018] Figure 8A schematic block diagram of units of the facial local area-based expression recognition device provided by the present invention;

[0019] Figure 9 A schematic block diagram of a computer device provided for an embodiment of the present invention;

[0020] Figure 10 95 key point diagrams of the facial expression recognition method based on local facial regions provided by the present invention;

[0021] Figure 11 The facial expression analysis image diagram is obtained by cutting the facial expression recognition method based on the local facial area provided by the present invention. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0023] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0024] It should also be understood that the terms used in this specification are only for the purpose of describing medical embodiments and are not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a", "an" and "the" are intended to include plural forms unless the context clearly indicates otherwise.

[0025] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0026] In order to solve the problem of redundant feature extraction in the traditional facial expression recognition method under the prior art, the present invention proposes an expression recognition method based on local facial regions. Figures 1 to 7 , Figure 10 and Figure 11 , the method comprising:

[0027] S110, acquiring image data containing a human face;

[0028] S120. Perform face detection on the image data to determine the target face region;

[0029] S130. Perform key point localization on the target face region and obtain the key point coordinates;

[0030] S140. Based on the preset key point combination corresponding to the selected facial expression, calculate the bounding box of the facial local region according to the key point coordinates, and crop to obtain the facial expression analysis image;

[0031] S150. Input the facial expression analysis image into the facial expression classification model processed by channel pruning, and output the facial expression recognition result.

[0032] First, before performing image data analysis and processing, source data collection is required. In the embodiments of the present invention, the original image data containing human faces is obtained through a mobile terminal camera, a surveillance camera, or an image acquisition device, and the original image contains the complete face region. Specifically, when collecting source data, the collection object makes different expressions according to requirements, taking into account various influencing factors such as human face angles, distances, and lighting, so that the data collection has diversity. Generally speaking, the larger the data volume, the better. In addition to manually collecting data, some open-source data can be used as auxiliary according to actual needs. Then, a face detection algorithm is used to process the acquired image data to identify and determine the target face region. Through precise facial localization, background interference is avoided, laying a foundation for subsequent key point localization and region cropping. After that, using the face key point detection technology, the system can accurately identify multiple key points on the face, such as the positions of eyes, eyebrows, mouth, etc., and obtain the corresponding coordinates. These key points provide key data for subsequent cropping of local regions and expression recognition. After that, according to the selected expression type, the system calculates the minimum bounding box of the facial local region according to the preset key point combination. For example, for a frowning expression, the key points in the eyebrow region are selected and the local region is cropped. The cropped region contains the key information of facial expression changes, removing the background region irrelevant to the expression, thereby reducing redundant features and improving the accuracy of recognition. After that, the cropped facial expression analysis image is then input into the expression classification model processed by channel pruning. The channel pruning technology can effectively reduce the computational complexity and storage requirements of the network, enabling this method to run on edge devices with relatively limited computing resources and ensuring the accuracy of real-time inference. Finally, through model inference, the finally output expression classification result is the recognized human face expression.

[0033] Compared with the prior art, in the embodiment of the present invention, the target face region is obtained from the image data, key point positioning is performed on the face region, a combination of key points related to specific facial expressions is selected, the bounding box of the facial local region is calculated, the facial expression analysis image is cropped, and the facial expression classification model processed by channel pruning is used to process the facial expression analysis image. Thereby, the redundancy of feature extraction is effectively reduced, the accuracy and efficiency of expression recognition are improved, and the model is made more lightweight and efficient.

[0034] In one embodiment, referring to Figure 2 , S120 includes:

[0035] S121. Detect multiple targets in the image data through a quantized-accelerated face detection model;

[0036] S122. Obtain the target face region from the multiple targets according to preset screening conditions.

[0037] Face detection is performed on the input image data by using an efficient quantized-accelerated face detection model to detect multiple targets in the image. The so-called quantization acceleration reduces the computational amount and memory occupancy of the model by quantizing the weights of the model, thereby improving the inference speed, and is applicable to environments with limited computing resources, such as mobile devices or edge devices. The goal of the face detection model is to identify the face region from the image and determine the position information of each target for subsequent processing.

[0038] After detecting multiple possible face regions, the system will obtain the target face region from the multiple targets according to preset screening conditions. The screening conditions usually include the size of the face bounding box and the angle and illumination conditions of the face. Among them, for the size of the face bounding box, usually the largest face region is selected as the target region because a larger region often represents a clearer and more accurate face. For the angle and illumination conditions of the face, in practical applications, since the face may have different angles or illumination conditions, the system will select the best face region that meets the conditions according to the frontal degree of the face and the illumination uniformity of the image. The above screening conditions ensure the accuracy of the face region and provide reliable input for subsequent key point positioning and expression recognition.

[0039] In actual use, the general object detection algorithm YOLO-v6-4.0 is used as the prototype, and the open-source face detection data widerface, etc. and some face detection data collected by oneself corresponding to the scene are used for training. To make the inference time of the algorithm model smaller, pruning and 8-bit quantization model acceleration methods are also used, and a method similar to yolov10 is used in the post-processing of the model to make the model inference speed faster.

[0040] In one embodiment, referring to Figure 3, S130 includes:

[0041] S131. Process the target face region through a lightweight human pose estimation model to locate each key point;

[0042] S132. Output the coordinates of the key points of multiple facial regions according to predefined rules. The facial regions include the eyebrow region, the eye region, and the mouth region, and each facial region is composed of the coordinates of multiple corresponding key points.

[0043] Refer to Figure 10 and Figure 11 , by using a lightweight human pose estimation model to process the target face region, the system can accurately locate each key point on the human face. These key points usually include important parts such as eyes, eyebrows, nose, and mouth. The lightweight human pose estimation model is chosen because it has high real-time performance and computational efficiency and is suitable for running on devices with limited computational resources. Specifically, the model extracts features from the input face region and outputs the coordinate information of 95 key points.

[0044] Based on the key point coordinates output by the lightweight human pose estimation model, according to predefined rules, extract the key point coordinates of multiple facial regions. These facial regions mainly include the eyebrow region, the eye region, and the mouth region. Refer to Figure 10 and Figure 11 , each facial region is composed of the coordinates of multiple corresponding key points. Specific key points such as 22, 31, 43, 53, etc. correspond to the upper-middle points of the left and right eyebrows and the upper-middle points of the left and right eyes respectively. These coordinate points may have some small-range jitters during expression changes, but it will not affect the overall recognition effect.

[0045] According to predefined rules, the system combines the coordinates of each key point to determine the specific position of each facial region. For example, refer to Figure 10 and Figure 11 , the eyebrow region can be determined by the upper-middle points of the left and right eyebrows, such as key points 22 and 31; the eye region is determined by the upper-middle points of the left and right eyes, such as key points 43 and 53. Through the combination of these key points, the system can crop the local region image related to expression recognition. The cropped local face region needs to be adjusted and normalized, for example, adjusted to a unified input size of 112×112 pixels, so as to facilitate the processing of the subsequent expression recognition model. The purpose of normalization processing is to make the input data consistent and improve the prediction accuracy of the expression recognition model.

[0046] In actual use, mainly based on the Tinypose algorithm model of Paddledetection and the Rtmpose algorithm model of Mmdetection, training was carried out on more than 40,000 pieces of 95-point data with existing annotations. Considering both speed and accuracy, finally Rtmpose was used as the final model. To improve the inference efficiency, the smallest version of Rtmpose was used.

[0047] In one embodiment, referring to Figure 4 , S140 includes:

[0048] S141. Establish a mapping relationship between facial expression types and the key point combinations, where each facial expression type is associated with at least one set of predefined key points;

[0049] S142. Select the corresponding key point combination from the mapping relationship according to the facial expression type to be recognized;

[0050] S143. Calculate the minimum bounding rectangle area of the selected key point combination on the image plane according to the key point coordinates;

[0051] S144. Crop the target facial expression analysis image according to the bounding rectangle area.

[0052] To ensure the accuracy and efficiency of facial expression recognition, the system needs to establish a mapping relationship between facial expression types and key point combinations. Each facial expression type is associated with at least one set of predefined key point position combinations. This mapping relationship ensures that the key point combinations required for different expressions can be effectively selected and provides an accurate basis for subsequent cropping. During the facial expression recognition process, according to the specific expression type to be recognized, the corresponding key point combination is selected from the preset mapping relationship. For example, if the expression to be recognized is "frowning", the system will select the key points related to frowning, such as the key point combination in the glabella area. For example, referring to Figure 10 and Figure 11, key points 22, 31, 43, 53. This process can be achieved by looking up the table corresponding to the key points or predefined expressions. According to the selected combination of key points, the system calculates the minimum bounding rectangle area of these key points on the image plane. The calculation method of the minimum bounding rectangle is: for the four selected key points, calculate their minimum values (xmin, ymin) and maximum values (xmax, ymax) on the X-axis and Y-axis respectively. Thus, the obtained upper left corner coordinates (xmin, ymin) and lower right corner coordinates (xmax, ymax) form the bounding box of this area. This bounding box is the required facial local area, providing a basis for subsequent image cropping. Based on the calculated minimum bounding rectangle bounding box, the system crops the target facial area from the original image. This area contains the key information required for expression analysis, such as the glabella area of the frowning expression. The cropped facial expression analysis image can be used as the input for the subsequent expression recognition model. The cropped facial expression analysis image needs to be data-annotated to provide standardized input for the training process. The annotation process includes classifying each expression type and associating it with the corresponding label, such as "frowning", "smiling", etc. After the annotation is completed, these annotated data will be used to train the expression recognition model and further optimize the model performance. Through the solution of this embodiment, the interference of irrelevant background information is effectively reduced, and the accuracy and robustness of expression recognition are improved.

[0053] In one embodiment, referring to Figure 5 , S150 includes:

[0054] S151. Perform pixel normalization processing on the facial expression analysis image to generate standardized input data:

[0055] S152. Perform forward inference on the standardized input data through the facial expression classification model to generate an initial probability distribution;

[0056] S153. Based on the initial probability distribution, perform dynamic weighted fusion in combination with the historical frame recognition results, and output the optimized expression recognition result.

[0057] Perform pixel normalization on the facial expression analysis image obtained in step S144. Through this process, factors such as the brightness and contrast of the image do not interfere with subsequent model inference, thereby improving the robustness of the model to input data. The normalized facial expression analysis image is input into a pre-trained facial expression classification model for forward inference. This step processes the image data through a deep neural network to generate a preliminary probability distribution, which represents the prediction probability of the model for each expression category. The so-called forward inference is to extract image features through a series of operations such as convolution, pooling, and fully connected layers, and finally output the probabilities of each category. After obtaining the initial probability distribution, to further improve the recognition accuracy, the system combines the recognition results of historical frames and adopts a dynamic weighted fusion strategy. This process uses the prediction information of previous frames to weight-adjust the prediction results of the current frame. Specifically, dynamic weighting is performed according to the reliability of historical frames and the prediction probability of the current frame to obtain a more stable and accurate final recognition result. The results of historical frames can usually provide supplementary information. Especially when the expression changes quickly or weakly in a real-time video stream, the recognition results of historical frames can provide important references for the recognition of the current frame. The result after dynamic weighted fusion is output as the optimized expression recognition result. This result is the final expression classification and may correspond to expression categories such as "smile", "frown", "surprise", etc. In addition, the dynamic weighted fusion strategy can effectively smooth the fluctuations of a single frame in a multi-frame video stream. Especially when encountering instantaneous expression changes, occlusion, or low-light environments, through the fusion of historical information, the system can make more accurate predictions. In addition, this strategy can improve the adaptability of the model to sudden expressions within a short period of time, improving real-time performance and accuracy. Through the solution of this embodiment, not only the accuracy of facial expression recognition is optimized, but also the stability of the recognition model and its adaptability to complex scenarios are improved.

[0058] Further, referring to Figure 6 , S152 includes:

[0059] S1521. Gradually extract hierarchical features through multiple convolutional layers in the pruned expression classification model, and the pruned multiple convolutional layers are pruned versions of MobileNet or ShuffleNet;

[0060] S1522. Input the hierarchical features of the final level into the fully connected layer to generate the initial probability distribution of each expression category.

[0061] In this embodiment, a pruned facial expression classification model is adopted. Specifically, a pruned version of a lightweight network structure such as MobileNet or ShuffleNet is used. The pruning technique reduces the computational amount and the number of parameters of the model by removing unimportant convolutional channels in the model, thereby improving the inference speed of the model and reducing the demand for computing resources. Although pruning reduces the complexity of the network, it still retains the effective facial expression recognition ability. During the inference process, the pruned facial expression classification model first gradually extracts hierarchical features through its multi-level convolutional layers. The convolutional layers can capture the detailed information of the facial image at different scales, especially the subtle changes in the local area. The multi-level convolutional layers enable the network to effectively extract image features from low-level to high-level by gradually enhancing the feature representation ability, such as the changes in regions such as eyes and mouth in facial expressions.

[0062] After the feature extraction by the multi-level convolutional layers, the network sends the last convolutional feature to the fully connected layer. The fully connected layer is responsible for mapping the high-dimensional features to the output of specific facial expression categories. During this process, the fully connected layer transforms the features extracted by the convolutional layer into the initial probability distribution of facial expression categories. This probability distribution represents the confidence of the model for each facial expression category and is the basis for subsequent decisions. The initial probability distribution output by the fully connected layer covers the prediction results of all preset facial expression categories. For example, it may include categories such as "smile", "frown", "surprise", etc., and each category corresponds to a probability value, reflecting the recognition confidence of the model for that facial expression. This initial probability distribution provides the basic data for the subsequent dynamic weighted fusion step.

[0063] In one embodiment, referring to Figure 7 , after S153, it further includes:

[0064] S154. Statistically analyze the facial expression recognition results of consecutive multiple frames, and determine the final facial expression category based on the results of the statistical analysis.

[0065] In order to make the results more accurate, the facial expression recognition results need to be further processed. By statistically analyzing the facial expression recognition results of consecutive multiple frames, the final facial expression category can be determined more precisely. The statistical analysis mainly includes methods such as frequency statistics and time series analysis of the recognition results. Through statistical analysis, the system can identify the stability of facial expression categories, thereby avoiding misjudging a certain facial expression in the case of unstable or large error in a single frame. Specifically, the facial expression recognition results of consecutive multiple frames will be used as input for frequency statistics. If the frequency of a certain facial expression category appears higher in consecutive multiple frames, then this facial expression category is considered as the current main facial expression. This process not only relies on the recognition results of a single frame, but also improves the accuracy of the final facial expression recognition through cross-frame information fusion.

[0066] Based on the statistical analysis of the multi-frame recognition results, the system will determine the final expression category. Suppose in a video sequence, a certain expression category dominates in most consecutive frames, then the final recognition result will tend to output that expression category. The key to this step is to comprehensively judge through the results at multiple time points, which can eliminate the influence of accidental misjudgments or expression changes within a short period. Through the above statistical analysis and fusion process, the system finally outputs the optimized expression recognition result and gives the corresponding expression category.

[0067] Figure 8 FIG. 4 is a schematic block diagram of an expression recognition device 600 based on a local facial region provided by an embodiment of the present invention. As Figure 8 shown, corresponding to the above expression recognition method based on a local facial region, the present invention also provides an expression recognition device 600 based on a local facial region. The expression recognition device 600 based on a local facial region includes units for executing the above expression recognition method based on a local facial region, and this device can be configured in terminals such as desktop computers, tablet computers, smart phones, etc.

[0068] Specifically, please refer to Figure 8 , the expression recognition device 600 based on a local facial region includes:

[0069] A source data acquisition unit 610, configured to acquire image data including a human face;

[0070] A target face acquisition unit 620, configured to perform face detection on the image data to determine a target face region;

[0071] A key point acquisition unit 630, configured to perform key point localization on the target face region and acquire key point coordinates;

[0072] An image cropping unit 640, configured to calculate a bounding box of a local facial region based on a preset key point combination corresponding to a selected facial expression, and crop to obtain a facial expression analysis image according to the key point coordinates;

[0073] An expression recognition unit 650, configured to input the facial expression analysis image into a facial expression classification model processed by channel pruning, and output a facial expression recognition result.

[0074] In one embodiment, the target face acquisition unit 620 includes:

[0075] A multi-target detection unit, configured to detect multiple targets in the image data through a quantized and accelerated face detection model;

[0076] A target face screening unit, configured to acquire a target face region from the multiple targets according to a preset screening condition.

[0077] In one embodiment, the key point acquisition unit 630 includes:

[0078] A key point localization unit, configured to process the target face region through a lightweight human pose estimation model to localize each of the key points;

[0079] A facial region localization unit, configured to output the key point coordinates of multiple facial regions according to predefined rules, where the facial regions include an eyebrow region, an eye region, and a mouth region, and each facial region is composed of multiple corresponding key point coordinates.

[0080] In one embodiment, the image cropping unit 640 includes:

[0081] A mapping relationship construction unit, configured to establish a mapping relationship between facial expression types and the key point combinations, where each facial expression type is associated with at least one group of predefined key points;

[0082] A key point combination selection unit, configured to select the corresponding key point combination from the mapping relationship according to the facial expression type to be recognized;

[0083] A cropping region localization unit, configured to calculate the minimum bounding rectangle region of the selected key point combination on the image plane according to the key point coordinates;

[0084] A target image cropping unit, configured to crop a target facial expression analysis image according to the bounding rectangle region.

[0085] In one embodiment, the expression recognition unit 650 includes:

[0086] A normalization unit, configured to perform pixel normalization processing on the facial expression analysis image to generate normalized input data:

[0087] A forward inference unit, configured to perform forward inference on the normalized input data through the facial expression classification model to generate an initial probability distribution;

[0088] A result optimization unit, configured to perform dynamic weighted fusion based on the initial probability distribution and combine historical frame recognition results to output an optimized expression recognition result.

[0089] Further, the forward inference unit includes:

[0090] A hierarchical feature extraction unit, configured to gradually extract hierarchical features through multiple convolutional layers in the pruned facial expression classification model, where the pruned multiple convolutional layers are pruned versions of MobileNet or ShuffleNet;

[0091] An initial probability generation unit, configured to input the hierarchical features of the final level into a fully connected layer to generate an initial probability distribution for each expression category.

[0092] In one embodiment, in the expression recognition unit 650, after the result optimization unit, there is further included:

[0093] A final expression determination unit, configured to perform statistical analysis on the expression recognition results of consecutive multiple frames, and determine a final expression category based on the results of the statistical analysis.

[0094] The above-mentioned expression recognition device 600 based on facial local regions can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 9 shown.

[0095] Please refer to Figure 9 , Figure 9 , which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a terminal or a server. Among them, the terminal can be an electronic device with communication functions such as a desktop computer, a tablet computer, a smart phone, etc. The server can be an independent server or a server cluster composed of multiple servers.

[0096] Refer to Figure 9 , this computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory can include a non-volatile storage medium 503 and an internal memory 504.

[0097] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, and when these program instructions are executed, the processor 502 can be made to execute an expression recognition method based on facial local regions.

[0098] The processor 502 is configured to provide computing and control capabilities to support the operation of the entire computer device 500.

[0099] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute an expression recognition method based on facial local regions.

[0100] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand, Figure 8The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. Specifically, the computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0101] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the steps of the above method.

[0102] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), and this processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program includes program instructions, and the computer program can be stored in a storage medium, and this storage medium is a computer-readable storage medium. These program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above method.

[0104] Therefore, the present invention also provides a storage medium. This storage medium may be a computer-readable storage medium. This storage medium stores a computer program, where the computer program includes program instructions. When these program instructions are executed by a processor, the processor is caused to execute the steps of the above method.

[0105] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, an optical disc, or other various computer-readable storage media that can store program codes.

[0106] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0107] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0108] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0109] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0110] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An expression recognition method based on a local facial region, characterized in that including: Obtain image data including a human face; Perform face detection on the image data to determine a target face region; Perform key point localization on the target face region and obtain key point coordinates; Based on a preset key point combination corresponding to a selected facial expression, calculate a bounding box for a local facial region according to the key point coordinates, and crop to obtain a facial expression analysis image; Input the facial expression analysis image into a facial expression classification model processed by channel pruning, and output a facial expression recognition result.

2. The method for facial expression recognition based on local facial regions according to claim 1, wherein The step of calculating a bounding box for a local facial region according to the key point coordinates based on a preset key point combination corresponding to a selected facial expression and cropping to obtain a facial expression analysis image includes: Establish a mapping relationship between facial expression types and the key point combinations, where each facial expression type is associated with at least one predefined set of the key points; Select the corresponding key point combination from the mapping relationship according to the facial expression type to be recognized; Calculate the minimum bounding rectangle region of the selected key point combination on the image plane according to the key point coordinates; Crop a target facial expression analysis image according to the bounding rectangle region.

3. The method for facial expression recognition based on a local facial region according to claim 2, wherein The step of inputting the facial expression analysis image into a facial expression classification model processed by channel pruning and outputting a facial expression recognition result includes: Perform pixel normalization processing on the facial expression analysis image to generate normalized input data: Perform forward inference on the normalized input data through the facial expression classification model to generate an initial probability distribution; Based on the initial probability distribution, perform dynamic weighted fusion in combination with historical frame recognition results, and output an optimized expression recognition result.

4. The method for facial expression recognition based on a local facial region according to claim 3, wherein The step of performing forward inference on the normalized input data through the facial expression classification model to generate an initial probability distribution includes: Gradually extract hierarchical features through multiple convolutional layers in the pruned expression classification model, and the pruned multiple convolutional layers are pruned versions of MobileNet or ShuffleNet; Input the hierarchical features of the final level into a fully connected layer to generate an initial probability distribution for each expression category.

5. The method for facial expression recognition based on a local facial region according to claim 3, wherein After the step of performing dynamic weighted fusion based on the initial probability distribution in combination with historical frame recognition results and outputting an optimized expression recognition result, it further includes: Perform statistical analysis on the expression recognition results of multiple consecutive frames, and determine the final expression category based on the results of the statistical analysis.

6. The method for facial expression recognition based on a local facial region according to claim 1, wherein The step of performing face detection on the image data to determine a target face region includes: Detect multiple targets in the image data through a quantized and accelerated face detection model; Obtain a target face region from the multiple targets according to preset screening conditions.

7. The method for facial expression recognition based on a local facial region according to claim 1, wherein The step of performing key point localization on the target face region and obtaining key point coordinates includes: Process the target face region through a lightweight human pose estimation model to locate each key point; Output the key point coordinates of multiple facial regions according to predefined rules, and the facial regions include an eyebrow region, an eye region, and a mouth region, and each facial region is composed of multiple corresponding key point coordinates.

8. An expression recognition device based on a local facial region, characterized in that, For performing the facial local-region-based expression recognition method according to any one of claims 1 to 7.

9. A computer device, characterized in that, The computer device includes a memory and a processor connected to the memory; the memory is used for storing a computer program; the processor is used for running the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program includes program instructions which, when executed by a processor, can implement the steps of the method according to any one of claims 1 to 7.