Method and device for training face recognition model, storage medium and electronic equipment
By training and optimizing the initial facial recognition model, the target loss function is used to improve the recognition accuracy and efficiency of the model, and the problem of poor facial expression recognition accuracy and efficiency in the prior art is solved.
Patent Information
- Application Number
- CN202510273565.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The prior art has the limitations of CNN capturing long-distance information in facial expression recognition and the high computational complexity of Transformer when processing ultra-high-definition images, resulting in poor recognition accuracy and efficiency.
By designing a method to train a facial recognition model, the initial facial recognition model is trained using the image dataset, including the backbone network module, feature processing module and object detection module, the model is optimized using the target loss function, including facial expression classification loss, facial detection positioning loss, target confidence loss and time consistency loss.
It realizes accurate recognition of facial expressions, improves the training effect and recognition efficiency of the model, and is suitable for efficient and real-time facial expression recognition scenarios.
Smart Images

Figure CN120220206A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image recognition. Specifically, it relates to a method, device, storage medium, and electronic device for training a facial recognition model. Background Art
[0002] Facial Expression Recognition is a bridge connecting facial feature analysis and high-level emotion understanding.
[0003] Currently, the implementation schemes of facial expression recognition in the prior art mainly rely on models based on Convolutional Neural Networks (CNNs) or Transformers. To some extent, these technical schemes can capture the changes in facial features, and then analyze and recognize different emotional states. However, due to the local receptive field limitation of CNNs in capturing long-distance information, it is difficult to effectively capture long-distance information in some cases, which may lead to poor results in tasks such as facial image segmentation. On the other hand, Transformers perform well in global modeling and can effectively capture long-distance dependencies, but the self-attention mechanism has a high complexity when dealing with large image sizes, especially when facing challenges in tasks such as ultra-high-definition image detection and small target detection.
[0004] Therefore, how to provide a technical scheme for accurate and efficient facial expression recognition has become an urgent technical problem to be solved. Summary of the Invention
[0005] Some embodiments of this application aim to provide a method, device, storage medium, and electronic device for training a facial recognition model. Through the technical scheme of the embodiments of this application, the accuracy of training the facial recognition model can be improved, and the accuracy and efficiency of facial expression recognition can be achieved, with high practicality.
[0006] First aspect, some embodiments of the present application provide a method for training a facial recognition model, which is used to train an initial facial recognition model to obtain a trained facial recognition model. The initial facial recognition model includes: a backbone network module, a feature processing module, and an object detection module. The method includes: inputting an image data set for training into the backbone network module, and outputting image feature information corresponding to the image data set; inputting the image feature information into the feature processing module, and outputting image fusion features of different sizes; inputting the image fusion features of different sizes into the object detection module, and outputting an image detection result; using the loss value of the image detection result determined by an object loss function to optimize the initial facial recognition model to obtain the trained facial recognition model, where the object loss function includes: a facial expression classification loss, a facial detection and localization loss, an object confidence loss, and a temporal consistency loss; the facial recognition model is used to recognize facial expressions in an image to be recognized.
[0007] Some embodiments of the present application train the model by inputting an image data set into the backbone network module, the feature processing module, and the object detection module in the initial facial model, and obtain an image detection result; calculate the loss value of the image detection result through an object loss function to optimize the initial facial model, and obtain a trained facial recognition model, so as to achieve accurate recognition of facial expressions in an image to be recognized, and the training effect of the facial recognition model is relatively high, and the recognition efficiency is relatively high.
[0008] In some embodiments, the image data set for training is continuously updated and obtained by the following method: using the i-th training data set to train the i-th initial facial recognition model to obtain the i-th facial recognition model training result; where i is a positive integer; confirming that the loss value of the i-th facial recognition model training result is greater than a set threshold, or i ≤ N, N is the number of iterations; updating the i-th training data set based on the i-th facial recognition model training result to obtain the (i + 1)-th training data set, where the (i + 1)-th training data set is the image data set.
[0009] Some embodiments of the present application can improve the accuracy of the training data set and further improve the training effect of the facial recognition model by screening and updating each training data set during the training process and then continuing to train the model.
[0010] In some embodiments, when i = 1, the first training dataset is obtained by the following method: cropping the collected original facial expression dataset to obtain a standard dataset; performing data augmentation on the expression categories in the standard dataset to obtain the first training dataset; wherein the data augmentation includes: mirror inversion, random region cropping, and random noise addition to the facial expression images in the standard dataset.
[0011] In some embodiments of the present application, after cropping and augmenting the original facial expression dataset, the first training dataset is obtained, improving the richness of the dataset.
[0012] In some embodiments, the backbone network module includes: a preprocessing module, a plurality of feature extraction modules, and a plurality of feature fusion modules; wherein, inputting the image dataset to be trained into the backbone network module and outputting the image feature information corresponding to the image dataset includes: the preprocessing module extracting features from the images in the image dataset to obtain facial features; the plurality of feature extraction modules obtaining the temporal dynamic information corresponding to the facial features and constructing a facial feature time series; the plurality of feature fusion modules mapping the facial feature time series to different dimensions to obtain the image feature information.
[0013] In some embodiments of the present application, the preprocessing module, the feature extraction module, and the feature fusion module in the backbone network module successively process the image dataset to obtain the image feature information, which can effectively process the data and provide data support for subsequent model training.
[0014] In some embodiments, inputting the image feature information into the feature processing module and outputting image fusion features of different sizes includes: the feature processing module performing sampling operations and convolutional operations on the image feature information in different dimensions to obtain the image fusion features of different sizes.
[0015] In some embodiments of the present application, the feature processing module processes the image feature information in different dimensions to obtain image fusion features of different sizes, which can provide data support for subsequent model training.
[0016] In some embodiments, inputting the image fusion features of different sizes into the target detection module and outputting an image detection result includes: respectively inputting the image fusion features of different sizes into the target detection layers in the target detection module to obtain the image detection result, wherein the target detection layers of different sizes are used to detect the image fusion features of different sizes.
[0017] In some embodiments, the target loss function is constructed as follows: Obtain the respective loss parameters and respective loss weights corresponding to the facial expression classification loss, the facial detection and localization loss, the target confidence loss, and the temporal consistency loss; perform weighted summation on the respective loss parameters and the respective loss weights to obtain the target loss function; the loss value of the image detection result is obtained by the following method: calculate different losses for the image detection result to obtain the respective loss values corresponding to the facial expression classification loss, the facial detection and localization loss, the target confidence loss, and the temporal consistency loss; use the respective loss values as the respective loss parameters and input them into the target loss function to obtain the loss value.
[0018] In some embodiments of the present application, the loss value of the image detection result is calculated by a target loss function composed of four parts, which can improve the accuracy and effect of model training and ensure that the recognition accuracy of the subsequent trained facial recognition model is relatively high.
[0019] In a second aspect, some embodiments of the present application provide an apparatus for training a facial recognition model. The apparatus is used to train an initial facial recognition model to obtain a trained facial recognition model. The initial facial recognition model includes: a backbone network module, a feature processing module, and a target detection module; the apparatus includes: an extraction module, configured to input an image data set for training into the backbone network module and output image feature information corresponding to the image data set; a fusion module, configured to input the image feature information into the feature processing module and output image fusion features of different sizes; a processing module, configured to input the image fusion features of different sizes into the target detection module and output an image detection result; an optimization module, configured to optimize the initial facial recognition model by using the loss value of the image detection result determined by the target loss function to obtain the trained facial recognition model, where the target loss function includes: a facial expression classification loss, a facial detection and localization loss, a target confidence loss, and a temporal consistency loss; and the facial recognition model is used to recognize facial expressions in an image to be recognized.
[0020] In a third aspect, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in any embodiment of the first aspect can be implemented.
[0021] In a fourth aspect, some embodiments of the present application provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any embodiment of the first aspect can be implemented.
[0022] Fifth aspect, some embodiments of the present application provide a computer program product, the computer program product includes a computer program, wherein when the computer program is executed by a processor, the method described in any embodiment of the first aspect can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the drawings required to be used in some embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 System diagram for training a face recognition model provided by some embodiments of the present application;
[0025] Figure 2 Method flowchart for training a face recognition model provided by some embodiments of the present application;
[0026] Figure 3 Block diagram of the device for training a face recognition model provided by some embodiments of the present application;
[0027] Figure 4 Schematic diagram of an electronic device provided by some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The technical solutions in some embodiments of the present application will be described below in conjunction with the drawings in some embodiments of the present application.
[0029] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0030] In the related art, currently, in the field of human-computer interaction, facial expression recognition plays a crucial role in tasks such as emotion analysis, user behavior prediction, and emotional robots; in the field of mental health, it also provides strong technical support for applications such as autism diagnosis and treatment of mood disorders. In addition, facial expression recognition is widely used in multimedia processing tasks such as video surveillance, driver fatigue detection, and virtual reality. Therefore, systematically studying facial expression recognition algorithms has profound application significance, which can not only improve the naturalness and accuracy of human-computer interaction but also provide a solid foundation for emotional intelligence in multiple fields. Currently, the CNN model extracts local and global features in facial images through multiple convolutional layers and pooling layers. These features include the shape, position, and texture information of key parts such as eyes, mouth, and eyebrows, and their change patterns in different emotional states provide the basis for model recognition. However, when dealing with complex facial expressions, the CNN model may be interfered by external factors such as lighting, occlusion, and angle changes, resulting in a decrease in recognition accuracy. The Transformer model uses the self-attention mechanism to capture the correlation information between regions in facial images. This mechanism regards the facial image as a sequence composed of pixels or feature vectors, and by calculating the influence weights of each element in the sequence on all other elements, it deeply explores potential facial expression features. When dealing with features with long-range dependencies, the Transformer model shows significant advantages. However, this method is also accompanied by high computational complexity and a large scale of model parameters, which may lead to challenges in the training process, including an extended training time and a significant demand for computing resources, thereby increasing resource consumption and implementation difficulty.
[0031] In view of this, some embodiments of the present application provide a method for training a facial recognition model. In this method, the designed initial facial recognition model is trained through an image data set to obtain an image detection result. Then, the loss of the image detection result is calculated through the target loss function constructed in the present application, which contains four parts of loss, so as to optimize the initial facial recognition model through the obtained loss value, and thus obtain a well-trained facial recognition model with high accuracy. Some embodiments of the present application can realize the all-round training and tuning of the initial facial recognition model by training the designed initial facial recognition model and using the constructed target loss function to optimize the initial facial recognition model, improve the effect and robustness of model training. Subsequently, facial expression recognition can be performed through the well-trained facial recognition model, with high accuracy and efficiency, thus improving the effect of facial expression recognition.
[0032] The following combines the attached Figure 1 Exemplarily elaborates the overall composition structure of the system for training a facial recognition model provided by some embodiments of the present application.
[0033] As Figure 1 shown, some embodiments of the present application provide a system diagram for training a facial recognition model. The system for training the facial recognition model may include: a terminal 100 and a server 200. Among them, the terminal 100 pre-arranges an image data set for model training, and the server 200 pre-deploys an initial facial recognition model. When the terminal 100 starts model training, it sends the image data set to the server 200. The server 200 performs multiple rounds of training on the initial facial recognition model through the image data set until the condition for stopping training is reached, and then outputs the trained facial recognition model. The facial recognition model is deployed in the server 200 and can subsequently perform facial expression recognition on the to-be-recognized images transmitted by the terminal 100.
[0034] In some embodiments of the present application, the initial facial recognition model may include: a backbone network module, a feature processing module, and an object detection module. The internal component module structure of the initial facial recognition model can be adjusted as needed. The terminal 100 can be a mobile terminal or a non-portable computer terminal, and the embodiments of the present application do not make specific limitations here.
[0035] Next, in combination with the attached Figure 2 drawings, an implementation process of training a facial recognition model executed by the server 200 provided by some embodiments of the present application is exemplarily described.
[0036] Please refer to the attached Figure 2 , Figure 2 which is a flowchart of a method for training a facial recognition model provided by some embodiments of the present application. Before executing the method for training the facial recognition model in the following S210 - S240, it is first necessary to prepare an image data set for training. Therefore, the process of obtaining the image data set is first exemplarily described below.
[0037] In some embodiments of the present application, the image data set for training is continuously updated and obtained through the following method:
[0038] S201, using the i-th training data set to train the i-th initial facial recognition model to obtain the i-th facial recognition model training result; where i is a positive integer;
[0039] For example, in the embodiments of the present application, during the process of model training, the model needs to be iteratively trained multiple times, and the training data set used in each training process is updated to improve the accuracy of model training. For example, the i-th initial facial model is trained using the i-th training data set to obtain the training result of this time. That is, when i = 1, the first initial facial model is trained using the first training data set to obtain the training result of the first facial recognition model. The first initial facial model is optimized using the target loss function in S240 through the training result of the first facial recognition model to obtain the second initial facial model (that is, the (i + 1)-th initial facial model).
[0040] In some embodiments of the present application, when i = 1, the first training data set is obtained by the following method: the collected original facial expression data set is cropped to obtain a standard data set; the expression categories in the standard data set are data-augmented to obtain the first training data set; wherein, the data augmentation includes: mirror inversion, random region cropping, and random noise addition to the facial expression images in the standard data set.
[0041] For example, in the embodiments of the present application, there are problems such as inconsistent specifications, sizes, and complex backgrounds in the collected original facial expression data set. In order to meet the unity of the input samples of the initial facial recognition model, the original facial expression data set needs to be preprocessed. For example, initially, the face detection tool (haarcascade_frontalface_alt.xml) of OpenCV is used to detect faces in all images in the original facial expression data set and extract the face regions, and the face data is cropped and unified to a size of 128×128 pixels to obtain a standard data set to adapt to the input requirements of the subsequent model. For the expression categories with a small amount of data in the standard data set, in order to balance the category distribution, the following data augmentation techniques are used to expand the data set: mirror flipping: horizontally flipping the image to increase data samples; random cropping: randomly selecting regions from the cropped image for adjustment to enhance data diversity; adding noise: adding random noise to the image to simulate the change in image quality under complex scenarios, etc. The first training data set for training is obtained through the above operations. It should be understood that the manner of processing the original facial expression data set can be adjusted, and the embodiments of the present application are not limited thereto.
[0042] S202, confirm that the loss value of the i-th facial recognition model training result is greater than the set threshold, or, i ≤ N, where N is the number of iterations;
[0043] For example, in the embodiments of the present application, when it is confirmed that the loss value of the training result of the i-th face recognition model does not meet the set threshold condition or the iteration times are not reached, the i-th training data set needs to be updated. If the loss value of the training result of the i-th face recognition model meets the set threshold condition, the i-th initial face model is used as the trained face recognition model. The set threshold can be A, and the loss value meeting the set threshold condition means that the loss value is less than or equal to A. On the contrary, the loss value not meeting the set threshold condition means that the loss value is greater than A. The values of A or N can be set according to the actual application scenario, and the embodiments of the present application do not make specific limitations here.
[0044] S203. Update the i-th training data set based on the training result of the i-th face recognition model to obtain the (i + 1)-th training data set, where the (i + 1)-th training data set is the image data set.
[0045] For example, in the embodiments of the present application, the iteration stage can be divided into two-layer iteration stages according to the model training stage, namely the first-layer iteration and the second-layer iteration, and a two-layer iteration strategy is adopted for further optimization. It should be noted that the update methods of the training data set are different in these two-layer iteration processes. In the first-layer iteration, based on the training result of the i-th face recognition model, the misclassified or low-confidence training data samples are automatically selected, re-annotated and added to the (i + 1)-th training data set, and the model is fine-tuned by continuing the training. In the second-layer iteration, new sample data is generated using the pre-trained model, and after screening out the high-confidence data from the i-th training data set, it is combined with the new sample data to form the (i + 1)-th training data set, expanding the scale of the data set and optimizing the global feature extraction ability of the face recognition model.
[0046] In order to improve the quality of the training data set, the present application is based on iterative semi-supervised annotation. Among them, high-quality annotated data is the premise and foundation for automatically recognizing facial expressions. To address this problem, an iterative semi-supervised annotation strategy (i.e., the above-introduced first-layer iteration and second-layer iteration) is used, and a pre-trained model is introduced in the annotation process. The model prediction is combined with a small amount of manual annotation, and the annotation of large-scale facial expressions is realized through continuous iteration, so as to continuously update the training data set during the training process. In the second-layer iteration, based on the recognized image detection results, training materials (i.e., the training data set) are added for fine-tuning. The goal of this stage is to increase the accuracy of future facial expression recognition for the initial face recognition model in more training materials.
[0047] Since the training image dataset in this application is continuously iteratively updated, the training process of S210 to S240 needs to be executed for the i-th training dataset until the loss value of the model in a certain iterative training is less than or equal to the set threshold, and then the training ends. The following describes the training process of a certain initial face recognition model with a certain training dataset.
[0048] In some embodiments of this application, the method for training a face recognition model may include:
[0049] S210, input the image dataset for training into the backbone network module, and output the image feature information corresponding to the image dataset.
[0050] For example, in some embodiments of this application, the i-th training dataset is input into the backbone network module of the i-th initial face recognition model. The backbone network module extracts and processes the features of the images in the i-th training dataset to obtain image feature information.
[0051] In some embodiments of this application, the backbone network module includes: a preprocessing module, multiple feature extraction modules, and multiple feature fusion modules; S210 may include: the preprocessing module extracts features from the images in the image dataset to obtain facial features; the multiple feature extraction modules obtain the temporal dynamic information corresponding to the facial features and construct a facial feature time series; the multiple feature fusion modules map the facial feature time series to different dimensions to obtain the image feature information.
[0052] For example, in the embodiments of this application, the feature extraction module in the backbone network module may adopt an ODSSBblock (Orthogonal Dynamic State Space Block) module or other types of feature processing modules, and the feature fusion module may adopt a VCM (Vision Clue Merge) module or other modules for fusing features. Specifically, the preprocessing module may first use a feature extraction algorithm to extract features (such as facial feature) from the images in the i-th training dataset to obtain facial features. Then, the feature extraction module may obtain the temporal information of these facial features and construct a facial feature time series in chronological order. Finally, the feature fusion module is used to map the facial feature time series to different spatial or temporal dimensions to obtain image feature information. Enhance local feature representation and subsequent detection accuracy.
[0053] S220, input the image feature information into the feature processing module, and output image fusion features of different sizes.
[0054] For example, in some embodiments of the present application, the feature processing module may process image feature information of different dimensions to obtain image fusion features at different sizes.
[0055] For example, in some embodiments of the present application, S220 may include: after the feature processing module performs sampling operations and convolution operations on the image feature information in different dimensions, the image fusion features of different sizes are obtained.
[0056] For example, in some embodiments of the present application, the feature processing module may perform upsampling and / or downsampling operations on the image feature information of different dimensions, perform convolution operations on the sampling results and the sampling results of other dimensions to obtain image fusion features. For example, there is image feature information in three dimensions. The result after performing downsampling operation on the image feature information of the first dimension is convolved with the result after performing downsampling operation on the image feature information of the second dimension to obtain image fusion features at a set size. For example, the downsampling operation can also be directly performed on the image feature information of the first dimension to output image fusion features of a set size. The set size can be arbitrarily set for the length, width, or height values according to actual needs, such as sizes of 80×80, 40×40, etc. It can be understood that which one or two or three of the upsampling, downsampling, and convolution operations on the image feature information can be flexibly set according to the actual application scenario, and the embodiments of the present application do not make specific limitations here.
[0057] S230, input the image fusion features of different sizes into the target detection module, and output an image detection result.
[0058] For example, in some embodiments of the present application, the target detection module includes sub-modules for detecting image fusion features of different sizes, one size corresponding to one sub-module, so as to obtain image detection results for image fusion features of different sizes.
[0059] In some embodiments of the present application, S230 may include: respectively inputting the image fusion features of different sizes into the target detection layers in the target detection module to obtain the image detection result, where the target detection layers of different sizes are used to detect image fusion features of different sizes.
[0060] For example, in some embodiments of the present application, by inputting image fusion features of different sizes into corresponding-sized sub-modules (as a specific example of the target detection layer), image detection results at corresponding sizes are obtained.
[0061] S240. Optimize the initial face recognition model using the loss value of the image detection result determined by the target loss function to obtain the trained face recognition model, where the target loss function includes: facial expression classification loss, face detection and localization loss, target confidence loss, and temporal consistency loss; the face recognition model is used to recognize facial expressions in the image to be recognized.
[0062] For example, in the embodiments of the present application, in order to reflect the accuracy of model training, the present application designs a target loss function containing multiple losses. Through the target loss function, the loss calculation can be performed on the image detection result and the face recognition samples in the training dataset to obtain the loss value, and the loss value is used to optimize the initial face recognition model for the i-th time, and so on until the trained face recognition model that meets the set threshold conditions is obtained. In the application stage of the face recognition model, the image to be recognized can be input into the face recognition model to output the facial expression recognition result; for example, the facial expression recognition result is classification labels such as happy, sad, melancholy, etc.
[0063] In order to achieve the accuracy of model optimization, in some embodiments of the present application, the target loss function is constructed in the following manner: obtain the respective loss parameters and respective loss weights corresponding to the facial expression classification loss, the face detection and localization loss, the target confidence loss, and the temporal consistency loss; perform weighted summation on the respective loss parameters and the respective loss weights to obtain the target loss function.
[0064] For example, in some embodiments of the present application, the target loss function is composed of four parts.
[0065] The first part is the facial expression classification loss L cls , which can use Cross-Entropy Loss or FocalLoss to handle the class imbalance problem. When calculating the multi-classification loss L cls of facial expressions, the following formula can be used:
[0066]
[0067] where N is the number of samples in the training dataset (usually the number of pixels, feature points, or other spatial dimensions in a video frame, etc.), C is the number of facial expression categories, y i,c is the true expression label of the i-th sample, is the predicted probability.
[0068] The second part is the face detection and localization loss L loc , which can use the CIoU loss of YOLOv8 for the face detection task, and the calculation formula is as follows:
[0069] Lloc = 1 - CIoU(b p , b t )
[0070] where b p is the predicted bounding box in face detection, and b t is the pre-annotated ground truth bounding box in face detection. The CIoU() function is used to calculate the similarity between b p and b t .
[0071] The third part is the object confidence loss L obj , which is used to measure whether the predicted bounding box contains a valid facial expression object (i.e., the difference between the confidence of the predicted bounding box and the true expression label). The calculation formula is as follows:
[0072]
[0073] where y i is the true expression label (ground truth label) of the i-th sample, taking values of 0 or 1: y i = 1 indicates that the sample is a positive sample (i.e., contains the object); y i = 0 indicates that the sample is a negative sample (i.e., does not contain the object). is the confidence (confidence score) of the i-th sample, representing the probability that the model believes the sample contains the object, with a value range of [0,1]. The log() function is used to calculate the difference between the predicted probability and the true label.
[0074] The fourth part is the temporal consistency loss L time , which can add temporal consistency loss for video or continuous frame expression recognition tasks in the training dataset to ensure the smoothness of the model's predictions for continuous frames. It is used to measure the difference between the model's outputs at adjacent time steps. The calculation formula is as follows:
[0075]
[0076] where f t is the output of the model at time step t (such as feature map, optical flow, depth map, etc.). f t-1 is the output of the model at time step t - 1. ||f t - f t-1 || 2 is the square of the Euclidean distance (L2 norm) between f t and f t-1 , representing the difference between the two.
[0077] Weighted sum the above four parts of losses as the final fusion loss function (as a specific example of the target loss function):
[0078] L total = λ loc L loc + λ cls L cls + λ obj L obj + λ time L time
[0079] Wherein, λ loc is the loss weight of the face detection and localization loss, λ cls is the loss weight of the facial expression classification loss, λ obj is the loss weight of the object confidence loss, λ time is the loss weight of the temporal consistency loss. The values of the loss weights can be set according to the actual situation, and the present application does not make specific limitations here.
[0080] In some embodiments of the present application, S240 may include: calculating different losses for the image detection result to obtain respective loss values corresponding to the facial expression classification loss, the face detection and localization loss, the object confidence loss, and the temporal consistency loss; inputting the respective loss values as the respective loss parameters into the target loss function to obtain the loss value.
[0081] For example, in some embodiments of the present application, by comparing the image detection result with the true face detection data in the training dataset, the above four parts of losses are calculated and input into the fusion loss function to obtain the loss value. Through this loss value, the i-th initial face recognition model can be optimized, and finally a trained face recognition model with high accuracy can be obtained.
[0082] It can be seen from some embodiments of the present application described above that the present application can improve the accuracy and efficiency of face recognition, and is particularly suitable for face expression recognition scenarios that require high-efficiency and real-time processing, with high practicality.
[0083] Please refer to Figure 3 , Figure 3 which shows a block diagram of the composition of the device for training a face recognition model provided by some embodiments of the present application. It should be understood that this device for training a face recognition model corresponds to the above method embodiments and can execute each step involved in the above method embodiments. The specific functions of this device for training a face recognition model can be referred to the description above. To avoid repetition, the detailed description is appropriately omitted here.
[0084] Figure 3The apparatus for training a facial recognition model comprises at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the apparatus for training a facial recognition model. The apparatus for training a facial recognition model is used to train an initial facial recognition model to obtain a trained facial recognition model, wherein the initial facial recognition model comprises: a backbone network module, a feature processing module and a target detection module; the apparatus comprises: an extraction module 310, which is used to input an image data set for training into the backbone network module and output image feature information corresponding to the image data set; a fusion module 320, which is used to input the image feature information into the feature processing module and output image fusion features of different sizes; a processing module 330, which is used to input the image fusion features of different sizes into the target detection module and output an image detection result; an optimization module 340, which is used to optimize the initial facial recognition model using the loss value of the image detection result determined by a target loss function to obtain the trained facial recognition model, wherein the target loss function comprises: facial expression classification loss, facial detection positioning loss, target confidence loss and time consistency loss; the facial recognition model is used to recognize facial expressions in images to be recognized.
[0085] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.
[0086] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations of the method corresponding to any of the above methods provided in the above embodiments.
[0087] Some embodiments of the present application further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.
[0088] like Figure 4 As shown, some embodiments of the present application provide an electronic device 400, which includes: a memory 410, a processor 420, and a computer program stored in the memory 410 and executable on the processor 420, wherein the processor 420 can implement a method as described in any of the above embodiments when reading the program from the memory 410 through a bus 430 and executing the program.
[0089] The processor 420 can process digital signals and can include various computing architectures, such as a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, the processor 420 can be a microprocessor.
[0090] The memory 410 can be used to store instructions executed by the processor 420 or data related to the instruction execution. These instructions and / or data can include code for implementing some or all of the functions of one or more modules described in the embodiments of the present application. The processor 420 of the embodiments of the present disclosure can be used to execute the instructions in the memory 410 to implement the methods shown above. The memory 410 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.
[0091] The above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0092] As mentioned above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0093] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
Claims
1. A method for training a facial recognition model, characterized in that: The method is used to train an initial facial recognition model to obtain a trained facial recognition model, wherein the initial facial recognition model includes: a backbone network module, a feature processing module and a target detection module; the method includes: Inputting an image data set for training into the backbone network module, and outputting image feature information corresponding to the image data set; Inputting the image feature information into the feature processing module, and outputting image fusion features of different sizes; Inputting the image fusion features of different sizes into the target detection module, and outputting the image detection result; The initial facial recognition model is optimized by using the loss value of the image detection result determined by the target loss function to obtain the trained facial recognition model, wherein the target loss function includes: facial expression classification loss, facial detection positioning loss, target confidence loss and time consistency loss; the facial recognition model is used to recognize facial expressions in the image to be recognized.
2. The method according to claim 1, characterized in that The image dataset used for training is continuously updated and acquired by the following method: Using the i-th training data set to train the i-th initial face recognition model, obtaining the i-th face recognition model training result; wherein i is a positive integer; Confirm that the loss value of the i-th facial recognition model training result is greater than a set threshold, or, i≤N, where N is the number of iterations; The i-th training data set is updated based on the i-th facial recognition model training result to obtain the i+1-th training data set, wherein the i+1-th training data set is the image data set.
3. The method according to claim 2, characterized in that When i=1, the first training data set is obtained by the following method: The collected original facial expression data set is cropped to obtain a standard data set; Data enhancement is performed on the expression categories in the standard data set to obtain the first training data set; wherein the data enhancement includes: mirror reversal, random area cropping and random noise addition to the facial expression images in the standard data set.
4. The method according to any one of claims 1 to 3, characterized in that The backbone network module includes: a preprocessing module, multiple feature extraction modules and multiple feature fusion modules; The step of inputting the image data set for training into the backbone network module and outputting image feature information corresponding to the image data set includes: The preprocessing module extracts features from the images in the image data set to obtain facial features; The multiple feature extraction modules obtain temporal dynamic information corresponding to the facial features and construct a facial feature time series sequence; The multiple feature fusion modules map the facial feature time series to different dimensions to obtain the image feature information.
5. The method according to claim 4, characterized in that The step of inputting the image feature information into the feature processing module and outputting image fusion features of different sizes includes: The feature processing module performs sampling operations and convolution operations on the image feature information in different dimensions to obtain image fusion features of different sizes.
6. The method according to any one of claims 1 to 3 and 5, characterized in that: The step of inputting the image fusion features of different sizes into the target detection module and outputting the image detection result comprises: The image fusion features of different sizes are respectively input into the target detection layers in the target detection module to obtain the image detection results, wherein the target detection layers of different sizes are used to detect image fusion features of different sizes.
7. The method according to any one of claims 1 to 3 and 5, characterized in that: The objective loss function is constructed as follows: Obtaining loss parameters and loss weights corresponding to the facial expression classification loss, the facial detection positioning loss, the target confidence loss, and the temporal consistency loss, respectively; Performing weighted summation on the loss parameters and the loss weights to obtain the target loss function; The loss value of the image detection result is obtained by the following method; Calculating different losses on the image detection result to obtain loss values corresponding to the facial expression classification loss, the facial detection positioning loss, the target confidence loss, and the time consistency loss respectively; The respective loss values are input into the target loss function as the respective loss parameters to obtain the loss values.
8. A device for training a facial recognition model, characterized in that: The device is used to train an initial facial recognition model to obtain a trained facial recognition model, wherein the initial facial recognition model includes: a backbone network module, a feature processing module and a target detection module; the device includes: An extraction module, used to input an image data set for training into the backbone network module, and output image feature information corresponding to the image data set; A fusion module, used for inputting the image feature information into the feature processing module and outputting image fusion features of different sizes; A processing module, used for inputting the image fusion features of different sizes into the target detection module and outputting the image detection result; An optimization module is used to optimize the initial facial recognition model using the loss value of the image detection result determined by a target loss function to obtain the trained facial recognition model, wherein the target loss function includes: facial expression classification loss, facial detection positioning loss, target confidence loss and time consistency loss; the facial recognition model is used to recognize facial expressions in the image to be recognized.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program executes the method according to any one of claims 1 to 7 when executed by a processor.
10. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program executes the method according to any one of claims 1 to 7 when being run by the processor.
Citation Information
Patent Citations
Method and device for correcting training set
CN109543713A
Classroom facial expression recognition method and device based on YOLOv4
CN116453178A
Face emotion recognition network model training method and device, equipment and medium
CN117315752A
Real-time monitoring system and method for high-risk animals around transformer substation based on deep learning
CN117649741A
Method for constructing training model for medical chest image disease classification based on Resnext integrated network
CN117911771A