Face recognition processing method based on multi-task cascaded convolutional network and storage medium

By constructing a multi-task cascaded convolutional network and optimizing the loss function, the problems of low recall and high computational complexity of multi-task neural network models in face recognition are solved, achieving efficient face recognition and detection.

CN120071412BActive Publication Date: 2025-11-18KUNMING RENLIANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411934269.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-11-18
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing multi-task neural network models have insufficient recall in face recognition and suffer from high computational complexity, increased model complexity and training difficulty, especially in deep learning networks, leading to performance degradation.

Method used

A multi-task cascaded convolutional network is constructed, which is divided into a three-level convolutional neural network. The model is optimized by using cross-entropy loss function, regression loss function and landmark location loss function to reduce unnecessary running parameters, unify candidate window size and filter overlap, thereby improving the efficiency of face detection and recognition.

Benefits of technology

The process of acquiring target candidate boxes has been optimized, which improves the speed and accuracy of face image recognition, simplifies the model structure, reduces computational complexity, and enables fast recognition and detection of multiple target faces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071412B_ABST
    Figure CN120071412B_ABST
Patent Text Reader

Abstract

The application provides a face recognition processing method based on a multi-task cascaded convolutional network and a storage medium, and the method comprises the following steps: constructing a multi-task cascaded convolutional network, dividing the multi-task cascaded convolutional network into three convolutional neural networks, the first cascaded network is P-Net, the second cascaded network and the third cascaded network are R-Net and O-Net respectively; acquiring a single face image sample set to train the multi-task cascaded convolutional network, and obtaining a converged multi-task cascaded convolutional network; acquiring a face image to be detected, preprocessing the face image to be detected, inputting the preprocessed face image to be detected into the first cascaded network, and obtaining a face candidate window image; screening the face candidate window image through the confidence of the second cascaded network; inputting the face candidate window image screened through the confidence into the third cascaded network, and outputting a final face image through the third cascaded network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a face recognition processing method and storage medium based on a multi-task cascaded convolutional network. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence face detection technology, convolutional neural network technology has played an increasingly important role. Among them, face image recognition, image processing, and automatic detection technologies are the most prominent. These technologies are ubiquitous in our lives. However, due to factors such as face angle, expression, and illumination, the loss function of convolutional neural networks has a huge impact on the calculation. Therefore, the recognition of facial attributes remains a huge challenge.

[0003] Research has found that multi-task learning through multi-task neural network models can complement each other, improve data utilization efficiency, and when there is correlation between different tasks, sharing data can help the model generalize better and improve model performance. Furthermore, by training on multiple tasks, the model can learn richer feature representations, thereby improving its robustness to input data.

[0004] However, researchers found that traditional multi-task neural network models were insufficient in recalling faces. Deep learning-based algorithms, on the other hand, use networks that are too deep, resulting in high computational complexity. Furthermore, with large model files, conflicts or competition between different tasks can lead to performance degradation. Additionally, multi-task neural network models typically involve more parameters, potentially increasing model complexity and training difficulty.

[0005] Therefore, there is an urgent need for a face image recognition method based on multi-task cascaded convolutional networks to improve the model's computation speed and detection efficiency. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and propose a face recognition processing method and storage medium based on a multi-task cascaded convolutional network. This method realizes the screening and selection of composite factors for target candidate boxes, and ultimately optimizes the acquisition process of target candidate boxes, thereby improving the recognition speed of face images.

[0007] This invention provides a face recognition processing method based on a multi-task cascaded convolutional network, comprising the following steps:

[0008] A multi-task cascaded convolutional network is constructed, which is divided into three convolutional neural networks. The first cascaded network is P-Net, the second cascaded network and the third cascaded network are R-Net and O-Net, respectively. A single face image sample set is collected to train the multi-task cascaded convolutional network to obtain the converged multi-task cascaded convolutional network.

[0009] Acquire a face image to be detected, preprocess the face image to be detected, and input the preprocessed face image to be detected into the first cascaded network to obtain an image containing face candidate windows;

[0010] The face candidate window image is filtered by the confidence level of the second cascaded network; the face candidate window image after confidence level filtering is input into the third cascaded network, and the third cascaded network processes and outputs the final image containing the face.

[0011] Preferably, the multi-task cascaded convolutional network is trained by acquiring a single-face image sample set to obtain a converged multi-task cascaded convolutional network. Specifically, obtaining the converged multi-task cascaded convolutional network includes:

[0012] For the single face image sample set, feature descriptors for face classification, candidate window regression, and face landmark localization are established respectively. Based on the feature descriptors for face classification, a cross-entropy loss function for face classification is constructed; based on the feature descriptors for candidate window regression, a regression loss function for candidate windows is constructed; and based on the feature descriptors for face landmark localization, a regression loss function for landmark locations is constructed.

[0013] The target total loss function of the multi-task concatenated convolutional network is constructed based on the cross-entropy loss function of face classification, the regression loss function of candidate windows, and the regression loss function of landmark location; when determining the target total loss function, the converged final multi-task concatenated convolutional network is obtained by solving.

[0014] Preferably, feature descriptors for face classification, candidate window regression, and face landmark localization are established for the single-face image sample set, respectively; a cross-entropy loss function for face classification is constructed based on the feature descriptors for face classification, a regression loss function for candidate windows is constructed based on the feature descriptors for candidate window regression, and a regression loss function for landmark locations is constructed based on the feature descriptors for face landmark localization, specifically including:

[0015] Classify facial features on a single-face image sample set, using the cross-entropy loss function for face classification:

[0016]

[0017] In the formula, pi represents the probability that the current face belongs to a human. The real label is used as the background.

[0018] Candidate windows are selected for prediction of facial features in the single-face image sample set. There is an offset between the predicted candidate windows and the actual candidate windows. The regression loss function for the candidate windows is calculated using Euclidean distance.

[0019]

[0020] In the formula, The position of the candidate window is predicted by a multi-task cascaded convolutional network. The position of the actual candidate window;

[0021] The Euclidean distance is calculated between the currently predicted facial landmark coordinates and the actual facial landmarks, and the regression loss function for the landmark location is calculated using the Euclidean distance.

[0022]

[0023] In the formula, The predicted facial landmark coordinates are obtained from a multi-task cascaded convolutional network, while These are the actual landmark coordinates of the face.

[0024] Preferably, a target total loss function for the multi-task cascaded convolutional network is constructed based on the cross-entropy loss function for face classification, the regression loss function for candidate windows, and the regression loss function for landmark locations. When determining the target total loss function, the converged final multi-task cascaded convolutional network is obtained by solving for it, specifically including:

[0025] Based on the aforementioned calculations of the cross-entropy loss function for face classification, the regression loss function for the candidate window, and the regression loss function for the landmark location, the target total loss function of the grid model is calculated, and its minimum value is obtained.

[0026]

[0027] In the formula, N is the number of training samples, αj represents the weights of the loss function for different network structures, βj is the sample label, and Lj is the loss function;

[0028] When determining the minimum value of the target total loss function, the optimized multi-task cascaded convolutional network parameters corresponding to the minimum value of the target total loss function are obtained; the multi-task cascaded convolutional network is trained based on the solved optimized multi-task cascaded convolutional network parameters to obtain the final multi-task cascaded convolutional network; that is, until the training network parameters are stable, the final multi-task cascaded convolutional network is saved after training is completed.

[0029] Preferably, the process involves acquiring a face image to be detected, preprocessing the face image, and inputting the preprocessed face image into the first cascaded network to obtain an image containing face candidate windows. Specifically, this includes:

[0030] The face image to be detected is preprocessed to obtain an image pyramid of the face image to be detected; the image pyramid includes multiple layers of preprocessed face images; the preprocessing refers to scaling the current face image to be detected multiple times, with each scaling yielding one layer of face image;

[0031] When performing sliding detection on the face image of each layer's scaled image, the step size of the associated adjacent layers is adjusted by varying the step size to obtain a sliding detection image containing face candidate windows.

[0032] Preferably, when performing sliding detection on the face image for each layer of scaled image, the step size associated with adjacent layers is adjusted by varying the step size to obtain a sliding detection image containing face candidate windows, specifically including:

[0033] During initialization, the scaling ratio of the current layer's scaled image relative to the original image is obtained; then, a single target face is detected on the scaled image of the current layer, and a standard step size is randomly set within the current layer.

[0034] On the scaled image of the current layer, slide detection of the first face target is performed according to the standard step size, and multiple candidate windows for the current first face target are output; single target detection is performed on the current first face target to obtain multiple candidate windows for the current first face target.

[0035] The process involves determining the confidence level of multiple candidate windows for the first face in the scaled image of the current layer; identifying windows with a confidence level greater than a standard confidence threshold as valid candidate windows; determining whether the number of valid candidate windows for the first face is greater than a candidate window count limit; if the number of valid candidate windows is greater than the candidate window count limit, it is considered that there are still many candidate windows for the selected first face, so the standard step size of the current layer is still small. Therefore, when detecting the second face in the scaled image of the current layer, the preset step size of the current layer is increased to achieve the detection of the second face. The above steps are repeated to determine whether the number of valid candidate windows for the second face is greater than the candidate window count limit. If the number of valid candidate windows is still greater than the candidate window count limit, the step size is increased again for the detection of the third face until all face detections in the current layer are completed or until the number of valid candidate windows for the nth face in the current layer is less than the candidate window count limit, at which point the standard step size is no longer increased; the final standard step size is determined as the target standard step size for the scaled image of the current layer.

[0036] When performing sliding detection on the scaled image of the next layer, first determine the relationship between the scaled image of the next layer and the scaled image of the current layer. If the scaled image of the next layer is a further scaled image of the scaled image of the current layer, then directly call the target standard step size of the scaled image of the previous layer as the standard step size of the scaled image of this layer to perform sliding detection of the 1-m face target objects on the scaled image of this layer.

[0037] Repeat the above steps to determine whether it is necessary to continue to perform sliding detection processing by adjusting the step size when detecting face targets of the first (m)th person in this layer with the target standard step size.

[0038] Preferably, before inputting the face candidate window image into the second cascaded network, the process of acquiring the face candidate window image further includes filtering the target candidate window in the face candidate window image that needs to be transformed into a square image candidate window.

[0039] Determine the position and size of the target candidate window that needs to be converted into a square; stretch the two longer sides of the candidate window outwards simultaneously so that the two shorter sides form a square with equal lengths to the two longer sides; crop the candidate window that has been stretched into a square, and scale the internal image of the cropped candidate window to 24×24.

[0040] Repeat the above steps: convert multiple candidate windows of the face candidate window image into square candidate windows until the internal images of all candidate windows have been converted into 24×24 square candidate windows;

[0041] Multiple square candidate windows are converted into 24×24 square face candidate window images and input into the second cascade network. The square candidate windows are sorted by confidence level, and the overlap between the square candidate window with the highest confidence level and the remaining square candidate windows with lower confidence levels is calculated. If the overlap between the square candidate window with the highest confidence level and the square candidate window with lower confidence levels is greater than the calculated overlap threshold, the low-confidence square candidate window is removed. The overlap between the remaining low-confidence square candidate windows is calculated until all square candidate windows are removed and the overlap is greater than the calculated overlap threshold. Multiple target face candidate window images with an overlap less than the calculated overlap threshold are obtained. The square candidate windows with an overlap less than the calculated overlap threshold are input into the third cascade network for subsequent operations.

[0042] Preferably, before the square candidate windows with an overlap less than the calculated overlap threshold are input into the third cascaded network, the following steps are also included:

[0043] The rectangular candidate window output by the second cascaded network is scaled to a square candidate window of size 48×48; the two longer sides of the candidate window are stretched outwards simultaneously so that the two shorter sides form a square with equal lengths to the two longer sides; the stretched square candidate window is cropped, and the internal image of the cropped candidate window is scaled to 48×48; all candidate windows are converted into 48×48 square candidate windows.

[0044] The candidate window is converted into a 48×48 square window, and the overlap of the confidence score is calculated to obtain the final target detection window.

[0045] Preferably, the third cascaded network processes and outputs the final image containing the face, specifically including: the third cascaded network outputs the final candidate box through the target detection window of the current face; determines the coordinates of the final candidate box, and obtains the final image containing the face.

[0046] The present invention provides an electronic device, including: a memory, a processor, and a face detection program of a multi-task cascaded convolutional network stored in the memory and running on the processor, wherein the face detection program of the multi-task cascaded convolutional network implements the steps of a face detection method of the multi-task cascaded convolutional network when executed by the processor.

[0047] This invention provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the face detection method of the multi-task cascaded convolutional network described above.

[0048] The face recognition processing method and storage medium based on multi-task cascaded convolutional networks provided by this invention have the following technical advantages:

[0049] Analysis of the face recognition processing method based on a multi-task cascaded convolutional network provided in this embodiment of the invention reveals that its main operations are: constructing a multi-task cascaded convolutional network; collecting a single-face image sample set and establishing feature descriptors for face classification, which can be used to compare and recognize faces in the image with known faces; feature descriptors for candidate window regression, which can be used to locate the face in the image; feature descriptors for face landmark localization, which can be used to determine the location of key points on the face, such as eyes and mouth; constructing a cross-entropy loss function for face classification based on the feature descriptors for face classification, thereby improving the accuracy of face classification; and constructing a regression loss function for candidate windows based on the feature descriptors for candidate window regression, thereby improving the accuracy of candidate window regression. Accuracy; Constructing a regression loss function for landmark location based on feature descriptors for face landmark localization can achieve more accurate face landmark localization; The role of these loss functions is to guide the model to learn correct feature representations and predictions during training, thereby improving the accuracy and robustness of face classification, localization, and landmark localization, and increasing the training speed of the multi-task cascaded network; By simultaneously optimizing the loss functions of these three components, the model can perform better in face detection and recognition tasks, and the target total loss function of the multi-task cascaded convolutional network is constructed using the loss functions of these three components; Solving this function yields the converged final multi-task cascaded convolutional network; The finally trained multi-task cascaded convolutional network can be used for real-time face detection and recognition tasks, achieving high accuracy and improving the speed of face image detection;

[0050] Furthermore, before inputting the face image to be detected into the multi-task cascaded network, the face image is preprocessed to obtain an image pyramid of the face image to be detected. Each layer of the image pyramid detects candidate windows by sliding a step size, and obtains the scaling ratio of the current layer's scaled image relative to the original image. By detecting a single target face on the scaled image of the current layer, the presence of targets in the current layer is determined. A standard step size is randomly set within the current layer. This helps the algorithm better explore the space of the current layer to find more targets. There may be multiple targets in the same image, and randomly setting the step size can, to some extent, prevent the algorithm from getting stuck in a local optimum and failing to find other targets. Face targets are detected by sliding a candidate window on the scaled image of the current layer according to the standard step size, and multiple candidate windows for the current face targets are output. These candidate windows are obtained by sliding a candidate window on the image with a fixed step size of the current layer to perform single target detection for the current face target. Multiple candidate windows for a target object can improve the accuracy and recall of face detection, finding more face targets. The confidence level of the face targets in multiple candidate windows of the current face target object in the scaled image of the current layer is determined. Window with a confidence level greater than a standard confidence threshold is considered a valid candidate window. Each candidate window can be evaluated to determine if it contains a face target and a confidence score is given. The confidence score indicates the presence of a face target. The number of valid candidate windows for the current face target object is checked against a threshold limit. The number of candidate windows determines the step size of the current layer. If the number of valid candidate windows exceeds the set threshold, the step size of the current layer needs to be increased to determine the number of subsequent face target candidate windows, ensuring the algorithm's performance and speed. This continues until the number of valid candidate windows for detected face targets is less than the threshold limit, at which point the step size is determined as the target standard step size for scaling the image of the current layer.

[0051] When performing sliding detection on the scaled image of the next layer, the relationship between the scaled image of the next layer and the scaled image of the current layer is first determined. If the scaled image of the next layer is a further scaled image of the current layer, the target standard step size of the scaled image of the previous layer is directly used as the standard step size of the scaled image of this layer to perform sliding detection of the 1-m face targets on the scaled image of this layer. Since the scaled image of the next layer is a further scaled image of the current layer, directly using the target standard step size of the previous layer as the standard step size of this layer can avoid repeated calculations, reduce the amount of computation, and improve the efficiency of the algorithm. Furthermore, by using the target standard step size of the previous layer, the results of the detection candidate windows of the previous layers and the information of the candidate windows on the image can be inherited.

[0052] Furthermore, after the face image to be detected is input into the multi-task cascaded network, the candidate windows of the face image to be detected are adjusted. By using a uniform input size, the number of parameters and computation can be reduced, thereby simplifying the model structure and reducing the risk of overfitting. At the same time, it avoids the problem of slow detection speed caused by inconsistent sizes. Moreover, by repeatedly calculating the overlap of confidence scores through candidate windows of uniform size, redundant candidate windows are filtered layer by layer to ensure that the final target detection box obtains the face image.

[0053] Therefore, the face recognition processing method based on multi-task cascaded convolutional networks adopted in this embodiment has been optimized and adjusted in many aspects, such as optimizing the network model, reducing unnecessary operating parameters, and quickly determining candidate windows. Ultimately, it ensures the rapid recognition and detection of multi-target faces in large-scale massive images and improves the efficiency of detection and processing. Attached Figure Description

[0054] Figure 1 This is a flowchart illustrating a face recognition processing method based on a multi-task cascaded convolutional network, as described in Embodiment 1.

[0055] Figure 2 The flowchart for constructing a multi-task cascaded convolutional network in a face recognition processing method based on a multi-task cascaded convolutional network is shown in Example 1.

[0056] Figure 3 Here is a flowchart illustrating the process of establishing feature descriptors for a single face image sample set in a face recognition processing method based on a multi-task cascaded convolutional network, as described in Example 1.

[0057] Figure 4 This is a flowchart illustrating the calculation of the target total loss function of a multi-task cascaded convolutional network in a face recognition processing method based on a multi-task cascaded convolutional network, as described in Example 1.

[0058] Figure 5 This is a flowchart illustrating the process of obtaining a face candidate window image in a face recognition processing method based on a multi-task cascaded convolutional network, as described in Embodiment 1.

[0059] Figure 6 This is a flowchart of the face image scaling and sliding detection process in a face recognition processing method based on a multi-task cascaded convolutional network, as described in Embodiment 1.

[0060] Figure 7 This is an example of a face recognition processing method based on a multi-task cascaded convolutional network in Embodiment 1, which involves scaling and sliding to detect face images.

[0061] Figure 8 This is a flowchart illustrating the process of obtaining a face candidate window image in a face recognition processing method based on a multi-task cascaded convolutional network, as described in Embodiment 1.

[0062] Figure 9 This is a square transformation image of the candidate window of the face image to be detected before it is input into the cascaded network in a face recognition processing method based on a multi-task cascaded convolutional network in Embodiment 1.

[0063] Figure 10 This is a schematic diagram of the structure of a storage medium for applying the above-described face recognition processing method based on a multi-task cascaded convolutional network in Embodiment 3.

[0064] Labels: Processor 1110; Communication interface 1120; Memory 1130; Computer storage medium 1140. Detailed Implementation

[0065] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0066] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings.

[0067] Example 1

[0068] like Figure 1 As shown, Embodiment 1 of the present invention provides a face recognition processing method based on a multi-task cascaded convolutional network, including the following operation steps:

[0069] S1. Construct a multi-task cascaded convolutional network, dividing the multi-task cascaded convolutional network into a three-level convolutional neural network. The first cascaded network is P-Net, the second cascaded network and the third cascaded network are R-Net and O-Net, respectively. Collect a single face image sample set to train the multi-task cascaded convolutional network to obtain the converged multi-task cascaded convolutional network.

[0070] S2. Obtain the face image to be detected, preprocess the face image to be detected, and input the preprocessed face image to be detected into the first cascaded network to obtain an image containing face candidate windows.

[0071] S3. The face candidate window image is filtered by the confidence of the second cascade network; the face candidate window image after confidence filtering is input into the third cascade network, and the third cascade network processes and outputs the final image containing the face.

[0072] It should be noted that the multi-task cascaded convolutional network is a three-level convolutional neural network, which includes: P-Net network: 12×12, for finding proposal boxes; R-Net network: 24×24, for refining proposal boxes; and O-Net network: 48×48, for outputting the final result.

[0073] Better, such as Figure 2 As shown, a single-face image sample set is collected to train the multi-task cascaded convolutional network, resulting in a converged multi-task cascaded convolutional network, specifically including:

[0074] S11. For the single face image sample set, establish feature descriptors for face classification, candidate window regression, and face landmark localization respectively; construct a cross-entropy loss function for face classification based on the feature descriptor for face classification, construct a regression loss function for candidate windows based on the feature descriptor for candidate window regression, and construct a regression loss function for landmark locations based on the feature descriptor for face landmark localization.

[0075] S12. Construct the target total loss function of the multi-task concatenated convolutional network based on the cross-entropy loss function of the face classification, the regression loss function of the candidate window, and the regression loss function of the landmark location; when determining the target total loss function, solve to obtain the converged final multi-task concatenated convolutional network.

[0076] It should be noted that the image size of the single face image sample set during training is the size specified by the network. The first cascaded network mentioned above is a convolutional neural network, which can be designed as a convolutional neural network without fully connected layers. Therefore, there are no size requirements during the training phase. The first cascaded network can predict N candidate windows and confidence scores for a multi-face sample set of arbitrary input size. Then, during training, candidate windows with high overlap can be filtered by thresholding. Because of its small structure, this network is very efficient.

[0077] In step S11, feature descriptors for classifying face features are established for each single face image sample set. These descriptors can be used to compare and identify faces in the image with known faces. Feature descriptors for candidate window regression can be used to locate the face in the image. Feature descriptors for face landmark localization can be used to determine the location of key points on the face, such as eyes and mouth. Based on the feature descriptors for face classification, a cross-entropy loss function for face classification is constructed to improve the accuracy of face classification. Based on the feature descriptors for candidate window regression, a regression loss function for candidate windows is constructed to improve the accuracy of candidate window regression. Based on the feature descriptors for face landmark localization, a regression loss function for landmark location is constructed to achieve more accurate face landmark localization. The role of these loss functions is to guide the model to learn the correct feature representation and prediction during training, so as to improve the accuracy and robustness of face classification, localization and landmark localization, and improve the training speed of multi-task cascaded networks.

[0078] In step S12, the target total loss function of the multi-task cascaded convolutional network is constructed based on the cross-entropy loss function of face classification, the regression loss function of candidate windows, and the regression loss function of landmark location; the final multi-task cascaded convolutional network after convergence is obtained by solving the solution; the finally trained multi-task cascaded convolutional network can be used for real-time face detection and recognition tasks, with high accuracy and improved speed when detecting face images.

[0079] Better, such as Figure 3 As shown, feature descriptors for face classification, candidate window regression, and face landmark localization are established for the single-face image sample set, respectively. A cross-entropy loss function for face classification is constructed based on the feature descriptors for face classification; a regression loss function for candidate windows is constructed based on the feature descriptors for candidate window regression; and a regression loss function for landmark locations is constructed based on the feature descriptors for face landmark localization. Specifically, this includes:

[0080] S111. Classify the faces of the single face image sample set (the current sample image i), and calculate the cross-entropy loss function for face classification using the single face image sample set:

[0081]

[0082] In the formula, pi represents the probability that the current face belongs to a human. The real label is used as the background.

[0083] S112. Candidate windows for predicting facial features in the single-face image sample set are selected by bounding boxes. There is an offset between the predicted candidate windows and the actual candidate windows. The regression loss function (i.e., offset loss) of the candidate windows is calculated using Euclidean distance.

[0084]

[0085] In the formula, The position of the candidate window is predicted by a multi-task cascaded convolutional network. This represents the position of the actual candidate window; the above position coordinates have four coordinates, including the top and bottom corners, height, and width of the candidate window. y represents the background coordinates predicted by a multi-task cascaded convolutional network; y represents the actual background coordinates.

[0086] S113. Calculate the Euclidean distance between the currently predicted facial landmark coordinates and the actual facial landmarks, and use the regression loss function (i.e., landmark location loss) calculated by the Euclidean distance to determine the landmark location.

[0087]

[0088] In the formula, The predicted facial landmark coordinates are obtained from a multi-task cascaded convolutional network, while These are the actual landmark coordinates of the face. Generally speaking, there are five facial locations: the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth. However, this embodiment of the invention preferably locks the facial landmark coordinates at four facial positions: the left eye, right eye, left corner of the mouth, and right corner of the mouth, thereby reducing the number of operating parameters. This is because this embodiment of the invention believes that once the positions of the left eye, right eye, left corner of the mouth, and right corner of the mouth are locked in a typical human face, the position of the nose is not of high importance.

[0089] Better, such as Figure 4 As shown, a target total loss function for a multi-task cascaded convolutional network is constructed based on the cross-entropy loss function for face classification, the regression loss function for candidate windows, and the regression loss function for landmark locations. When determining the target total loss function, the converged final multi-task cascaded convolutional network is obtained by solving for it, specifically including:

[0090] S121. Based on the aforementioned calculated cross-entropy loss function for face classification, regression loss function for candidate windows, and regression loss function for landmark locations, calculate the target total loss function (total loss function) of the grid model, and find its minimum value based on the target total loss function:

[0091]

[0092] Each convolutional neural network uses different tasks, and each stage of the convolutional neural network has a different training objective. Therefore, the entire training and learning process can be transformed into finding the process of minimizing the total loss function mentioned above.

[0093] In the formula, N is the number of training samples, αj represents the weights of the loss functions for det, box, and landmark in different network structures, βj is the sample label, Lj is the loss function, and i represents the sample image.

[0094] In the formula, αj represents the weights of the loss functions for det, box, and landmark in different network structures. Clearly, the weights of the loss functions for det, box, and landmark differ for different network structures; and j∈{det,box,landmark}. For example, P-Net and R-Net use (αdet=1, αbox=0.5, αlandmark=0.5), while O-Net uses (αdet=1, αbox=0.5, αlandmark=1) to more accurately locate facial landmarks. Because O-Net aims to output accurate landmarks, the landmark loss weights in O-Net are larger than those in the previous two stages of the network.

[0095] Research has found that an excessive number of samples leads to low model computational efficiency: if the number of samples in a single face image sample set is large, a large amount of image and labeled data needs to be processed, which increases the overall computation time. For each sample, operations such as forward propagation, loss function calculation, and backpropagation are required, and the computation time increases exponentially with the increase in the number of samples. To address this, S113 reduces the number of landmark location annotations, thereby reducing computational parameters and increasing computational efficiency.

[0096] S122. When determining the minimum value of the target total loss function, obtain the optimized multi-task cascaded convolutional network parameters corresponding to the minimum value of the target total loss function; train the multi-task cascaded convolutional network based on the solved optimized multi-task cascaded convolutional network parameters to obtain the final multi-task cascaded convolutional network; that is, until the training network parameters are stable, after training is completed, save the final multi-task cascaded convolutional network.

[0097] The multi-task concatenated convolutional network used in this embodiment of the invention employs a lightweight model structure, reducing the number of network layers (the first grid does not have a fully connected layer) and the number of parameters to reduce computational complexity. Through the above optimization methods, the computational efficiency of the above processing can be significantly improved, thereby enabling more efficient performance of tasks such as face gender classification, candidate window regression, and face landmark localization.

[0098] Better, such as Figure 5 As shown, researchers found that when using a P-Net network to slide across the image with a stride of 2 during detection, the small stride can cause a face to be bounded multiple times. This is because each scaling operation creates a new scaled image, and preprocessing involves scaling the current face image multiple times. After multiple scaling operations, the current 12×12 candidate window needs to be used for stride-based sliding detection on the scaled image at that layer. If different scaling ratios result in the same sliding distance for a given 12×12 candidate window on each image layer, it can lead to either too many candidate windows (too small a stride) or too few candidate windows (too large a stride). Therefore, this invention suggests the need for effective and reasonable control over the current stride-based sliding detection.

[0099] In the above technical solution, acquiring a face image to be detected, preprocessing the face image to be detected, and inputting the preprocessed face image to be detected into the first cascaded network to obtain an image containing face candidate windows specifically includes:

[0100] S21. The face image to be detected is preprocessed to obtain an image pyramid of the face image to be detected; the image pyramid includes multiple layers of preprocessed face images; the preprocessing refers to scaling the current face image to be detected multiple times, with each scaling yielding one layer of face image;

[0101] S22. When performing sliding detection on the face image of each layer's scaled image, the step size of the associated adjacent layers is adjusted by changing the step size to obtain a sliding detection image containing face candidate windows.

[0102] In step S21, the face image to be detected is preprocessed by scaling the current face image to be detected multiple times, and each scaling yields a face image of one layer, so as to obtain the image pyramid of the face image to be detected; a series of images at different scales can be generated to perform face detection at different scales.

[0103] In step S22, when performing sliding detection of face images on the scaled images of each layer, the step size of the adjacent layers is adjusted by changing the step size to obtain a window image containing face candidate images. Using a larger step size at a larger image scale can reduce the amount of computation and speed up the detection process; while using a smaller step size at a smaller image scale can improve the detection accuracy of small-scale faces.

[0104] Better, such as Figure 6 , 7As shown, when performing sliding detection on the face image for each layer of scaled image, the step size associated with adjacent layers is adjusted by varying the step size to obtain an image containing face candidate windows, specifically including:

[0105] S221. During initialization, obtain the scaling ratio of the current layer's scaled image relative to the original image; then detect a single target face on the scaled image of the current layer, and randomly set a standard step size within the current layer.

[0106] S222. Perform sliding detection of the first face target on the scaled image of the current layer according to the standard step size, and output multiple candidate windows for the current first face target; perform single target detection on the current first face target to obtain multiple candidate windows for the current first face target;

[0107] S223. Determine the confidence level of the first face target among multiple candidate windows of the current first face target in the scaled image of the current layer; determine that a confidence level greater than the standard confidence threshold is a valid candidate window; determine whether the multiple valid candidate windows of the current first face target are greater than the candidate window number limit threshold; if the number of valid candidate windows is greater than the candidate window number limit threshold, it is considered that there are still many candidate windows of the currently selected first face target, so it is determined that the standard step size of the current layer is still small. Therefore, when detecting the second face target in the scaled image of the current layer (the scenario assumed here is that the face sizes of several people in the same scaled image will not differ too much), For images with large numbers of faces, especially those with multiple faces in the same close-up position, increase the preset step size of the current layer to detect the second face. Repeat the above steps to determine whether the number of valid candidate windows for the second face is greater than the candidate window number limit threshold. If the number of valid candidate windows is still greater than the candidate window number limit threshold, continue to increase the step size for the third face detection until all face targets in the current layer have been detected, or until the number of valid candidate windows for the nth face detection in the current layer is less than the candidate window number limit threshold. Then, stop increasing the standard step size. Determine the final standard step size as the target standard step size for scaling the image in the current layer.

[0108] S224. When performing sliding detection on the scaled image of the next layer, first determine the relationship between the scaled image of the next layer and the scaled image of the current layer. If the scaled image of the next layer is a further scaled image of the scaled image of the current layer (i.e., the next layer image is the next layer after the current layer image, then the two layers are considered to be adjacent upper and lower layers); then directly call the target standard step size of the scaled image of the previous layer as the standard step size of the scaled image of this layer to perform sliding detection of the 1-m face target objects on the scaled image of this layer.

[0109] Repeat step S223 above to determine whether it is necessary to continue to perform sliding detection processing by adjusting the step size when the target standard step size is used to detect the face target of the 1-m person in this layer.

[0110] When S224 is executed, the target standard step size of the current layer's scaled image is directly called as the target standard step size of the next layer's scaled image to slide and detect the 1-m face targets on the scaled image of the next layer. The above steps S222-S223 are repeated to perform the standard step size increase judgment processing. Through the above processing, the number of effective candidate windows can be reasonably limited to a large extent, avoiding the low detection efficiency caused by too many effective candidate windows.

[0111] It should be noted that since each scaling results in a new scaled image, preprocessing refers to scaling the current face image to be detected multiple times. After multiple scaling processes, the current 12×12 candidate window needs to be used to perform step-size sliding detection on the scaled image of that layer. Due to different scaling ratios, if the sliding distance of a certain current 12×12 candidate window is the same on each layer of the image, there will be too many candidate windows (i.e., too small a step size) or too few candidate windows (e.g., too large a step size). Therefore, this embodiment of the invention believes it is necessary to implement effective and reasonable control over the current step-size sliding detection.

[0112] In step S221, during initialization, the scaling ratio of the current layer's scaled image relative to the original image is obtained; then, a single target face is detected on the scaled image of the current layer to determine which targets exist in the current layer; a standard step size is randomly set within the current layer; this can help the algorithm better explore the space of the current layer to find more targets. There may be multiple targets in the same image, and by randomly setting the step size, the algorithm can avoid getting stuck in a local optimum and failing to find other targets to a certain extent.

[0113] In step S222, the first face target is detected by sliding on the scaled image of the current layer according to the standard step size, and multiple candidate windows for the current first face target are output. These candidate windows are obtained by sliding a candidate window on the image with a fixed step size of the current layer to perform single target detection for the current first face target, thereby improving the accuracy and recall of face detection and finding more face targets.

[0114] In step S223, the confidence level of the first face target in multiple candidate windows of the first face target in the scaled image of the current layer is determined; a candidate window with a confidence level greater than the standard confidence threshold is determined to be a valid candidate window; each candidate window can be evaluated to determine whether it contains a face target and a confidence score is given, which can be used to determine whether a face target exists; it is determined whether the multiple valid candidate windows of the first face target are greater than the candidate window number limit threshold; the step size of the current layer can be determined by the number of candidate windows, and if the number of generated valid candidate windows exceeds the set threshold, the step size of the current layer needs to be increased to determine the number of candidate windows of the second face target, so as to ensure the performance and speed of the algorithm, until the number of valid candidate windows of the detected face target is less than the candidate window number limit threshold, then the step size at this time is determined as the target standard step size of the scaled image of the current layer;

[0115] In step S224, when performing sliding detection on the scaled image of the next layer, the relationship between the scaled image of the next layer and the scaled image of the current layer is first determined. If the scaled image of the next layer is a further scaled image of the current layer, then the target standard step size of the scaled image of the previous layer is directly used as the standard step size of the scaled image of this layer to perform sliding detection of the 1-m face targets on the scaled image of this layer. Since the scaled image of the next layer is a further scaled image of the current layer, directly using the target standard step size of the previous layer as the standard step size of this layer can avoid repeated calculations, reduce the amount of calculation, and improve the efficiency of the algorithm. Furthermore, by using the target standard step size of the previous layer, the results of the detection candidate windows of the previous layer and the information of the candidate windows on the image can be inherited.

[0116] Better, such as Figure 8 As shown, in the specific execution process of this embodiment of the invention, after each layer of the image pyramid is scaled through a sliding candidate window, multiple candidate windows are generated for each layer. Each candidate window of each layer of the scaled image has a confidence level. Each layer of the scaled image is input into the first cascaded network, and the overlap of the confidence levels of the candidate windows of each layer is calculated. The candidate windows of the layer are sorted by confidence level, and the overlap of the confidence level of the candidate window with the highest confidence level with the confidence level of the other candidate windows is calculated. If the overlap of the confidence level of the candidate window with the highest confidence level with the other candidate windows is greater than a preset threshold, it is removed. If the overlap of the confidence level of the other candidate windows is less than the preset threshold, face candidate window images of multiple candidate windows are obtained.

[0117] like Figure 9As shown, the study found that although a 12×12 candidate window was used for step-size sliding detection on the scaled image of this layer, most of the face bounding boxes predicted by the P-Net network were rectangular candidate windows. When passed to the R-Net network, the internal image of the candidate window needed to be scaled to 24x24 (note that converting the 12×12 image to a 24x24 image can also be understood as scaling). Further research revealed that before inputting the face candidate window image into the second cascaded network, each candidate window of the selected face candidate window image needed to be scaled to avoid distortion during image scaling. The following processing method was adopted for this purpose.

[0118] Preferably, before inputting the face candidate window image into the second cascaded network, the method further includes filtering target candidate windows in the face candidate window image that need to undergo square transformation of the image candidate window;

[0119] S31. Determine the position and size of the target candidate window that needs to be converted into a square; stretch the two longer sides of the candidate window outwards simultaneously so that the lengths of the two shorter sides and the two longer sides form a square of equal length; crop the candidate window that has been stretched into a square, and scale the internal image of the cropped candidate window to 24×24.

[0120] Repeat the above steps: convert multiple candidate windows of the face candidate window image into square candidate windows until the internal images of all candidate windows have been converted into 24×24 square candidate windows;

[0121] It should be noted that since most of the face candidate windows predicted by the first-level network are rectangular boxes, the images of the face candidate windows need to be scaled to 24x24 when they are transmitted to the second-level network. If the images of the rectangular candidate windows are scaled directly, the faces will be distorted, resulting in poor performance. If the images of the rectangular candidate windows are cropped, converted to squares, and then scaled, black borders will appear at the edges of the images. The final approach is to convert the rectangular boxes to squares according to the maximum side length, then crop the images from the original images according to the square windows, scale them to 24x24, and then pass them into the second-level network. This improves the efficiency of the second-level network in further eliminating overlapping candidate windows.

[0122] The above process involves converting the rectangular candidate window into a square according to the maximum side length, then cropping the image from the original image into a square candidate window and scaling it to 24x24 before feeding it into the second cascaded network. This improves the efficiency of the second cascaded network in further eliminating overlapping candidate windows. Therefore, the above processing operation also has the technical effect of improving the computational efficiency of multi-task cascaded convolutional networks.

[0123] S32. Convert multiple square candidate windows into 24×24 square face candidate window images and input them into the second cascade network. Sort the square candidate windows by confidence level, and calculate the overlap between the square candidate window with the highest confidence level and the remaining square candidate windows with lower confidence levels. If the overlap between the square candidate window with the highest confidence level and the square candidate window with lower confidence levels is greater than the calculated overlap threshold, then the square candidate window with lower confidence level is removed. Calculate the overlap between the remaining square candidate windows with lower confidence levels until all square candidate windows with overlap levels greater than the calculated overlap threshold are removed. Obtain multiple target face candidate window images with overlap levels less than the calculated overlap threshold. Input the square candidate windows with overlap levels less than the calculated overlap threshold into the third cascade network for subsequent operations.

[0124] After the square candidate box is output through the overlap calculation of the second cascade network, the overlap is calculated again by the third cascade network, and facial key point feature identification is performed simultaneously. When obtaining the square candidate window through the output of the second cascade network, some rectangular candidate windows will also appear. These rectangular candidate windows need to be scaled according to the rectangular candidate window in S31 above to form a 48×48 square candidate window. Therefore, the square candidate window also includes the following image processing method before being input into the third cascade network:

[0125] S33. Before inputting the square candidate windows with an overlap degree less than the calculated overlap degree threshold into the third cascaded network, the rectangular candidate windows output by the second cascaded network are scaled into square candidate windows of size 48×48; the two longer sides of the candidate window are stretched outwards simultaneously so that the lengths of the two shorter sides and the two longer sides form a square of equal length; the stretched square candidate windows are cropped, and the internal image of the cropped candidate windows is scaled into 48×48; all candidate windows are converted into 48×48 square candidate windows.

[0126] S34. Convert the candidate window into a 48×48 square window and calculate the overlap of the confidence level to obtain the final target detection window.

[0127] The third cascaded network processes and outputs the final image containing the face, specifically including: the third cascaded network outputs the final candidate box (predicted box) through the target detection window of the current face; determines the coordinates of the final candidate box, and obtains the final image containing the face.

[0128] It should be noted that the labels of the final output image containing the face obtained above need to be verified (the output results include the labels and the confidence and offset of the final candidate box, etc., which will not be elaborated here). If the pass rate meets the standard, it proves that the above model has fully converged and meets the standard, which will not be elaborated here.

[0129] The aforementioned MTCNN is a cascaded network model comprising three networks: P-Net, R-Net, and O-Net. Researchers found that the computational results of these three networks progressively increase in accuracy, with the latter two networks (R-Net and O-Net) further refining the results of the preceding network. When generating the dataset, previous networks are invoked for inference, and their inference results are used to generate the dataset for the next network. While the network training process is complex, the embodiments described in this application significantly reduce computational load and improve operational efficiency.

[0130] The O-Net network, a third-level concatenated network, calculates the overlap of candidate windows and then determines the final target detection window. The O-Net network consists of a 6-layer network structure (the first 4 layers are convolutional layers, and the last 2 layers are fully connected layers). It can not only filter the overlap of candidate windows but also predict the position of the final facial key points. During the training phase, the predicted facial landmark coordinates are compared with the real facial landmarks using Euclidean distance. The regression loss function (landmark position loss) calculated using Euclidean distance is used to predict and determine the position of the key points of the face (i.e., the positions of 5 facial features, including the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth).

[0131] In step S31, the two longer sides of the candidate window are stretched outwards simultaneously to avoid slight changes in the candidate window during stretching, which could cause target shift in the image inside the candidate window. After stretching the candidate window into a square, the square candidate window is cropped, and the internal image of the cropped candidate window is scaled 24×24 to ensure that all candidate windows have the same size. This eliminates the detection result deviation caused by size differences, ensures that the algorithm has the same detection capability for face targets of different sizes, and also reduces the detection speed affected by size differences.

[0132] In step S32, the converted multiple 24×24 square candidate windows are input into the second cascaded network to calculate the overlap of the confidence scores of the square candidate windows. By calculating the overlap of the confidence scores, the square candidate windows are filtered to reduce redundant target candidate windows and improve the speed of subsequent detection.

[0133] In step S33, the filtered square candidate windows and some rectangular candidate windows generated when the second cascaded network is output are transformed into 48×48 candidate windows. The larger size helps to preserve key details and texture information in the image, thereby better capturing the features of the face target. This can improve the effect of subsequent feature extraction algorithms and enhance the representation ability of the face target.

[0134] In step S33, the candidate windows are converted into 48×48 squares and input into the third cascade network for the final calculation of the confidence overlap. All the redundant square candidate windows are then filtered out to obtain the final target detection window and the face image is obtained through the target detection window.

[0135] Example 2

[0136] This second embodiment provides an electronic device, including: a memory, a processor, and a face detection program of a task-cascaded convolutional network stored in the memory and running on the processor. When the face detection program of the task-cascaded convolutional network is executed by the processor, it implements the steps of the face recognition processing method based on a multi-task cascaded convolutional network described in the first embodiment.

[0137] Example 3

[0138] like Figure 10 As shown, on the other hand, this third embodiment, based on the face recognition processing method based on a multi-task cascaded convolutional network provided in the first embodiment of the invention, also provides a computer storage medium 1140 (hereinafter referred to as the storage medium). A schematic diagram of the computer storage medium structure framework provided in the third embodiment of the invention is shown, which includes:

[0139] Memory 1130 is used to store computer programs;

[0140] The communication interface 1120 is used to connect the memory 1130 to the processor 1110;

[0141] Processor 1110 is configured to execute a computer program to implement a face recognition processing method based on a multi-task cascaded convolutional network, as disclosed in any of the above-described embodiments.

[0142] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits, digital signal processors, digital signal processing devices, programmable logic devices, field-programmable gate arrays, general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.

[0143] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0144] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0145] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; those skilled in the art can modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A face recognition processing method based on a multi-task cascaded convolutional network, characterized in that, The following steps are included: A multi-task cascaded convolutional network is constructed, which is divided into three convolutional neural networks. The first cascaded network is P-Net, the second cascaded network and the third cascaded network are R-Net and O-Net, respectively. A single face image sample set is collected to train the multi-task cascaded convolutional network to obtain the converged multi-task cascaded convolutional network. Acquire a face image to be detected, preprocess the face image to be detected, and input the preprocessed face image to be detected into the first cascaded network to obtain an image containing face candidate windows; The candidate image containing the face is filtered by the confidence level of the second cascaded network; the filtered candidate image is then input into the third cascaded network, which processes it and outputs the final image containing the face. Acquire a face image to be detected, preprocess the face image to be detected, and input the preprocessed face image to be detected into the first cascaded network to obtain an image containing face candidate windows, specifically including: The face image to be detected is preprocessed to obtain an image pyramid of the face image to be detected; the image pyramid includes multiple layers of preprocessed face images; the preprocessing refers to scaling the current face image to be detected multiple times, with each scaling yielding one layer of face image; When performing sliding detection on the face image of each scaled image layer, the step size of the associated adjacent layers is adjusted by varying the step size to obtain an image containing face candidate windows, specifically including: During initialization, the scaling ratio of the current layer's scaled image relative to the original image is obtained; then, a single target face is detected on the scaled image of the current layer, and a standard step size is randomly set within the current layer. On the scaled image of the current layer, slide detection of the first face target is performed according to the standard step size, and multiple candidate windows for the current first face target are output; single target detection is performed on the current first face target to obtain multiple candidate windows for the current first face target. The process involves determining the confidence level of multiple candidate windows for the first face in the scaled image of the current layer; identifying windows with a confidence level greater than a standard confidence threshold as valid candidate windows; determining whether the number of valid candidate windows for the first face is greater than a candidate window count limit; if the number of valid candidate windows is greater than the candidate window count limit, it is considered that there are still many candidate windows for the selected first face, so the standard step size of the current layer is still small. Therefore, when detecting the second face in the scaled image of the current layer, the preset step size of the current layer is increased to achieve the detection of the second face. The above steps are repeated to determine whether the number of valid candidate windows for the second face is greater than the candidate window count limit. If the number of valid candidate windows is still greater than the candidate window count limit, then the detection of the third face is initiated. The step size is increased until multiple face targets in the current layer are detected, or until the number of valid candidate windows for the nth face target in the current layer is less than the candidate window limit threshold. Then the standard step size is no longer increased. The final standard step size is determined as the target standard step size of the scaled image in the current layer. When performing sliding detection on the scaled image in the next layer, the relationship between the scaled image in the next layer and the scaled image in the current layer is first determined. If the scaled image in the next layer is a further scaled image of the scaled image in the current layer, the target standard step size of the scaled image in the previous layer is directly used as the standard step size for sliding detection of the (1-m)th face targets on the scaled image in this layer. The above steps are repeated until it is determined whether adjusting the step size is necessary for sliding detection when performing sliding detection on the (1-m)th face targets in this layer using the target standard step size.

2. The face recognition processing method based on a multi-task cascaded convolutional network according to claim 1, characterized in that, The multi-task cascaded convolutional network is trained by acquiring a single-face image sample set to obtain a converged multi-task cascaded convolutional network. Specifically, the process includes: For the single face image sample set, feature descriptors for face classification, candidate window regression, and face landmark localization are established respectively. Based on the feature descriptors for face classification, a cross-entropy loss function for face classification is constructed; based on the feature descriptors for candidate window regression, a regression loss function for candidate windows is constructed; and based on the feature descriptors for face landmark localization, a regression loss function for landmark locations is constructed. The target total loss function of the multi-task concatenated convolutional network is constructed based on the cross-entropy loss function of face classification, the regression loss function of candidate windows, and the regression loss function of landmark location; when determining the target total loss function, the converged final multi-task concatenated convolutional network is obtained by solving.

3. The face recognition processing method based on a multi-task cascaded convolutional network according to claim 2, characterized in that, For the single-face image sample set, feature descriptors for face classification, candidate window regression, and face landmark localization are established respectively. Based on the face classification feature descriptors, a cross-entropy loss function for face classification is constructed; based on the candidate window regression feature descriptors, a candidate window regression loss function is constructed; and based on the face landmark localization feature descriptors, a landmark location regression loss function is constructed. Specifically, this includes: Classify facial features on a single-face image sample set, using the cross-entropy loss function for face classification: ; In the formula, pi represents the probability that the current face belongs to a human. Real labels as background; Candidate windows are selected for prediction of facial features in the single-face image sample set. There is an offset between the predicted candidate windows and the actual candidate windows. The regression loss function for the candidate windows is calculated using Euclidean distance. ; In the formula, The position of the candidate window is predicted by a multi-task cascaded convolutional network. The position of the actual candidate window; The Euclidean distance is calculated between the currently predicted facial landmark coordinates and the actual facial landmarks, and the regression loss function for the landmark location is calculated using the Euclidean distance. ; In the formula, The predicted facial landmark coordinates are obtained from a multi-task cascaded convolutional network, while These are the actual landmark coordinates of the face.

4. The face recognition processing method based on a multi-task cascaded convolutional network according to claim 3, characterized in that, Based on the cross-entropy loss function for face classification, the regression loss function for candidate windows, and the regression loss function for landmark locations, a target total loss function for a multi-task cascaded convolutional network is constructed. When determining the target total loss function, the converged final multi-task cascaded convolutional network is obtained by solving for it, specifically including: Based on the aforementioned calculations of the cross-entropy loss function for face classification, the regression loss function for the candidate window, and the regression loss function for the landmark location, the target total loss function is calculated, and its minimum value is obtained. ; ; In the formula, N is the number of training samples, αj represents the weights of the loss function for different network structures, βj is the sample label, and Lj is the loss function; When determining the minimum value of the target total loss function, the optimized multi-task cascaded convolutional network parameters corresponding to the minimum value of the target total loss function are obtained; the multi-task cascaded convolutional network is trained based on the solved optimized multi-task cascaded convolutional network parameters to obtain the final multi-task cascaded convolutional network; that is, until the training network parameters are stable, the final multi-task cascaded convolutional network is saved after training is completed.

5. The face recognition processing method based on a multi-task cascaded convolutional network according to claim 4, characterized in that, To acquire the face image to be detected, before inputting the face candidate window image into the second cascaded network, the method further includes filtering the target candidate windows in the face candidate window image that need to undergo square transformation of the image candidate window; Determine the position and size of the target candidate window that needs to be converted into a square; stretch the two longer sides of the candidate window outwards simultaneously so that the two shorter sides form a square with equal lengths to the two longer sides; crop the candidate window that has been stretched into a square, and scale the internal image of the cropped candidate window to 24×24. Repeat the above steps: convert multiple candidate windows of the face candidate window image into square candidate windows until the internal images of all candidate windows have been converted into 24×24 square candidate windows; Multiple square candidate windows are converted into 24×24 square face candidate window images and input into the second cascade network. The square candidate windows are sorted by confidence, and the overlap between the square candidate window with the highest confidence and the remaining square candidate windows with lower confidence is calculated. If the overlap between the highest confidence square candidate window and the low confidence square candidate window is greater than the calculated overlap threshold, then the low confidence square candidate window is removed; the overlap of the remaining low confidence square candidate windows is calculated until the overlap of all removed square candidate windows is greater than the calculated overlap threshold. Multiple target face candidate window images with an overlap degree less than the calculated overlap degree threshold are obtained. The square candidate windows with an overlap degree less than the calculated overlap degree threshold are input into the third cascade network for subsequent operations.

6. The face recognition processing method based on a multi-task cascaded convolutional network according to claim 5, characterized in that, Before the square candidate windows with an overlap less than the calculated overlap threshold are input into the third cascade network, the following steps are also included: The rectangular candidate window output by the second cascaded network is scaled to a square candidate window of size 48×48; the two longer sides of the candidate window are stretched outwards simultaneously so that the two shorter sides form a square with equal lengths to the two longer sides; the stretched square candidate window is cropped, and the internal image of the cropped candidate window is scaled to 48×48; all candidate windows are converted into 48×48 square candidate windows. The candidate window is converted into a 48×48 square window, and the overlap of the confidence score is calculated to obtain the final target detection window.

7. The face recognition processing method based on a multi-task cascaded convolutional network according to claim 6, characterized in that, The third cascaded network processes and outputs the final image containing the face, specifically including: The third cascaded network outputs the final candidate box through the target detection window of the current face; determines the coordinates of the final candidate box, and obtains the final image containing the face.

8. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the face recognition processing method based on a multi-task cascaded convolutional network as described in any one of claims 1-7.