Model branch system for improving robustness of face recognition model and training method thereof

By introducing a model branch system at the feature space level into the face recognition model, and using the residual compensation branch and feature pyramid model training module, the problem of insufficient robustness of the existing face recognition model in complex scenarios is solved, and the robustness of the model and the simplification of the optimization process is achieved.

CN120220199APending Publication Date: 2025-06-27ZHEJIANG SUNNY INTELLIGENT OPTICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311811004.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing face recognition models have insufficient robustness when dealing with low-quality blurred faces, large-angle faces, and cross-RGB-IR modal recognition, and existing solutions usually require retraining the model, which is cumbersome.

Method used

A model branch system based on feature space is proposed, and the robustness of the face recognition model is improved by introducing residual compensation branches and feature pyramid model training modules. The system can train a small branch model separately without retraining the face feature model to correct fuzzy, angle and cross-modal problems.

Benefits of technology

It has achieved robustness improvements in the face recognition model in a variety of complex scenarios, including low-quality blurred faces, large-angle faces and cross-RGB-IR modal recognition, and there is no need to retrain the face feature model, simplifying the model optimization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220199A_ABST
    Figure CN120220199A_ABST
Patent Text Reader

Abstract

The invention provides a model branch system for improving robustness of a face recognition model and a training method thereof, the model branch system can be connected behind a feature layer of a trained face feature model, and the input of the model branch system is a feature output by the trained face feature model. After the target is corrected by the model branch system, features after target correction are output, the model branch system is a compensation branch based on residual errors, and the model branch system comprises n full connection layers connected front and back, ReLU and a residual error module, and the residual module compensates and corrects the original features in combination with a calculation result of the data evaluation scale and an output result activated by the full connection layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition, and in particular to a model branch system for improving the robustness of a face recognition model and a training method thereof. Background Art

[0002] With the emergence and rise of deep learning, face recognition technology has also flourished and been widely applied in various fields.

[0003] Today, face recognition models with good performance already have good robustness for the recognition of high-quality face data of the same modality, such as clear and small-angle RGB face images. However, in the actual application process, the following pain points still exist: for example, due to reasons such as the need to reduce the cost of imaging equipment, long recognition distance, or motion blur, the quality of the input image of the network is low and blurred, resulting in recognition failure; for example, cross RGB-IR modality recognition between human-certificate RGB images and infrared camera IR images in the field of security access control is also a challenge; the recognition accuracy of the model for faces at large angles is not satisfactory, etc.

[0004] For the above several problems, existing solutions often optimize and solve them one by one in different directions. For example:

[0005] 1) To improve the recognition robustness for low-quality blurred faces. By adding image noise, image downsampling and then upsampling and other schemes to generate low-quality blurred data, but the actual effect still cannot align with the actual image blur effect.

[0006] 2) To improve the cross RGB-IR modality recognition performance. It is often necessary to render the corresponding IR dataset based on the existing public RGB face dataset through the existing data synthesis algorithm, which requires a long time and a large amount of computing resources.

[0007] 3) To improve the recognition rate for faces at large angles. The corresponding frontal face image can be generated by straightening through the GAN network, but the details of the generated image are poor. Without the supervision of an extremely accurate identification model, it cannot be guaranteed that the generated face image belongs to the same person's sample.

[0008] In addition, for new data input at the image level, the model needs to be retrained, and it cannot be based on the current existing model. Retraining and fitting the model to select the best is a very long process. Summary of the Invention

[0009] A main advantage of the present invention is to provide a model branch system for improving the robustness of a face recognition model and a training method thereof, wherein the model branch system can solve the problems of low-quality blurred faces, large-angle faces, and difficult cross RGB-IR modality recognition in the current face recognition field.

[0010] Another advantage of the present invention is that it provides a model branching system and a training method for improving the robustness of face recognition models, wherein the model branching system optimizes problems such as blur, large angle, and cross-modality in face recognition based on the feature space level rather than the image level, and there is no need to retrain the face feature model.

[0011] Another advantage of the present invention is that it provides a model branching system and a training method thereof for improving the robustness of a face recognition model, wherein the model branching system is based on a well-performing model that has been trained and only requires a small branch to be trained separately.

[0012] Another advantage of the present invention is that it provides a model branch system and a training method for improving the robustness of face recognition models, wherein the model branch system is small, lightweight, simple and effective, and can be added behind the feature output layer of any model.

[0013] Another advantage of the present invention is that it provides a model branching system and a training method for improving the robustness of face recognition models, wherein the model branching system has a certain universality in improving the robustness of face recognition models in various dimensions such as angle, blur, different modalities or brightness and darkness.

[0014] Another advantage of the present invention is that it provides a model branching system and a training method for improving the robustness of face recognition models, wherein the model branching system introduces a feature pyramid, introduces the feature map difference information of the to-be-corrected class data and the target class data at different stages into the loss function of the model training, and brings the features of the two at different stages closer, so that the branch training is easier to converge and the effect is better.

[0015] According to one aspect of the present application, the present application provides a model branch system for improving the robustness of a face recognition model, which can be connected to the feature layer of a trained face feature model, wherein the input of the model branch system is the feature output by the trained face feature model, and after the model branch system corrects the target, the output is the feature after the target correction, wherein the model branch system is a compensation branch based on residuals, wherein the model branch system includes n fully connected layers and ReLU connected forward and backward, and a residual module, and the residual module combines the calculation results of the data evaluation scale and the output results after activation of the fully connected layer to compensate and correct the original features.

[0016] According to one embodiment of the present application, the target correction is a face correction item, including clarifying a blurred face, straightening a face at a large angle, and cross-modal conversion.

[0017] According to an embodiment of the present application, the model branch system is obtained by selecting relevant face data for training according to a correction target. For example, if the correction target is the blurred face recognition rate, it is obtained by training with clear and blurred face pairs; if the correction target is the large-angle face recognition rate, it is trained with large-angle and frontal face data pairs; if the RGB-IR cross-modal recognition performance is to be improved, it is trained with RGB and IR face data pairs.

[0018] According to an embodiment of the present application, if it is branch training for improving blurred face recognition, the training data of the model branch system is implemented as follows: for blurred branch training, each sample ID includes blurred face image data and clear face data.

[0019] According to an embodiment of the present application, if it is for large pitch angle branch training, each sample ID includes large pitch angle and small-angle frontal face data.

[0020] According to an embodiment of the present application, for a certain face image P i , let the original output feature of the model be f i , and the calculation result of the image evaluation scale be m i The sigmoid function is S(x), and the output result after activation by the fully connected layer is W(f i ), and the finally corrected output feature is

[0021]

[0022] Based on and the category label, through softmax or other commonly used face feature loss functions for output, combined with cross-entropy to calculate Loss, and backpropagation is used to train and update the branch parameters.

[0023] According to an embodiment of the present application, it further includes a correction scale control module for controlling the correction size.

[0024] According to an embodiment of the present application, the correction scale control module includes a data evaluation scale calculation unit. For the blurred training set, the gray variance or gradient sum is used to describe the blur degree of the image. For the angle training set, based on the face key point information, the corresponding face angle value is directly calculated and used as the evaluation scale θ.

[0025] According to an embodiment of the present application, the correction scale control module includes a correction scale control unit, and the correction scale control unit uses the sigmoid function to control the correction scale.

[0026]

[0027] A is used to control the amplitude, a controls the threshold at which the correction starts to take effect, and k controls the average rate of change. If you want the correction effect to change sharply in a certain range, k takes a large value; if you want a slow change, take a smaller k value.

[0028] According to an embodiment of the present application, a feature pyramid model training module is further included, which introduces feature map difference information of the to-be-corrected class data and the target class data at different stages into the loss function of the model training, and brings the features of the two at different stages closer.

[0029] According to one embodiment of the present application, the feature pyramid model training module includes: building a feature pyramid, each feature_map represents a level of the pyramid, generally selecting the feature_map of the previous layer before the feature resolution size changes; adjusting the feature_map of each level of the pyramid to have the same number of channels; starting from the feature_map with the smallest resolution at the bottom layer, up-sampling to the feature_map of the previous layer, and adding it to the feature_map of the previous layer to perform feature fusion.

[0030] According to an embodiment of the present application, the feature pyramid of all the to-be-corrected class data and the target class data in each batch is calculated and extracted, and the feature pyramid is recorded. and They respectively represent the kth face data of the target class data of the i-th ID and the jth data of the class data to be corrected.

[0031] According to another aspect of the present application, the present application further provides a training method for improving the robustness of a face recognition model, wherein the training method comprises the following steps:

[0032] Designing a model branch system based on the residual, wherein the model branch system is connected after the feature layer of the trained face feature model; and

[0033] The features output by the trained facial feature model are input into the model branch system, which processes the features and outputs the features after a certain target correction, wherein the target correction is a face correction item, including clarifying a blurred face, straightening a face at a large angle, and cross-modal conversion.

[0034] According to one embodiment of the present application, the model branch system selects relevant facial data for training according to the correction target. If the correction target is the blurred face recognition rate, it is obtained by training with clear and blurred face pairs; if the correction target is the wide-angle face recognition rate, it is trained with wide-angle and frontal face data pairs; if the correction target is to improve the RGB-IR cross-modal recognition performance, it is trained with RGB and IR face data pairs.

[0035] According to an embodiment of the present application, the model branch system is a residual-based compensation branch, wherein the model branch system includes n fully connected layers connected in series and ReLU.

[0036] According to an embodiment of the present application, assuming that the branch training is for improving the recognition of blurred faces and large-angle faces, the training data of the model branch system is as follows: If the blurred branch training is performed, each sample ID includes blurred face image data and clear face data, preferably n pairs of blurred and clear face data pairs, where the blurred face corresponds to one clear face; if the angle branch training is performed, each sample ID includes large pitch angle and frontal small angle face data.

[0037] According to an embodiment of the present application, a branch structure is added after the feature layer of the trained face feature model, the parameters of the face feature model are frozen, and only the branch is trained; or the features corresponding to the training set are inferred and saved in advance based on the trained face feature model, and the features are directly read later to train the branch.

[0038] According to an embodiment of the present application, for a certain face image P i , let the original output feature of the model be f i , and the calculation result of the image evaluation scale be m i The sigmoid function is S(x), and the output result after activation by the fully connected layer is W(f i ), and the finally corrected output feature is

[0039]

[0040] Based on and the category label, through softmax, or other commonly used face feature loss functions for output, the Loss is calculated by combining cross entropy, and the branch parameters are trained and updated by backpropagation.

[0041] According to an embodiment of the present application, it further includes controlling the correction scale, where controlling the correction scale includes data evaluation size calculation and correction scale control. For the blurred training set, the gray variance or gradient sum is used to describe the blur degree of the image. For the angle training set, based on the face key point information, the corresponding face angle value is directly calculated and obtained as the evaluation scale θ.

[0042] According to an embodiment of the present application, the sigmoid function is used to control the correction scale as follows:

[0043]

[0044] A is used to control the amplitude, a controls the threshold at which the correction starts to take effect, and k controls the average rate of change. If you want the correction effect to change sharply in a certain range, k takes a large value; if you want a slow change, take a smaller k value.

[0045] According to one embodiment of the present application, it further includes feature pyramid model training. Through the feature pyramid model training, the feature map difference information of the to-be-corrected class data and the target class data at different stages is introduced into the loss function of the model training, so as to bring the features of the two at different stages closer.

[0046] According to one embodiment of the present application, feature pyramid model training includes:

[0047] 1) Build a feature pyramid. Each feature_map represents a level of the pyramid. Generally, the feature_map of the previous layer before the feature resolution size changes is selected;

[0048] 2) Adjust the feature_map of each level of the pyramid to have the same number of channels;

[0049] 3) Starting from the feature_map with the smallest resolution at the bottom layer, upsample to the feature_map of the previous layer and add it to the feature_map of the previous layer for feature fusion.

[0050] According to an embodiment of the present application, the feature pyramid of all the to-be-corrected class data and the target class data in each batch is calculated and extracted, and the feature pyramid is recorded. and They respectively represent the kth face data of the target class data of the i-th ID and the jth data of the class data to be corrected.

[0051] According to one embodiment of the present application, during training, in each mini-batch, n IDs are sampled, and each ID has p and q pieces of target dimension class data and to-be-corrected dimension class data, and the average features of the target dimension and to-be-corrected dimension of the i-th ID are calculated as follows:

[0052]

[0053]

[0054] Finally, the LOSSmse (mean square error loss) corresponding to each mini-batch is obtained by the following formula:

[0055]

[0056] Further objects and advantages of the present invention will be fully apparent from an understanding of the following description and the accompanying drawings.

[0057] These and other objects, features, and advantages of the present invention will be fully embodied by the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The technical solution of the present invention will be further described in detail below in conjunction with the drawings and embodiments. In the drawings, unless otherwise specified, the same reference numerals are used to represent the same components. Among them:

[0059] Figure 1 is a schematic diagram of a model branch system for enhancing the robustness of a face recognition model according to a first preferred embodiment of the present invention.

[0060] Figure 2 is a schematic diagram of a model branch system for enhancing the robustness of a face recognition model according to a first preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] It should be noted that the embodiments shown in the drawings are only examples for specifically and vividly explaining and illustrating the concept of the present invention. In terms of size and structure, they are neither necessarily drawn to scale nor constitute a limitation to the concept of the present invention.

[0062] Those skilled in the art should understand that the term "one" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of this element can be multiple. The term "one" cannot be understood as a limitation on the number.

[0063] Referring to FIGS. Figure 1 and Figure 2 shown, a model branch system for enhancing the robustness of a face recognition model and its training method according to a first preferred embodiment of the present application will be elucidated in the following description. The model branch system for enhancing the robustness of the face recognition model is a model branch based on the residual. For the sake of convenience of description, it is hereinafter simply referred to as the model branch system. The model branch system can be connected after the feature layer of the trained face feature model. The input of the model branch system is the feature output by the trained face feature model. After the target is corrected by the model branch system, the output is the feature after the target correction. The target correction of the model branch system is a face correction term, including clarifying the correction of blurred faces, straightening large-angle faces, and cross-modal conversion, such as performing RGB and IR modal conversion, that is, RGB and IR face cross-modal style transfer. Those skilled in the art can understand that the cross-modal conversion can also be implemented as other types of cross-modal conversions, and no limitation is imposed thereon.

[0064] The model branch system is obtained by selecting relevant face data for training according to the correction target. For example, if the correction target is the recognition rate of blurred faces, it is obtained by training with clear and blurred face pairs; if the correction target is the recognition rate of large-angle faces, it is trained with large-angle and frontal face data pairs; if the performance of RGB-IR cross-modal recognition is to be improved, it is trained with RGB and IR face data pairs.

[0065] It can be understood that the model branch system of the preferred embodiment of the present application optimizes problems such as blur, large angle, and cross-modal in face recognition based on the feature space level rather than the image level, and does not require retraining the face feature model. Based on the well-trained model with good performance, only a small branch needs to be trained separately, which can simplify the model. The model branch system of the preferred embodiment of the present application is very small, lightweight, simple and effective compared with the face recognition model, and can be added behind the feature output layer of any model. In addition, the model branch system of the preferred embodiment of the present application has a certain universality for improving the robustness of the face recognition model in various dimensions such as angle, blur, different modalities, or brightness and darkness.

[0066] In the present application, the pre-trained face feature model can be, but is not limited to, a model of any network such as Resnet or MolileFaceNet. The model branch system is a compensation branch based on residuals, where the model branch system includes n fully connected layers connected in sequence and ReLU (Rectified Linear Unit, also known as the rectified linear unit), and its actual application quantity can be adjusted according to the training effect, and generally 2 or 3 are used. The model branch system further includes a residual module, which combines the calculation result of the data evaluation scale and the output result after activation by the fully connected layer to compensate and correct the original features.

[0067] Specifically, in the preferred embodiment of the present application, assuming it is the branch training for improving the recognition of blurred faces and large-angle faces, the training data of the model branch system is implemented as follows: for the blurred branch training, each sample ID includes blurred face image data and clear face data. Preferably, there are n pairs of blurred and clear face data pairs, with one blurred face corresponding to one clear face. For the angle branch training, the head pose of the face includes three angles: pitch angle, yaw angle, and roll angle. Since the roll angle does not need to be considered for face alignment, the large pitch angle and yaw angle need to be solved by two branches. Therefore, when preparing the data, it also includes according to the corresponding angle dimension. For example, for the large pitch angle branch training, each sample ID includes large pitch angle and frontal small angle face data. Preferably, there are n pairs of large pitch angle and frontal small angle face data pairs.

[0068] In this preferred embodiment of the present application, after adding a branch structure to the feature layer of the trained face feature model, the parameters of the face feature model are frozen, and only the branch is trained. It is also possible to infer and save the features of the corresponding training set in advance based on the trained face feature model, and then directly read the features to train the branch.

[0069] Taking one inference forward pass as an example, for a certain face image P i , let the original output feature of the model be f i , and the calculation result of the image evaluation scale be m i The sigmoid function is S(x), and the output result after activation by the fully connected layer is W(f i ), and the finally corrected output feature is

[0070]

[0071] Based on and the category label, through softmax, or other commonly used face feature loss functions for output, such as ArcFace, CosFace, etc., combined with cross-entropy to calculate Loss, and backpropagation is used to train and update the branch parameters.

[0072] The model branch system further includes a correction scale control module for controlling the correction size. The correction scale control module includes a data evaluation scale calculation unit and a correction scale control unit. Among them, for the fuzzy training set, the data evaluation scale calculation unit uses an evaluation parameter to describe the blurring degree of the image, and can use the gray variance, gradient sum, etc. Taking the gray variance s as an example,

[0073]

[0074]

[0075] where, w x , h y respectively represent the size of the face image, the network input is 112*112, and f(x,y) represents the corresponding pixel gray value.

[0076] For the angle training set, the data evaluation scale calculation unit can directly calculate and obtain the corresponding face angle value as the evaluation scale θ based on the face key point information.

[0077] It should be noted that for the RGB and IR cross-modal training sets, since there is no intuitively quantifiable evaluation scale for the cross-modal dimension, the scale evaluation and correction control module cannot take effect at this time, but the training branch can still obtain a certain improvement effect.

[0078] Since the original model has strong robustness for clear images, slightly blurred faces, frontal faces, and small-angle faces, it is necessary to control the correction scale. When the face is slightly blurred or the face angle is small, the original features are slightly corrected or not corrected. When it is greater than a certain preset threshold, the correction effect is increased. The correction scale control unit of the model branch system in the preferred embodiment of the present application uses the sigmoid function to control the correction scale as follows:

[0079]

[0080] A is used to control the amplitude, a controls the threshold at which the correction starts to take effect, and k can control the average change rate. If it is desired that the correction effect changes steeply in a certain interval, k takes a large value; if it is desired to change slowly, a smaller k value is taken.

[0081] The model branch system further includes a feature pyramid model training module to make the model or the model branch system converge more easily. Through the feature pyramid model training module, the difference information of the feature maps at different stages of the data to be corrected and the target data is introduced into the loss function of the model training, and the features of the two at different stages are brought closer.

[0082] Specifically, in the preferred embodiment of the present application, the feature pyramid model training module includes: building a feature pyramid, where each feature_map represents a level of the pyramid. Generally, the feature_map of the previous layer before the change in the feature resolution size is selected; adjusting the feature_maps of each level of the pyramid to the same number of channels; starting from the feature_map with the smallest resolution at the bottom layer, upsampling to the previous layer of the feature_map and adding it to the previous layer of the feature_map for feature fusion. Repeat this process until each face image obtains a feature_map that fuses multiple scales.

[0083] Furthermore, calculate and extract the feature pyramids of all the data to be corrected and the target data in each batch, denoted as and respectively represent the k-th face data of the target data of the i-th ID and the j-th data of the data to be corrected.

[0084] It can be understood that during the training process, in each mini-batch, n IDs are sampled, and for each ID, there are p and q pieces of data in the target dimension class and the data to be corrected dimension class respectively. The average features of the target dimension and the data to be corrected dimension of the i-th ID are calculated as follows:

[0085]

[0086]

[0087] Finally, the loss mse (mean squared error loss) corresponding to each mini-batch is obtained by the following formula:

[0088]

[0089] It is worth mentioning that, different from the function of enabling targets of different sizes to have appropriate feature representations at corresponding scales in object detection algorithms, the model branch system in the preferred embodiment of the present application introduces the difference information of feature maps at different stages of the data to be corrected and the target class data into the loss function of model training through the feature pyramid model training module, bringing the features of the two at different stages closer. The branch training is more likely to converge and has better effects.

[0090] According to another aspect of the present application, the present application further provides a training method for improving the robustness of a face recognition model, where the training method includes the following steps:

[0091] Design a model branch system based on residuals, and the model branch system is connected after the feature layer of the trained face feature model; and

[0092] Input the features output by the trained face feature model into the model branch system, and the model branch system processes the features and outputs the features after a certain target correction, where the target correction is a correction item for the face, including clarifying the blurred face, straightening the large-angle face, and cross-modal conversion. As an example, in a specific example of the present application, the cross-modal conversion may include but is not limited to the modal conversion between RGB and IR, that is, the cross-modal style transfer between RGB and IR faces.

[0093] In the training method of the preferred embodiment of the present application, the model branch system is obtained by selecting relevant face data for training according to the correction target. For example, if the correction target is the recognition rate of blurred faces, it is obtained by training with clear and blurred face pairs; if the correction target is the recognition rate of large-angle faces, it is trained with large-angle and frontal face data pairs; if the cross-modal recognition performance between RGB and IR is improved, it is trained with RGB and IR face data pairs.

[0094] In the training method of this preferred embodiment of the present application, the trained face feature model can be, but is not limited to, a model of any network such as Resnet, MobileFaceNet, etc. Among them, the model branch system is a compensation branch based on residuals. The model branch system includes n fully connected layers connected in sequence and ReLU, and the actual number used can be adjusted according to the training effect, generally 2 or 3 are used. The model branch system further includes a residual module, and the residual module combines the calculation result of the data evaluation scale and the output result after activation by the fully connected layer to compensate and correct the original features.

[0095] Specifically, in the training method of this preferred embodiment of the present application, assuming that it is branch training for improving the recognition of blurred faces and large-angle faces, the training data of the model branch system is as follows: If it is blurred branch training, each sample ID includes blurred face image data and clear face data, preferably n pairs of blurred and clear face data pairs, where one blurred face corresponds to one clear face; If it is angle branch training, the head pose of the face includes three angles: pitch angle, yaw angle, and roll angle. Since the roll angle does not need to be considered for face alignment, the large pitch angle and yaw angle need to be solved by two branches. Therefore, when preparing the data, it also includes according to the corresponding angle dimension. For example, for large pitch angle branch training, each sample ID includes large pitch angle and frontal face small angle face data. Preferably, there are n pairs of large pitch angle and frontal face small angle face data.

[0096] In the training method of this preferred embodiment of the present application, after adding the branch structure to the feature layer of the trained face feature model, freeze the parameters of the face feature model and only train the model of the branch; or based on the already trained face feature model, infer and save the features of the corresponding training set in advance, and then directly read the features to train the branch.

[0097] Taking one inference forward pass as an example, for a certain face image P i , let the original output feature of the model be f i , the calculation result of the image evaluation scale be m i , the sigmoid function be S(x), the output result after activation by the fully connected layer be W(f i ), and the finally corrected output feature be

[0098]

[0099] Based on and the category label, through softmax, or other commonly used face feature loss functions for output, such as ArcFace, CosFace, etc., combine cross-entropy to calculate Loss, and perform backpropagation to update the branch parameters.

[0100] The training method in the preferred embodiment of the present application further includes controlling the correction scale, wherein controlling the correction scale includes data evaluation size calculation and correction scale control, wherein for the fuzzy training set, an evaluation parameter is used to describe the blur degree of the image, and grayscale variance, gradient sum, etc. can be used. Taking grayscale variance s as an example,

[0101]

[0102]

[0103] Among them, w x ,h y The table represents the size of the face image, the network input is 112*112, and f(x,y) represents the corresponding pixel grayscale value.

[0104] Data evaluation scale calculation For the angle training set, the corresponding angle value of the face can be directly calculated based on the facial key point information as the evaluation scale θ.

[0105] It should be noted that for the RGB and IR cross-modal training sets, since there is no intuitive and quantifiable evaluation scale for the cross-modal dimension, the scale evaluation and correction control module cannot take effect at this time, but the training branch can still achieve a certain improvement effect.

[0106] Since the original model is more robust to clear images, slightly blurred faces, frontal faces and faces at small angles, when the face is slightly blurred or at a small angle, the original features are slightly corrected or not corrected, and when it is greater than a preset threshold, the correction effect is increased. The training method of this preferred embodiment of the present application uses a sigmoid function to control the correction scale, as follows:

[0107]

[0108] A is used to control the amplitude, a controls the threshold at which the correction starts to take effect, and k controls the average rate of change. If you want the correction effect to change sharply in a certain range, k takes a large value; if you want a slow change, take a smaller k value.

[0109] In the training method of this preferred embodiment of the present application, feature pyramid model training is further included to make the model or the model branch system converge more easily. Through the feature pyramid model training, the feature map difference information of the data to be corrected and the target data at different stages is introduced into the loss function of the model training, so that the features of the two at different stages are brought closer.

[0110] Specifically, in the preferred embodiment of the present application, the feature pyramid model training includes:

[0111] 1) Build a feature pyramid. Each feature map represents a level of the pyramid. Generally, select the feature map of the previous layer before the change in the size of the feature resolution.

[0112] 2) Adjust the feature maps of each level of the pyramid to have the same number of channels.

[0113] 3) Starting from the feature map with the smallest resolution at the bottom layer, upsample it to the feature map of the upper layer and add it to the feature map of the upper layer for feature fusion. Repeat this process until each face image obtains a feature map that fuses multiple scales.

[0114] 4) Further, calculate and extract the feature pyramids of all the data to be corrected and the target data in each batch. Denote and as the k-th face data of the target data and the j-th data of the data to be corrected for the i-th ID, respectively.

[0115] During training, in each mini-batch, sample n IDs, with p and q images of the target dimension data and the data to be corrected dimension data for each ID respectively. The average features of the target dimension and the data to be corrected dimension for the i-th ID are calculated as follows:

[0116]

[0117]

[0118] Finally, the LOSSmse (mean squared error loss) corresponding to each mini-batch is obtained through the following formula:

[0119]

[0120] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The object of the present invention has been fully and effectively achieved. The function and structural principle of the present invention have been demonstrated and explained in the embodiments. Without departing from the said principle, the embodiments of the present invention can have any deformation or modification.

[0121] The technical scope of the present invention is not limited to the content described above. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of the present invention, and these deformations and modifications all fall within the protection scope of the present invention.

Claims

1. A model branch system for enhancing the robustness of a face recognition model, which can be connected after the feature layer of a trained face feature model. The input of the model branch system is the feature output by the trained face feature model, and after the model branch system corrects the target, the output is the feature after target correction. It is characterized in that The model branch system is a compensation branch based on residuals, wherein the model branch system includes n fully connected layers and ReLU connected forward and backward, and a residual module. The residual module combines the calculation results of the data evaluation scale and the output results after activation of the fully connected layer to compensate and correct the original features.

2. The model branching system according to claim 1, wherein the target correction is a correction item for the face, including clarifying correction for blurred faces, straightening faces at large angles, and cross-modal conversion.

3. The model branching system according to claim 2, wherein the model branching system selects relevant facial data for training according to the correction target. If the correction target is the blurred face recognition rate, the model branching system is trained through clear and blurred face pairs; if the correction target is the wide-angle face recognition rate, the model branching system is trained through wide-angle and frontal face data pairs; if the correction target is to improve the RGB-IR cross-modal recognition performance, the model branching system is trained through RGB and IR face data pairs.

4. The model branching system according to claim 3, wherein if it is branch training for improving fuzzy face recognition, the training data of the model branching system is implemented as follows: fuzzy branch training is performed, and each sample ID includes fuzzy face image data and clear face data.

5. The model branching system according to claim 3, wherein if training is performed for a large pitch angle branch, each sample ID includes large pitch angle and frontal small angle face data.

6. The model branch system according to claim 1, wherein for a certain face image P i , let the original output feature of the model be f i , and the calculation result of the image evaluation scale be m i The sigmoid function is S(x), and the output result after activation by the fully connected layer is W(f i ), and the finally corrected output feature is Based on and category labels, calculate the Loss by combining cross-entropy with a softmax or other commonly used face feature loss function for output, and perform backpropagation to train and update the branch parameters.

7. The model branching system according to claim 1 or 3 further comprises a correction scale control module for controlling the correction size.

8. The model branching system according to claim 7, wherein the correction scale control module comprises a data evaluation size calculation unit, wherein the data evaluation scale calculation unit adopts the grayscale variance or gradient and the blur degree of the image to describe the fuzzy training set, and directly calculates and obtains the corresponding angle value of the face as the evaluation scale θ based on the facial key point information for the angle training set.

9. The model branching system according to claim 7, wherein the correction scale control module comprises a correction scale control unit, wherein the correction scale control unit uses a sigmoid function to control the correction scale. A is used to control the amplitude, a controls the threshold at which the correction starts to take effect, and k controls the average rate of change. If you want the correction effect to change sharply in a certain range, k takes a large value; if you want a slow change, take a smaller k value.

10. The model branching system according to claim 1 further comprises a feature pyramid model training module, which introduces feature map difference information of the to-be-corrected class data and the target class data at different stages into the loss function of the model training, and brings the features of the two at different stages closer.

11. The model branch system according to claim 10, wherein the feature pyramid model training module comprises: Build a feature pyramid. Each feature_map represents a level of the pyramid. Generally, select the feature_map of the previous layer before the change in the feature resolution size. Adjust the feature_maps of each level of the pyramid to the same number of channels. Starting from the feature_map with the smallest resolution at the bottom layer, upsample it to the feature_map of the upper layer and add it to the feature_map of the upper layer for feature fusion.

12. The model branch system according to claim 11, wherein the feature pyramid of all data of classes to be corrected and target class data in each batch is calculated and extracted, denoted as and respectively represent the k-th face data of the target class data of the i-th ID and the j-th data of the data of the class to be corrected.

13. A training method for improving the robustness of a face recognition model, wherein the training method comprises the following steps: Design a model branch system based on residuals, and the model branch system is connected after the feature layer of the trained face feature model; and Input the features output by the trained face feature model into the model branch system. The model branch system processes the features and outputs the features after a certain target correction, where the target correction is a face correction item, including clarifying the blurred face, straightening the large-angle face, and cross-modal conversion.

14. The training method according to claim 13, wherein the model branch system is obtained by training with relevant face data selected according to the correction target. For example, if the correction target is the recognition rate of blurred faces, it is trained with pairs of clear and blurred face data; if the correction target is the recognition rate of large-angle faces, it is trained with pairs of large-angle and frontal face data; if the RGB-IR cross-modal recognition performance is improved, it is trained with pairs of RGB and IR face data.

15. The training method according to claim 14, wherein the model branch system is a compensation branch based on residuals, and the model branch system comprises n fully connected layers and ReLU connected in sequence.

16. The training method according to claim 15, wherein assuming that the branch training is for improving the recognition of blurred faces and large-angle faces, the training data of the model branch system is as follows: If the blurred branch training is performed, each sample ID includes blurred face image data and clear face data, preferably n pairs of blurred and clear face data pairs, where one blurred face corresponds to one clear face; if the angle branch training is performed, each sample ID includes large pitch angle and small frontal face angle data.

17. The training method according to claim 15, wherein a branch structure is added after the feature layer of the trained face feature model, the parameters of the face feature model are frozen, and only the branch is trained; or based on the already trained face feature model, the features of the corresponding training set are inferred and saved in advance, and then the features are directly read to train the branch.

18. The training method according to claim 15, wherein for a certain face image P i , let the original output feature of the model be f i , and the calculation result of the image evaluation scale be m i The sigmoid function is S(x), and the output result after activation by the fully connected layer is W(f i ), and the finally corrected output feature is Based on and category labels, calculate the Loss by combining cross-entropy with softmax or other loss functions commonly used for outputting facial features, and perform backpropagation to train and update the branch parameters.

19. The training method according to claim 18 further comprises controlling the correction scale, wherein controlling the correction scale includes calculating the data evaluation size and controlling the correction scale. For the blurred training set, the gray variance or the gradient sum is used to describe the blur degree of the image. For the angle training set, based on the face key point information, the corresponding face angle value is directly calculated as the evaluation scale θ.

20. The training method according to claim 19, wherein a sigmoid function is used to control the correction scale, as follows: A is used to control the amplitude, a controls the threshold at which the correction starts to take effect, and k controls the average rate of change. If you want the correction effect to change sharply in a certain range, k takes a large value; if you want a slow change, take a smaller k value.

21. The training method according to claim 19 further comprises feature pyramid model training, through which feature pyramid model training is used to introduce feature map difference information of the data to be corrected and the target data at different stages into the loss function of model training, so as to bring the features of the two data at different stages closer.

22. The training method according to claim 21, wherein the feature pyramid Model training includes: 1) Build a feature pyramid. Each feature_map represents a level of the pyramid. Generally, the feature_map of the previous layer before the feature resolution size changes is selected; 2) Adjust the feature_map of each level of the pyramid to have the same number of channels; 3) Starting from the feature_map with the smallest resolution at the bottom layer, upsample to the feature_map of the previous layer and add it to the feature_map of the previous layer for feature fusion.

23. The training method according to claim 22, wherein the feature pyramids of all the data of the classes to be corrected and the target class data in each batch are calculated and extracted, denoted as and respectively represent the k-th face data of the target class data of the i-th ID and the j-th data of the class to be corrected.

24. The training method according to claim 22, wherein during training, in each mini-batch, n IDs are sampled, each ID has p, q pieces of target dimension class data and to-be-corrected dimension class data, and the average features of the target dimension and to-be-corrected dimension of the i-th ID are calculated as follows: Finally, the LOSSmse corresponding to each mini-batch is obtained by the following formula: