Face Cross-Age Recognition Method, Device and Storage Medium
By using a hybrid feature extraction network based on the Transformer model and a decorrelation adversarial learning algorithm in face recognition, the recognition robustness problem under the influence of face age changes is solved, and efficient and accurate cross-age face recognition is achieved.
Patent Information
- Application Number
- CN202210355976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-06
AI Technical Summary
The prior art is difficult to effectively handle face age changes in facial recognition, resulting in poor robustness of cross-age face recognition, and the traditional convolutional neural network model is highly complex and the recognition process takes a long time.
A hybrid feature extraction network based on the Transformer model is adopted, and the traditional convolutional neural network is replaced by the T2T-ViT model, mixed features in face images are extracted, combined with the residual factor decomposition module and the decorrelation adversarial learning algorithm, decoupling identity information and age information is improved, and the robustness and efficiency of the identification model are improved.
The number of parameters and calculation complexity of the model is reduced, the speed and accuracy of face recognition are improved, and the ability to recognize faces across ages is significantly enhanced.
Smart Images

Figure CN114863512B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method, device and storage medium for cross-age face recognition. Background Art
[0002] For many years, face recognition has always been a research hotspot in the field of computer vision. In recent years, with the rapid development of artificial intelligence, face recognition algorithms based on deep learning have achieved excellent results and have been applied in various fields of life. Although general face recognition has achieved remarkable success, as people age, the appearance of the face also changes drastically, and the changes between people are also different; moreover, the face appearances of different identities have similar age-related information. For example, the differences between different periods of the same person are generally greater than the differences between children of different identities. How to minimize the impact of age changes is a long-term difficulty for current face recognition systems to correctly recognize faces in many practical applications. For example, when looking for missing children, as the age increases, the appearance of the face changes, which brings great difficulties to finding missing children. Therefore, solving the problem of age-invariant face recognition is of great significance.
[0003] Cross-age face recognition is a type of face recognition. Different from general identity-based face recognition algorithms, cross-age face recognition takes into account the face age information on the basis of identity information, and by decoupling the identity information and age information, a more generalizable face recognition model is obtained.
[0004] Among the cross-age face recognition algorithms based on deep learning, there are mainly generative methods and discriminant methods. The generative method assists face recognition by synthesizing face images of different ages. For example, a GAN model is used to improve the quality of the generated aged faces. However, accurately simulating the aging process is difficult and complex, and the unstable artifacts in the synthesized faces will significantly affect the performance of face recognition. The discriminant method is to decompose the age information and use the model to judge the identity information of the face. On the premise that it is assumed that the face information can be well modeled by the decomposed components, feature decomposition plays a key role in the invariant learning of features. In the discriminant method, for the feature extraction of faces, the current mainstream neural network is the convolutional neural network.
[0005] In the related art, a patent application for an invention with the application number 202010675730.7 discloses a face recognition method based on a deep convolutional neural network. This method first uses the MTCNN model to detect face photos; then performs alignment processing on the detected face images, crops the processed images to a size of 112 * 112; then uses a 100-layer deep convolutional neural network ResNet as the backbone network. For identity recognition, it matches face feature vectors and obtains the corresponding identity recognition information of the face feature vectors according to the similarity; for age recognition, based on the face feature vectors and multiple age classification age recognition models, it obtains the probability corresponding to each age classification of the age recognition model, and obtains the recognized age according to the age classification and the probability.
[0006] However, the problems it has are as follows: on the one hand, it does not remove the age information in the face, has poor robustness to cross-age faces, and the face recognition effect is poor; on the other hand, using a 100-layer ResNet as the backbone network, its parameter quantity and MACs are relatively large, reducing the performance of face recognition.
[0007] The literature Wang H, Gong D, Li Z, et al. "Decorrelated Adversarial Learning for Age-Invariant Face Recognition" [J]. 2019 proposed a decorrelated adversarial learning algorithm (DAL) based on linear feature decomposition. This algorithm adversarially minimizes the correlation between a person's identity information and age information. Through adversarial training, the person's identity information and age information can be made sufficiently uncorrelated, and the age information in the identity information can be significantly reduced.
[0008] However, it uses a traditional convolutional neural network, the parameter quantity and MACs of the model are relatively large, the model is relatively complex, and the parallel optimization ability is poor, and the face recognition process takes a long time. Summary of the Invention
[0009] The technical problem to be solved by the present invention is how to increase the robustness to face age information, reduce the complexity of the model and the time consumption in the face recognition process.
[0010] The present invention realizes the solution of the above technical problems through the following technical means:
[0011] On the one hand, the present invention proposes a face cross-age recognition method, and the method includes the following steps:
[0012] Obtain the face image to be detected;
[0013] Extract features from the face image using a hybrid feature extraction network based on the Transformer model to obtain hybrid features, which are a combination of face age features and face identity features;
[0014] Process the hybrid features using a residual factorization module to obtain the face identity features;
[0015] Compare the face identity features with the features in the face feature library to obtain the identity information of the face image.
[0016] The present invention uses a hybrid feature extraction network based on the Transformer model to extract the hybrid features of face age features and face identity features in a face image. By using the T2T-ViT model in the Transformer to replace the traditional convolutional neural network, the number of model parameters (Params) and MACs can be reduced, and the complexity of the model and the time consumption in the face recognition process can be reduced.
[0017] Further, the obtaining of any face image includes:
[0018] Obtain an arbitrary image and perform face detection on the image using a face detection model;
[0019] If a face part is detected, align the images of all detected face parts, and the alignment result is used as the face image to be detected;
[0020] If no face part is detected, obtain the image again.
[0021] Further, the hybrid feature extraction network includes: a T2T model and a Face Age model. The input of the T2T model is the face image, and the output is connected to the input of the Face Age model. The hybrid features output by the Face Age model are used as the input of the residual factorization module;
[0022] Among them, the T2T model uses the T2T-ViT network model to extract the face features of the face image. The Face Age model uses a feature recombination module, including a first Linear class, a View() function, a convolutional layer, a first activation function, a normalization layer, a Drop Out layer, a Flatten layer, a second Linear class, a second activation function, and an L2 Norm layer connected in sequence. The input of the first Linear class is connected to the output of the T2T model, and the output of the L2 Norm layer is connected to the input of the residual factorization module.
[0023] Further, the hybrid feature extraction network is used to extract features from the face image, and the formula representation of the obtained hybrid features is as follows:
[0024] X = E(F) = X id + X age
[0025] where X is the hybrid feature, E(F) is the feature obtained by extracting the face image using the hybrid feature extraction network, X id is the identity feature of the face, and X age is the age feature of the face.
[0026] Further, the residual factor decomposition module is used to process the hybrid feature to obtain the face identity feature X id The formula representation is as follows:
[0027]
[0028] where represents the result after convolutional processing of EFM(X), RFM(X) represents the result of processing the hybrid feature X using the residual factor decomposition module, and X age is the age feature of the face.
[0029] Further, before obtaining the face image to be detected, it further includes:
[0030] Performing decorrelation adversarial learning on the face age feature and the face identity feature to obtain the correlation between the face age feature and the face identity feature;
[0031] Calculating an optimized total loss function as the model training loss function according to the loss between the face age feature and the true age label value, the loss between the face identity feature and the true identity label value, and the correlation.
[0032] Further, the performing decorrelation adversarial learning on the face age feature and the face identity feature to obtain the correlation between the face age feature and the face identity feature includes:
[0033] Using a linear normalization mapping module to normalize the face identity feature X id and the face age feature X age to V id and V age :
[0034]
[0035]
[0036] Among them, and are neural network model parameters;
[0037] Calculate the correlation ρ between V id and V age :
[0038]
[0039] Among them, μ id and are the mean and variance of V id respectively, μ age and are the mean and variance of V age respectively. ∈ is a constant for maintaining numerical stability. Cov() represents the correlation operation, and Var() represents the variance operation.
[0040] Furthermore, the total optimized loss function is:
[0041] L = L id + αL age + βρ
[0042] Among them, L id is the loss between X id and the true identity label value, and L age is the loss between X age and the true age label value of the training set. α and β are balance ratio coefficients.
[0043] In addition, the present invention also proposes a face cross-age recognition device, and the device includes:
[0044] An acquisition module, configured to acquire a face image to be detected;
[0045] A hybrid feature extraction module, configured to use a hybrid feature extraction network based on a Transformer model to extract features from the face image to obtain hybrid features, where the hybrid features are a mixture of face age features and face identity features;
[0046] A face identity feature extraction module, configured to process the hybrid features by using a residual factorization module to obtain the face identity features;
[0047] A comparison module, configured to compare the face identity features with the features in the face feature library to obtain the identity information of the face image.
[0048] In addition, the present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method is implemented.
[0049] The advantages of the present invention are as follows:
[0050] (1) The present invention uses a hybrid feature extraction network based on the Transformer model to extract the hybrid features of the face age feature and the face identity feature in the face image. By using the T2T-ViT model in the Transformer to replace the traditional convolutional neural network, the number of model parameters (Params) and MACs can be reduced, and the complexity of the model and the time consumption of the face recognition process are reduced.
[0051] (2) The present invention adopts the decorrelation adversarial learning algorithm (DAL) to perform adversarial learning on the face age feature and the face identity feature, remove the age information in the face, and adversarially minimize the correlation between the person's identity information and age information. Through adversarial training, the person's identity information and age information can be made fully uncorrelated, and the age information in the identity information can be significantly reduced.
[0052] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0053] Figure 1 is the flowchart of the face cross-age recognition method in the present invention;
[0054] Figure 2 is the flowchart of face image detection in the present invention;
[0055] Figure 3 is the network structure diagram of the face recognition model in the present invention;
[0056] Figure 4 is the network structure diagram of the hybrid feature extraction network T2T-Face Age in the present invention;
[0057] Figure 5 is the network structure diagram of the residual factor decomposition module RFM in the present invention;
[0058] Figure 6 is the structure diagram of the face cross-age recognition device in the present invention. Detailed Embodiments
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] Referring to Figure 1 , an age-crossing face recognition method is proposed in an embodiment of the present invention. The method includes the following steps:
[0061] S10. Obtain a face image to be detected;
[0062] It should be noted that in this embodiment, a traditional face detection model can be used to detect face images in any input image.
[0063] S20. Use a hybrid feature extraction network based on a Transformer model to extract features from the face image to obtain hybrid features, where the hybrid features are a combination of face age features and face identity features;
[0064] It should be noted that in this embodiment, the hybrid feature extraction network based on the Transformer model is used to extract the hybrid features of face age features and face identity features in the face image. By using the T2T-ViT model in the Transformer instead of the traditional convolutional neural network, the number of model parameters (Params) and MACs can be reduced, the complexity of the model can be reduced, and the time consumption in the face recognition process can be reduced at the same time.
[0065] S30. Use a residual factorization module to process the hybrid features to obtain the face identity features;
[0066] S40. Compare the face identity features with the features in the face feature library to obtain the identity information corresponding to the face image.
[0067] It should be noted that the features stored in the face feature library are face features corresponding to the identities of each user.
[0068] In this embodiment, a face recognition model is used to recognize a face image. The face recognition model includes a hybrid feature extraction network and a residual factor decomposition module. The hybrid feature extraction network based on the Transformer model is used to extract the hybrid features of the face age feature and the face identity feature in the face image. The residual factor decomposition module is used to process the hybrid features to obtain the face identity feature and compare it with the features in the face feature library to obtain the identity information of the target face. By using the T2T-ViT model in the Transformer to replace the traditional convolutional neural network, the number of model parameters (Params) and MACs can be reduced, and the complexity of the model and the time consumption of the face recognition process can be reduced.
[0069] In one embodiment, referring to Figure 2 , step S10 includes:
[0070] S11. Obtain any image I
[0071] S12. Use a face detection model to perform face detection on the image I, and determine whether a face image is detected. If so, execute step S13; if not, execute step S11.
[0072] S13. Align the images of all detected face parts, and the alignment result is used as the face image F to be detected.
[0073] It should be noted that in this embodiment, the RetinaFace face detection model can be used to detect any image I.
[0074] In one embodiment, referring to Figure 3 , the face recognition model includes a hybrid feature extraction network and a residual decomposition factor module. The structure of the hybrid feature extraction network T2T-Face Age refers to Figure 4 : The hybrid feature extraction network includes a T2T model and a Face Age model. The input of the T2T model is the face image, and the output is connected to the input of the Face Age model. The hybrid features output by the Face Age model are used as the input of the residual factor decomposition module.
[0075] Among them, the T2T model adopts the T2T-ViT network model to extract the face features of the face image. The Face Age model adopts a feature recombination module, which includes a first Linear class, a View() function, a convolutional layer, a first normalization layer, an activation function, a second normalization layer, a Drop Out layer, a Flatten layer, a second Linear class, a third normalization layer, and an L2 Norm layer connected in sequence. The input of the first Linear class is connected to the output of the T2T model, and the output of the L2 Norm layer is connected to the input of the residual factorization module. The first Linear class mainly transforms features, converting the sequence features into The View() function mainly processes the sequence form into an image form in the spatial dimension, that is, converting the features into where: C = S, L2 = H × W; the combined action of the convolutional layer, the first normalization layer, and the activation function is to convert the number of features into 512 dimensions, that is, converting the features into The combined action of the second normalization layer, the Drop Out layer, the Flatten layer, the second Linear class, the third normalization layer, and the L2 Norm layer is to output the final 512-dimensional features, that is, converting the features into
[0076] It should be noted that the T2T-ViT network model is the T2T-ViT network model in the Transformer model, which is a basic network model and can be used for tasks such as classification and detection. In this embodiment, the T2T-ViT network model is used to extract face features from face images.
[0077] The feature recombination model is used to recombine the serialized face features in the spatial dimension into an image form to retain the feature information of the image blocks, which can better perform face feature decomposition.
[0078] It should be noted that in this embodiment, the face recognition model is used for human identity recognition, and the result contains n categories. Compared with the traditional face recognition for distinguishing true and false (the result has only two categories), it is more difficult. The T2T model in this embodiment adopts the T2T-ViT network model, which is an improvement of the ViT model and has better recognition effects. And in this embodiment, a Face Age module, that is, a feature recombination module, is added to recombine the serialized face features in the spatial dimension into an image form to retain the feature information of the image blocks, which can better perform face feature decomposition and further improve the face recognition effect and accuracy.
[0079] In one embodiment, in step S20, the hybrid feature extraction network is used to extract features from the face image, and the formula representation of the obtained hybrid features is as follows:
[0080] X = E(F) = X id +X age
[0081] where X is the hybrid feature, E(F) is the feature obtained by using the hybrid feature extraction network to extract the face image, X id is the identity feature of the face, and X age is the age feature of the face.
[0082] Referring to Figure 5 , the residual factor decomposition module RFM is a module in the T2T-ViT network model, including a third Linear class, a third activation function, a fourth Linear class, and a fourth activation function connected in sequence;
[0083] The formula for processing the hybrid feature by using the residual factor decomposition module to obtain the face identity feature X id is as follows:
[0084]
[0085] where represents the result after convolutional processing of RFM(X), RFM(X) represents the result of processing the hybrid feature X by using the residual factor decomposition module, and X age is the age feature of the face.
[0086] In one embodiment, before step S10, the method further includes: training the face recognition model, specifically:
[0087] S100. Perform decorrelation adversarial learning on the face age feature and the face identity feature to obtain the correlation between the face age feature and the face identity feature;
[0088] S200. Calculate the optimized total loss function as the model training loss function according to the loss between the face age feature and the true age label value, the loss between the face identity feature and the true identity label value, and the correlation.
[0089] It should be noted that in this embodiment, the decorrelation adversarial learning algorithm (DAL) is used to perform decorrelation adversarial learning on the face age feature and the face identity feature, removing the age information in the face and adversarially minimizing the correlation between the person's identity information and age information. Through adversarial training, the person's identity information and age information can be made sufficiently uncorrelated, and the age information in the identity information can be significantly reduced.
[0090] In one embodiment, step S100 includes the following steps:
[0091] Using a linear normalization mapping module, normalize the face identity feature X id and the face age feature X age to V id and V age :
[0092]
[0093]
[0094] where and are neural network model parameters;
[0095] Calculate the correlation ρ between V id and V age :
[0096]
[0097] where μ id and are the mean and variance of V id respectively, μ age and are the mean and variance of V age respectively, ∈ is a constant for maintaining numerical stability, Cov() represents the correlation operation, and Var() represents the variance operation.
[0098] In one embodiment, the optimized total loss function in step S200 is:
[0099] L = L id + αL age + βρ
[0100] where L id is the loss between X id and the true identity label value, L age is the loss between X age and the true age label value of the training set, and α, β are balance ratio coefficients.
[0101] It should be noted that the true identity tag value and the true age tag value are the losses of the training set tags for training the face recognition model.
[0102] In this embodiment, based on the T2T-ViT model in Transformer, a model for extracting face features is designed to replace the traditional convolutional neural network, reducing the number of model parameters (Params) and MACs, and improving the performance of face recognition; and the decorrelation adversarial learning algorithm (DAL) is used to remove the age information in the face, increasing the robustness to the face age information and improving the recognition accuracy of cross-age faces.
[0103] And on some publicly available datasets, it is verified that the cross-age face recognition method of this embodiment achieves the optimal effect of cross-age face recognition. The test environment is as follows: Window10 (Intel(R) Core(TM) i7-9750H CPU @ 2.60GHz 2.59GHz, 16.0G memory), Python3, Pytorch1.8.1; graphics card: NVIDIA GeForce RTX 2060, CUDA: 11.4. The test results are shown in Table 1 and Table 2:
[0104] Table 1 Test results on publicly available datasets
[0105]
[0106] Table 2 Test MACs and Params
[0107]
[0108] In addition, referring to Figure 6 , an embodiment of the present invention also proposes a cross-age face recognition device, which includes:
[0109] An acquisition module 10, configured to acquire a face image to be detected;
[0110] A hybrid feature extraction module 20, configured to extract features from the face image by using a hybrid feature extraction network based on a Transformer model to obtain hybrid features, where the hybrid features are a mixture of face age features and face identity features;
[0111] A face identity feature extraction module 30, configured to process the hybrid features by using a residual factorization module to obtain the face identity features;
[0112] A comparison module 40, configured to compare the face identity features with the features in a face feature library to obtain the identity information of the face image.
[0113] In this embodiment, a hybrid feature extraction network based on the Transformer model is used to extract the hybrid features of the face age feature and the face identity feature in the face image. By using the T2T-ViT model in the Transformer to replace the traditional convolutional neural network, the number of model parameters (Params) and MACs can be reduced, and the complexity of the model and the time consumption of the face recognition process can be reduced.
[0114] In one embodiment, the device further includes: a face detection module, specifically used for:
[0115] Obtain any image I, and use a face detection model to perform face detection on the image I to determine whether a face image is detected;
[0116] If a face part is detected, align the images of all detected face parts, and the alignment result is used as the face image F to be detected;
[0117] If no face part is detected, obtain the image I again.
[0118] In one embodiment, the structure of the hybrid feature extraction network T2T-Face Age refers to Figure 4 : The hybrid feature extraction network includes: a T2T model and a Face Age model. The input of the T2T model is the face image, and the output is connected to the input of the Face Age model. The hybrid features output by the Face Age model are used as the input of the residual factorization module;
[0119] Among them, the T2T model uses the T2T-ViT network model, and the Face Age model uses a feature recombination module. The feature recombination module includes a first Linear class, a View() function, a convolutional layer, a first normalization layer, an activation function, a second normalization layer, a Drop Out layer, a Flatten layer, a second Linear class, a third normalization layer, and an L2 Norm layer connected in sequence. The input of the first Linear class is connected to the output of the T2T model, and the output of the L2 Norm layer is connected to the input of the residual factorization module. The first Linear class is mainly used to transform features and convert sequence features into The View() function is mainly used to process the sequence form into an image form in the spatial dimension, that is, to convert features into where: C = S, L2 = H×W; the convolutional layer, the first normalization layer, and the activation function together act to convert the number of features into 512 dimensions, that is, to convert features into The combined effect of the second normalization layer, Drop Out layer, Flatten layer, second Linear class, third normalization layer, and L2 Norm layer is to output the final 512-dimensional feature, that is, to convert the feature into
[0120] In one embodiment, the hybrid feature extraction network is used to extract features from the face image, and the formula representation of the obtained hybrid feature is as follows:
[0121] X = E(F) = X id + X age
[0122] where X is the hybrid feature, E(F) is the feature after extracting the face image using the hybrid feature extraction network, X id is the identity feature of the face, and X age is the age feature of the face.
[0123] In one embodiment, the residual factor decomposition module is used to process the hybrid feature to obtain the face identity feature X id The formula representation is as follows:
[0124]
[0125] where represents the result after performing convolution processing on RFM(X), RFM(X) represents the result after processing the hybrid feature X using the residual factor decomposition module, and X age is the age feature of the face.
[0126] In one embodiment, the device further includes a training module, specifically including:
[0127] The decorrelation adversarial training unit is used to perform decorrelation adversarial learning on the face age feature and the face identity feature to obtain the correlation between the face age feature and the face identity feature;
[0128] The loss function calculation unit is used to calculate the optimized total loss function as the model training loss function according to the loss between the face age feature and the true age label value, the loss between the face identity feature and the true identity label value, and the correlation.
[0129] It should be noted that in this embodiment, the decorrelation adversarial learning algorithm (DAL) is used to perform adversarial learning on the facial age feature and the facial identity feature, removing the age information in the face, and adversarially minimizing the correlation between the identity information and the age information of a person. Through adversarial training, the identity information and the age information of a person can be made sufficiently uncorrelated, and the age information in the identity information can be significantly reduced.
[0130] In one embodiment, the decorrelation adversarial training unit is specifically configured to:
[0131] Use a linear normalization mapping module to normalize the facial identity feature X id and the facial age feature X age to V id and V age :
[0132]
[0133]
[0134] Wherein, and are neural network model parameters;
[0135] Calculate the correlation ρ between V id and V age :
[0136]
[0137] Wherein, μ id and are the mean and variance of V id respectively, μ age and are the mean and variance of V age respectively, ∈ is a constant to maintain numerical stability, Cov() represents the correlation operation, and Var() represents the variance operation.
[0138] In one embodiment, the optimized total loss function is:
[0139] L = L id + αL age + βρ
[0140] Wherein, L id is the loss between X id and the true identity label value, L age is the loss between X age and the true age label value of the training set, and α, β are balance ratio coefficients.
[0141] It should be noted that other embodiments or implementation methods of the face cross-age recognition device of the present invention can refer to the above-mentioned method embodiments, and will not be repeated here.
[0142] In addition, an embodiment of the present invention also discloses a computer-readable medium, on which computer-readable instructions are stored, and the computer-readable instructions can be executed by a processor to implement the face cross-age recognition method as described above.
[0143] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0144] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0145] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0146] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0147] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A cross-age face recognition method, characterized in that, The method includes: Obtaining a face image to be detected; Using a hybrid feature extraction network based on a Transformer model to extract features from the face image, obtaining hybrid features, where the hybrid features are a combination of face age features and face identity features; Using a residual factorization module to process the hybrid features to obtain the face identity features; Comparing the face identity features with the features in the face feature library to obtain the identity information corresponding to the face image; Among them, the hybrid feature extraction network includes: a T2T model and a Face Age model. The input of the T2T model is the face image, and the output is connected to the input of the Face Age model. The hybrid features output by the Face Age model are used as the input of the residual factorization module; Among them, the T2T model uses a T2T-ViT network model, and the Face Age model uses a feature recombination module. The feature recombination module includes a first Linear class, a View() function, a convolutional layer, a first activation function, a normalization layer, a Drop Out layer, a Flatten layer, a second Linear class, a second activation function, and an L2 Norm layer connected in sequence. The input of the first Linear class is connected to the output of the T2T model, and the output of the L2 Norm layer is connected to the input of the residual factorization module; Among them, before obtaining the face image to be detected, it further includes: Performing decorrelation adversarial learning on the face age features and the face identity features to obtain the correlation between the face age features and the face identity features; Calculating an optimized total loss function as the model training loss function according to the loss between the face age features and the true age label value, the loss between the face identity features and the true identity label value, and the correlation; 2. The face cross-age recognition method according to claim 1, characterized in that The obtaining of the face image to be detected includes: Obtaining an arbitrary image and using a face detection model to detect faces in the image; If a face part is detected, aligning the images of all detected face parts, and the alignment result is used as the face image to be detected; If no face part is detected, re-obtaining the image.
3. The face cross-age recognition method according to claim 1, wherein Using the hybrid feature extraction network to extract features from the face image, and the formula representation of the obtained hybrid features is as follows: Among them, is the mixed feature, is the feature extracted from the face image using the mixed feature extraction network, is the identity feature of the face, is the age feature of the face.
4. The face cross-age recognition method according to claim 1, wherein Processing the mixed features by using the residual factorization module to obtain face identity features The formula representation is as follows: Among them, represents the result after performing convolution processing, represents the result after processing the mixed feature using the residual factorization module, which is the age feature of the face.
5. The face cross-age recognition method according to claim 1, characterized in that The performing of decorrelation adversarial learning on the face age features and the face identity features to obtain the correlation between the face age features and the face identity features includes: Using a linear normalization mapping module, normalize the face identity feature and the face age feature by performing normalization mapping to and : Among them, and are neural network model parameters; Calculation and the correlation between : Among them, and are the mean and variance of respectively, and are the mean and variance of respectively, is a constant that holds the numerical stability, .
6. The face cross-age recognition method according to claim 5, wherein The optimized total loss function is: Among them, is the loss between the true identity tag value is the loss with the true age tag value of the training set and is the balance ratio coefficient.
7. A cross-age face recognition device, characterized in that, The device includes: An obtaining module, configured to obtain a face image to be detected; A hybrid feature extraction module, configured to use a hybrid feature extraction network based on a Transformer model to extract features from the face image, obtaining hybrid features, where the hybrid features are a combination of face age features and face identity features; A face identity feature extraction module, configured to use a residual factorization module to process the hybrid features to obtain the face identity features; A comparison module, configured to compare the face identity features with the features in the face feature library to obtain the identity information of the face image; Wherein, the hybrid feature extraction network includes: a T2T model and a Face Age model. The input of the T2T model is the face image, and the output is connected to the input of the Face Age model. The hybrid feature output by the Face Age model is used as the input of the residual factor decomposition module; Wherein, the T2T model adopts a T2T-ViT network model, and the Face Age model adopts a feature recombination module. The feature recombination module includes a first Linear class, a View() function, a convolutional layer, a first activation function, a normalization layer, a Drop Out layer, a Flatten layer, a second Linear class, a second activation function, and an L2 Norm layer connected in sequence. The input of the first Linear class is connected to the output of the T2T model, and the output of the L2 Norm layer is connected to the input of the residual factor decomposition module; The device further includes a training module, specifically including: A decorrelation adversarial training unit, configured to perform decorrelation adversarial learning on the face age feature and the face identity feature to obtain the correlation between the face age feature and the face identity feature; A loss function calculation unit, configured to calculate an optimized total loss function as the model training loss function according to the loss between the face age feature and the true age label value, the loss between the face identity feature and the true identity label value, and the correlation; 8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Face Recognition Method and System Based on Deep Convolutional Neural Networks
CN111985323B
Cross-age face recognition method, system and device and storage medium
CN111881722A
Face recognition method and device and electronic equipment
CN112597941A