Cross-age identity authentication method and system based on attention mechanism

By adopting an attention mechanism-based method in cross-age identity authentication, the YOLOv11 detection model and Resnet-50 backbone network are improved, and combining multi-task learning and cross-age domain adversarial learning, the problems of decreased recognition rate and insufficient feature separation in the existing technology are solved, and high accuracy and robust cross-age identity authentication is achieved.

CN119939559APending Publication Date: 2025-05-06SOUTH CHINA UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510005011.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing cross-age identity authentication methods have a decrease in recognition rate when facing the appearance changes caused by age changes, and the generative model is prone to artifacts and high computing power requirements. The discriminant model needs to fully separate identity characteristics and age characteristics, and insufficient decomposition affects the recognition rate.

Method used

Using an interage age identity authentication method based on attention mechanism, multi-task learning is performed through the improved YOLOv11 detection model and Resnet-50 backbone network, combining channel and spatial attention modules, to separate age-related features and identity-related features, and to encourage features to decorrelate them through cross-age domain confrontation learning.

Benefits of technology

It improves the accuracy of cross-age identity authentication, enhances the accuracy and robustness of face detection, fully separates identity characteristics and age characteristics, effectively resists age interference, and improves recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939559A_ABST
    Figure CN119939559A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-age identity authentication method and system based on an attention mechanism. A video input device is used for capturing a current picture and transmitting the current picture to a background for processing; the improved YOLOv11 detection model is used for detecting and intercepting a face image; and establishing a cross-age identity authentication network based on an attention mechanism as a recognition model, and performing identity authentication on the intercepted face image. Aiming at the problem that the identity authentication accuracy is reduced due to age change, the cross-age identity authentication network based on the attention mechanism is constructed, the identity-related features and the age-related features in the face image are effectively separated, and the accuracy of cross-age identity authentication is improved. In addition, the improved YOLOv11 network is adopted as a detection network, and the robustness and the detection precision of face detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning and computer vision, and specifically relates to a cross-age identity authentication method and system based on an attention mechanism. Background Art

[0002] Identity authentication has achieved remarkable success, but the appearance changes caused by age changes often reduce the recognition rate of faces, and cross-age identity authentication has important applications in tracking missing children and identifying absconding fugitives. Therefore, how to minimize the impact of age changes is a lingering challenge for current identity authentication systems to correctly identify identities in many practical applications.

[0003] Existing research methods for cross-age identity authentication can be roughly summarized into two categories: generative models and discriminative models. The generative model converts faces of different age groups into the same age group after a given face image in order to minimize the impact of age changes on identity authentication. Guo et al. proposed a new method for generating aging faces with identity preservation based on multi-scale attention and multi-attribution fusion of genetic networks (Identity-Preserving_Face_Aging_With_Multi-Attribution_Fusion_and_Multi-Scale_Attention, Yingchun Guo o, Mingyuan Lio, and Gang Yan), which has good performance in identity preservation and accurate age conversion, but the generative method is prone to artifacts, the generated image quality is unstable, and the computing power requirements are high. The discriminative model extracts age-invariant features by disentangling identity-related information from mixed facial information, so that the facial recognition system only uses identity-related information for recognition. This method does not require high computing power, the results are relatively stable, and it is easy to deploy and implement, but it requires sufficient separation of identity features and age features to achieve a high accuracy rate. Liu Cheng et al. proposed a cross-age face recognition method based on Transformer (A cross-age face recognition method based on Transformer, Liu Cheng, Cao Liangcai, Jin Ye, Wang Haowei, Yin Songfeng, Tsinghua University Hefei Public Security Institute, Hefei, Anhui; Tsinghua University State Key Laboratory of Precision Measurement Technology and Instrumentation; Hefei Public Security Bureau Criminal Police Detachment), which extracts face age and identity mixed features through an improved model, and then obtains face age features and identity features through residual factor decomposition. This method has a relatively lightweight model and uses simple modules to decompose identity features and age features, but it also has the problem of insufficient decomposition affecting the recognition rate.

[0004] Therefore, a stable, highly accurate cross-age identity authentication system that can fully separate identity features and age features and effectively resist age interference is of great significance to real life. Summary of the invention

[0005] In order to solve the problem of decreased accuracy of identity authentication due to existing age changes, the present invention proposes a cross-age identity authentication method and system based on an attention mechanism, which can improve the accuracy of cross-age identity authentication.

[0006] The purpose of the present invention is achieved by at least one of the following technical solutions.

[0007] A cross-age identity authentication method based on an attention mechanism includes the following steps:

[0008] S1. Use a video input device to capture images. The video input device captures the current image into several frames of images and inputs the images into a face detection model.

[0009] S2, using the face detection model to detect the face, and extracting the face image to be recognized according to the face coordinates;

[0010] S3, preprocessing the captured face image;

[0011] S4, input the preprocessed image into the trained cross-age identity authentication network for identity authentication. Further, in step S2, the face detection model uses an improved YOLOv11 detection model to detect faces. The improved YOLOv11 detection model is an improvement on the C3K2 module in the YOLOv11 backbone network and the neck network. The improved C3K2 module includes adding multiple bottleneck layer branches and configuring a weight generator, which is used to generate respective weights for each bottleneck layer branch according to the current image features; when the image features enter each bottleneck layer branch, they are first multiplied by the weight of the corresponding branch, and then enter the branch, and more effective image features are screened out after weighted processing of multiple branches.

[0012] Furthermore, in step S3, the face image preprocessing specifically includes:

[0013] Data cleaning: Filter out blurred face images or misdetected non-face images from the captured images;

[0014] Image resizing and pixel normalization: The image size is uniformly scaled to 112×112, and the pixels are uniformly normalized to the range of [-1, 1].

[0015] Furthermore, in step S4, cross-age identity authentication network content:

[0016] Backbone network: mainly used to extract input facial image features and convert them into facial information feature maps that are a mixture of age-related features and identity-related features;

[0017] Attention module: extracts features of mixed facial features in channel and spatial dimensions;

[0018] Multi-task learning module: mainly used to calculate the loss of age prediction, the loss of identity prediction, and the loss of the correlation between the two features of identity and age, and to remove the correlation between the two features.

[0019] Furthermore, in step S4, the backbone network and the attention module specifically include:

[0020] The backbone network uses the Resnet-50 network model to extract mixed face feature maps from images;

[0021] The attention module is divided into a channel attention module and a spatial attention module. The channel attention module adopts a compressed excitation network. The compressed excitation network output feature map includes the following steps:

[0022] First, a compression operation is performed, and the W×H×C feature map is compressed into 1×1×C using channel global average pooling, where W is the width of the feature map, H is the height of the feature map, and C is the number of channels of the feature map; channel average pooling is as follows:

[0023]

[0024] Among them, Z c represents the cth element of the compressed feature vector Z, U c Represents the cth channel of the input feature map U;

[0025] Then, two fully connected layers are used to perform the extraction operation to capture the compressed feature information. The first fully connected layer compresses the C channels of the compressed feature vector Z into C / r channels to reduce the amount of calculation, where r is the compression ratio, and then passes through a ReLU nonlinear activation layer; the second fully connected layer restores the number of channels to C channels, and then passes through the Sigmoid function to obtain a 1×1×C weight S, which is used to characterize the weight of each channel in the input feature map U. The formula for the extraction operation is as follows:

[0026] S = σ(W2δ(W1z));

[0027] Where δ(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function, and W1 and W2 represent the weights of the two fully connected layers respectively;

[0028] Finally, a weighted operation is performed to multiply the weight S by the input feature map to obtain the output feature map after channel attention weighting. The weighted process formula is as follows:

[0029]

[0030] in Represents the output feature map The cth channel component, U c represents the cth channel of the input feature map U, S c represents the c-th component of weight S.

[0031] Furthermore, the spatial attention module uses a combination of maximum pooling and average pooling to extract spatial hierarchical features: first, the input feature map U of size W×H×C is subjected to global maximum pooling and global average pooling in the channel dimension to obtain two feature maps of size H×W×1. The formulas of the above global maximum pooling and global average pooling are as follows:

[0032] Z max [i,j]=Max(U[i,j,i],U[i,j,2],...,U[i,j,C];

[0033]

[0034] Among them, Z max [i, j] indicates that the width dimension of the maximum pooling output feature map is i, the height dimension is the element at j, C indicates the number of channels in the input feature map, U[i, j, c] indicates that the width dimension is i, the height dimension is j, and the channel dimension is c in the input feature map, Z avg [i, j] represents the element at the width dimension i and the height dimension j of the average pooling output feature map;

[0035] Then the feature map Z obtained by maximum pooling and average pooling is max and Z avg By channel splicing, we get a feature map of H×W×2, and perform a convolution operation on the splicing result to get a feature map of H×W×1. Then, we pass the Sigmoid function to get the spatial attention weight matrix M. The formula of this process is as follows:

[0036] M=σ(f(Z max ; Z avg ));

[0037] Among them, σ(·) represents the Sigmoid activation function and f(·) represents the convolution operation.

[0038] The obtained spatial attention weight matrix M is multiplied by the input feature map U for weighting, and the formula is as follows:

[0039]

[0040] in, represents the weighted output feature map, U[i, j] represents the element at the input feature map with width dimension i and height dimension j, and M[i, j] represents the element at the weight matrix with width dimension i and height dimension j. Further, in step S4, the multi-task learning framework for cross-age identity authentication network training includes the following contents:

[0041] For the age prediction task, linear regression is used to estimate age, with mean square error as the loss function. The formula is as follows:

[0042]

[0043] Among them, L age represents the age prediction loss, F age (.) represents the representation function of the age prediction module, represents the age feature vector of the i-th training set sample, z i The age label of the i-th training set sample;

[0044] For the identity recognition task, the ArcFace function is used as the loss function of identity recognition. The formula is as follows:

[0045]

[0046] Among them, L id represents the identity prediction loss, N is the size of the training batch, that is, the number of input images; k is the number of categories, that is, the number of identities trained; y represents the true identity category label; Represents the true category y i The cosine value of the angle after the offset processing; is the angle of the target category, i.e. the true label; m is the weighted angle, which is used to adjust the interval of the target angle and enhance the discrimination between features; s is a scaling factor, which is used to control the learning speed of the model, cosθ j is the cosine similarity between each category and the feature vector; for the correlation loss between age features and identity features, a cross-age domain adversarial learning is introduced to encourage the de-correlation of age features and identity features through continuous domain adaptation of the gradient reversal layer, and the correlation loss L is obtained. cor , the final total loss function formula of the multi-task learning framework is as follows:

[0047] L=L id +αL age +βL cor ;

[0048] Among them, α and β are weight parameters that control the balance of different loss terms.

[0049] Furthermore, in step S4, the cross-age identity authentication network performs identity authentication, including: the face image to be identified passes through the backbone network to obtain a mixed feature map C, and the mixed feature map is weighted by the attention module to obtain the age-related feature X age , subtract the age-related features from the mixed feature map to obtain the identity-related features X id , the formula is as follows:

[0050] X id =XX age ;

[0051] The identity-related feature X id A 512-dimensional feature vector V is obtained through the linear layer, and the feature vector V is input into the background storage device for storage; the cosine similarity between the feature vector V and the other saved feature vectors is calculated one by one, and the formula is as follows:

[0052]

[0053] Where n is the total number of feature vectors stored in the background storage device, D is the feature vector stored in the background storage device, V i and D i Represents the i-th component of the feature vector. If the similarity is greater than the threshold, the two are considered to be the same person and the identity authentication result is returned.

[0054] A system for implementing the cross-age identity authentication method based on the attention mechanism includes: a data acquisition module, which is used to collect image data and input the image data into a face detection module;

[0055] The face detection module extracts the features of the image, cuts out the face image to be identified according to the face coordinates, and outputs the face image to be identified;

[0056] The identity authentication module is used to authenticate the preprocessed face image to be identified.

[0057] A computer device comprises: a memory and a processor and a computer program stored in the memory. When the computer program is executed on the processor, a cross-age identity authentication method based on an attention mechanism is implemented.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] 1. The present invention improves the structure of the YOLOv11 detection network in the face detection stage, adds multiple bottleneck layer branches to the C3K2 module of YOLOv11, and generates respective weights for each bottleneck layer branch using a weight generator according to the current image features. After weighted processing of multiple branches, more effective image features can be screened out for subsequent processing, thereby improving the detection accuracy of face images and the robustness of face detection.

[0060] 2. When constructing a cross-age identity authentication network, the present invention enhances the network's ability to capture image features by constructing an attention module. Among them, the channel attention module uses the SE network to solve the channel dependency problem through excitation and compression operations, and obtains the weight of each channel to weight the feature map on the channel; the spatial attention module uses a combination of maximum pooling and average pooling to better extract the spatial weights to weight the feature map spatially. By weighting the feature map, age-related features and identity-related features can be more fully separated.

[0061] 3. When constructing a cross-age identity authentication network, the present invention constructs a more effective multi-task learning framework to fully separate age-related features and identity-related features, thereby improving recognition accuracy. The mean square error and ArcFace function are used as loss functions for age prediction and identity recognition tasks, respectively, which is simple and efficient. On this basis, a cross-age domain adversarial learning is introduced, which encourages the de-correlation of age features and identity features through continuous domain adaptation of the gradient reversal layer. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 This is a flow chart of the cross-age identity authentication system based on the attention mechanism of this embodiment;

[0063] Figure 2 This is a schematic diagram of the YOLOv11 model structure in this embodiment;

[0064] Figure 3 Schematic diagram of the structure of the C3K2 module after improvement of the YOLOv11 model in this embodiment;

[0065] Figure 4 This is a schematic diagram of the framework structure of the cross-age identity authentication network in this embodiment;

[0066] FIG5( a ) is a schematic diagram of the identity prediction module structure of the cross-age identity authentication network model implemented in this embodiment;

[0067] FIG5( b ) is a schematic diagram of the structure of the age prediction module of the cross-age identity authentication network model implemented in this embodiment;

[0068] FIG5(c) is a schematic diagram of the structure of the identity age correlation prediction module of the cross-age identity authentication network model implemented in this embodiment. DETAILED DESCRIPTION

[0069] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0070] like Figure 1 As shown, this embodiment is a cross-age identity authentication method based on an attention mechanism, using an improved YOLOv11 as a face detection network, and a cross-age identity authentication system based on an attention mechanism to separate age-related features and identity-related features for identity authentication. The recognition accuracy of the identity authentication system can be effectively improved and the interference of age factors on identity authentication can be effectively counteracted. The authentication method specifically includes the following steps:

[0071] S1. Image acquisition stage based on video input device: The video input device captures the current picture into several frames of images, which are then input into the data storage device, and the background processing device reads the images and performs subsequent processing.

[0072] For example, when an RGB camera is used to capture the current picture, each frame of the current picture will be input into the computer, read by the system running in the background, and then processed.

[0073] S2, face detection stage based on improved YOLOv11 detection model: use the face detection model to detect the face, and extract the face image to be recognized according to the face coordinates;

[0074] In order to improve the accuracy of face detection, the following improvements have been made to the YOLOv11 detection model structure: mainly the C3K2 module in the YOLOv11 backbone network and the neck network has been improved, including adding multiple bottleneck layer branches and configuring a weight generator to generate respective weights for each bottleneck layer branch according to the current image features. When the image features enter each bottleneck layer branch, they will first be multiplied by the weight of the branch and then enter the branch for processing. After weighted processing of multiple branches, more effective image features can be screened out for subsequent processing, thereby improving the detection accuracy of face images.

[0075] The overall network framework of the YOLOv11 detection model is as follows Figure 2As shown in the figure. The backbone network of YOLOv11 extracts and outputs the features of the image. The neck network (Neck) strengthens the expressiveness of the features and performs multi-scale feature fusion. Finally, the head network (Head) outputs the final detection results including the face frame coordinates and confidence. The face image to be recognized is intercepted according to the face coordinates.

[0076] The YOLOv11 model includes Conv and C3K2 modules, where Conv is a convolutional layer, and the C3K2 module is divided into two states: C3K=False and C3K=True. When C3K=False, the input features pass through the convolutional layer, and then a branch is drawn here to cross the bottleneck layer group, and the main branch features pass through the segmentation layer, and part of them pass through the bottleneck layer group, and the other part is processed by the bottleneck layer group, and then recombined with the above-mentioned branch that crosses the bottleneck layer and the part that crosses the bottleneck layer after segmentation, and then passes through the convolutional layer to obtain the output; C3K=True replaces the bottleneck layer group in the above process with the C3K module group. The structure of the C3K module is as follows: after the input features pass through the convolutional layer, a branch is drawn out, which passes through the convolutional layer and then crosses the bottleneck layer group, and the main branch is processed by the bottleneck layer group and recombined with the branch, and then passes through the convolutional layer to obtain the output.

[0077] This embodiment improves the state of C3K=False. The improved C3K2 module is as follows: Figure 3 As shown in the figure, the C3K2 module in this state adds 3 bottleneck layer branches, and input features are input into the weight generator. The structure of the weight generator is as follows: the W×H×C feature map is compressed into 1×1×C after average pooling, where W is the width of the feature map, H is the height of the feature map, and C is the number of channels of the feature map; it is then compressed into C / 4 channels through the layer fc1 to reduce the amount of calculation, and the second fully connected layer fc2 converts the number of channels into 3 channels. Finally, the Softmax layer generates the weights (k1, k2, k3) of each branch, where k1, k2, and k3 are the weights of the three bottleneck layer branches respectively. The input features are first multiplied by the weights and then input into each branch, and then reorganized after being processed by the bottleneck layer group of each branch, and finally the output is obtained through the convolution layer. The YOLOv11 detection model will continuously read each incoming frame for detection. When a face appears in the picture, the backbone network (Backbone) of YOLOv11 extracts the features of the image and outputs them. The neck network (Neck) will enhance the expressiveness of the features and perform multi-scale feature fusion. Finally, the head network (Head) outputs the final detection results including the face frame coordinates and confidence level, and then extracts the face image to be identified based on the face coordinates.

[0078] S3. Preprocess the face image to be recognized:

[0079] Data cleaning: Filter out unclear face images or misdetected non-face images from the captured images; perform data cleaning first to filter out face images with low confidence.

[0080] Then the images are uniformly scaled to a size of 112×112, and the pixels of each face image to be recognized are normalized to the range of [-1, 1].

[0081] S4. Input the preprocessed image into the trained cross-age identity authentication network for identity authentication.

[0082] The framework of the cross-age identity authentication network is as follows Figure 4 As shown in FIG. 1 , the network framework specifically includes the following contents: Backbone network part: mainly used to extract the features of the input face image and convert it into a face information feature map mixed with age-related features and identity-related features;

[0083] Attention module part: mainly extracts mixed facial features in channel and spatial dimensions to obtain a better feature screening weight grid, thereby screening out age-related features and identity-related features;

[0084] Multi-task learning module: mainly used to calculate the loss of age prediction, the loss of identity prediction and the loss of the correlation between identity and age features, so that the network can better learn age features and identity features and remove the correlation between the two features.

[0085] As one of the embodiments, the backbone network of this embodiment adopts Resnet-50 residual network to extract the features of the input mixed face image and convert it into a face information feature map X that is a mixture of age-related features and identity-related features. In this embodiment, the feature map size is (512,7,7).

[0086] The attention module is divided into a channel attention module and a spatial attention module. The channel attention module uses a squeeze-and-excitation (SE) network. In order to solve the channel dependency problem, the network first performs a squeeze operation. First, the W×H×C feature map is compressed into 1×1×C using channel global average pooling, where W is the width of the feature map, H is the height of the feature map, and C is the number of channels of the feature map. The principle of channel average pooling is as follows:

[0087]

[0088] Among them, Z c represents the cth element of the compressed feature vector Z, U c Represents the c-th channel of the input feature map U.

[0089] The network then uses two fully connected layers to perform an extraction operation to capture the compressed feature information. The first fully connected layer compresses the C channels of the compressed feature vector Z into C / r channels to reduce the amount of calculation, where r is the compression ratio, and then passes through a ReLU nonlinear activation layer. The second fully connected layer restores the number of channels to C channels, and then passes through the Sigmoid function to obtain a 1×1×C weight S, which is used to characterize the weight of each channel in the input feature map U. The formula for the extraction operation is as follows:

[0090] S = σ(W2δ(W1z));

[0091] Where δ(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function, and W1 and W2 represent the weights of the two fully connected layers, respectively.

[0092] Finally, the network performs a weighted operation, multiplying the weight S by the input feature map to obtain the output feature map after channel attention weighting. The weighted process formula is as follows:

[0093]

[0094] in Represents the output feature map The cth channel component, U c represents the cth channel of the input feature map U, S c represents the c-th component of weight S.

[0095] The spatial attention module uses a combination of maximum pooling and average pooling to extract spatial hierarchical features. First, the input feature map U of size W×H×C is subjected to global maximum pooling and global average pooling in the channel dimension to obtain two feature maps of size H×W×1. The formulas for the above global maximum pooling and global average pooling are as follows:

[0096] Z max [i,j]=Max(U[i,j,i], U[i,j,2],...,U[i,j,C]);

[0097]

[0098] Among them, Z max [i, j] indicates the element at position j with width dimension of the maximum pooling output feature map, C indicates the number of channels in the input feature map, and U[i, j, c] indicates the element at position c with width dimension i, height dimension j, and channel dimension in the input feature map. avg [i, j] represents the element at the width dimension i and the height dimension j of the average pooling output feature map.

[0099] Then the feature map Z obtained by maximum pooling and average pooling is max and Z avg By channel splicing, we get a feature map of H×W×2, and perform a convolution operation on the splicing result to get a feature map of H×W×1. Then, we pass the Sigmoid function to get the spatial attention weight matrix M. The formula of this process is as follows:

[0100] M=σ(f(Z max ; Z avg ));

[0101] Among them, σ(·) represents the Sigmoid activation function and f(·) represents the convolution operation.

[0102] The obtained spatial attention weight matrix M is multiplied by the input feature map U for weighting, and the formula is as follows:

[0103]

[0104] in, Represents the weighted output feature map, U[i, j] represents the element at the width dimension i and the height dimension j of the input feature map, and M[i, j] represents the element at the width dimension i and the height dimension j in the weight matrix.

[0105] The multi-task learning framework for cross-age identity authentication network training specifically includes the following:

[0106] For the age prediction task, linear regression is used to estimate age, and the mean square error is used as the loss function. The specific formula is as follows:

[0107]

[0108] Among them, L age represents the age prediction loss, F age (.) represents the representation function of the age prediction module, represents the age feature vector of the i-th training set sample, z i The age label of the i-th training set sample. The training set can be obtained from the public face dataset published on the Internet;

[0109] For the identity recognition task, the ArcFace function is used as the loss function of identity recognition. The specific formula is as follows:

[0110]

[0111] Among them, L idrepresents the identity prediction loss, N is the size of the training batch, that is, the number of input images; k is the number of categories, that is, the number of identities trained; y represents the true identity category label; The cosine value of the angle representing the true category yi after the offset processing; is the angle of the target category, i.e. the true label; m is the weighted angle, which is used to adjust the interval of the target angle and enhance the discrimination between features; s is a scaling factor, which is used to control the learning speed of the model, cosθ j is the cosine similarity between each category and the feature vector.

[0112] For the correlation loss between age features and identity features, a cross-age domain adversarial learning is introduced. The continuous domain adaptation of the gradient reversal layer (GRL) is used to encourage the de-correlation of age features and identity features, and the correlation loss L is obtained. cor The final total loss function formula of the multi-task learning framework is as follows:

[0113] L=L id +αL age +βL cor ;

[0114] Among them, α and β are weight parameters that control the balance of different loss terms.

[0115] The trained cross-age identity authentication network is used for identity authentication, which specifically includes the following: the face image to be identified passes through the backbone network to obtain a mixed feature map X, a compressed excitation network is used as the channel attention module, and a combination of maximum pooling and average pooling is used as the spatial attention module. The output weights of the channel attention module and the output weights of the spatial attention module are combined into an attention weight grid W, and then the feature map is multiplied by the attention grid to obtain the age-related feature X age The feature map minus the age-related features to obtain the identity-related features X id , the specific process is as follows:

[0116]

[0117] in, Represents the multiplication of corresponding elements.

[0118] The multi-task learning module uses the ArcFace function as the loss function for identity recognition. id Input the identity prediction module to obtain the identity feature vector V id , V id Input the ArcFace function to calculate the identity recognition error L id ; Using mean square error as the loss function for age prediction, X age Input the age prediction module to get the predicted age Zpred , Z pred With the label Age Z label Calculate the mean square error to get the age prediction loss L age . Use X id After the gradient reversal layer (GRL) is input into the age prediction module, the correlation loss L can be obtained. cor . The network framework of each prediction module is shown in Figure 5(a), Figure 5(b), and Figure 5(c). The size of the input feature map is W×H×C, where W is the width of the feature map, H is the height of the feature map, and C is the number of channels of the feature map. The identity prediction module is shown in Figure 5(a). The input feature map passes through the BN layer, then passes through the fully connected layer to obtain a 512-dimensional feature vector, and then passes through a BN layer to obtain the output. The age prediction module is shown in Figure 5(b). The input feature map passes through the BN layer, then passes through the fully connected layer to obtain a 512-dimensional feature vector, and then passes through a fully connected layer to obtain a 101-dimensional output feature vector. The identity feature and age feature correlation prediction module is shown in Figure 5(c). The input feature map is processed by the gradient reversal layer (GRL) and then input into the same structure as the age prediction module to obtain a 101-dimensional output feature vector.

[0119] The final total loss is as follows:

[0120] L=L id +αL age +βL cor ;

[0121] In this embodiment, α is set to 0.01 and β is set to 0.02.

[0122] The identity-related feature X id The 512-dimensional feature vector V is obtained through the linear layer. The feature vector V can be selected to be input into the background storage device for storage; if there are already saved feature vectors in the background storage device, you can choose to save the feature vector V or calculate the cosine similarity with the other saved feature vectors in the background storage device one by one. The formula is as follows:

[0123]

[0124] Where n is the total number of feature vectors stored in the background storage device, D is the feature vector stored in the background storage device, V i and D i Represents the i-th component of the feature vector. If the similarity is greater than the threshold, the two are considered to be the same person and the identity authentication result is returned.

[0125] When the calculated cosine similarity is greater than the set threshold, it is considered that the current face to be identified and the face stored in the database belong to the same person, and the identification result is returned. As one embodiment, the threshold of this embodiment is set to 0.7.

[0126] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well.

Claims

1. A cross-age identity authentication method based on attention mechanism, characterized in that: The following steps are involved: S1. Use a video input device to capture images. The video input device captures the current image into several frames of images and inputs the images into a face detection model. S2, using the face detection model to detect the face, and extracting the face image to be recognized according to the face coordinates; S3, preprocessing the captured face image; S4. Input the preprocessed image into the trained cross-age identity authentication network for identity authentication.

2. According to the cross-age identity authentication method based on the attention mechanism of claim 1, it is characterized in that: In step S2, the face detection model uses an improved YOLOv11 detection model to detect faces. The improved YOLOv11 detection model is an improvement on the C3K2 module in the YOLOv11 backbone network and the neck network. The improved C3K2 module includes adding multiple bottleneck layer branches and configuring a weight generator. The weight generator is used to generate respective weights for each bottleneck layer branch according to current image features. When the image features enter each bottleneck layer branch, they are first multiplied by the weight of the corresponding branch, and then enter the branch. After weighted processing of multiple branches, more effective image features are screened out.

3. According to the cross-age identity authentication method based on the attention mechanism of claim 1, it is characterized in that: In step S3, the face image is preprocessed, specifically including: Data cleaning: Filter out blurred face images or misdetected non-face images from the captured images; Image resizing and pixel normalization: The image size is uniformly scaled to 112×112, and the pixels are uniformly normalized to the range of [-1,1].

4. According to the cross-age identity authentication method based on the attention mechanism of claim 1, it is characterized in that: In step S4, cross-age identity authentication network content: Backbone network: mainly used to extract input facial image features and convert them into facial information feature maps that are a mixture of age-related features and identity-related features; Attention module: extracts features of mixed facial features in channel and spatial dimensions; Multi-task learning module: mainly used to calculate the loss of age prediction, the loss of identity prediction, and the loss of the correlation between the two features of identity and age, and to remove the correlation between the two features.

5. According to the cross-age identity authentication method based on the attention mechanism of claim 4, it is characterized in that: In step S4, the backbone network and the attention module specifically include: The backbone network uses the Resnet-50 network model to extract mixed face feature maps from images; The attention module is divided into a channel attention module and a spatial attention module. The channel attention module adopts a compressed excitation network. The compressed excitation network output feature map includes the following steps: First, a compression operation is performed, and the W×H×C feature map is compressed into 1×1×C using channel global average pooling, where W is the width of the feature map, H is the height of the feature map, and C is the number of channels of the feature map; channel average pooling is as follows: Among them, Z c represents the cth element of the compressed feature vector Z, U c Represents the cth channel of the input feature map U; Then two fully connected layers are used to perform the extraction operation to capture the compressed feature information. The first fully connected layer compresses the C channels of the compressed feature vector Z into C / r channels to reduce the amount of calculation, where r is the compression ratio, and then passes through a ReLU nonlinear activation layer; the second fully connected layer restores the number of channels to C channels, and then passes through the Sigmoid function to obtain a 1×1×C weight S, which is used to characterize the weight of each channel in the input feature map U. The formula for the extraction operation is as follows: S = σ(W2δ(W1z)); Where δ(·) represents the ReLU activation function, σ(·) represents the Sigmoid activation function, and W1 and W2 represent the weights of the two fully connected layers respectively; Finally, a weighted operation is performed to multiply the weight S by the input feature map to obtain the output feature map after channel attention weighting. The weighted process formula is as follows: in Represents the output feature map The cth channel component, U c represents the cth channel of the input feature map U, S c represents the c-th component of weight S.

6. A cross-age identity authentication method based on attention mechanism according to claim 5, characterized in that: The spatial attention module uses a combination of maximum pooling and average pooling to extract spatial hierarchical features: first, the input feature map U of size W×H×C is subjected to global maximum pooling and global average pooling in the channel dimension to obtain two feature maps of size H×W×1. The formulas for the above global maximum pooling and global average pooling are as follows: Z max [i,j]=Max(U[i,j,i],U[i,j,2],...,U[i,j,C]; Among them, Z max [i,j] indicates that the width dimension of the maximum pooling output feature map is i, the height dimension is the element at j, C indicates the number of channels in the input feature map, U[i,j,c] indicates that the width dimension is i, the height dimension is j, and the channel dimension is the element at c in the input feature map, Z avg [i,j] represents the element at the width dimension i and the height dimension j of the average pooling output feature map; Then the feature map Z obtained by maximum pooling and average pooling is max and Z avg By channel splicing, we get a feature map of H×W×2, and perform a convolution operation on the splicing result to get a feature map of H×W×1. Then, we pass the Sigmoid function to get the spatial attention weight matrix M. The formula of this process is as follows: M=σ(f(Z max ;WITH avg )); Among them, σ(·) represents the Sigmoid activation function, f(·) represents the convolution operation; the obtained spatial attention weight matrix M is multiplied with the input feature map U for weighting, and the formula is as follows: in, Represents the weighted output feature map, U[i,j] represents the element at the input feature map with width dimension i and height dimension j, and M[i,j] represents the element at the weight matrix with width dimension i and height dimension j.

7. According to the cross-age identity authentication method based on the attention mechanism of claim 1, it is characterized in that: In step S4, the multi-task learning framework for cross-age identity authentication network training includes the following contents: For the age prediction task, linear regression is used to estimate age, with mean square error as the loss function. The formula is as follows: Among them, L age represents the age prediction loss, F age (.) represents the representation function of the age prediction module, represents the age feature vector of the i-th training set sample, z i The age label of the i-th training set sample; For the identity recognition task, the ArcFace function is used as the loss function of identity recognition. The formula is as follows: Among them, L id represents the identity prediction loss, N is the size of the training batch, that is, the number of input images; k is the number of categories, that is, the number of identities trained; y represents the true identity category label; Represents the true category y i The cosine value of the angle after the offset processing; is the angle of the target category, i.e. the true label; m is the weighted angle, which is used to adjust the interval of the target angle and enhance the distinction between features; s is a scaling factor, which is used to control the learning speed of the model, cosθ j is the cosine similarity between each category and the feature vector; For the correlation loss between age features and identity features, a cross-age domain adversarial learning is introduced. The continuous domain adaptation of the gradient reversal layer is used to encourage the de-correlation of age features and identity features, and the correlation loss L is obtained. cor , the final total loss function formula of the multi-task learning framework is as follows: L=L id +αL age +βL cor ; Among them, α and β are weight parameters that control the balance of different loss terms.

8. According to the cross-age identity authentication method based on the attention mechanism of claim 1, it is characterized in that: In step S4, the cross-age identity authentication network performs identity authentication, including: The face image to be recognized is passed through the backbone network to obtain a mixed feature map X, and the mixed feature map is weighted by the attention module to obtain the age-related feature X age , subtract the age-related features from the mixed feature map to obtain the identity-related features X id , the formula is as follows: X i d =X-X age ; The identity-related feature X id A 512-dimensional feature vector V is obtained through the linear layer, and the feature vector V is input into the background storage device for storage; the cosine similarity between the feature vector V and the other saved feature vectors is calculated one by one, and the formula is as follows: Where n is the total number of feature vectors stored in the background storage device, D is the feature vector stored in the background storage device, V i and D i Represents the i-th component of the feature vector. If the similarity is greater than the threshold, the two are considered to be the same person and the identity authentication result is returned.

9. A system for implementing a cross-age identity authentication method based on an attention mechanism as claimed in any one of claims 1 to 8, characterized in that: include: A data acquisition module is used to collect image data and input the image data into the face detection module; The face detection module extracts the features of the image, cuts out the face image to be identified according to the face coordinates, and outputs the face image to be identified; The identity authentication module is used to authenticate the preprocessed face image to be identified.

10. A computer device, characterized in that: It comprises: a memory and a processor and a computer program stored in the memory. When the computer program is executed on the processor, a cross-age identity authentication method based on an attention mechanism as described in any one of claims 1 to 8 is implemented.