Cross-Age Face Recognition Method
Through deep learning-based image processing technology and dual-channel contrast learning mechanism, the cross-age face images are processed and analyzed, which solves the problem of the reduction in the accuracy of face recognition technology in cross-age scenarios, and improves the accuracy and robustness of recognition.
Patent Information
- Application Number
- CN202411657825.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-20
AI Technical Summary
The existing facial recognition technology has decreased the recognition accuracy in cross-age scenarios, making it difficult to effectively deal with changes in facial details and texture caused by skin sagging and changes in skin tone.
The image processing technology based on deep learning is used to process the face images across ages, and the facial structure feature representation is mined, and the age gap is predicted through the dual-channel comparison learning mechanism to determine whether the face object is the same object.
It improves the accuracy and robustness of facial recognition, effectively solving the problem of the decrease in recognition accuracy in cross-age scenarios.
Smart Images

Figure CN119206836B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of face recognition technology, and more specifically, to a cross-age face recognition method. Background Art
[0002] Face recognition technology has demonstrated high application value in multiple fields, such as access control security, security permissions, etc. However, over time, the appearance of the face will change, especially in the facial detail texture part, with problems such as skin laxity and skin color change.
[0003] In the existing face recognition technology, training face classification is the main means, expecting to learn the similarities and differences in the face distribution during the training process, so that the model has the ability to cluster the same faces and distinguish different faces. When facing cross-age face recognition, due to the large differences in face data under the labels, it is difficult for the model to be optimized in the desired direction during the learning process, resulting in a decrease in the face recognition accuracy rate during application.
[0004] Therefore, a cross-age face recognition method is expected. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. The embodiments of this application provide a cross-age face recognition method, which uses deep learning-based image processing technology to process two cross-age face images, extracts the face structure feature representations of both, and based on this, predicts their ages. At the same time, a differential comparison analysis of the face structure features of the cross-age face images is carried out, and further, the age gap is predicted based on the face difference features. Thus, based on the prediction results of the two channels, it is determined whether the face objects in the cross-age face images are the same object. Through this two-channel contrast learning mechanism, the problem of the decrease in the recognition accuracy rate of traditional face recognition technology in cross-age scenarios can be effectively solved, and the accuracy and robustness of face recognition can be improved.
[0006] Correspondingly, according to one aspect of this application, a cross-age face recognition method is provided, which includes:
[0007] Obtain a first face image to be recognized and a second face image to be recognized;
[0008] Perform age prediction and difference calculation on the first face image to be recognized and the second face image to be recognized to obtain an age gap anchor value;
[0009] Perform face feature difference measurement on the first face image to be recognized and the second face image to be recognized to obtain a first-second face semantic differential feature map of the objects to be recognized;
[0010] Perform attention enhancement on the face semantic differential feature maps of the first and second objects to be recognized to obtain face semantic differential enhanced feature maps of the first and second objects to be recognized;
[0011] Generate an age difference decoding value based on face semantics based on the face semantic differential enhanced feature maps of the first and second objects to be recognized;
[0012] In response to the difference between the age difference decoding value based on face semantics and the age gap anchor value being within a preset range, determine that the objects in the first face image to be recognized and the second face image to be recognized are the same object.
[0013] In the above cross-age face recognition method, performing attention enhancement on the face semantic differential feature maps of the first and second objects to be recognized includes: performing adaptive attenuation modulation based on spatial distribution guidance on the pixel-level semantic association structure of the face semantic differential feature maps of the first and second objects to be recognized to obtain a semantic association weight map between spatially attenuated pixel granularities of the face semantic differential feature space of the object to be recognized; performing feature modulation on the face semantic differential feature maps of the first and second objects to be recognized based on the semantic association weight map between spatially attenuated pixel granularities of the face semantic differential feature space of the object to be recognized to obtain the face semantic differential enhanced feature maps of the first and second objects to be recognized.
[0014] In the above cross-age face recognition method, performing age prediction and difference calculation on the first face image to be recognized and the second face image to be recognized to obtain an age gap anchor value includes: respectively inputting the first face image to be recognized and the second face image to be recognized into a face feature extractor based on a dilated convolutional neural network model to obtain a face feature map of the first object to be recognized and a face feature map of the second object to be recognized; respectively inputting the face feature map of the first object to be recognized and the face feature map of the second object to be recognized into an age regression predictor based on a decoder to obtain a first age prediction value and a second age prediction value; calculating the difference between the first age prediction value and the second age prediction value to obtain the age gap anchor value.
[0015] In the above cross-age face recognition method, performing face feature difference measurement on the first face image to be recognized and the second face image to be recognized to obtain face semantic differential feature maps of the first and second objects to be recognized includes: inputting the face feature map of the first object to be recognized and the face feature map of the second object to be recognized into a face semantic difference feature measurement network to obtain the face semantic differential feature maps of the first and second objects to be recognized, where the face semantic difference feature measurement network is used to calculate the position-wise difference between the face feature map of the first object to be recognized and the face feature map of the second object to be recognized to obtain the face semantic differential feature maps of the first and second objects to be recognized.
[0016] In the above cross - age face recognition method, an adaptive attenuation modulation based on spatial distribution guidance is performed on the pixel - level semantic association structure of the first - second face semantic differential feature maps of the objects to be recognized to obtain a semantic association weight map between the spatially attenuated pixel granularities of the face semantic differential feature space of the objects to be recognized, including: performing feature fine - grained spatial decoupling along the channel dimension on the first - second face semantic differential feature maps of the objects to be recognized to obtain a set of face pixel granularity semantic differential feature vectors of the objects to be recognized; inputting any two face pixel granularity semantic differential feature vectors in the set of face pixel granularity semantic differential feature vectors of the objects to be recognized into a semantic association score metric network to obtain a set of semantic association score vectors between the differential feature pixels of the face of the objects to be recognized; performing semantic association adaptive attenuation modulation on the set of semantic association score vectors between the differential feature pixels of the face of the objects to be recognized based on the spatial span between any two face pixel granularity semantic differential feature vectors in the set of face pixel granularity semantic differential feature vectors of the objects to be recognized to obtain a set of semantic association score vectors between the spatially attenuated pixel granularities of the face differential features of the objects to be recognized; performing dimension reconstruction and weight assignment on the set of semantic association score vectors between the spatially attenuated pixel granularities of the face differential features of the objects to be recognized to obtain the semantic association weight map between the spatially attenuated pixel granularities of the face semantic differential feature space of the objects to be recognized.
[0017] In the above cross - age face recognition method, inputting any two face pixel granularity semantic differential feature vectors in the set of face pixel granularity semantic differential feature vectors of the objects to be recognized into a semantic association score metric network to obtain a set of semantic association score vectors between the differential feature pixels of the face of the objects to be recognized, including: concatenating and fusing any two face pixel granularity semantic differential feature vectors in the set of face pixel granularity semantic differential feature vectors of the objects to be recognized, multiplying by a weight parameter matrix, and then performing a dot product with a bias vector to obtain the set of semantic association score vectors between the differential feature pixels of the face of the objects to be recognized.
[0018] In the above cross-age face recognition method, semantic association adaptive attenuation modulation is performed on the set of semantic association score vectors between semantic differential feature pixels of the face of the object to be recognized based on the spatial span between any two semantic differential feature vectors of the face pixels of the object to be recognized at the pixel granularity, so as to obtain a set of semantic association score vectors between spatially attenuated pixels of the face differential features of the object to be recognized, including: performing spatial position annotation on each semantic differential feature vector of the face pixels of the object to be recognized at the pixel granularity to obtain a set of spatial position data of the face differential features of the object to be recognized; calculating the Euclidean distance between the spatial position data of any two semantic differential feature vectors of the face pixels of the object to be recognized at the pixel granularity as the spatial span value of the two semantic differential feature vectors of the face pixels of the object to be recognized to obtain a set of spatial span values of the face differential features of the object to be recognized at the pixel granularity; based on each spatial span value of the face differential features of the object to be recognized at the pixel granularity in the set of spatial span values of the face differential features of the object to be recognized at the pixel granularity, performing semantic association adaptive attenuation modulation on each semantic association score vector between semantic differential feature pixels of the face of the object to be recognized at the pixel granularity in the set of semantic association score vectors between semantic differential feature pixels of the face of the object to be recognized to obtain the set of semantic association score vectors between spatially attenuated pixels of the face differential features of the object to be recognized.
[0019] In the above cross-age face recognition method, dimension reconstruction and weight assignment are performed on the set of semantic association score vectors between spatially attenuated pixels of the face differential features of the object to be recognized to obtain a semantic association weight map between spatially attenuated pixels of the semantic differential features of the face of the object to be recognized, including: after arranging the set of semantic association score vectors between spatially attenuated pixels of the face differential features of the object to be recognized into a semantic association score feature map between spatially attenuated pixels of the face differential space of the object to be recognized, inputting it into a feature dimension modulation layer based on a point convolutional layer and a weight assignment network based on a Sigmoid function to obtain the semantic association weight map between spatially attenuated pixels of the semantic differential features of the face of the object to be recognized.
[0020] In the above cross-age face recognition method, feature modulation is performed on the first-second semantic differential feature map of the face of the object to be recognized based on the semantic association weight map between spatially attenuated pixels of the semantic differential features of the face of the object to be recognized to obtain a first-second semantic differential enhanced feature map of the face of the object to be recognized, including: calculating the point-by-position multiplication between the semantic association weight map between spatially attenuated pixels of the semantic differential features of the face of the object to be recognized and the first-second semantic differential feature map of the face of the object to be recognized to obtain the first-second semantic differential enhanced feature map of the face of the object to be recognized.
[0021] In the above cross-age face recognition method, based on the first and second face semantic differential enhanced feature maps of the objects to be recognized, generating an age difference decoding value based on face semantics includes: inputting the first and second face semantic differential enhanced feature maps of the objects to be recognized into an age gap regression prediction module based on a decoder to obtain the age difference decoding value based on face semantics.
[0022] Compared with the prior art, the cross-age face recognition method provided by this application uses image processing technology based on deep learning to process two cross-age face images, extracts the face structure feature representations of the two, and predicts their ages based on this. At the same time, it conducts a differential comparison analysis of the face structure features of the cross-age face images, further predicts the age gap based on the face difference features, and thus determines whether the face objects in the cross-age face images are the same object based on the prediction results of the two channels. Through this two-channel contrast learning mechanism, it can effectively solve the problem of the decline in recognition accuracy of traditional face recognition technology in cross-age scenarios, and improve the accuracy and robustness of face recognition. Description of the Drawings
[0023] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. They are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0024] Figure 1 It is a flowchart of the cross-age face recognition method according to an embodiment of the present application.
[0025] Figure 2 It is a schematic diagram of data flow of the cross-age face recognition method according to an embodiment of the present application.
[0026] Figure 3 It is a flowchart of step S120 in the cross-age face recognition method according to an embodiment of the present application.
[0027] Figure 4 It is a flowchart of step S140 in the cross-age face recognition method according to an embodiment of the present application. Detailed Embodiments
[0028] Next, exemplary embodiments according to the present application will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described here.
[0029] Regarding the problem of cross-age face recognition, the present application proposes a cross-age face recognition method based on dual-channel face contrast learning, which can effectively improve the face recognition performance in cross-age scenarios. Specifically, it uses deep learning-based image processing technology to process two cross-age face images, extracts the face structure feature representations of both, and predicts their ages based on this. At the same time, it conducts a differential contrast analysis of the face structure features of the cross-age face images, and further predicts the age gap based on the face difference features. Thus, based on the prediction results of the dual channels, it determines whether the face objects in the cross-age face images are the same object. Through this dual-channel contrast learning mechanism, it can effectively solve the problem of the decline in recognition accuracy of traditional face recognition technology in cross-age scenarios, and improve the accuracy and robustness of face recognition.
[0030] Figure 1 FIG. is a flowchart of the cross-age face recognition method according to an embodiment of the present application. Figure 2 FIG. is a schematic diagram of data flow of the cross-age face recognition method according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the cross-age face recognition method according to an embodiment of the present application includes the steps: S110, obtaining a first face image to be recognized and a second face image to be recognized; S120, performing age prediction and difference calculation on the first face image to be recognized and the second face image to be recognized to obtain an age gap anchor value; S130, performing face feature difference measurement on the first face image to be recognized and the second face image to be recognized to obtain a first-second face semantic differential feature map of the objects to be recognized; S140, performing attention enhancement on the first-second face semantic differential feature map of the objects to be recognized to obtain a first-second face semantic differential enhanced feature map of the objects to be recognized; S150, generating an age difference decoding value based on the face semantics based on the first-second face semantic differential enhanced feature map of the objects to be recognized; S160, in response to the difference between the age difference decoding value based on the face semantics and the age gap anchor value being within a preset range, determining that the objects in the first face image to be recognized and the second face image to be recognized are the same object.
[0031] In the above cross-age face recognition method, in step S110, a first face image to be recognized and a second face image to be recognized are obtained. In the technical solution of this application, in order to ensure that the obtained first face image to be recognized and the second face image to be recognized have a certain age span to meet the requirements of cross-age face recognition, first, a face recognition data set applicable to the cross-age scenario is constructed. Specifically, first, age pseudo-labels are assigned to the face recognition data set. In the face recognition data set, there are n images of the same face under the same face ID. An open-source face age prediction model is used to label the age of each image to obtain n labels, and finally, the average value is used to assign it to the face ID as the age pseudo-label of the face ID, representing the age situation of the face. Next, cross-age face generation is performed on each image with an age pseudo-label. Different age intervals will be set, and 5 age intervals are set for 0-100. If the pseudo-label falls within the interval of 0-20, face generation will be performed in the other four intervals to simulate the face images in the cross-age scenario. At the end of this stage, each face ID has face images in different age intervals. Then, each face ID is logically divided into two subsets according to the age size to ensure that the two pictures selected each time have a certain age span during subsequent model training to adapt to the cross-age face recognition scenario. Based on this, the cross-age scenario data set is made. Then, the first face image to be recognized and the second face image to be recognized are extracted from the above cross-age scenario data set, where the first face image to be recognized and the second face image to be recognized correspond to different subsets respectively to ensure that the two have a certain age span.
[0032] In the above cross-age face recognition method, in step S120, age prediction and difference calculation are performed on the first face image to be recognized and the second face image to be recognized to obtain an age gap anchor value. Among them, Figure 3 FIG. is a flowchart of step S120 in the cross-age face recognition method according to an embodiment of the present application. As Figure 3 shown, step S120 includes: S121, respectively inputting the first face image to be recognized and the second face image to be recognized into a face feature extractor based on a dilated convolutional neural network model to obtain a first face feature map of the object to be recognized and a second face feature map of the object to be recognized; S122, respectively inputting the first face feature map of the object to be recognized and the second face feature map of the object to be recognized into an age regression predictor based on a decoder to obtain a first age prediction value and a second age prediction value; S123, calculating the difference between the first age prediction value and the second age prediction value to obtain the age gap anchor value.
[0033] Specifically, in step S121, the first face image to be identified and the second face image to be identified are respectively input into a face feature extractor based on a dilated convolutional neural network model to obtain a face feature map of the first object to be identified and a face feature map of the second object to be identified. It should be understood that face images of different age groups have different skin textures, facial contours, facial expressions, etc. Therefore, in order to fully learn the facial features in the image and use them as a basis for cross-age face recognition, the present application adopts a dilated convolutional neural network model to construct a face feature extractor, and processes the first face image to be identified and the second face image to be identified respectively. By introducing a dilated convolution operation, the dilated convolutional neural network model can expand the receptive field while maintaining the image resolution, capture richer contextual information, and take into account the detailed features and overall facial structure features in the face image, thereby providing a more accurate feature representation for subsequent cross-age face recognition.
[0034] Specifically, in step S122, the first face feature map of the object to be identified and the second face feature map of the object to be identified are respectively input into the decoder-based age regression predictor to obtain the first age prediction value and the second age prediction value. That is, in order to improve the accuracy of cross-age face recognition, the present application adds the task of face age regression in addition to the face comparison task, and predicts the age of the first face feature map of the object to be identified and the second face feature map of the object to be identified, and uses this as prior information to provide important prior knowledge and reference basis for the subsequent face comparison task. In a specific example of the present application, the age regression predictor adopts a multi-layer perceptron structure, and after flattening the input face feature map, it is nonlinearly transformed layer by layer through the fully connected layer to decouple the age information contained therein, and then outputs the corresponding age prediction value through the output layer. In an embodiment of the present application, during the training process of the decoder-based age regression predictor, the L2 loss function is used to optimize the network to minimize the difference between the predicted age value and the true age value.
[0035] Specifically, in step S123, the difference between the first age prediction value and the second age prediction value is calculated to obtain the age difference anchor value. That is, by calculating the difference between the first age prediction value and the second age prediction value, the age difference between the two objects to be identified is quantitatively represented, thereby providing age information reference for subsequent facial feature comparison analysis. For example, if the age difference between the two images is small, but the feature map difference is large, they may belong to different facial objects; conversely, if the age difference is large, but the feature map difference is small, they may belong to different age groups of the same facial object.
[0036] In the above cross-age face recognition method, in step S130, face feature difference measurement is performed on the first face image to be recognized and the second face image to be recognized to obtain a first-second face semantic differential feature map of the objects to be recognized. In a specific example of the present application, the face feature map of the first object to be recognized and the face feature map of the second object to be recognized are input into a face semantic difference feature measurement network to obtain the first-second face semantic differential feature map of the objects to be recognized. Among them, the face semantic difference feature measurement network is used to calculate the position-by-position difference between the face feature map of the first object to be recognized and the face feature map of the second object to be recognized to reveal the feature differences between the two, such as the position differences and shape differences of facial organs such as eyes and noses, so as to obtain the first-second face semantic differential feature map of the objects to be recognized.
[0037] In the above cross-age face recognition method, in step S140, attention enhancement is performed on the first-second face semantic differential feature map of the objects to be recognized to obtain a first-second face semantic differential enhanced feature map of the objects to be recognized. In particular, considering that in the first-second face semantic differential feature map of the objects to be recognized, different feature regions may have different importance. For example, the feature differences of key facial organs such as eyes and noses are more important for the recognition process. Therefore, in order to improve the model's ability to capture key difference features in the first-second face semantic differential feature map of the objects to be recognized, the present application proposes a feature enhancement method based on spatial context awareness. By learning the pixel-level context semantic association structure of the first-second face semantic differential feature map of the objects to be recognized, it dynamically adjusts its feature weights and performs adaptive focusing on key features, thereby enhancing the model's sensitivity to the feature differences of key facial organs.
[0038] Figure 4 It is a flowchart of step S140 in the cross-age face recognition method according to an embodiment of the present application. As Figure 4 shown, step S140 includes: S141, performing adaptive attenuation modulation based on spatial distribution guidance on the pixel-level semantic association structure of the first-second face semantic differential feature map of the objects to be recognized to obtain a semantic association weight map between spatial attenuation pixel granularities of the face semantic differential feature space of the objects to be recognized; S142, performing feature modulation on the first-second face semantic differential feature map of the objects to be recognized based on the semantic association weight map between spatial attenuation pixel granularities of the face semantic differential feature space of the objects to be recognized to obtain the first-second face semantic differential enhanced feature map of the objects to be recognized.
[0039] Specifically, the step S141 includes: performing feature fine-grained spatial decoupling on the first-second face semantic differential feature map of the object to be recognized along the channel dimension to obtain a set of face pixel granularity semantic differential feature vectors of the object to be recognized; inputting any two face pixel granularity semantic differential feature vectors in the set of face pixel granularity semantic differential feature vectors of the object to be recognized into a semantic association score metric network to obtain a set of semantic association score vectors between face differential feature pixels of the object to be recognized; performing semantic association adaptive attenuation modulation on the set of semantic association score vectors between face differential feature pixels of the object to be recognized based on the spatial span between any two face pixel granularity semantic differential feature vectors in the set of face pixel granularity semantic differential feature vectors of the object to be recognized to obtain a set of semantic association score vectors between spatially attenuated face differential feature pixels of the object to be recognized; performing dimension reconstruction and weight assignment on the set of semantic association score vectors between spatially attenuated face differential feature pixels of the object to be recognized to obtain a semantic association weight map between spatially attenuated face semantic differential features of the object to be recognized.
[0040] In a specific example of the present application, inputting any two face pixel granularity semantic differential feature vectors in the set of face pixel granularity semantic differential feature vectors of the object to be recognized into a semantic association score metric network to obtain a set of semantic association score vectors between face differential feature pixels of the object to be recognized includes: concatenating and fusing any two face pixel granularity semantic differential feature vectors in the set of face pixel granularity semantic differential feature vectors of the object to be recognized, multiplying by a weight parameter matrix, and then performing a dot product with a bias vector to obtain a set of semantic association score vectors between face differential feature pixels of the object to be recognized.
[0041] That is, first, perform fine-grained spatial decoupling on the first-second face semantic differential feature map of the object to be recognized along the channel dimension, decompose it into multiple feature vectors of pixel granularity, and form a set of face pixel granularity semantic differential feature vectors of the object to be recognized, thereby enhancing the model's recognition ability for local detail features. Then, considering that local features belonging to the same facial organ usually have stronger semantic associations, therefore, the present application further estimates the semantic connection strength between the two by performing semantic association measurement on any two face pixel granularity semantic differential feature vectors in the set, and generates a set of semantic association score vectors between face differential feature pixels of the object to be recognized, thereby revealing the semantic association structure in the first-second face semantic differential feature map of the object to be recognized.
[0042] In a specific example of the present application, performing semantic association adaptive attenuation modulation on the set of semantic association score vectors between the differential feature pixel granularities of the face of the object to be recognized includes: performing spatial position annotation on each semantic differential feature vector of the face pixel granularity of the object to be recognized in the set of semantic differential feature vectors of the face pixel granularity of the object to be recognized to obtain a set of spatial position data of the differential feature pixels of the face of the object to be recognized; calculating the Euclidean distance between the spatial position data of any two semantic differential feature vectors of the face pixel granularity of the object to be recognized in the set of spatial position data of the differential feature pixels of the face of the object to be recognized as the spatial span value between the two semantic differential feature vectors of the face pixel granularity of the object to be recognized to obtain a set of spatial span values of the differential feature pixels of the face of the object to be recognized; based on each spatial span value of the differential feature pixels of the face of the object to be recognized in the set of spatial span values of the differential feature pixels of the face of the object to be recognized, performing semantic association adaptive attenuation modulation on each semantic association score vector between the differential feature pixel granularities of the face of the object to be recognized in the set of semantic association score vectors between the differential feature pixel granularities of the face of the object to be recognized to obtain a set of semantic association score vectors between the spatially attenuated pixel granularities of the differential features of the face of the object to be recognized.
[0043] In a specific example of the present application, performing semantic association adaptive attenuation modulation on each semantic association score vector between the differential feature pixel granularities of the face of the object to be recognized in the set of semantic association score vectors between the differential feature pixel granularities of the face of the object to be recognized includes: dividing the spatial span value of the differential feature pixels of the face of the object to be recognized by a preset scaling factor to obtain a spatially scaled span value of the differential feature pixels of the face of the object to be recognized; calculating the exponential function value with e as the base and the opposite of the spatially scaled span value of the differential feature pixels of the face of the object to be recognized as the exponent to obtain a semantic association spatial attenuation coefficient; calculating the semantic association score vector between the differential feature pixel granularities of the face of the object to be recognized divided by the semantic association spatial attenuation coefficient to obtain the semantic association score vector between the spatially attenuated pixel granularities of the differential features of the face of the object to be recognized.
[0044] In a specific example of the present application, performing dimension reconstruction and weight assignment on the set of semantic association score vectors between the spatially attenuated pixel granularities of the differential features of the face of the object to be recognized includes: arranging the set of semantic association score vectors between the spatially attenuated pixel granularities of the differential features of the face of the object to be recognized into a semantic association score feature map between the spatially attenuated pixels of the face of the object to be recognized, and then inputting it into a feature dimension modulation layer based on a point convolution layer and a weight assignment network based on the Sigmoid function to obtain a semantic association weight map between the spatially attenuated pixel granularities of the semantic differential features of the face of the object to be recognized.
[0045] Here, considering that in an image, adjacent pixel points often have more similar semantic information, therefore, the present application further performs spatial position annotation on the semantic differential feature vectors of each face pixel granularity of the objects to be recognized, adds their exact position information in the original feature map, and calculates the spatial span between any two semantic differential feature vectors of the face pixel granularity of the objects to be recognized based on this, so as to reveal the relative position relationship between the local features of each pixel granularity, in order to filter out interfering terms that are seemingly similar but actually irrelevant during the subsequent feature modulation process. Then, an adaptive attenuation modulation is performed on the semantic association score vector between the two based on the spatial span value between any two semantic differential feature vectors of the face pixel granularity of the objects to be recognized. When two local features of pixel granularity show both strong semantic correlation and are close to each other, the importance of both will be maximized; conversely, if they have a high semantic similarity but are far apart, their influence will be appropriately weakened, so as to accurately reflect the most noteworthy feature part in the first-second face semantic differential feature map of the objects to be recognized. Then, after arranging the set of semantic association score vectors between the face differential feature pixels of the objects to be recognized after spatial attenuation modulation into a feature map form, feature dimension modulation and weight assignment are performed on it through point convolution and the Sigmoid function to generate the corresponding semantic differential feature spatial attenuation pixel granularity semantic association weight map of the face of the object to be recognized.
[0046] Specifically, step S142 includes: calculating the first-second face semantic differential enhanced feature map by performing element-wise multiplication between the semantic differential feature spatial attenuation pixel granularity semantic association weight map of the face of the object to be recognized and the first-second face semantic differential feature map of the objects to be recognized. That is, through element-wise multiplication operation, weighted modulation is performed on the first-second face semantic differential feature map to selectively amplify important information and suppress irrelevant background information or non-critical feature regions, highlighting the expression of key face difference features, thereby obtaining the first-second face semantic differential enhanced feature map.
[0047] Correspondingly, step S140 includes: processing the first-second face semantic differential feature map with the following feature enhancement formula to obtain the first-second face semantic differential enhanced feature map, where the feature enhancement formula is:
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061] Among them, is the face semantic difference feature map of the first and second objects to be recognized, and , where , and are the number of channels, height, and width of the face semantic difference feature map of the first and second objects to be recognized respectively, represents the feature decoupling operation, is a set of face pixel granularity semantic difference feature vectors of the object to be recognized, is the face pixel granularity semantic difference feature vector of the object to be recognized, and is the length of the face pixel granularity semantic difference feature vector of the object to be recognized, is the spatial position data of the face pixel granularity semantic difference feature vector of the object to be recognized, where , is the spatial position coordinate of the face pixel granularity semantic difference feature vector of the object to be recognized, represents a set of spatial position data of face difference feature pixels of the object to be recognized, and represent any two face pixel granularity semantic difference feature vectors in the set of face pixel granularity semantic difference feature vectors of the object to be recognized, and represent the weight parameter matrix and the bias vector respectively, represents the semantic association score vector between face difference feature pixels of the object to be recognized, is the set of semantic association score vectors between face difference feature pixels of the object to be recognized, is the spatial span value between the semantic differential feature vectors of the face pixels of any two objects to be recognized, is the set of spatial span values of the face differential feature pixel granularity of the object to be recognized, is the preset scaling factor, is the semantic association score vector between the spatially attenuated pixels of the face differential feature of the object to be recognized, is the set of semantic association score vectors between the spatially attenuated pixels of the face differential feature of the object to be recognized, represents the semantic association score feature map between the spatially attenuated pixels of the face differential space of the object to be recognized, represents the point convolution operation, and 𝐴' is the semantic association score feature map between the spatially attenuated pixels of the face differential space of the object to be recognized after feature dimension modulation, is the Sigmoid function, represents the semantic association weight map between the spatially attenuated pixels of the face semantic differential feature space of the object to be recognized, represents element-wise multiplication by position, represents the first-second face semantic differential enhancement feature map of the object to be recognized.
[0062] In the above cross-age face recognition method, in step S150, based on the first-second face semantic differential enhancement feature map, an age difference decoding value based on face semantics is generated. In a specific example of the present application, step S150 includes: inputting the first-second face semantic differential enhancement feature map into an age gap regression prediction module based on a decoder to obtain the age difference decoding value based on face semantics. In the specific solution of the present application, the age gap regression prediction module based on the decoder also adopts a multi-layer perceptron structure, which performs feature analysis on the first-second face semantic differential enhancement feature map through layer-by-layer non-linear transformation, and maps the face difference information contained therein into an age difference decoding value. By predicting the age gap decoding based on the face semantic difference features, the feature differences between two objects to be recognized are quantitatively represented. In the embodiment of the present application, during the training process of the age gap regression prediction module based on the decoder, the InfoNCE (Information Noise Contrastive Estimation) loss function is used for network optimization.
[0063] Preferably, inputting the first-second face semantic differential enhanced feature map of the object to be recognized into the age gap regression prediction module based on the decoder to obtain the age difference decoding value based on the face semantics includes: determining the number of zero eigenvalues in the first-second face semantic differential enhanced feature map of the object to be recognized, and respectively calculating the reciprocal of the logarithm to the base 2 of the number of zero eigenvalues and the exponential value to the base of the natural constant of the reciprocal of the number of zero eigenvalues to obtain the first face semantic differential representation value of the object to be recognized and the second face semantic differential representation value of the object to be recognized; expanding the first-second face semantic differential enhanced feature map of the object to be recognized into the first-second face semantic differential enhanced feature vector; calculating the power function of each eigenvalue of the first-second face semantic differential enhanced feature vector with the reciprocal of the number of zero eigenvalues as the exponent to obtain the first-second face semantic differential structure state vector; performing dot multiplication on the first-second face semantic differential structure state vector and the first-second face semantic differential representation value to obtain the first-second face semantic differential information representation vector; performing dot multiplication on the autocorrelation matrix of the first-second face semantic differential structure state vector with the second first-second face semantic differential representation value and the weight hyperparameter respectively to obtain the first-second face semantic differential regression understanding matrix; performing matrix multiplication on the first-second face semantic differential information representation vector and the first-second face semantic differential regression understanding matrix to obtain the optimized first-second face semantic differential enhanced feature vector; inputting the optimized first-second face semantic differential enhanced feature vector into the age gap regression prediction module based on the decoder to obtain the age difference decoding value based on the face semantics.
[0064] Here, the first face semantic differential representation value of the object to be recognized and the second face semantic differential representation value of the object to be recognized are expressed as:
[0065]
[0066]
[0067] Among them, is the number of zero eigenvalues in the first-second face semantic differential enhanced feature map of the object to be recognized, is the first face semantic differential representation value of the object to be recognized, the second face semantic differential representation value of the object to be recognized, represents the logarithmic function to the base 2, represents the natural exponential function.
[0068] And the optimized first-second face semantic differential enhanced feature vector of the object to be recognized is expressed as:
[0069]
[0070]
[0071]
[0072] wherein, the first-second face semantic differential enhanced feature vector of the object to be recognized is a row vector, represents a power function with the reciprocal of the number of zero eigenvalues as the exponent for each eigenvalue of the calculated feature vector, represents the transpose of a vector, represents matrix multiplication operation, is a weight hyperparameter, represents the first-second face semantic differential regression understanding matrix of the object to be recognized, represents the first-second face semantic differential information representation vector of the object to be recognized.
[0073] Here, when the first face feature map of the object to be recognized and the second face feature map of the object to be recognized respectively represent the face image semantic encoding features of the first face image to be recognized and the second face image to be recognized, during face semantic difference feature measurement and feature attention enhancement based on spatial modulation, the cross-domain face semantic differential features of the first-second face semantic differential enhanced feature map will have an imbalance in remote class regression understanding representation due to spatial flexible modulation, affecting the accuracy of the decoding result.
[0074] Based on this, the present application realizes effective long-range modeling of the first-second face semantic differential enhanced feature map at the field linear complexity by modeling the zero dimension of the vector field of the first-second face semantic differential enhanced feature vector obtained after unfolding the first-second face semantic differential enhanced feature map as the global field redundancy dependence, so as to perform structured correlation self-distillation between information representation and class understanding in the selective state space based on feature distribution during the feature regression process of the first-second face semantic differential enhanced feature map, thereby promoting the remote dynamic class understanding balance of its feature convergence while maintaining the overall information complexity of the feature distribution of the first-second face semantic differential enhanced feature map, so as to improve the decoding understanding consistency between the decoding regression and the feature extraction process, and improve the accuracy of the face semantic-based age difference decoding value obtained by inputting the first-second face semantic differential enhanced feature map into the age gap regression prediction module based on the decoder.
[0075] In the above cross-age face recognition method, in step S160, in response to the difference between the age difference decoding value based on face semantics and the age gap anchor value being within a preset range, it is determined that the objects in the first face image to be recognized and the second face image to be recognized are the same object. That is, based on the comprehensive judgment of facial feature similarity and age factor verification, the final confirmation of face recognition is completed. For example, when the difference is less than the set threshold, it indicates that the facial difference degree between the two images conforms to the natural change range of the same face object at different age groups, and thus can be confirmed as the same person. On the contrary, if the difference exceeds the threshold, it indicates that the facial difference degree between the two images exceeds the natural change range of the same face object at different age groups, and therefore can be determined as different individuals.
[0076] In summary, the cross-age face recognition method according to the embodiments of the present application is elucidated. It uses image processing technology based on deep learning to process two cross-age face images, extracts the face structure feature representations of both, and predicts their ages based on this. At the same time, it conducts a differential comparison analysis of the face structure features of the cross-age face images, further predicts the age gap based on the face difference features, and thus determines whether the face objects in the cross-age face images are the same object based on the prediction results of the two channels. Through this two-channel contrast learning mechanism, the problem of the decline in recognition accuracy of traditional face recognition technology in cross-age scenarios can be effectively solved, and the accuracy and robustness of face recognition can be improved.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A cross-age face recognition method, characterized in that: include: Acquire a first face image to be recognized and a second face image to be recognized; Performing age prediction and difference calculation on the first face image to be recognized and the second face image to be recognized to obtain an age difference anchor value; Performing facial feature difference measurement on the first to-be-recognized face image and the second to-be-recognized face image to obtain a first-second to-be-recognized face semantic difference feature map; Performing attention enhancement on the first and second face semantic difference feature maps of the objects to be identified to obtain the first and second face semantic difference enhanced feature maps of the objects to be identified; Based on the semantic difference enhanced feature map of the first and second objects to be identified, generating an age difference decoding value based on facial semantics; In response to a difference between the age difference decoding value based on facial semantics and the age difference anchor value being within a preset range, determining that the objects in the first to-be-recognized face image and the second to-be-recognized face image are the same object; Among them, the attention is strengthened on the first and second semantic differential feature maps of the faces of the objects to be identified, including: performing fine-grained spatial decoupling of the features of the first and second semantic differential feature maps of the faces of the objects to be identified along the channel dimension to obtain a set of semantic differential feature vectors of the face pixel granularity of the objects to be identified; inputting any two semantic differential feature vectors of the face pixel granularity of the objects to be identified in the set of semantic differential feature vectors of the face pixel granularity of the objects to be identified into a semantic association score measurement network to obtain a set of semantic association score vectors between the pixel granularity of the face differential features of the objects to be identified; based on the semantic differential feature vectors of the face pixel granularity of the objects to be identified, the semantic association score measurement network is used to measure the semantic association score between the pixel granularity of the face differential features of the objects to be identified. The spatial span performs semantic association adaptive attenuation modulation on the set of semantic association score vectors between pixel granularities of face differential features of the to-be-recognized object to obtain a set of semantic association score vectors between spatial attenuation pixel granularities of face differential features of the to-be-recognized object; the set of semantic association score vectors between spatial attenuation pixel granularities of face differential features of the to-be-recognized object is dimensional reconstructed and weighted to obtain a semantic association weight map between spatial attenuation pixel granularities of semantic differential features of the to-be-recognized object; based on the semantic association weight map between spatial attenuation pixel granularities of semantic differential features of the to-be-recognized object, feature modulation is performed on the first and second semantic differential feature maps of faces of the to-be-recognized object to obtain the first and second semantic differential enhanced feature maps of faces of the to-be-recognized object; The method of generating an age difference decoding value based on facial semantics based on the first and second facial semantics difference enhanced feature maps of the objects to be identified comprises: The first and second facial semantic difference enhanced feature maps of the objects to be identified are input into the decoder-based age difference regression prediction module to obtain the facial semantic-based age difference decoding value.
2. The cross-age face recognition method according to claim 1, characterized in that: Performing age prediction and difference calculation on the first face image to be recognized and the second face image to be recognized to obtain an age difference anchor value, including: Inputting the first face image to be recognized and the second face image to be recognized into a face feature extractor based on a dilated convolutional neural network model to obtain a first face feature map of an object to be recognized and a second face feature map of an object to be recognized; Inputting the first face feature map of the object to be identified and the second face feature map of the object to be identified into the decoder-based age regression predictor to obtain a first age prediction value and a second age prediction value; The difference between the first age prediction value and the second age prediction value is calculated to obtain the age difference anchor value.
3. The cross-age face recognition method according to claim 2, characterized in that: Measuring the difference of facial features of the first to-be-recognized facial image and the second to-be-recognized facial image to obtain a semantic difference feature map of the first to-be-recognized object's facial features, including: The first face feature map of the object to be identified and the second face feature map of the object to be identified are input into a face semantic difference feature measurement network to obtain the first-second face semantic difference feature map of the object to be identified, wherein the face semantic difference feature measurement network is used to calculate the positional difference between the first face feature map of the object to be identified and the second face feature map of the object to be identified to obtain the first-second face semantic difference feature map of the object to be identified.
4. The cross-age face recognition method according to claim 3, characterized in that: Inputting any two semantic difference feature vectors of the face pixel granularity of the object to be identified from the set of semantic difference feature vectors of the face pixel granularity of the object to be identified into the semantic association score measurement network to obtain a set of semantic association score vectors between the face difference feature pixel granularity of the object to be identified, including: Any two semantic differential feature vectors of the pixel granularity of the faces of the objects to be identified in the set of semantic differential feature vectors of the pixel granularity of the faces of the objects to be identified are cascaded and fused, multiplied by a weight parameter matrix, and then dotted with a bias vector to obtain a set of semantic association score vectors between the pixel granularity of the differential features of the objects to be identified.
5. The cross-age face recognition method according to claim 4, characterized in that: Based on the spatial span between any two semantic differential feature vectors of the pixel granularity of the face of the object to be identified in the set of the semantic differential feature vectors of the pixel granularity of the face of the object to be identified, the set of semantic association score vectors between pixel granularity of the face differential feature of the object to be identified is subjected to semantic association adaptive attenuation modulation to obtain a set of semantic association score vectors between pixel granularity of the face differential feature of the object to be identified, including: Performing spatial position annotation on each of the face pixel granularity semantic differential feature vectors of the to-be-identified object in the set of face pixel granularity semantic differential feature vectors of the to-be-identified object to obtain a set of face differential feature pixel spatial position data of the to-be-identified object; Calculating the Euclidean distance between the spatial position data of any two face pixel granularity semantic differential feature vectors of the objects to be identified in the set of face differential feature pixel spatial position data of the objects to be identified as the spatial span value of the two face pixel granularity semantic differential feature vectors of the objects to be identified to obtain a set of face differential feature pixel granularity spatial span values of the objects to be identified; Based on each face differential feature pixel granularity spatial span value of the object to be identified in the set of face differential feature pixel granularity spatial span values of the object to be identified, semantic association adaptive attenuation modulation is performed on each face differential feature pixel granularity semantic association score vector of the object to be identified in the set of face differential feature pixel granularity semantic association score vectors to obtain the set of face differential feature spatial attenuated pixel granularity semantic association score vectors of the object to be identified.
6. The cross-age face recognition method according to claim 5, characterized in that: The set of semantic association score vectors of the space attenuation pixel granularity of the face difference feature of the object to be identified is reconstructed in dimension and weighted to obtain a semantic association weight map of the space attenuation pixel granularity of the face semantic difference feature of the object to be identified, including: After arranging the set of semantic association score vectors between spatial attenuation pixel granularities of the face differential features of the object to be identified into a semantic association score feature map between spatial attenuation pixel granularities of the face differential features of the object to be identified, it is input into a feature dimension modulation layer based on a point convolution layer and a weight allocation network based on a Sigmoid function to obtain a semantic association weight map between spatial attenuation pixel granularities of the semantic differential features of the object to be identified.
7. The cross-age face recognition method according to claim 6, characterized in that: The method comprises: performing feature modulation on the first and second to-be-recognized object face semantic difference feature maps based on the semantic association weight map between the attenuated pixel granularity of the face semantic difference feature space of the to-be-recognized object to obtain the first and second to-be-recognized object face semantic difference enhanced feature maps, including: Calculate the semantic association weight map between the spatial attenuation pixel granularity of the semantic difference feature of the face of the object to be identified and multiply the first and second semantic difference feature maps of the faces of the objects to be identified by the position points to obtain the first and second semantic difference enhanced feature maps of the faces of the objects to be identified.
Citation Information
Patent Citations
Cross-age face recognition method, system and device and storage medium
CN111881721A
Age recognition model training method, face age recognition method and related device
CN113065525A