A nonparametric face age estimation method based on feature aggregation
Patent Information
- Application Number
- CN202410597428.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-14
AI Technical Summary
但通过人脸图像准确确定年龄仍然是一项具有挑战性的任务,其中一个问题就是面部特征空间分布不均匀,同一年龄不同个体的面部外貌存在很大的差异
[0043]采用上述技术方案所产生的有益效果在于:本发明提供的基于特征聚合的非标态人脸年龄估计方法,使用残差网络和图注意力网络进行特征聚合,捕捉人脸图像年龄特征之间的内在相关性,解决面部特征空间分布不均匀问题;另外通过人脸对齐、亮度归一化和光照增强方法解决了实际应用场景中人脸姿态不标准、光照不均匀的难题,在未成年网络防沉迷方面具有极高的应用价值。
Smart Images

Figure CN118486058B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biometric recognition technology, and in particular to a non-standard face age estimation method based on feature aggregation. Background Technology
[0002] Facial features convey important, perceptible information related to individual characteristics, such as personal identity, facial expressions, gender, and age. Currently, with the widespread use of electronic products, an increasing number of minors are becoming addicted to the internet, making the management of internet addiction among minors a pressing issue. Automatically estimating age on the device to control internet addiction among minors has significant social implications for protecting their healthy development.
[0003] During human development, facial shape, texture, skin color, blemishes, skin laxity, and hairline all change to varying degrees. These changes can be used to extract facial age features for age estimation. With the development of deep learning, most current age estimation methods use various convolutional neural networks to extract features from facial images and then estimate the age value through classification or regression. However, accurately determining age from facial images remains a challenging task. One problem is the uneven distribution of facial features, resulting in significant differences in facial appearance between individuals of the same age. Additionally, end-users may be lying flat or on their side, causing the facial images captured by the terminal camera to be not in a standard position and potentially resulting in low brightness. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to address the shortcomings of the prior art by providing a non-standard face age estimation method based on feature aggregation. This method uses a graph attention network to aggregate age-related face image features, captures the intrinsic correlation between features, and thus estimates the age of the end user.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0006] A non-standard face age estimation method based on feature aggregation specifically includes the following steps:
[0007] Step 1: Image data preprocessing;
[0008] Acquire face images, use a face detection model to detect faces in the images and perform face alignment, and then perform brightness normalization; for images with low brightness, use a lighting enhancement algorithm to enhance the lighting.
[0009] Step 2: Extract features from the image;
[0010] We use ResNet, a convolutional neural network, as the feature extraction model. During training, we use Adaptive Mean-Residue Loss as the loss function to improve the robustness of the model in the age estimation task through distribution learning.
[0011] Step 3: Construct a graph based on the extracted features;
[0012] Features before the ResNet fully connected layer are extracted from all training images. The features of one image correspond to a node in the graph. For each training image, the output is obtained after passing through the complete ResNet. The output is then input into the softmax layer to obtain the predicted probability for each label of the image. The predicted probability values are sorted in descending order, and the classifications corresponding to the top five probability values are taken as the prediction set for that node. The number of intersections of the prediction sets between any two nodes is used as the weight of the edge between the two nodes, thereby constructing the adjacency matrix of the graph.
[0013] Step 4: Based on the graph attention network, perform feature aggregation on the nodes of the graph obtained in Step 3, capture the intrinsic correlation between features, then classify the nodes, aggregate age-related features, and save the feature corresponding to the node with the largest degree under each label as a prototype.
[0014] Step 5: For the test image, first extract the features before the fully connected layer using the trained ResNet, calculate the probability of each label based on the cosine similarity between the feature and the prototype of each label, and obtain the final predicted age value.
[0015] Furthermore, the face detection model in step 1 specifically uses a three-layer cascaded network, referred to as the P-network, R-network, and O-network. The P-network quickly generates coarse candidate boxes, the R-network filters to obtain high-precision candidate boxes, and the O-network generates bounding boxes and key point coordinates. After the face image is cropped out through the three-layer cascaded network, it is uniformly sized to 256×256 and its brightness is normalized.
[0016] Furthermore, in the three-layer cascaded network, the P-network includes three convolutional layers, which ultimately produce three results: a 1×1×2 face classification result, a 1×1×4 bounding box regression result, and a 1×1×10 face keypoint location. For an image of size H×W×3, after passing through the P-network, a 5-channel S×S feature map is output. That is, after the sliding window operation of the P-network, the image yields S×S proposal boxes, each corresponding to one confidence score and four offsets. Then, the NMS algorithm is used to retain the proposal boxes corresponding to the confidence scores greater than the set threshold of 0.6, and their corresponding bounding box offsets are subjected to bounding box regression to obtain the coordinate information in the original image.
[0017] Image data is obtained from the original image based on the bounding box coordinates of the P-network and input into the second layer of the face detection model, the R-network. The R-network adds a fully connected layer with 128 neurons on top of the three convolutional layers to filter out more errors and remove data with low scores.
[0018] The face candidate boxes output by the R-network are input into the third layer of the face detection model, the O-network, for further detail processing. The O-network consists of four convolutional layers. The fourth convolutional layer collects more face features and removes incorrect candidate boxes. At this point, the output branch includes three categories: whether there is a face in the image; the horizontal and vertical coordinates of the starting point or center point of the regressed box and the length and width of the bounding box; and information on five key points of the face, including the position of the left eye, the position of the right eye, the position of the nose, the left position of the mouth, and the right position of the mouth. Each key position is represented using two dimensions.
[0019] Then, based on the obtained facial landmark locations and the manually defined standard facial landmark positions, an affine transformation is performed on the outlined facial image, and the aligned facial image is adjusted to a size of 256×256.
[0020] Furthermore, the illumination enhancement algorithm first decomposes the input image into sub-images of different scales, and then performs logarithmic transformation and Gaussian blur on each sub-image at each scale to obtain a logarithmic domain image L(x,y); for each scale of the logarithmic domain image, the standard deviation of the surrounding pixels is calculated to determine the gain coefficient of that pixel, i.e.: Among them G i,j σ represents the gain coefficient of pixel (i,j). i,j Let L' represent the standard deviation of the pixels surrounding pixel (i,j), where k is a constant used to control the magnitude of the gain coefficient; then, multiplicative enhancement is performed on the logarithmic domain image at each scale, i.e.: L′ i,j =G i,j *L i,j , where L′ i,j For the enhanced pixel value, L i,j The values are the pixel values of the original logarithmic domain image. Finally, the enhanced logarithmic domain image at each scale is subjected to an inverse logarithmic transformation to obtain the enhanced sub-image. All the enhanced sub-images are then combined to obtain the final enhanced image.
[0021] Furthermore, in step 2, the loss function, Adaptive Mean-Residue Loss, is specifically:
[0022] L = L s +λ1L m+λ2L r
[0023] Among them, L s The loss is softmax. It is the average loss, representing the variance between the mean of the age distribution and the true age; The residual loss represents the tail residual error in the predicted age distribution after the top-K operation. It is calculated by sorting the predicted probability values in descending order, taking the labels corresponding to the top K values, and calculating the loss function for the predicted values of the other NK labels.
[0024] Furthermore, the graph attention network in step 4 is composed of stacked graph attention layers. For a single graph attention layer, the input is a set of node features, h = {h1, h2, ..., h...} N},h i ∈R F Where N is the number of nodes and F is the number of features for each node, a graph attention layer generates a new set of node features h′={h'1,h'2,…,h' N},h' i ∈R F′ ;
[0025] The calculation of the graph attention layer is divided into two steps:
[0026] First, calculate the attention coefficient; for vertex i, calculate the attention coefficient for itself and its neighboring nodes j∈N one by one. i similarity coefficient e ij e ij =a([Wh i ||Wh j ]),j∈N i , where N i Wh represents all neighboring nodes of vertex i. i This indicates that the shared W vector learned by the model itself is used to perform a dimensionality transformation on the original feature vector. [·||·] indicates that the transformed features of nodes i and j are concatenated. a(·) indicates that the concatenated high-dimensional features are mapped to a real number. After obtaining the correlation coefficient, it is normalized to obtain the attention coefficient.
[0027]
[0028] Then, the features of the neighboring nodes are weighted and summed based on the calculated attention coefficients:
[0029]
[0030] Where, h′ i The output of each node i in the attention layer incorporates new features from its neighboring nodes, and σ(·) is the activation function.
[0031] To improve the performance of the aggregator, the graph attention network uses a multi-head attention mechanism, that is, it uses K independent attention mechanisms, i.e., different values of a and W, and then concatenates the results:
[0032]
[0033] Here, || represents the concatenation operation. This represents the weight coefficients calculated by the k-th attention mechanism, but it will lead to h′ i Since it has a higher dimension (1, kh'), the multi-head attention in the intermediate layer is spliced together, while the aggregation method of the output layer is to average the results obtained from the K attention mechanisms:
[0034]
[0035] Then, for each label, calculate the node with the highest degree under that label, and save the features output by that node as the prototype proto for that label. i 'i' represents the corresponding tag value, proto i It is a one-dimensional vector with a channel size of C.
[0036] Furthermore, the final age prediction of the test image in step 5 specifically involves:
[0037] First, the features f before the fully connected layers are extracted using the trained ResNet. Feature f is a one-dimensional vector with channel size C. Then, the cosine similarity s between each feature and the prototype of each saved label is calculated. i :
[0038]
[0039] Where i represents the corresponding label value; based on the calculated cosine similarity, the probability a that the age of the test image might be labeled i is calculated. i :
[0040]
[0041] The predicted age of the test image is:
[0042]
[0043] The beneficial effects of adopting the above technical solution are as follows: The non-standard face age estimation method based on feature aggregation provided by this invention uses residual networks and graph attention networks for feature aggregation, captures the intrinsic correlation between age features of face images, and solves the problem of uneven spatial distribution of facial features; in addition, it solves the problems of non-standard face pose and uneven illumination in practical application scenarios through face alignment, brightness normalization and illumination enhancement methods, and has extremely high application value in preventing minors from becoming addicted to the Internet. Attached Figure Description
[0044] Figure 1 This is a diagram illustrating the overall architecture of the non-standard face age estimation method based on feature aggregation provided in this embodiment of the invention.
[0045] Figure 2 The image preprocessing effect diagram provided in the embodiment of the present invention;
[0046] Figure 3 A visualization of the training set composition diagram provided in an embodiment of the present invention;
[0047] Figure 4 A schematic diagram illustrating the attention mechanism provided in an embodiment of the present invention;
[0048] Figure 5 A visualization of facial age features provided for an embodiment of the present invention. Detailed Implementation
[0049] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0050] Facial age estimation faces the problem of uneven facial feature space, where individuals of the same age have significant differences in facial appearance. Furthermore, in real-world scenarios, facial image pose and lighting are not in a standard state. This embodiment first aligns the faces of the images and standardizes their size, then normalizes the brightness and enhances the lighting for images with excessively low light. The preprocessed facial images are then input into a convolutional neural network for feature extraction. The extracted features are then used to construct a graph attention network to train a graph attention network, capturing the intrinsic correlation between age features in facial images. During prediction, the features of the test image are input into the graph attention network to obtain the final age features, thus completing the facial age estimation.
[0051] In this embodiment, the software environment is a Windows 10 system, and the simulation environment is PyCharm 2022.3.2 x64. The collection and use of personal information and privacy information (including facial recognition data) in this invention comply with relevant national laws and regulations.
[0052] In this embodiment, the overall architecture for facial age recognition is designed as follows: Figure 1 As shown in the diagram, according to the architecture diagram, a non-standard face age estimation method based on feature aggregation includes the following steps:
[0053] Step 1: Image Data Preprocessing; Acquire face images, use a face detection model to detect faces in the images and perform face alignment, then normalize the brightness of the obtained images. For low-brightness images, use the MSRCR algorithm for illumination enhancement, such as... Figure 2 As shown.
[0054] The face detection model consists of three layers: P-network, R-network, and O-network. First, the face image is input into the first layer, the P-network.
[0055] The P-network consists of three convolutional layers, ultimately producing three results: a 1×1×2 face classification result, a 1×1×4 bounding box regression result, and a 1×1×10 facial landmark location result. For example, an image of size H×W×3, after passing through the P-network, outputs a 5-channel F×F feature map. This means that the image, after the sliding window operation of the P-network, yields F×F proposal boxes, each corresponding to one confidence score and four offsets. Then, using Non-Maximum Suppression (NMS), the proposal boxes with confidence scores greater than a set threshold of 0.6 are retained, and their corresponding bounding box offsets are used for bounding box regression to obtain their coordinate information in the original image.
[0056] Image data is obtained from the original image based on the bounding box coordinates of the P-network and then input into the second layer, the R-network. The R-network adds a fully connected layer with 128 neurons to the three convolutional layers, which can filter out more erroneous information and remove data with low scores.
[0057] The face candidate boxes output by the R-network are input into the third layer, the O-network, for further detail processing. The O-network contains four convolutional layers. The fourth convolutional layer collects more face features and removes incorrect candidate boxes. At this point, the output branch contains three categories: whether there is a face in the image, the horizontal and vertical coordinates of the starting point (or center point) of the face box and the length and width of the border, and information on five key points of the face (left eye position, right eye position, nose position, left mouth position, and right mouth position). Each key position is represented using two dimensions.
[0058] Based on the obtained facial landmark localization and manually defined standard facial landmark positions, an affine transformation is performed on the outlined face image, and the aligned face image is resized to 256×256. Considering that illumination intensity affects the accuracy of facial age estimation, brightness normalization is performed on each RGB value of each pixel in the image. For low-brightness images, the MSRCR algorithm is used for illumination enhancement.
[0059] The MSRCR algorithm is based on Retinex theory, which posits that an image consists of an illumination image and a reflectance image. The illumination image refers to the information of the incident component of an object, denoted by L(x,y), while the reflectance image refers to the reflected portion of the object, denoted by R(x,y). The final image observed or received by the camera is denoted by I(x,y), where I(x,y) = R(x,y)·L(x,y). Illumination enhancement aims to obtain the reflectance component R(x,y). Since the logarithmic form most closely resembles the human perception of brightness, the above process is transformed into the logarithmic domain, thus converting complex multiplication into addition: log[I(x,y)] = log[R(x,y)] + log[L(x,y)].
[0060] The MSRCR algorithm first decomposes the input image into sub-images of different scales, and then performs a logarithmic transformation on each scale sub-image to obtain a logarithmic domain image. For each scale's logarithmic domain image, the standard deviation of all surrounding pixels is calculated to determine the gain coefficient for that pixel, i.e.: Among them G i,j σ represents the gain coefficient of pixel (i,j). i,j Let L' represent the standard deviation of pixels surrounding pixel (i,j), and k be a constant used to control the magnitude of the gain coefficient; then, multiplicative enhancement is performed on the logarithmic domain image at each scale, i.e.: L′ i,j =G i,j *L i,j , where L′ i,j For the enhanced pixel value, L i,j The values are the pixel values of the original logarithmic domain image. Finally, the enhanced logarithmic domain image at each scale is subjected to an inverse logarithmic transformation to obtain the enhanced sub-image. Finally, all the enhanced sub-images are synthesized to obtain the enhanced image.
[0061] Step 2: Extract features from the image; use ResNet as the feature extraction model, and during training use Adaptive Mean-Residue Loss (L = L s +λ1L m +λ2L r ) as the loss function, where L sThe commonly used softmax loss is used. It is the average loss, representing the variance between the predicted mean and the actual age of the age distribution. The residual loss represents the tail residual error in the predicted age distribution after the top-K operation. It is calculated by sorting the predicted probability values in descending order, taking the labels corresponding to the top K values, and calculating the loss function for the other (NK) labels.
[0062] Step 3: Construct a graph based on the extracted features; extract features from all training images before the ResNet fully connected layer. Each image's feature corresponds to a node in the graph, and each feature has a dimension of (1, num_fc). For each image, pass through the complete ResNet to obtain the output. Input the output into the softmax layer to obtain the predicted probability for each label of the image, with a dimension of (1, num_classes). Sort the obtained probability values in descending order and take the classification corresponding to the top five probability values as the prediction set for that node. The number of elements in the intersection of the prediction sets between any two nodes is used as the weight of the edge between those two nodes. This is used to construct the adjacency matrix of the graph. The visualization is shown below. Figure 3 As shown. For example, if num_classes = 6, the corresponding label value is (1,2,3,4,5,6). The prediction set obtained by node i after passing through the ResNet and softmax layers is {2,3,4,5,6}, and the prediction set of node j is {5,6,7,8,9}. The intersection of the prediction sets of node i and node j is {5,6}, and the number of elements in the intersection is 2. Therefore, there is an edge between node i and node j, and the weight of the edge is 2.
[0063] Step 4: Feature aggregation based on graph attention network; Since the task of this invention belongs to inductive node classification, that is, the nodes in the test set are unknown and have not appeared in the training set, a graph attention network is used to aggregate node features. The graph attention network is composed of stacked graph attention layers. For a single graph attention layer, the input is a set of node features, h = {h1, h2, ..., h...} N},h i ∈R F Where N is the number of nodes and F is the number of features for each node, a graph attention layer generates a new set of node features h′={h′1,h′2,…,h′ N},h′ i ∈R F′ The core idea of graph attention layers is to learn the attention weights of a node and its neighbors for each node, and then perform a weighted average of the features of each node based on these attention weights to obtain a new feature representation for that node, such as... Figure 4 As shown. Therefore, the calculation of the graph attention layer involves two steps:
[0064] First, calculate the attention coefficient; for node i, calculate the attention coefficient for itself and its neighboring nodes j∈N one by one. i (N i The similarity coefficient e of all neighboring nodes of vertex i ij e ij =a([Wh i ||Wh j ]),j∈N i , of which Wh i This indicates that the shared W vector learned by the model itself is used to perform a dimensionality transformation on the original feature vector. [·||·] indicates that the transformed features of nodes i and j are concatenated. a(·) maps the concatenated high-dimensional features to a real number. After obtaining the correlation coefficient, it is normalized to obtain the attention coefficient.
[0065] Then, the features of neighboring nodes are weighted and summed based on the calculated attention coefficients: Where h′ i The output of node i from the attention layer incorporates new features from its neighboring nodes, and σ(·) is the activation function. To improve the performance of the aggregator, the graph attention network uses a multi-head attention mechanism, that is, it uses K independent attention mechanisms (with different a and W), and then concatenates the results again: Where || represents the concatenation operation. These are the weight coefficients calculated by the k-th attention mechanism, but they will cause h′ to... i Since it has a higher dimension (1, kh'), the multi-head attention in the intermediate layer is spliced together, while the aggregation method of the output layer is to average the results obtained from the K attention mechanisms:
[0066] For each label, calculate the node with the highest degree under that label, and save the features output by that node as the prototype proto for that label. i (i represents the corresponding label value), proto i It is a one-dimensional vector with a channel size of C.
[0067] Step 5: For the test image, first extract the features f before the fully connected layer using the trained ResNet. f is a one-dimensional vector with channel size C. Calculate the cosine similarity s between each feature and the prototype of each saved label. i (i represents the corresponding label value):
[0068]
[0069] Based on the calculated cosine similarity, calculate the probability 'a' that the age of the test image might be labeled i. i :
[0070]
[0071] The predicted age of the test image is:
[0072]
[0073] The visualization of facial age features after applying the method in this embodiment is as follows: Figure 5 As shown.
[0074] In summary, the non-standard face age estimation method based on feature aggregation of the present invention meets the high accuracy requirement of face age estimation and solves the problems of non-standard face pose and uneven lighting in practical application scenarios.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A non-standard face age estimation method based on feature aggregation, characterized in that: The method includes the following steps: Step 1: Image data preprocessing; Acquire face images, use a face detection model to detect faces in the images and perform face alignment, and then perform brightness normalization; for images with low brightness, use a lighting enhancement algorithm to enhance the lighting. Step 2: Extract features from the image; We use ResNet, a convolutional neural network, as the feature extraction model. During training, we use Adaptive Mean-ResidueLoss as the loss function to improve the robustness of the model in the age estimation task through distribution learning. Step 3: Construct a graph based on the extracted features; Features before the ResNet fully connected layer are extracted from all training images. The features of one image correspond to a node in the graph. For each training image, the output is obtained after passing through the complete ResNet. The output is then input into the softmax layer to obtain the predicted probability for each label of the image. The predicted probability values are sorted in descending order, and the classifications corresponding to the top five probability values are taken as the prediction set for that node. The number of intersections of the prediction sets between any two nodes is used as the weight of the edge between the two nodes, thereby constructing the adjacency matrix of the graph. Step 4: Based on the graph attention network, perform feature aggregation on the nodes of the graph obtained in Step 3, capture the intrinsic correlation between features, then classify the nodes, aggregate age-related features, and save the feature corresponding to the node with the largest degree under each label as a prototype. Step 5: For the test image, first extract the features before the fully connected layer using the trained ResNet, calculate the probability of each label based on the cosine similarity between the feature and the prototype of each label, and obtain the final predicted age value.
2. The non-standard face age estimation method based on feature aggregation according to claim 1, characterized in that: The face detection model in step 1 is as follows: a three-layer cascaded network is used, referred to as the P-network, R-network and O-network respectively. The P-network quickly generates coarse candidate boxes, the R-network filters to obtain high-precision candidate boxes, and the O-network generates bounding boxes and key point coordinates. After the face image is cropped out by the three-layer cascaded network, it is uniformly sized to 256×256 and its brightness is normalized.
3. The non-standard face age estimation method based on feature aggregation according to claim 2, characterized in that: In the three-layer cascaded network, the P-network includes three convolutional layers, which ultimately yield three results: a 1×1×2 face classification result, a 1×1×4 bounding box regression result, and a 1×1×10 face landmark location. For an image of size H×W×3, after passing through a P-network, a 5-channel S×S feature map is output. That is, after the sliding window operation of the P-network, the image yields S×S proposal boxes, each corresponding to 1 confidence score and 4 offsets. Then, the NMS algorithm is used to retain the proposal boxes corresponding to the confidence scores greater than the set threshold of 0.6, and the corresponding bounding box offsets are subjected to bounding box regression to obtain the coordinate information in the original image. Image data is obtained from the original image based on the bounding box coordinates of the P-network and input into the second layer of the face detection model, the R-network. The R-network adds a fully connected layer with 128 neurons on top of the three convolutional layers to filter out more errors and remove data with low scores. The face candidate boxes output by the R-network are input into the third layer of the face detection model, the O-network, for further detail processing. The O-network consists of four convolutional layers. The fourth convolutional layer collects more face features and removes incorrect candidate boxes. At this point, the output branch includes three categories: whether there is a face in the image; the horizontal and vertical coordinates of the starting point or center point of the regressed box and the length and width of the bounding box; and information on five key points of the face, including the position of the left eye, the position of the right eye, the position of the nose, the left position of the mouth, and the right position of the mouth. Each key position is represented using two dimensions. Then, based on the obtained facial landmark locations and the manually defined standard facial landmark positions, an affine transformation is performed on the outlined facial image, and the aligned facial image is adjusted to a size of 256×256.
4. The non-standard face age estimation method based on feature aggregation according to claim 1, characterized in that: The illumination enhancement algorithm first decomposes the input image into sub-images of different scales, and then performs logarithmic transformation and Gaussian blur on each sub-image at each scale to obtain a logarithmic domain image L(x, y). For each scale of the logarithmic domain image, the standard deviation of the surrounding pixels is calculated to determine the gain coefficient of that pixel. Among them G i,j σ represents the gain coefficient of pixel (i, j). i,j Let L' represent the standard deviation of the pixels surrounding pixel (i, j), and k be a constant used to control the magnitude of the gain coefficient; then, multiplicative enhancement is performed on the logarithmic domain image at each scale, i.e.: L′ i,j =G i,j *L i,j , where L′ i,j For the enhanced pixel value, L i,j The values are the pixel values of the original logarithmic domain image. Finally, the enhanced logarithmic domain image at each scale is subjected to an inverse logarithmic transformation to obtain the enhanced sub-image. All the enhanced sub-images are then combined to obtain the final enhanced image.
5. The non-standard face age estimation method based on feature aggregation according to claim 1, characterized in that: In step 2, the loss function, Adaptive Mean-Residue Loss, is specifically as follows: L=L s +λ1L m +λ2L r Among them, L s The loss is softmax. It is the average loss, representing the variance between the mean of the age distribution and the true age; The residual loss represents the tail residual error in the predicted age distribution after the top-K operation. It is calculated by sorting the predicted probability values in descending order, taking the labels corresponding to the top K values, and calculating the loss function for the predicted values of the other NK labels.
6. The non-standard face age estimation method based on feature aggregation according to claim 1, characterized in that: The graph attention network in step 4 is composed of stacked graph attention layers. For a single graph attention layer, the input is a set of node features, h = {h1, h2, ..., h...}. N }, h i ∈R F Where N is the number of nodes and F is the number of features for each node, a graph attention layer generates a new set of node features h′={h′1,h′2,…,h′ N }, h′ i ∈R F′ ; The calculation of the graph attention layer is divided into two steps: First, calculate the attention coefficient; For vertex i, calculate the relationship between it and its neighbors j∈N one by one. i similarity coefficient e ij e ij =a([Wh i ||Wh j ]), j∈N i , where N i Wh represents all neighboring nodes of vertex i. i This indicates that the shared W vector learned by the model itself is used to perform a dimensionality transformation on the original feature vector. [·||·] indicates that the transformed features of nodes i and j are concatenated. a(·) indicates that the concatenated high-dimensional features are mapped to a real number. After obtaining the correlation coefficient, it is normalized to obtain the attention coefficient. Then, the features of the neighboring nodes are weighted and summed based on the calculated attention coefficients: Where, h′ i The output of each node i in the attention layer incorporates new features from its neighboring nodes, and σ(·) is the activation function. To improve the performance of the aggregator, the graph attention network uses a multi-head attention mechanism, that is, it uses K independent attention mechanisms, i.e., different values of a and W, and then concatenates the results: Here, || represents the concatenation operation. This represents the weight coefficients calculated by the k-th attention mechanism, but it will lead to h′ i Since it has a higher dimension (1, kh'), the multi-head attention in the intermediate layer is spliced together, while the aggregation method of the output layer is to average the results obtained from the K attention mechanisms: Then, for each label, calculate the node with the highest degree under that label, and save the features output by that node as the prototype proto for that label. i 'i' represents the corresponding tag value, proto i It is a one-dimensional vector with a channel size of C.
7. The non-standard face age estimation method based on feature aggregation according to claim 1, characterized in that: The final age prediction of the test image in step 5 is specifically as follows: First, the features f before the fully connected layers are extracted using the trained ResNet. Feature f is a one-dimensional vector with channel size C. Then, the cosine similarity s between each feature and the prototype of each saved label is calculated. i : Where i represents the corresponding label value; based on the calculated cosine similarity, the probability a that the age of the test image might be labeled i is calculated. i : The predicted age of the test image is:
Citation Information
Patent Citations
Face age automatic estimation method based on weak supervision
CN111274882A
Cross-age face recognition method and system based on ternary constraints
CN113705383A