Age prediction method, device, electronic device and storage medium
By building an age prediction model of the convolution module, visual memory module and fully connected module, age characteristics with time dimension information are extracted and model training is carried out, the problem of insufficient face age prediction accuracy in the prior art is solved, and a higher age prediction accuracy is achieved.
Patent Information
- Application Number
- CN202111609601.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The existing face age prediction algorithm has insufficient recognition accuracy and cannot effectively learn the changes in face age characteristics, resulting in a large deviation between predicted age and actual age.
An age prediction method is proposed. By constructing an age prediction model including a convolutional module, a visual memory module and a fully connected module, the visual memory module extracts age features with time dimension information, and combines features through a fully connected module, and combines the model with a preset age constraint loss function.
It effectively reduces the deviation between predicting age and real age, improves the accuracy of face age prediction, and enables the age prediction model to better learn the aging situation of the same character with age.
Smart Images

Figure CN114267071B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of image processing, and in particular, to an age prediction method, device, electronic device, and storage medium. Background Art
[0002] A face image often contains a lot of face feature information. Among them, age, as an important feature information, has been widely used in the field of face recognition. Especially on the mobile APP side, face age prediction is actually a challenging task. For each face image, as age increases, the aging of the face is actually a slow process. At the same age stage, there are some similar features among different people.
[0003] Currently, age is usually regarded as a single label information. During the model training process, a one-to-one correspondence between a face image and age is established. Due to the problem of collecting face aging images of the same person in different years, the existing age estimation model cannot learn the age feature changes of the face with time memory. As a result, in actual applications, the predicted age deviates greatly from the actual age.
[0004] In the process of implementing the embodiments of the present application, the inventors found that the current technical solutions have at least the following technical problems: the recognition accuracy of the existing face age prediction algorithm is insufficient. Summary of the Invention
[0005] The main technical problem to be solved by the embodiments of the present application is to provide an age prediction method, device, electronic device, and storage medium to improve the accuracy of face age prediction.
[0006] In a first aspect, an age prediction method is provided in the embodiments of the present application, including:
[0007] Obtain a plurality of face image sets, where the face image set includes at least real face images of the same person corresponding to several years;
[0008] According to the real face images, construct an age prediction model, where the age prediction model includes a convolutional module, a visual memory module, and a fully connected module. Among them, the convolutional module is used to input multiple real face images of at least the same person, and perform feature extraction on each real face image to output a first feature map corresponding to each real face image; the visual memory module is used to receive the first feature map output by the convolutional module, and extract age features with time dimension information to output a second feature map; the fully connected module is used to combine the second feature map output by the visual memory module;
[0009] Train the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model;
[0010] Based on the trained age prediction model, perform age prediction on the target image to obtain the age prediction value of the person in the target image.
[0011] In some embodiments,
[0012] The convolution module includes multiple convolutional layers, and the multiple convolutional layers are used to extract the first feature maps of multiple real face images. Among them, each convolutional layer is used to extract the first feature map of one real face image;
[0013] The visual memory module includes multiple visual memory layers. Each visual memory layer is connected to a convolutional layer in a one-to-one correspondence. Each visual memory layer is used to receive the first feature map output by the convolutional layer corresponding to it and the first feature map output by another convolutional layer to extract the feature information of the two first feature maps in the time dimension;
[0014] The fully connected module includes multiple fully connected layers. Each fully connected layer is connected to a visual memory layer in a one-to-one correspondence and is used to combine the second feature maps output by multiple visual memory layers to generate a combined feature map;
[0015] The age prediction model further includes:
[0016] Multiple classification layers. Each classification layer is connected to a fully connected layer in a one-to-one correspondence and is used to predict the age prediction value corresponding to the combined feature map output by the fully connected layer;
[0017] Among them, each real face image corresponds to a channel, and each channel corresponds to a convolutional layer, a visual memory layer, a fully connected layer, and a classification layer.
[0018] In some embodiments, the age prediction model includes:
[0019] Multiple groups of visual memory layers. One visual memory layer in each group of visual memory layers is connected in series with one visual memory layer in another group of visual memory layers to form a series structure;
[0020] Among them, one visual memory layer in the next group of visual memory layers is used to obtain the second feature maps output by at least two visual memory layers in the previous group of visual memory layers, so as to further process the second feature maps output by the at least two visual memory layers, obtain the second feature maps and input them into the next group of visual memory layers, and so on until the feature maps are output to the last group of visual memory layers.
[0021] In some embodiments, the age prediction model includes three channels, namely the upper channel, the middle channel, and the lower channel. Among them, each channel corresponds to a real face image, and the real ages corresponding to the real face images in different channels are different. Among them, the age prediction value output by the classification layer corresponding to the middle channel is used as the age prediction value output by the age prediction model.
[0022] In some embodiments, the age constraint loss function is:
[0023]
[0024] where α is the weight parameter of the middle channel, β is the weight parameter of the upper channel, γ is the weight parameter of the lower channel, N is the total number of samples, is the true age value of the middle channel of the j-th sample, is the predicted age value of the middle channel of the j-th sample, is the true age value of the upper channel of the j-th sample, is the predicted age value of the upper channel of the j-th sample, is the true age value of the lower channel of the j-th sample, is the predicted age value of the lower channel of the j-th sample.
[0025] In some embodiments, based on the preset age constraint loss function, the age prediction model is trained to obtain the trained age prediction model, including:
[0026] Iteratively train the age prediction model based on the age constraint loss function;
[0027] If the number of iterations is greater than the first number threshold, or the loss of the age prediction model is less than the first loss threshold, stop the iterative training to obtain the trained age prediction model.
[0028] In some embodiments, the method further includes:
[0029] Preprocess the real face images in the face image set, including:
[0030] According to the face key point algorithm, obtain the center coordinates of the left and right eyeballs in the real face image;
[0031] Calculate the angle between the line connecting the center coordinates of the left and right eyeballs in the real face image and the horizontal direction;
[0032] Taking the center coordinates of the left and right eye eyeballs in the real face image as the base points, rotate the real face image by the angle;
[0033] Crop the face region in the rotated real face image and resize it to the preset resolution to obtain the preprocessed real face image.
[0034] In a second aspect, an age prediction device provided by an embodiment of the present application includes:
[0035] A set construction unit for obtaining several sets of face images, where each set of face images includes at least real face images of the same person corresponding to several years;
[0036] A model construction unit for constructing an age prediction model based on real face images. The age prediction model includes a convolutional module, a visual memory module, and a fully connected module. The convolutional module is used to input multiple real face images of at least the same person, and perform feature extraction on each real face image to output a first feature map corresponding to each real face image; the visual memory module is used to receive the first feature maps output by the convolutional module, and extract age features with time dimension information to output a second feature map; the fully connected module is used to combine the second feature maps output by the visual memory module;
[0037] A model training unit for training the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model;
[0038] An age prediction unit for predicting the age of a target image based on the trained age prediction model to obtain an age prediction value of the person in the target image.
[0039] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0040] A memory and one or more processors. The one or more processors are used to execute one or more computer programs stored in the memory. When the one or more processors execute the one or more computer programs, the electronic device implements the method as in the first aspect.
[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the method as in the first aspect.
[0042] Beneficial effects of the embodiments of the present application: Different from the prior art, an age prediction method provided by the embodiments of the present application includes: obtaining a plurality of sets of face images, where the set of face images includes at least real face images of the same person corresponding to several years; constructing an age prediction model according to the real face images, the age prediction model including a convolution module, a visual memory module, and a fully connected module, where the convolution module is used to input multiple real face images of at least the same person, and perform feature extraction on each real face image to output a first feature map corresponding to each real face image; the visual memory module is used to receive the first feature map output by the convolution module, and extract age features with time dimension information to output a second feature map; the fully connected module is used to combine the second feature map output by the visual memory module; training the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model; and predicting the age of a target image based on the trained age prediction model to obtain an age prediction value of the person in the target image.
[0043] On the one hand, by constructing an age prediction model, which includes a convolution module, a visual memory module, and a fully connected module, the visual memory module extracts features from the first feature map output by the convolution module to extract age features with time dimension information, and the fully connected module combines the second feature map output by the visual memory module, enabling the age prediction model to learn the aging situation of the same person as they age, effectively reducing the deviation between the predicted age and the real age. On the other hand, by training the age prediction model with an age constraint loss function and predicting the age of a target image based on the trained age prediction model, the embodiments of the present application can improve the accuracy of face age prediction. Description of the Drawings
[0044] One or more embodiments are illustrated by way of example in the accompanying drawings, and these exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a scale limitation.
[0045] Figure 1 is a schematic diagram of the application environment of an age prediction method provided by the embodiments of the present application;
[0046] Figure 2 is a schematic flowchart of an age prediction method provided by the embodiments of the present application;
[0047] Figure 3 is a schematic diagram of an age prediction model provided by the embodiments of the present application;
[0048] Figure 4It is a schematic diagram of another age prediction model provided by an embodiment of the present application;
[0049] Figure 5 It is a schematic flowchart of iterative training of an age prediction model provided by an embodiment of the present application;
[0050] Figure 6 It is a schematic structural diagram of an age prediction device provided by an embodiment of the present application;
[0051] Figure 7 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0052] The present application will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present application. These all belong to the protection scope of the present application.
[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] It should be noted that if there is no conflict, the various features in the embodiments of the present application can be combined with each other, and all are within the protection scope of the present application. In addition, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. In addition, the terms "first", "second", "third", etc. used herein do not limit the data and execution order, but only distinguish the same items or similar items with basically the same function and role.
[0055] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in this specification in the description of the present application are only for the purpose of describing specific embodiments and are not used to limit the present application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.
[0056] In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0057] Before elaborating on the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations:
[0058] (1) A neural network, also simply referred to as neural networks (NNs) or called a connection model, is a mathematical model of an algorithm that mimics the behavioral characteristics of animal neural networks and performs distributed parallel information processing. A neural network relies on the complexity of the system and adjusts the relationships between a large number of internal nodes to achieve the purpose of processing information. Specifically, a neural network can be composed of neural units, which can be specifically understood as a neural network with an input layer, hidden layers, and an output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the intermediate layers are all hidden layers. Among them, a neural network with many hidden layers is called a deep neural network (DNN). The operation of each layer in a neural network can be described by the mathematical expression y = a(W·x + b). Physically, the operation of each layer in a neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / dimensionality reduction; 2. Enlargement / shrinkage; 3. Rotation; 4. Translation; 5. "Bending". Among them, operations 1, 2, and 3 are completed by "W·x", operation 4 is completed by "+b", and operation 5 is implemented by "a()". The reason for using the word "space" here is that the objects to be classified are not individual things but a class of things. Space refers to the set of all individuals of this class of things. Among them, W is the weight matrix of each layer of the neural network, and each value in this matrix represents the weight value of a neuron in this layer. This matrix W determines the space transformation from the above-mentioned input space to the output space, that is, W of each layer of the neural network controls how to transform the space. The purpose of training a neural network is to finally obtain the weight matrices of all layers of the trained neural network. Therefore, the training process of a neural network is essentially to learn the way of controlling space transformation, and more specifically, to learn the weight matrix.
[0059] It should be noted that in the embodiments of the present application, the models adopted based on machine learning tasks are essentially neural networks. Common components in a neural network include a convolutional layer, a pooling layer, a normalization layer, a transposed convolutional layer, etc. By assembling these common components in a neural network, a model is designed. When the model error satisfies a preset condition or the number of adjusted model parameters reaches a preset threshold when determining the model parameters (weight matrices of each layer), the model converges.
[0060] Among them, the convolutional layer is configured with multiple convolutional kernels, and each convolutional kernel is set with a corresponding stride to perform a convolutional operation on the image. The purpose of the convolutional operation is to extract different features of the input image. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners. Deeper convolutional layers can iteratively extract more complex features from the low-level features.
[0061] The transposed convolutional layer is used to map a low-dimensional space to a high-dimensional space while maintaining their connection relationship / pattern (here the connection relationship refers to the connection relationship during convolution). The transposed convolutional layer is configured with multiple convolutional kernels, and each convolutional kernel is set with a corresponding stride to perform a transposed convolutional operation on the image. Generally, the upsample() function is built into the framework library (such as the PyTorch library) used to design neural networks. By calling the upsample() function, the mapping from a low-dimensional space to a high-dimensional space can be achieved.
[0062] The pooling layer is used to reduce the dimension of data or represent the image with higher-level features by imitating the human visual system. Common operations of the pooling layer include max pooling, average pooling, stochastic pooling, median pooling, and combined pooling, etc. Generally speaking, pooling layers are periodically inserted between the convolutional layers of a neural network to achieve dimensionality reduction.
[0063] The normalization layer is used to perform a normalization operation on all neurons in the intermediate layer to prevent gradient explosion and gradient disappearance.
[0064] (2) Loss function refers to a function that maps the values of a random event or its related random variables to non - negative real numbers to represent the "risk" or "loss" of the random event. A loss function is a non - negative real - valued function used to quantify the difference between the predicted label and the true label in model prediction. In applications, the loss function is usually associated with the learning criterion and the optimization problem, that is, the model is solved and evaluated by minimizing the loss function. For example, it is used in parametric estimation in statistics and machine learning. During the process of training a neural network, since we hope that the output of the neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the true target value, and then update the weight matrix of each layer of the neural network according to the difference between them (of course, there is usually an initialization process before the first update, that is, parameters are pre - configured for each layer in the neural network). For example, if the predicted value of the network is too high, we adjust the weight matrix to make it predict lower, and keep adjusting until the neural network can predict the true target value. Therefore, it is necessary to pre - define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the neural network becomes a process of minimizing this loss as much as possible.
[0065] The technical solution of the present application will be specifically described below with reference to the accompanying drawings of the specification.
[0066] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the application environment of an age prediction method provided by an embodiment of the present application;
[0067] As Figure 1 shown, the application environment 100 includes: an electronic device 101 and a server 102. The electronic device 101 and the server 102 communicate through wired or wireless communication means.
[0068] Among them, the electronic device 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. A client can be provided in the electronic device 101. The client can be a video client, a browser client, an online shopping client, an instant messaging client, etc. The present application does not limit the type of the client.
[0069] The electronic device 101 and the server 102 can be directly or indirectly connected through wired or wireless communication means, and this application does not limit this. The electronic device 101 can obtain a target image, predict the age prediction value of the target image, and display the target image and its age prediction value on a visualization interface. Among them, the target image can be an image stored in the memory of the electronic device 101 or an image received from other devices.
[0070] Alternatively, the electronic device 101 can receive the age prediction value of the target image sent by the server 102, and display the target image and its age prediction value on the visualization interface. The user can browse the face images stored in the electronic device, and trigger the age prediction instruction for the face image by triggering the age prediction button corresponding to any face image. The electronic device can respond to the age prediction instruction, obtain a face image through an image acquisition device, and use the face image as the target image. Among them, the image acquisition device can be built into the electronic device 101 or externally connected to the electronic device 101, and this application does not limit this.
[0071] The electronic device 101 can send the target image to the server 102, and receive the age prediction value of the target image returned by the server 102, and then display the target image and its age prediction value on the visualization interface so that the user can understand the prediction result of the target image.
[0072] It can be understood that the electronic device 101 can generally refer to one of multiple electronic devices, and this application only takes the electronic device 101 as an example for illustration. Those skilled in the art can know that the number of the above-mentioned electronic devices can be more or less. For example, the above-mentioned electronic device can be only one, or the above-mentioned electronic devices can be dozens or hundreds, or more in number, and this application does not limit the number and type of the electronic devices.
[0073] Among them, the server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0074] The server 102 and the electronic device 101 can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here. The server 102 can maintain a face image database for storing multiple face images. The server 102 can receive an age prediction instruction and a target image sent by the electronic device 101, and based on the age prediction instruction, perform age prediction on the target image based on the target image to obtain an age prediction value of the target image, and then send the age prediction value of the target image to the electronic device 101.
[0075] It can be understood that the number of the above-mentioned servers 102 can be more or less, and the embodiments of this application do not limit this. Of course, the server 102 can also include other functional servers to provide more comprehensive and diverse services.
[0076] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an age prediction method provided by an embodiment of this application;
[0077] Among them, this age prediction method is applied to the above-mentioned electronic device. Specifically, the execution subject of this age prediction method is one or more processors of the electronic device.
[0078] As Figure 2 shown, this age prediction method includes:
[0079] Step S201: Obtain a plurality of face image sets, where the face image set includes at least real face images of the same person corresponding to several years;
[0080] Specifically, this face image set is composed of real face images of the same person within N years. For example: collect real face images of the same person within continuous or discontinuous N years, and label the age label of each real face image. For example: use the one-hot method for labeling. Preferably, the embodiments of this application collect real face images of the same person within continuous several years. For example: face images of the same person within continuous 5 years, so as to better learn the changes in face aging features and make the age prediction value output by the age prediction model more accurate.
[0081] Furthermore, a plurality of face image sets can also be divided into a face age training set and a face age test set. For example: divide the face image set according to the ratio of the face age training set to the face age test set of 8:2, so as to better train the age prediction model based on the face age training set and test the age prediction model based on the face age test set, thereby better generating the age prediction model.
[0082] It can be understood that the set of face images can be color ID photos or color selfies collected by an image acquisition device, etc. It can also be understood that the set of face images can be data in an existing open-source face library, where the open-source face library can be the FERET face database, the CMU Multi-PIE face database, or the YALE face database, etc. Here, the source of the image samples is not restricted, as long as the image is a color image including a face, such as a face image in RGB format.
[0083] Specifically, the age prediction method further includes: preprocessing the real face images in the set of face images, including:
[0084] According to the face key point algorithm, obtain the center coordinates of the left and right eyeballs in the real face image;
[0085] Calculate the angle between the line connecting the center coordinates of the left and right eyeballs in the real face image and the horizontal direction;
[0086] Taking the center coordinates of the left and right eyeballs in the real face image as the base points, rotate the real face image by this angle;
[0087] Intercept the face region in the rotated real face image and adjust the size to the preset resolution to obtain the preprocessed real face image.
[0088] Specifically, use the face key point algorithm to obtain the center coordinates of the left and right eyeballs in the real face image; calculate the angle between the line connecting the center coordinates of the left and right eyeballs in the real face image and the horizontal direction; taking the center coordinates of the left and right eyeballs in the real face image as the base points, rotate the real face image by the above angle; intercept the face region in the rotated real face image and adjust the size to the preset resolution to obtain the preprocessed real face image.
[0089] According to the face key point algorithm, locate several key points on the face of the real face image, including points in areas such as eyebrows, eyes, nose, mouth, and face contour. Then, from these several key points, obtain the center coordinates of the left and right eyeballs in the real face image, and calculate the angle θ between the line connecting the center coordinates of the left and right eyeballs and the horizontal direction.
[0090] It can be understood that this angle θ is the angle by which the face deviates from the frontal face. In order to adjust the face in the real face image to the frontal face, taking the center coordinates of the left and right eyeballs in the real face image as the base points, rotate the real face image by the above angle θ to obtain the frontal face.
[0091] Specifically, the following formula can be used to calculate the rotated real face image:
[0092]
[0093] Among them, (x, y) are the two-dimensional coordinates of the pixel points in the real face image before rotation, and (x', y') are the two-dimensional coordinates of the pixel points in the real face image after rotation.
[0094] Based on the fact that the face in the real face image after rotation is a frontal face, the effective face region in the real face image after rotation is intercepted and resized to a preset resolution, that is, a normalization operation is performed on the intercepted effective face region in the real face image after rotation. For example, the size of the effective face region is unified to 256*256 to obtain the preprocessed real face image. Thus, the size of the preprocessed real face image is the preset resolution, and the preprocessed real face image only includes the frontal face and does not include other background pixels.
[0095] In the embodiment of the present application, preprocessing the real face image is beneficial for the age prediction model to better learn features and can help the age prediction model converge better.
[0096] Step S202: Construct an age prediction model according to the real face image. The age prediction model includes a convolution module, a visual memory module, and a fully connected module. Among them, the convolution module is used to input at least multiple real face images of the same person and extract features from each real face image to output a first feature map corresponding to each real face image; the visual memory module is used to receive the first feature map output by the convolution module and extract age features with time dimension information to output a second feature map; the fully connected module is used to combine the second feature map output by the visual memory module;
[0097] Specifically, please refer to Figure 3 , Figure 3 which is a schematic diagram of an age prediction model provided by the embodiment of the present application;
[0098] As Figure 3 shown, the age prediction model 300 includes a convolution module 301, a visual memory module 302, and a fully connected module 303. Among them, the convolution module 301 is connected to the visual memory module 302, and the visual memory module 302 is connected to the fully connected module 303.
[0099] It can be understood that for the preprocessed real face images, since the prediction range of the existing age estimation model is between plus or minus five years, and coupled with the changes in the shooting environment (lighting, brightness) and facial expressions, the prediction range changes even more. Therefore, in order to better learn the changes in facial age features between the ages of five before and after, and make the predicted age value of the age prediction model more accurate, when constructing the age prediction model, multiple input interfaces are adopted. The multiple input interfaces use face images that change between the ages of five before and after to provide age feature information in the time dimension for the age prediction model.
[0100] Among them, the convolution module 301 is used to input multiple real face images of the same person, and extract features from each real face image to output the first feature map corresponding to each real face image. That is, the input of the convolution module 301 is the real face image in the face image set. Specifically, the input of the convolution module 301 is the real face images of the same person corresponding to several years. In order to better learn the changes in facial features, preferably, in the embodiments of the present application, the input of the convolution module 301 is the real face images of the same person corresponding to several years. For example, the real face images of the same person within five years are used to more accurately extract more detailed age features. In the embodiments of the present application, the convolution module 301 adopts a convolution feature extraction structure with multiple input channels.
[0101] Specifically, the convolution module includes multiple convolutional layers. The multiple convolutional layers are used to extract the feature maps of multiple real face images. Among them, each convolutional layer is used to extract the first feature map of one real face image. For example, the convolution module adopts a three-channel input structure, that is, three groups of standard convolutional layers are used to extract the feature information of three different real face images. For example, the size of the input real face image is 256*256, and the convolutional layer of the convolution module uses a 3*3 convolutional kernel, the number is set to 16, and the stride is 2 to obtain three groups of first feature maps with a size of 128*128*16.
[0102] Among them, the visual memory module 302 (ConvGRU), connected to the convolution module 301, is used to receive the first feature map output by the convolution module 301, and extract the age features with time dimension information to output the second feature map. Specifically, the input of the visual memory module (ConvGRU) is the first feature map output by the convolution module. Among them, if the convolution module includes three convolutional layers, the first feature map output by the convolution module is input into the visual memory module in a pairwise combination manner, so that the visual memory module extracts the age feature information of the aging of the face in the front and back age ranges.
[0103] In an embodiment of the present application, the visual memory module 302 (ConvGRU) includes a gated recurrent unit (GRU) network. It can be understood that the gated recurrent unit (GRU) network is a recurrent neural network that is simpler than the LSTM network. The GRU network introduces a gating mechanism to control the way of information update. Moreover, the GRU does not introduce additional memory units. Specifically, the GRU network introduces an update gate to control how much information from the historical state needs to be retained in the current state (without passing through a non-linear transformation) and how much new information needs to be received from the candidate state. That is to say, the GRU network directly uses a gate to control the balance between input and forgetting.
[0104] Among them, the fully connected module 303 is connected to the visual memory module 302 and is used to combine the second feature maps extracted by the visual memory module 302. Specifically, the input of the fully connected module 303 is multiple second feature maps output by the visual memory module 302, and the multiple second feature maps are combined to achieve feature fusion of the multiple feature maps, and a feature map after feature fusion is output.
[0105] In an embodiment of the present application, the convolution module includes multiple convolution layers, and the multiple convolution layers are used to extract feature maps of multiple real face images. Among them, each convolution layer is used to extract a feature map of one real face image;
[0106] The visual memory module includes multiple visual memory layers. Each visual memory layer is connected to a convolution layer in a one-to-one correspondence. Each visual memory layer is used to receive the first feature map output by the corresponding convolution layer and the first feature map output by another convolution layer to extract the feature information of the two first feature maps in the time dimension;
[0107] The fully connected module includes multiple fully connected layers. Each fully connected layer is connected to a visual memory layer in a one-to-one correspondence and is used to combine the second feature maps output by the multiple visual memory layers to generate a combined feature map;
[0108] The age prediction model further includes:
[0109] Multiple classification layers. Each classification layer is connected to a fully connected layer in a one-to-one correspondence and is used to predict the age prediction value corresponding to the combined feature map output by the fully connected layer;
[0110] Among them, each real face image corresponds to a channel, and each channel corresponds to a convolution layer, a visual memory layer, a fully connected layer, and a classification layer.
[0111] Among them, the age prediction model includes:
[0112] Multiple groups of visual memory layers, where one visual memory layer in each group of visual memory layers is connected in series with one visual memory layer in another group of visual memory layers to form a series structure;
[0113] Among them, one visual memory layer in the next group of visual memory layers is used to obtain the second feature maps output by at least two visual memory layers in the previous group of visual memory layers, so as to further process the second feature maps output by the at least two visual memory layers, obtain the second feature maps and input them into the next group of visual memory layers, and so on until the second feature maps are output to the last group of visual memory layers.
[0114] In the embodiment of the present application, the visual memory layer, that is, the ConvGRU structure, is embedded in the age prediction network structure. Since the ConvGRU structure has a memory function, the age prediction model in the present application can not only extract aging features in terms of depth and breadth, but also extract aging feature information from the time dimension, so as to learn the different degrees of aging changes of the same person with age, and further effectively reduce the situation where the result of the age prediction model deviates greatly from the actual age, and improve the applicability and accuracy of the age prediction model.
[0115] Specifically, please refer to Figure 4 , Figure 4 which is a schematic diagram of another age prediction model provided by the embodiment of the present application;
[0116] As Figure 4 shown, first, the input of the convolutional layer is three real face images of the same person at different ages, corresponding to three different channels, arranged in chronological order. The basic feature extraction structures of the three channels are the same. For example: the three channels are the upper channel, the middle channel, and the lower channel. Among them, each channel corresponds to a real face image, and the real ages corresponding to the real face images of different channels are different. Among them, the age prediction value output by the classification layer corresponding to the middle channel is used as the age prediction value output by the age prediction model.
[0117] Taking the middle channel as an example, the convolutional layer performs a convolutional operation on a real face image, extracts features from the original image of size 256*256*3, and obtains a first feature map of size 128*128*16. Then, through the visual memory layer (ConvGRU layer), at this time, since the visual memory layer requires two image inputs, the features extracted by the convolutional layer of the upper channel and the features extracted by the convolutional layer of the middle channel need to be used as inputs to obtain a second feature map of size 64*64*32.
[0118] Similarly, the second feature map output by the previous visual memory layer of the upper channel and the second feature map output by the visual memory layer of the middle channel are input into the next visual memory layer, and so on. Finally, a second feature map with a size of 8*8*256 is obtained. Then, the second feature maps of the three channels are concatenated to obtain a feature map with a size of 8*8*768. Then, convolution operation is performed for feature fusion to obtain a feature with a size of 8*8*512. Then, a fully connected layer is connected to obtain a feature with a size of 1*1024. Then, through the softmax layer, the probability of the corresponding age value is predicted.
[0119] In the embodiment of the present application, through the way of cascading multiple groups of visual memory layers, the changing feature information of human face aging is continuously extracted from three aspects: network depth, width, and time dimension. Finally, through the splicing method, the second feature maps extracted from the three channels are combined together, and then feature fusion is realized through a convolutional layer with a size of 3*3, and then input into the classification structure to predict the age value corresponding to the image. The present application can improve the accuracy of human face age prediction.
[0120] Step S203: Based on a preset age constraint loss function, train the age prediction model to obtain a trained age prediction model;
[0121] Specifically, the age loss function is:
[0122]
[0123] where α is the weight parameter of the middle channel, β is the weight parameter of the upper channel, γ is the weight parameter of the lower channel, N is the total number of samples, is the true age value of the middle channel of the jth sample, is the predicted age value of the middle channel of the jth sample, is the true age value of the upper channel of the jth sample, is the predicted age value of the upper channel of the jth sample, is the true age value of the lower channel of the jth sample, is the predicted age value of the lower channel of the jth sample.
[0124] Specifically, N represents the total number of training samples, represents the true age value of the middle channel of the jth sample, represents the predicted age value of the middle channel of the jth sample. Similarly and respectively represent the true age value and the predicted age value of the upper channel of the jth sample. and They respectively represent the true age value and the predicted age value of the channel under the j-th sample. α, β, and γ are respectively the weight parameters of the middle channel, the upper channel, and the lower channel. It can be understood that in order to reflect that the middle channel is the main one, the value of the weight parameter α of the middle channel is greater than any one of the weight parameter β of the upper channel and the weight parameter γ of the lower channel.
[0125] In the embodiment of the present application, the values of α, β, and γ can be set according to specific needs, and the present application does not limit this. Preferably, the values of α, β, and γ are α = 20, β = 1, and γ = 1.
[0126] It can be understood that in the age prediction network structure in the embodiment of the present application, the features extracted by the middle channel are the main ones, and the features extracted by the upper channel and the lower channel are externally connected to the fully connected layer to predict the age corresponding to the image. On the one hand, the age range predicted by the middle channel can be restricted. On the other hand, the upper channel and the lower channel provide information on the aging characteristics in the time dimension for the middle channel, thereby improving the accuracy of face age prediction.
[0127] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of the iterative training of an age prediction model provided by the embodiment of the present application;
[0128] As Figure 5 shown, the process of the iterative training of the age prediction model includes:
[0129] Step S501: Construct an age constraint loss function;
[0130] Specifically, the age constraint loss function is:
[0131]
[0132] where α is the weight parameter of the middle channel, β is the weight parameter of the upper channel, γ is the weight parameter of the lower channel, N is the total number of samples, is the true age value of the middle channel of the j-th sample, is the predicted age value of the middle channel of the j-th sample, is the true age value of the upper channel of the j-th sample, is the predicted age value of the upper channel of the j-th sample, is the true age value of the lower channel of the j-th sample, is the predicted age value of the lower channel of the j-th sample.
[0133] Step S502: Iteratively train the age prediction model based on the age constraint loss function;
[0134] Specifically, the parameters of the age prediction model are optimized by the age constraint loss function.
[0135] Step S503: Whether the number of iterations is greater than the first number threshold;
[0136] Specifically, the embodiments of the present application adopt the Adam algorithm (Adaptive Moment Estimation Algorithm) to optimize the model parameters. For example: the number of iterations is set to 500 times, the initial learning rate is set to 0.001, the weight decay is set to 0.0005, and every 50 iterations, the learning rate decays to 1 / 10 of the original.
[0137] It can be understood that the Adam algorithm (Adaptive Moment Estimation Algorithm) can be regarded as the combination of the momentum method and the RMSprop algorithm, which not only uses momentum as the parameter update direction but also can adaptively adjust the learning rate.
[0138] Specifically, it is judged whether the number of iterations is greater than the first number threshold, and the first number threshold is preset, for example: set to 500 times. If the number of iterations is greater than the first number threshold, go to step S505: Training is completed; if the number of iterations is not greater than the first number threshold, go to step S504: Whether the loss of the age prediction model is less than the first loss threshold.
[0139] It can be understood that the first number threshold is specifically set according to specific needs and is not limited here.
[0140] Step S504: Whether the loss of the age prediction model is less than the first loss threshold;
[0141] Specifically, it is judged whether the loss of the age prediction model is less than the first loss threshold. If so, go to step S505: Training is completed; if not, return to step S502: Iteratively train the age prediction model based on the age constraint loss function.
[0142] In the embodiments of the present application, it is judged whether the loss of the age prediction model is less than the first loss threshold, that is, it is judged whether the loss calculated by the age constraint loss function is less than the first loss threshold, so as to determine whether to end the iterative process in advance when the number of iterations is less than the first number threshold, that is, stop iterative training to quickly obtain the trained age prediction model.
[0143] In the embodiments of the present application, the first loss threshold can be set to 0.0005, 0.001. It can be understood that the first loss threshold is specifically set according to specific needs and is not limited here.
[0144] Step S505: Training is completed;
[0145] Specifically, after the training is completed, a trained age prediction model is obtained. At this time, the trained age prediction model can be called to predict the age prediction value of the target image.
[0146] Step S204: Based on the trained age prediction model, perform age prediction on the target image to obtain the age prediction value of the target image.
[0147] Specifically, after the trained age prediction model is obtained through training, based on the trained age prediction model, age prediction is performed on the input target image to obtain the age prediction value of the target image.
[0148] In an embodiment of the present application, an age prediction method is provided, including: obtaining a plurality of face image sets, where the face image sets include at least real face images of the same person corresponding to several years; constructing an age prediction model according to the real face images, the age prediction model including a convolution module, a visual memory module, and a fully connected module, where the convolution module is used to input multiple real face images of at least the same person, and perform feature extraction on each real face image to output a first feature map corresponding to each real face image; the visual memory module is used to receive the first feature map output by the convolution module, and extract age features with time dimension information to output a second feature map; the fully connected module is used to combine the second feature map output by the visual memory module; training the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model; based on the trained age prediction model, performing age prediction on the target image to obtain the age prediction value of the person in the target image.
[0149] On the one hand, by constructing an age prediction model, the age prediction model includes a convolution module, a visual memory module, and a fully connected module. The visual memory module performs feature extraction on the first feature map output by the convolution module to extract age features with time dimension information, and the fully connected module combines the second feature map output by the visual memory module, so that the age prediction model can learn the aging situation of the same person as the age changes, effectively reducing the deviation between the predicted age and the real age. On the other hand, by training the age prediction model with an age constraint loss function, and based on the trained age prediction model, performing age prediction on the target image, the embodiment of the present application can improve the accuracy of face age prediction.
[0150] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an age prediction device provided by an embodiment of the present application;
[0151] Among them, the age prediction device is applied to an electronic device. Specifically, the age prediction device is applied to one or more processors of the electronic device.
[0152] As Figure 6 shown, the age prediction device 60 includes:
[0153] A set construction unit 601, configured to obtain a plurality of face image sets, where each face image set includes at least real face images of the same person corresponding to several years;
[0154] A model construction unit 602, configured to construct an age prediction model according to the real face images. The age prediction model includes a convolutional module, a visual memory module, and a fully connected module. Among them, the convolutional module is configured to input multiple real face images of at least the same person, and perform feature extraction on each real face image to output a first feature map corresponding to each real face image; the visual memory module is configured to receive the first feature maps output by the convolutional module, and extract age features with time dimension information to output a second feature map; the fully connected module is configured to combine the second feature maps output by the visual memory module;
[0155] A model training unit 603, configured to train the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model;
[0156] An age prediction unit 604, configured to predict the age of a target image based on the trained age prediction model to obtain an age prediction value of the person in the target image.
[0157] In the embodiments of the present application, the age prediction device can also be built by hardware devices. For example, the age prediction device can be built by one or more than two chips, and each chip can work in coordination with each other to complete the age prediction method described in the above embodiments. For another example, the age prediction device can also be built by various logic devices, such as built by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0158] The age prediction device in the embodiments of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0159] The age prediction device in the embodiments of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0160] The age prediction device provided in the embodiments of the present application can implement Figure 2 each of the implemented processes. To avoid repetition, they will not be elaborated here.
[0161] It should be noted that the above-mentioned age prediction device can execute the age prediction method provided in the embodiments of the present application and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in the embodiments of the age device, reference can be made to the age prediction method provided in the embodiments of the present application.
[0162] In an embodiment of the present application, by providing an age prediction device, including: a set construction unit, configured to obtain a plurality of face image sets, where the face image sets include at least real face images of the same person corresponding to several years; a model construction unit, configured to construct an age prediction model according to the real face images, and the age prediction model includes a convolution module, a visual memory module, and a fully connected module, where the convolution module is configured to input multiple real face images of at least the same person, and perform feature extraction on each real face image to output a first feature map corresponding to each real face image; the visual memory module is configured to receive the first feature map output by the convolution module, and extract age features with time dimension information to output a second feature map; the fully connected module is configured to combine the second feature map output by the visual memory module; a model training unit, configured to train the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model; an age prediction unit, configured to perform age prediction on a target image based on the trained age prediction model to obtain an age prediction value of the person in the target image.
[0163] On the one hand, by constructing an age prediction model, the age prediction model includes a convolution module, a visual memory module, and a fully connected module. The visual memory module performs feature extraction on the first feature map output by the convolution module to extract age features with time dimension information, and the fully connected module combines the second feature map output by the visual memory module, so that the age prediction model can learn the aging situation of the same person as the age changes, effectively reducing the deviation between the predicted age and the real age. On the other hand, by training the age prediction model with an age constraint loss function, and performing age prediction on a target image based on the trained age prediction model, the embodiment of the present application can improve the accuracy of face age prediction.
[0164] The embodiment of the present application also provides an electronic device. Please refer to Figure 7 , Figure 7 which is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application;
[0165] As Figure 7 shown, the electronic device 70 includes at least one processor 701 and a memory 702 that are communicatively connected ( Figure 7 taking bus connection and one processor as an example).
[0166] Among them, the processor 701 is used to provide computing and control capabilities to control the electronic device 70 to perform corresponding tasks. For example, it controls the electronic device 70 to execute the age prediction method in any of the above method embodiments, including: obtaining a plurality of face image sets, where the face image sets include at least real face images of the same person corresponding to several years; constructing an age prediction model according to the real face images, and the age prediction model includes a convolutional module, a visual memory module, and a fully connected module. Among them, the convolutional module is used to input multiple real face images of at least the same person, and perform feature extraction on each real face image to output a first feature map corresponding to each real face image; the visual memory module is used to receive the first feature maps output by the convolutional module, and extract age features with time dimension information to output second feature maps; the fully connected module is used to combine the second feature maps output by the visual memory module; based on a preset age constraint loss function, train the age prediction model to obtain a trained age prediction model; based on the trained age prediction model, perform age prediction on the target image to obtain the age prediction value of the person in the target image.
[0167] On the one hand, by constructing an age prediction model, which includes a convolutional module, a visual memory module, and a fully connected module, the visual memory module extracts features from the feature maps output by the convolutional module to extract age features with time dimension information, and the fully connected module combines the feature maps output by the visual memory module, enabling the age prediction model to learn the aging situation of the same person as they age, effectively reducing the deviation between the predicted age and the real age. On the other hand, by training the age prediction model with an age constraint loss function, and based on the trained age prediction model, performing age prediction on the target image, the embodiments of the present application can improve the accuracy of face age prediction.
[0168] The processor 601 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a hardware chip, or any combination thereof; it may also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0169] The memory 602, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the age prediction method in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory 602, the processor 601 can implement the age prediction method in any of the following method embodiments. Specifically, the memory 602 may include a volatile memory (VM), such as a random access memory (RAM); the memory 602 may also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or other non-transitory solid-state storage devices; the memory 502 may also include a combination of the above types of memories.
[0170] In the embodiments of the present application, the memory 602 may further include a memory remotely set relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0171] In an embodiment of the present application, the electronic device 60 may further include components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The electronic device 60 may further include other components for implementing the functions of the device, which will not be elaborated here.
[0172] An embodiment of the present application also provides a computer-readable storage medium, such as a memory including program code. The above program code can be executed by a processor to complete the age prediction method in the above embodiment. For example, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0173] An embodiment of the present application also provides a computer program product, which includes one or more program codes stored in a computer-readable storage medium. The processor of the electronic device reads the program code from the computer-readable storage medium, and the processor executes the program code to complete the method steps of the age prediction method provided in the above embodiment.
[0174] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiment can be completed by hardware, or can be completed by hardware related to program code. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0175] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disc, a Read-Only Memory (ROM), or a Random Access Memory (RAM), etc.
[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other changes in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An age prediction method, characterized in that, Including: Obtaining a plurality of sets of face images, where the set of face images includes at least real face images corresponding to the same person in several years; Constructing an age prediction model according to the real face images, the age prediction model including a convolution module, a visual memory module, and a fully connected module, where the convolution module is used to input multiple real face images of at least the same person, and perform feature extraction on each real face image to output a first feature map corresponding to each real face image; the visual memory module is used to receive the first feature map output by the convolution module, and extract age features with time dimension information to output a second feature map; the fully connected module is used to combine the second feature maps output by the visual memory module; The convolution module includes a plurality of convolutional layers, and the plurality of convolutional layers are used to extract first feature maps of multiple real face images, where each convolutional layer is used to extract a first feature map of one real face image; The visual memory module includes a plurality of visual memory layers, each visual memory layer is connected to a convolutional layer in a one-to-one correspondence, and each visual memory layer is used to receive the first feature map output by the convolutional layer corresponding to it and the first feature map output by another convolutional layer to extract the feature information of the two first feature maps in the time dimension; The fully connected module includes a plurality of fully connected layers, each fully connected layer is connected to a visual memory layer in a one-to-one correspondence, and is used to combine the second feature maps output by the plurality of visual memory layers to generate a combined feature map; The age prediction model further includes: A plurality of classification layers, each classification layer is connected to a fully connected layer in a one-to-one correspondence, and is used to predict the age prediction value corresponding to the combined feature map output by the fully connected layer; Wherein, each real face image corresponds to a channel, and each channel corresponds to a convolutional layer, a visual memory layer, a fully connected layer, and a classification layer; The age prediction model includes: Multiple groups of visual memory layers, and one visual memory layer in each group of visual memory layers is connected in series with one visual memory layer in another group of visual memory layers to form a series structure; Wherein, one visual memory layer in the next group of visual memory layers is used to obtain the second feature maps output by at least two visual memory layers in the previous group of visual memory layers, so as to further process the second feature maps output by the at least two visual memory layers, obtain the second feature maps and input them into the next group of visual memory layers, and so on until the second feature maps are output to the last group of visual memory layers; Training the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model; Performing age prediction on a target image based on the trained age prediction model to obtain an age prediction value of the person in the target image.
2. The method according to claim 1, characterized in that, The age prediction model includes three channels, namely an upper channel, a middle channel, and a lower channel, where each channel corresponds to a real face image, and the real ages corresponding to the real face images of different channels are different, and the age prediction value output by the classification layer corresponding to the middle channel is used as the age prediction value output by the age prediction model.
3. The method according to claim 1, characterized in that, The age constraint loss function is: Among them, is the weight parameter of the middle channel, is the weight parameter of the upper channel, is the weight parameter of the lower channel, is the total number of samples, is the true age value of the middle channel of the j-th sample, is the predicted age value of the middle channel of the j-th sample, is the true age value of the upper channel of the j-th sample, is the predicted age value of the upper channel of the j-th sample, is the true age value of the lower channel of the j-th sample, is the predicted age value of the lower channel of the j-th sample.
4. The method according to claim 1, characterized in that, Training the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model includes: Performing iterative training on the age prediction model based on the age constraint loss function; If the number of iterations is greater than a first number threshold, or the loss of the age prediction model is less than a first loss threshold, stop the iterative training to obtain a trained age prediction model.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: Preprocessing the real face images in the face image set, including: Obtaining the central coordinates of the left and right eyeballs in the real face image according to the face key point algorithm; Calculating the angle between the line connecting the central coordinates of the left and right eyeballs in the real face image and the horizontal direction; Rotating the real face image by the angle with the central coordinates of the left and right eyeballs in the real face image as the base points; Cropping the face area in the rotated real face image and resizing it to a preset resolution to obtain a preprocessed real face image.
6. An age prediction device, characterized in that, Including: A set construction unit for obtaining a plurality of face image sets, where each face image set includes real face images corresponding to at least the same person in several years; A model construction unit for constructing an age prediction model according to the real face images. The age prediction model includes a convolution module, a visual memory module, and a fully connected module. The convolution module is used to input multiple real face images of at least the same person and extract features from each real face image to output a first feature map corresponding to each real face image. The visual memory module is used to receive the first feature maps output by the convolution module and extract age features with time dimension information to output second feature maps. The fully connected module is used to combine the second feature maps output by the visual memory module; The convolution module includes a plurality of convolutional layers for extracting first feature maps of multiple real face images, where each convolutional layer is used to extract a first feature map of one real face image; The visual memory module includes a plurality of visual memory layers, each visual memory layer is connected to a convolutional layer in a one-to-one correspondence. Each visual memory layer is used to receive the first feature map output by its corresponding convolutional layer and the first feature map output by another convolutional layer to extract the feature information of the two first feature maps in the time dimension; The fully connected module includes a plurality of fully connected layers, each fully connected layer is connected to a visual memory layer in a one-to-one correspondence, and is used to combine the second feature maps output by multiple visual memory layers to generate a combined feature map; The age prediction model further includes: A plurality of classification layers, each classification layer is connected to a fully connected layer in a one-to-one correspondence, and is used to predict the age prediction value corresponding to the combined feature map output by the fully connected layer; Wherein, each real face image corresponds to a channel, and each channel corresponds to a convolutional layer, a visual memory layer, a fully connected layer, and a classification layer; The age prediction model includes: Multiple groups of visual memory layers, and one visual memory layer in each group of visual memory layers is connected in series with one visual memory layer in another group of visual memory layers to form a series structure; Among them, one visual memory layer in the next group of visual memory layers is used to obtain the second feature maps output by at least two visual memory layers in the previous group of visual memory layers, so as to further process the second feature maps output by the at least two visual memory layers, obtain the second feature maps and input them into the next group of visual memory layers, and so on until the second feature maps are output to the last group of visual memory layers; A model training unit, configured to train the age prediction model based on a preset age constraint loss function to obtain a trained age prediction model; An age prediction unit, configured to predict the age of a target image based on the trained age prediction model to obtain an age prediction value of the person in the target image.
7. An electronic device, characterized in that, including: A memory and one or more processors, where the one or more processors are configured to execute one or more computer programs stored in the memory, and when the one or more processors execute the one or more computer programs, the electronic device implements the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method according to any one of claims 1-5.
Citation Information
Patent Citations
Target model training method, face image generation method and related device
CN113221645A
Training method of image processing model, image processing model and computer equipment
CN113762117A