Age prediction model training method and device, equipment, medium and product
Through the improved cross entropy loss function and total loss function, combined with classification and regression loss, the problem of large age prediction deviation in the prior art is solved, the accuracy and convergence speed of the age prediction model are improved, and it is suitable for the continuity problem of age or age group.
Patent Information
- Application Number
- CN202411795207.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, the deviation caused by the direct regression age prediction method is large, and because the age or age group is continuity, it is not suitable for age prediction of classification models.
The parameters of the age prediction model are updated through the total loss function using an improved cross-entropy loss function, combining classification and regression losses, taking into account the continuity of age or age groups.
The accuracy and convergence speed of the age prediction model are improved, and are suitable for the continuity problems of age or age groups.
Smart Images

Figure CN119964212A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to an age prediction model training method, device, equipment, medium and product. Background Art
[0002] With the continuous development of deep learning, deep learning of facial images can predict the age corresponding to the face.
[0003] In related technologies, most studies focus on improving the optimization effect of regression networks and the improvement of loss functions, and directly regressing age prediction based on face images. However, due to factors such as image quality, direct regression age prediction may result in large deviations. In addition, whether it is age or age range, it is continuous in itself and is also not suitable for age prediction by classification models. Summary of the invention
[0004] In view of this, the present invention provides an age prediction model training method, device, equipment, medium and product to solve the problems in the related art of large deviation caused by direct regression age prediction based on face images, and the problem that classification models are not applicable due to the continuity of age or age groups.
[0005] In a first aspect, the present invention provides an age prediction model training method, the method comprising:
[0006] Input the face image samples with age labels and classification labels into the target age prediction model to obtain the age classification results and age prediction results corresponding to the face image samples;
[0007] Apply the target classification loss function to the age classification results and classification labels corresponding to the face image samples to obtain the classification loss corresponding to the face image samples;
[0008] Applying the target regression loss function to the age prediction result and age label corresponding to the face image sample to obtain the regression loss corresponding to the face image sample;
[0009] Among them, the target classification loss function is:
[0010]
[0011] Among them, y ic =-s(1-tanh(|cy true |)); N is the total number of face image samples; c is the category; M is the number of categories; i is the face image sample; p ic is the predicted probability that face image sample i belongs to category c; y ic is the true probability that face image sample i belongs to category c; ytrue is the true category to which the face image sample i belongs; s is an adjustable scaling factor, and 0 <s≤1;
[0012] According to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample, the total loss corresponding to the face image sample is obtained, and the total loss corresponding to the face image sample is used to update the parameters of the target age prediction model.
[0013] In an optional implementation, before inputting the face image samples carrying the age label and the classification label into the target age prediction model, the method further includes:
[0014] Obtain multiple original face image samples, and assign an age label corresponding to each original face image sample according to the real age corresponding to the face in each original face image sample;
[0015] Perform face alignment on each original face image sample;
[0016] According to the age labels, the plurality of original face image samples after face alignment are classified according to age groups to obtain a plurality of category intervals;
[0017] According to the true probability that each original face image sample after face alignment belongs to each category, a corresponding classification label is assigned to each original image sample, and multiple face image samples carrying age labels and classification labels are obtained.
[0018] In an optional implementation, inputting the face image samples carrying the age labels and classification labels into the target age prediction model to obtain the age classification results and age prediction results corresponding to the face image samples includes:
[0019] Input the face image samples with age labels and classification labels into the classification model to obtain the predicted probability that the face image samples belong to each category;
[0020] According to the weight and prediction probability corresponding to each category, the age prediction result corresponding to the face image sample is obtained.
[0021] In an optional implementation, the weight corresponding to each category is the middle value of each category interval.
[0022] In an optional implementation, obtaining the age prediction result corresponding to the face image sample according to the weight and prediction probability of each category includes:
[0023] Multiply the weight and prediction probability corresponding to each category and sum them up to obtain the age prediction result corresponding to the face image sample.
[0024] In an optional implementation, obtaining the total loss corresponding to the face image sample according to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample includes:
[0025] Loss = αL 分类 +(1-α)L 回归
[0026] Among them, L 分类 is the classification loss corresponding to the face image sample, L 回归 is the regression loss corresponding to the face image sample, α is a constant and 0<α<1.
[0027] In a second aspect, the present invention provides an age prediction model training device, the device comprising:
[0028] The age prediction module is used to input the face image samples carrying age labels and classification labels into the target age prediction model to obtain the age classification results and age prediction results corresponding to the face image samples;
[0029] A classification loss function calculation module is used to apply the target classification loss function to the age classification result and classification label corresponding to the face image sample to obtain the classification loss corresponding to the face image sample;
[0030] A regression loss function calculation module, used for applying the target regression loss function to the age prediction result and the age label corresponding to the face image sample to obtain the regression loss corresponding to the face image sample;
[0031] Among them, the target classification loss function is:
[0032]
[0033] Among them, y ic =-s(1-tanh(|cy true |)); N is the total number of face image samples; c is the category; M is the number of categories; i is the face image sample; p ic is the predicted probability that face image sample i belongs to category c; y ic is the true probability that face image sample i belongs to category c; y true is the true category to which the face image sample i belongs; s is an adjustable scaling factor, and 0 <s≤1;
[0034] The total loss calculation module is used to obtain the total loss corresponding to the face image sample according to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample. The total loss corresponding to the face image sample is used to update the parameters of the target age prediction model.
[0035] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0036] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to cause a computer to execute the method of the first aspect or any corresponding embodiment thereof.
[0037] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the method of the first aspect or any corresponding embodiment thereof.
[0038] The embodiment of the present invention adopts an improved cross entropy loss function to take into account the continuity of age or age group itself, and also takes the prediction of adjacent ages or age groups into account in the loss value calculation, that is, adjacent ages or age groups are also considered to have a positive contribution to the correct prediction, which not only accelerates the convergence speed of the age prediction model, but also improves the accuracy of the age prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 is a flow chart of an age prediction model training method according to an embodiment of the present invention;
[0041] Figure 2 is a structural block diagram of an age prediction model training device according to an embodiment of the present invention;
[0042] Figure 3 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0044] As an important aspect of describing personnel portraits, age attributes are widely used, such as screening suspicious individuals based on age and recommending advertisements to customer groups based on age.
[0045] With the continuous development of deep learning, deep learning of facial images can predict the age corresponding to the face.
[0046] In related technologies, most studies focus on improving the optimization effect of regression networks and the improvement of loss functions, and directly regressing age prediction based on face images. However, as people age, their faces behave differently. In childhood, it manifests as changes in face shape, and in adulthood, it manifests as changes in facial texture. Due to the differences in the imaging quality of facial images, low-quality faces may "look" like changes in face shape (serious camera distortion) or texture changes (a lot of noise and lines on the photo), which will lead to large age deviations in direct regression age prediction.
[0047] In addition, since the output of age prediction is an integer age, the problem can be considered as a classification problem. For example, the age range of 0-99 can be divided into 100 categories. However, when directly classifying age, its continuity is often ignored. This is because in conventional classification tasks, the categories are completely independent, such as cat and dog classification. Conventional classification tasks generally use softmax function (normalized exponential function) + cross entropy loss function for classification tasks. Suppose that the three categories of cat, dog, and bird are classified, and their corresponding one-hot encodings are
[100] ,
[010] , and
[001] respectively. In the current training round, a picture of a cat is predicted. Suppose the prediction result P1 is [0.2 0.6 0.2] or P2 is [0.2 0.20.6]. Obviously, both prediction results are wrong, and the loss values calculated by the cross entropy loss function are the same. However, if we use the above method to directly classify age, assuming that they correspond to 0, 1, and 2 years old respectively, in fact, 1 year old is closer to 0 years old than 2 years old, but the losses caused by 1 and 2 years old are the same. This is because the softmax function only emphasizes the maximization of the difference between categories, but ignores the continuous correlation of the age problem itself.
[0048] Similarly, after age groups are divided, there is continuity between the age groups, which is also not suitable for classification models.
[0049] In view of this, the present invention provides an age prediction model training method. By adopting an improved cross-entropy loss function, the continuity of age or age group itself can be taken into account, and the prediction of adjacent ages or age groups is also taken into account in the loss value calculation, that is, adjacent ages or age groups are also considered to have a positive contribution to the correct prediction, which not only accelerates the convergence speed of the age prediction model, but also improves the accuracy of the age prediction model.
[0050] According to an embodiment of the present invention, an embodiment of an age prediction model training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0051] Figure 1 FIG. 1 is a flow chart of an age prediction model method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0052] Step S101, inputting a face image sample carrying an age label and a classification label into a target age prediction model to obtain an age classification result and an age prediction result corresponding to the face image sample.
[0053] In step S101, the age classification result is an intermediate result obtained after a face image sample carrying an age label and a classification label is input into a target age prediction model and in the process of performing age prediction according to the target age prediction model.
[0054] The age prediction result is the output result of the target age prediction model after the face image sample carrying the age label and the classification label is input into the target age prediction model.
[0055] Among them, the age labels and classification labels carried by face image samples can be manually labeled or machine labeled.
[0056] Step S102, applying the target classification loss function to the age classification result and classification label corresponding to the face image sample to obtain the classification loss corresponding to the face image sample.
[0057] In step S102, the classification loss is calculated based on the age classification result corresponding to the face image sample and the classification label carried by the face image sample. The classification loss is obtained by bringing the age classification result corresponding to the face image sample and the classification label carried by the face image sample into the target classification loss function.
[0058] Specifically, the target classification loss function is:
[0059]
[0060] Among them, y ic =-s(1-tanh(|cy true |)); N is the total number of face image samples; c is the category; M is the number of categories; i is the face image sample; p ic is the predicted probability that face image sample i belongs to category c; y ic is the true probability that face image sample i belongs to category c; y true is the true category to which the face image sample i belongs; s is an adjustable scaling factor, and 0 <s≤1。
[0061] It should be noted that the classification labels can be classified by age or by age group.
[0062] The embodiment of the present invention improves the target classification loss function, taking into account the continuity of age or age group, and converts y ic A symbol function with a value of 0 or 1 is used instead. If the true category y of the face image sample i true is equal to c, then y ic Take 1 if the true category y of face image sample i true is not equal to c, then y ic Take 0. And, when y true When it is not equal to c, by adjusting the scaling factor s, the prediction probability of adjacent ages or adjacent age groups can be considered, that is, falling on the surrounding ages or age groups, especially the adjacent ages or age groups, are also considered to be approximately correctly classified, that is, the prediction of the adjacent ages or age groups is also taken into account in the calculation of the loss value of the face image sample i, that is to say, the adjacent ages or age groups are also considered to have a positive contribution to the correct prediction, which not only accelerates the convergence speed of the age prediction model, but also improves the accuracy of the age prediction model.
[0063] Step S103, applying the target regression loss function to the age prediction result and age label corresponding to the face image sample to obtain the regression loss corresponding to the face image sample.
[0064] In step S103, the regression loss is calculated based on the age prediction result corresponding to the face image sample and the age label carried by the face image sample. The regression loss is obtained by bringing the age prediction result corresponding to the face image sample and the age label carried by the face image sample into the target regression loss function.
[0065] It should be noted that the target regression loss function can be selected according to the specific situation and is not specifically limited here.
[0066] Step S104, obtaining the total loss corresponding to the face image sample according to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample, and the total loss corresponding to the face image sample is used to update the parameters of the target age prediction model.
[0067] Among them, the total loss is determined based on the classification loss and regression loss. In the training age prediction model, the parameters of the target age prediction model are updated by obtaining the total loss. When the total loss reaches the expected threshold, the training ends.
[0068] In some optional implementations, before inputting the face image samples carrying the age labels and the classification labels into the target age prediction model, the age prediction model training method further includes:
[0069] Step a: obtaining a plurality of original face image samples, and assigning an age label corresponding to each original face image sample according to the real age corresponding to the face in each original face image sample.
[0070] Among them, the embodiments of the present invention can collect open source face datasets such as AFAD (Asian Face) dataset, FG-Net dataset, MORPH dataset and FairFace dataset with pre-annotated age, and can also collect data by itself in combination with usage scenarios, and collect data based on the age on the ID card or the age entered by the collection object.
[0071] Step b: perform face alignment on each original face image sample.
[0072] Among them, face alignment is to adjust the position and orientation of the face in each original face image so that the face in each original face image sample maintains the same position and the specific positions such as the eyes, nose, mouth, etc. are aligned. By performing face alignment on each original face image sample, the embodiment of the present invention can solve the problems of scale change, rotation change and posture change in the original face image sample, thereby improving the accuracy and stability of the age prediction model.
[0073] As an example, use the pre-trained face detector and face key point detector to perform face detection on each original face image, and obtain the face region coordinate frame and face key point coordinates (x src ,y src ) The key points of the face are generally the center of the left and right eyes, the tip of the nose, and the left and right corners of the mouth, that is, five sets of coordinates. At the same time, a standard-sized face is preset, and the coordinates of the standard face key points are (x tgt ,y tgt), and there are also five sets of coordinates. Among them, the face detector and the face key point detector can be selected according to specific needs, and are not specifically limited here.
[0074] According to the coordinates of the facial key points (x src ,y src ), and the preset standard facial key point coordinates (x tgt ,y tgt ), the affine transformation matrix between points is:
[0075]
[0076] According to the five sets of coordinates, the transformation matrix M can be obtained:
[0077]
[0078] Among them, a1 is the scaling factor in the horizontal direction, b1 is the shear factor in the horizontal direction, t1 is the translation in the horizontal direction, a2 is the shear factor in the vertical direction, b2 is the scaling factor in the horizontal direction, and t2 is the translation in the vertical direction.
[0079] Therefore, according to the standard facial key point coordinates (x tgt ,y tgt ) The coordinates of the facial key points of the original face image are obtained through the transformation matrix M (x src ,y src ), and use the pixel values corresponding to the facial key point coordinates to update the pixel values at the standard facial key point coordinate positions to complete face alignment.
[0080] Step c: According to the age labels, the multiple original face image samples after face alignment are classified according to age groups to obtain multiple category intervals.
[0081] Specifically, the age division length is set to STEP, and STEP can be fixed or variable. The embodiment of the present invention takes the fixed STEP as an example to illustrate the classification of multiple categories.
[0082] First, count the age ranges corresponding to all faces in all original face image samples [L min , L max ], and then divided into categories according to age groups, that is, [L min , L min+STEP ), [L min+STEP , L min+2*STEP ),…,[L min+n*STEP , L min+(n+1)*STEP ), ..., are respectively categories 0, 1, 2, ..., n, .... Among them, L max into the last category.
[0083] Step d: according to the true probability that each original face image sample after face alignment belongs to each category, a corresponding classification label is assigned to each original image sample, thereby obtaining multiple face image samples carrying age labels and classification labels.
[0084] In some optional implementations, a face image sample carrying an age label and a classification label is input into a target age prediction model to obtain an age classification result and an age prediction result corresponding to the face image sample, including:
[0085] The face image samples carrying age labels and classification labels are input into the classification model to obtain the predicted probability that the face image samples belong to each category.
[0086] Among them, the classification model can be set according to the specific deployment scenario, which can be resnet50 network, mobilenetV3, mobilenetV3-small, etc., and is not specifically limited here.
[0087] According to the weight and prediction probability corresponding to each category, the age prediction result corresponding to the face image sample is obtained.
[0088] In some optional implementations, the weight corresponding to each category is the middle value of each category interval.
[0089] In some optional implementations, obtaining the age prediction result corresponding to the face image sample according to the weight and prediction probability of each category includes:
[0090] Multiply the weight and prediction probability corresponding to each category and sum them up to obtain the age prediction result corresponding to the face image sample.
[0091] As an example, suppose the predicted probability of class c is y pc The age prediction result is the weighted sum of the prediction probabilities of each category, that is, the prediction probability corresponding to each category is multiplied by the corresponding age weight and then added together to obtain the final age prediction result Y est .
[0092] Specific:
[0093]
[0094] Among them, y wc is the age weight of category c. Without loss of generality, the median of the category interval can be used as the age weight. Take any age interval as an example: [L min+n*STEP , L min+(n+1)*STEP ), whose weight is L min+n*STEP+1 / 2*STEP , where n is an integer greater than or equal to 0.
[0095] In some optional implementations, the total loss corresponding to the face image sample is obtained according to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample, including:
[0096] Loss = αL 分类 +(1-α)L 回归
[0097] Among them, L 分类 is the classification loss corresponding to the face image sample, L 回归 is the regression loss corresponding to the face image sample, α is a constant and 0<α<1.
[0098] As an example, first, using a fixed STEP = 5 as an example, we divide the age into multiple categories, and get a total of 15 categories, numbered from '0' to '14'. Secondly, we select mobilenetV3-small as the backbone network, set the number of categories M to 15, and select Loss = αL as the loss function. 分类 +(1-α)L 回归 , α is set to 0.6. epch is set to 50, model training is started, and the final model training accuracy is 96%. Finally, the predicted probabilities of each age category are weighted summed and the final age prediction value is regressed. At this time, the regression loss is about 3.
[0099] In summary, the embodiment of the present invention converts the direct regression age into a classification + regression problem by dividing the age into groups, and at the same time proposes an improved cross entropy loss function calculation, taking the interval continuity into account in the age prediction model, thereby improving the training convergence speed of the age prediction model and the accuracy of age prediction. At the same time, the total loss function of classification + regression is used to evaluate the age prediction model, which further improves the training convergence speed of the age prediction model. In addition, the grouping step size can be flexibly set to fine-tune multiple classifications, control the number of categories, and adapt to different usage scenarios, and the scaling factor of the segmented continuous interval can be flexibly set to improve the training effect of the age prediction model on different face image samples.
[0100] In the present embodiment, a device for training an age prediction model is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions thereof will not be repeated. As used below, the term "module" may implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and contemplated.
[0101] This embodiment also provides an age prediction model training device, such as Figure 2 As shown, including:
[0102] The age prediction module 201 is used to input the face image samples carrying the age label and the classification label into the target age prediction model to obtain the age classification result and the age prediction result corresponding to the face image samples.
[0103] The classification loss function calculation module 202 is used to apply the target classification loss function to the age classification results and classification labels corresponding to the face image samples to obtain the classification loss corresponding to the face image samples.
[0104] The regression loss function calculation module 203 is used to apply the target regression loss function to the age prediction result and age label corresponding to the face image sample to obtain the regression loss corresponding to the face image sample.
[0105] Among them, the target classification loss function is:
[0106]
[0107] Among them, y ic =-s(1-tanh(|cy true |)); N is the total number of face image samples; c is the category; M is the number of categories; i is the face image sample; p ic is the predicted probability that face image sample i belongs to category c; y ic is the true probability that face image sample i belongs to category c; y true is the true category to which the face image sample i belongs; s is an adjustable scaling factor, and 0 <s≤1。
[0108] The total loss calculation module 204 is used to obtain the total loss corresponding to the face image sample according to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample. The total loss corresponding to the face image sample is used to update the parameters of the target age prediction model.
[0109] In some optional implementations, the age prediction model training device also includes a face image sample acquisition module 205.
[0110] The face image sample acquisition module 205 further comprises:
[0111] The original face image acquisition unit 2051 is used to acquire multiple original face image samples, and assign an age label corresponding to each original face image sample according to the real age corresponding to the face in each original face image sample.
[0112] The pre-processing unit 2052 is used to perform face alignment on each original face image sample.
[0113] The age segmentation unit 2053 is used to classify the plurality of original face image samples after face alignment according to age groups according to age labels to obtain a plurality of category intervals.
[0114] The face image sample confirmation unit 2054 is used to assign a corresponding classification label to each original image sample according to the true probability that each original face image sample belongs to each category after face alignment, and obtain multiple face image samples carrying age labels and classification labels.
[0115] In some optional implementations, the age prediction module 201 further includes:
[0116] The category prediction probability acquisition unit 2011 is used to input the face image samples carrying the age label and the category label into the classification model to obtain the prediction probability that the face image samples belong to each category.
[0117] The age prediction result obtaining unit 2012 is used to obtain the age prediction result corresponding to the face image sample according to the weight and prediction probability corresponding to each category.
[0118] In some optional implementations, the weight corresponding to each category is the middle value of each category interval.
[0119] In some optional implementations, the age prediction result acquisition unit 2012 is further configured to multiply the weight and prediction probability corresponding to each category and sum them up to obtain the age prediction result corresponding to the face image sample.
[0120] In some optional implementations, the total loss calculation module 204 further includes:
[0121] Loss = αL 分类 +(1-α)L 回归
[0122] Among them, L 分类 is the classification loss corresponding to the face image sample, L 回归 is the regression loss corresponding to the face image sample, α is a constant and 0<α<1.
[0123] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0124] The age prediction model training device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0125] The embodiment of the present invention also provides a computer device, such as Figure 3 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface).
[0126] In some optional embodiments, if desired, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 3 A processor 10 is taken as an example.
[0127] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0128] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0129] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0130] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0131] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 3 The example of connecting through bus is taken in the following.
[0132] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0133] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0134] A portion of the embodiments of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in a computer-readable medium includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which the computer program instructions are executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0135] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for training an age prediction model, characterized in that: The method comprises: Input the face image samples with age labels and classification labels into the target age prediction model to obtain the age classification results and age prediction results corresponding to the face image samples; Apply the target classification loss function to the age classification results and classification labels corresponding to the face image samples to obtain the classification loss corresponding to the face image samples; Applying the target regression loss function to the age prediction result and age label corresponding to the face image sample to obtain the regression loss corresponding to the face image sample; Among them, the target classification loss function is: Among them, y ic =-s(1-tanh(cy true |)); N is the total number of face image samples; c is the category; M is the number of categories; i is the face image sample; p ic is the predicted probability that face image sample i belongs to category c; y ic is the true probability that face image sample i belongs to category c; y true is the true category to which the face image sample i belongs; s is an adjustable scaling factor, and 0 <s≤1; According to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample, the total loss corresponding to the face image sample is obtained, and the total loss corresponding to the face image sample is used to update the parameters of the target age prediction model.
2. The method according to claim 1, characterized in that Before inputting the face image samples carrying the age labels and classification labels into the target age prediction model, the method further includes: Obtain multiple original face image samples, and assign an age label corresponding to each original face image sample according to the real age corresponding to the face in each original face image sample; Perform face alignment on each original face image sample; According to the age labels, the plurality of original face image samples after face alignment are classified according to age groups to obtain a plurality of category intervals; According to the true probability that each original face image sample after face alignment belongs to each category, a corresponding classification label is assigned to each original image sample, and multiple face image samples carrying age labels and classification labels are obtained.
3. The method according to claim 2, characterized in that The face image samples carrying age labels and classification labels are input into the target age prediction model to obtain age classification results and age prediction results corresponding to the face image samples, including: Input the face image samples with age labels and classification labels into the classification model to obtain the predicted probability that the face image samples belong to each category; According to the weight and prediction probability corresponding to each category, the age prediction result corresponding to the face image sample is obtained.
4. The method according to claim 3, characterized in that The weight corresponding to each category is the middle value of each category interval.
5. The method according to claim 3, characterized in that: The age prediction result corresponding to the face image sample is obtained according to the weight and prediction probability of each category, including: Multiply the weight and prediction probability corresponding to each category and sum them up to obtain the age prediction result corresponding to the face image sample.
6. The method according to claim 1, characterized in that The total loss corresponding to the face image sample is obtained according to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample, including: Loss=ααL 分类 +(1-α)L 回归 Among them, L 分类 is the classification loss corresponding to the face image sample, L 回归 is the regression loss corresponding to the face image sample, α is a constant and 0<α<1.
7. An age prediction model training device, characterized in that: The device comprises: The age prediction module is used to input the face image samples carrying age labels and classification labels into the target age prediction model to obtain the age classification results and age prediction results corresponding to the face image samples; A classification loss function calculation module is used to apply the target classification loss function to the age classification result and classification label corresponding to the face image sample to obtain the classification loss corresponding to the face image sample; A regression loss function calculation module, used for applying the target regression loss function to the age prediction result and the age label corresponding to the face image sample to obtain the regression loss corresponding to the face image sample; Among them, the target classification loss function is: Among them, y ic =-s(1-tanh(cy true |)); N is the total number of face image samples; c is the category; M is the number of categories; i is the face image sample; p ic is the predicted probability that face image sample i belongs to category c; y ic is the true probability that face image sample i belongs to category c; y true is the true category to which the face image sample i belongs; s is an adjustable scaling factor, and 0 <s≤1; The total loss calculation module is used to obtain the total loss corresponding to the face image sample according to the classification loss corresponding to the face image sample and the regression loss corresponding to the face image sample. The total loss corresponding to the face image sample is used to update the parameters of the target age prediction model.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the method according to any one of claims 1 to 6.