A Generative Feature Order Regularization Method for Deep Ordinal Regression Models

Through the generative feature order regularization method PnP-FOR, the problem that feature representation contains irrelevant information in deep order regression is solved, and the performance of image classification tasks is improved, especially in scenarios such as face age estimation, medical image classification and historical image age classification.

CN114494785BActive Publication Date: 2025-07-08FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210116188.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2025-07-08
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

现有深度有序回归方法在学习特征表示时,包含了太多与有序标签无关的信息,损害了输入空间中内在的一维有序关系的学习。

Method used

A generative feature order regularization method PnP-FOR is proposed. By calculating the Kullback-Leibler divergence between the feature probability distribution and the ordered label distribution in the embedded space, it constrains the intrinsic dimensions of the feature representation, adopts batch training method and is compatible with the existing methods, and does not change the network architecture.

Benefits of technology

The performance of image classification tasks is improved, especially in scenarios such as face age estimation, medical image classification and historical image age classification. The feature representation retains orderly information, and improves the learning effect of classification accuracy and orderly relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494785B_ABST
    Figure CN114494785B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image classification, and specifically relates to a generative feature order regularization method for a deep ordered regression model. The present invention maps an input image into a low-dimensional feature representation through a deep convolutional neural network, and calculates the distances between the low-dimensional feature representations of all samples in a batch in the corresponding low-dimensional space; then calculates the distances between the ordered labels of all samples in the batch; normalizes the obtained feature representation distance vectors and label distance vectors between all samples in the batch and other samples respectively; calculates the divergence of the normalized feature vectors and label distance vectors to constrain the distribution of features to be consistent with the distribution of ordered labels in the embedding space, that is, to ensure the orderliness of features; the final loss function of the model includes an ordered regression loss and this KL divergence loss; the method of the present invention helps to improve the classification performance in various task scenarios (such as face age estimation, medical image classification, historical image age classification, etc.).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image classification, and particularly relates to an image classification method with ordered labels. Background Art

[0002] Ordinal regression (OR) is a classic problem in machine learning, specifically for predicting data with ordered labels. Typical applications include face age estimation, where the image label is the age of the face, from small to large; rating a movie, such as from one star to ten stars; disease diagnosis based on medical images, etc. Due to the orderliness of the labels, OR is an intermediate problem between classification and regression [1][2].

[0003] In the past few years, deep neural networks have promoted the development of OR methods [1][2][3][4][5][6], namely deep ordinal regression methods. These methods mainly focus on modeling the mapping from the feature representation obtained from the input data to the ordered label space. For example, a popular method for modeling ordered information is to model ordinal regression as multiple binary classification problems [1][2][3][7][8], because the ordered labels are monotonically distributed. Some methods impose constraints on the output logits, which often force the output of the model to follow a predefined distribution, such as the binomial distribution, Poisson distribution, or Beta distribution [8][9]. In addition, by converting the ordered labels into label / ordered distributions [6]

[10] , traditional classification models can be easily adapted to the deep ordinal regression task. However, through experiments, it is found that the feature representation learned from the input images contains too much information irrelevant to the ordered labels, seriously damaging the essential goal of existing deep ordinal regression methods, that is, learning one-dimensional ordered relationships.

[0004] In the present invention, a novel generative perspective is proposed for deep ordinal regression. Different from the current deep ordinal regression learning the mapping from the input to the features and then to the ordered label space, the generative view of the present invention assumes that the input image can be generated from the ordered labels. In other words, in the context of ordinal regression, even in a high-dimensional space, the input image should have only one inherent dimension - ordered label information.

[0005] Therefore, the present invention proposes a plug-and-play feature order regularization method PnP-FOR for depth order regression, which enables the feature representations learned by deep learning models to retain order information. Specifically, the goal is to calculate the Kullback-Leibler divergence between the constructed feature probability distribution and the ordered label distribution in the embedding space. The proposed PnP-FOR has the advantages of being batch-training based, easy to implement, and compatible with existing methods without modifying the network architecture. The experimental results in three scenarios, namely face age estimation, medical image classification, and historical image age classification, verify the effectiveness of the PnP-FOR proposed by the present invention when combined with existing depth order regression methods. Summary of the Invention

[0006] The object of the present invention is to propose a generative feature order regularization method for a depth order regression model to improve the performance of image classification tasks in multiple label order scenarios.

[0007] Depth order regression aims to predict the ordered labels of given samples using a deep neural network. Existing methods mainly focus on modeling the mapping from the input feature representation to the ordered label space. However, the feature representations in the embedding space learned by current methods contain too much information unrelated to the ordered labels, which seriously impairs the learning of the intrinsic one-dimensional order relationship in the input space. Therefore, the present invention proposes a novel generative perspective for depth ordinal regression and a plug-and-play feature order regularization method PnP-FOR to extract the order information in the embedding space. More specifically, from the perspective of a generative model, the input samples can be generated from the ordered labels, and even in a high-dimensional space, the feature representations should have an intrinsic one-dimensional order dimension. To this end, the present invention extracts the order information in the embedding space and constrains the intrinsic dimension of the feature representation by measuring the Kullback-Leibler divergence between the probability distributions constructed in the embedding space and the ordered label space. The proposed PnP-FOR is batch-training based, easy to implement, and can be well compatible with existing methods without modifying the network architecture. A large number of experimental results in three different scenarios - face age estimation, disease progression prediction based on medical images, and historical image age prediction - show that the proposed PnP-FOR improves the existing depth order regression methods. More importantly, the visualization of the PnP-FOR feature representations highlights the dominant position of the order information in the embedding space.

[0008] The generative feature order regularization method for a depth order regression model provided by the present invention includes:

[0009] First, map the input image into a low-dimensional feature representation through a deep convolutional neural network, and calculate the distances between the low-dimensional feature representations of all samples in a batch in the corresponding low-dimensional space;

[0010] Then, after obtaining the distances between samples in a batch during the forward propagation process, calculate the distances between the ordered labels of all samples in the batch at the same time; normalize the feature representation distance vectors and label distance vectors between all samples in a batch obtained respectively;

[0011] Then, calculate the (KL) divergence for the normalized feature vectors and label distance vectors to constrain the distribution of features to be consistent with the distribution of ordered labels in the embedding space, that is, to ensure the orderliness of features;

[0012] The final loss function of the model includes the general ordered regression loss and this KL divergence loss;

[0013] The generative feature ordered regularization method for a deep ordered regression model provided by the present invention specifically includes the following steps:

[0014] (1) Select a discriminative ordered regression model;

[0015] The existing deep ordered regression method is based on the following discriminant model:

[0016] p(Y|X) = g(f θ (X)); (1)

[0017] This also constructs a Markov chain:

[0018] X → Z = f θ (X) → Y; (2)

[0019] Among them, both X and Y are random variables, representing a sample point in the input space and the class label corresponding to this sample point respectively, f θ (·) is a feature extraction network, which is a convolutional neural network with parameters θ in the present invention, and g(·) is a label mapping function. Specifically, if there are N independent samples {(x1, y1), (x2, y2),..., (x N , y N )}, then the labels satisfy:

[0020] y ∈ {r1, r2,..., r K}, and r1 ≤ rx ≤ … ≤ r K ;

[0021] It can be obtained that the goal of the ordered regression model is to learn a mapping from to , where K represents the number of ordered labels;

[0022] (2) Model the ordinal regression from the perspective of the generative model;

[0023] The feature representation learned by the discriminant ordinal regression model in step (1) has many features irrelevant to the ordinal information, that is, it cannot guarantee the ordinal relationship of features in the latent space; while regarding X as being generated by the label Y, the generative model can be expressed as:

[0024] p(X, Y) = p(Y)p(X|Y), (3)

[0025] The conditional probability p(X|Y) indicates that the input variable X in the high-dimensional space and the ordinal label Y have the same intrinsic dimension in the latent space, that is, a one-dimensional vector; then making the feature vector Z in the latent space have the same dimension as Y will simplify the learning objective;

[0026] Regarding Y as the low-dimensional representation of X embedded in the high-dimensional manifold, then Y is the latent variable of X with an intrinsic dimension of 1. Assuming that the intrinsic dimension of the latent variable Z is also 1, the following Markov chain is obtained:

[0027] Y → Z → X, (4)

[0028] From the criterion of information theory, the information bottleneck criterion can be used to describe the relevant information between X and Y, that is, the mutual information (X; Y); Y implicitly determines the features in X that are relevant or irrelevant to the ordinal information, and the following Markov chain is obtained:

[0029] Y → X → Z, (5)

[0030] For an X, the optimal feature representation Z should capture the features most relevant to the category Y and make the prediction of X unaffected by irrelevant information; therefore, the mutual information between Y and Z should be maximized to maintain the ordinal information in Z; even so, the learned feature representation still contains a large amount of information unrelated to the order;

[0031] (3) Construct the probability distribution of the ordinal label;

[0032] Inspired by the stocastic neighbor embedding (SNE)

[23] , the present invention intends to reduce the dimension of the feature representation of the sample to align it with the one-dimensional distribution of the ordinal label. First, map the ordinal label to a probability distribution:

[0033]

[0034] Here, P ii= 0, the ordered label is converted into a probability, representing the distance between any two ordered labels, where y i , y j , y k , y l respectively represent the ordered labels corresponding to different samples in a batch;

[0035] (4) Construct the probability distribution of the feature representation;

[0036] Similar to the label probability distribution, the probability distribution of the feature representation is:

[0037]

[0038] Here, Q ii = 0, z = f θ (x), where z i , z j , z k , z l respectively represent the feature representations of different samples in a batch;

[0039] (5) Ordered feature regularization;

[0040] So far, the probability distribution representations of the sample feature expressions and the ordered labels are obtained; then the KL (Kullback-Leibler) divergence can be used to measure the distance between the probability distribution of the feature representation and the probability distribution of the label:

[0041]

[0042] Formula (8) aligns the feature distribution and the label distribution in the latent space; since the label space is globally ordered, the batch-training-based method can learn the global ordered information; this feature ordered regularization method can be flexibly embedded into the existing deep ordered regression methods without changing their structures;

[0043] (6) Construct the loss function;

[0044] The loss function used in the present invention is:

[0045]

[0046] where λ is a hyperparameter; represents the loss function involved in the existing deep ordered regression method, generally the cross-entropy loss between the output vector of the fully connected layer of the neural network and the label, or it may vary depending on the specific ordered regression method. For example, in the SORD method, the cross-entropy loss is calculated through the constructed soft label and the output vector of the fully connected layer [6]. To prevent overfitting, the present invention adds a Dropout module

[22] to the fully connected layer.

[0047] The feature extraction network used in the present invention is specifically all convolutional layer modules of the VGG-16 network

[17] ; a linear layer and a ReLU activation layer are connected after the last convolutional layer; in order to prevent overfitting, Dropout

[22] is added; the output of Dropout here is used as the feature representation of the input image and is used to calculate D in formula (9). KL Finally, this feature representation is input into a classifier, which is a linear layer with an input dimension of 4096 and an output dimension of the number K of ordered labels.

[0048] From the perspective of the generative model, the present invention constrains the feature representation in the deep learning-based ordinal regression method, which can keep the input image in a one-dimensional ordered distribution in the latent space, thereby improving the performance of the existing deep ordinal regression method; during the training process of the neural network, for the low-dimensional feature representation obtained from a batch of samples, that is, the output of the last convolutional layer, calculate the distances between all samples in the batch, then construct a probability distribution for each sample, which represents the probability that the sample selects other samples in the batch as neighbors; similarly, calculate the corresponding probability distribution for the ordered labels corresponding to the samples in the same batch, and this probability distribution represents the global ordered information; normalize the feature probability distribution and the label probability distribution respectively, and set the distance calculated for the same sample itself to 0; in order to keep the features ordered, minimize the KL divergence between the feature probability distribution and the label probability distribution to make the features keep ordered in the latent space; the method of the present invention can be well applied to the currently popular batch-based neural network training framework, and since the label probability distribution represents the global ordered information, the samples in each batch can also learn the global ordered information; in the tasks of face age estimation, medical image classification, and historical image age classification, the effectiveness of the proposed method is verified; the visualization results obtained by t-SNE can clearly show the one-dimensional ordered distribution of the sample points, which is significantly improved compared with the existing ordinal regression methods

[11] .

[0049] The method of the present invention is a "plug-and-play" regularization method, that is, it can be flexibly embedded into all existing deep learning-based ordinal regression methods without changing the model structure of the existing methods; the method of the present invention helps to improve the classification performance of the existing deep ordinal regression methods in various task scenarios (such as face age estimation, medical image classification, historical image age classification, etc.). Brief Description of the Drawings

[0050] Figure 1 is the overall framework of the present invention.

[0051] Figure 2 Shows the differences between the non-ordinal regression method, the existing ordinal regression method, and the PnP-FOR method.

[0052] Figure 3 Comparison of the existing ordinal regression method and the t-SNE visualization results after its combination with PnP-FOR.

[0053] Figure 4 Effect of different batch sampling strategies on the PnP-FOR method. Detailed implementation manner

[0054] The present invention will be further introduced below through simulation examples, and the classification results on real data and the comparison of their visualization results will be shown.

[0055] The present invention uses three face age estimation datasets: MORPH

[12] , FG-NET

[13] , and Adience[3]; three medical image classification datasets: LIDC-IDRI

[14] , BUSI

[15] , and Diabetic Retinopathy (DR)[8]; and one historical image age classification dataset: HID

[16] ; the image labels in all three types of datasets are ordinal labels. In the experiment, VGG-16

[17] is used as the feature extraction network, and the last fully connected layer is modified to a linear layer and a ReLU layer, and Dropout

[22] is added. The number of neurons in the linear layer is 4096, that is, the feature vector is 4096-dimensional, and finally this feature vector is input into the classifier. During training, all images are first scaled to 256×256 in size and then randomly cropped to 224×224 in size, and data augmentation methods such as random rotation and random inversion are adopted.

[0056] The hyperparameter settings for all experiments are as follows: the learning rate is 0.0001, it decays by 0.1 every 20 epochs, and the total number of epochs is 50; the batch sizes for training are 32 respectively. The optimizer is Adam, and the weight decay is 0.0001. The hyperparameter λ is set to 10. All experiments are implemented under the PyTorch framework and trained using an NVIDIA A100 GPU.

[0057] During the experiment, classification accuracy (Accuracy), cumulative score (CS), and mean absolute error (MAE) are used to measure the quantization performance of different datasets:

[0058]

[0059] Among them, TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative respectively;

[0060]

[0061] The CS score is mainly used to evaluate the age estimation dataset, where N is the number of test data, N l represents the samples with a prediction error not exceeding l years. In the experiment, l = 5;

[0062]

[0063] where, represents the predicted value of the i-th sample, y i is the true label of this sample.

[0064] Experimental Example 1: Comparison of quantization performance under different scenarios

[0065] Table 1 Performance comparison of different methods in face age estimation

[0066]

[0067] Table 1 Performance comparison of different methods in medical image classification

[0068]

[0069] For quantitative comparison, experimental comparisons were made between some state-of-the-art ordinal regression methods and their combinations with our PnP-FOR. In Table 1, we mainly compared four baseline methods and their combinations with PnP-FOR: Mean-Variance [4], Poisson [9], SORD [6], and POE [5]. It is worth noting that the proposed PnP-FOR improves over the existing methods in terms of MAE, CS, and accuracy. Mean-Variance and POE focus on learning the uncertainty of the neural network output vector and the latent space features respectively. Mean-Variance calculates the mean loss and variance loss of the output vector to control the uncertainty in the output space. PnP-FOR+Mean-Variance achieved the best MAE value on the MORPH and Adience datasets, which verifies that PnP-FOR makes the learned mean and variance consistent with the values in the label space. For the FG-Net dataset with expression and illumination variations, PnP-FOR+SORD outperformed other methods. This indicates that better ordinal representations can be more suitable for the soft labels constructed by SORD. More importantly, PnP-FOR significantly improves the CS score or classification accuracy, indicating that the feature representation retains the ordinal information beneficial to the ordinal classification task.

[0070] The essence of medical disease progression is distinguished by presenting ordinal information of disease progression. Table 2 compares PnP-FOR and four baseline methods for age estimation. A similar conclusion can be drawn that the proposed PnP-FOR can improve the classification performance of existing deep ordinal regression methods on medical data.

[0071] Table 3 presents the results of the historical image age classification task. A similar conclusion can be drawn that the proposed PnP-FOR improves the performance of the ordinal regression model. What differentiates HID from the previous two tasks is that there is no fixed ordinal information in the content of images of different categories or ages. The discriminative information for different ages includes hue, architectural style, etc. Therefore, POE models uncertainty by learning the distribution of latent space features. Although the uncertainty of POE can model the implicit information of HID, it only uses the triplet loss that focuses on local ordinal constraints to maintain the ordinal relationship. In contrast, PnP-FOR can cover the global ordinal relationship. Therefore, the feature ordinality of PnP-FOR further improves the performance of HID.

[0072] Table 3 Performance comparison of different methods on the HID dataset

[0073]

[0074] Among these four baseline methods, SORD seems to be a good choice for combination with PnP-FOR. SORD is easier to learn the ordinal relationship, especially when the feature space is regularized by PnP-FOR.

[0075] Experimental Example 2: Comparison of t-SNE visualization results of different methods

[0076] To further highlight the superiority of PnP-FOR, we Figure 3 visualized the feature representations of four baseline methods with and without PnP-FOR by t-SNE. It can be seen that the traditional OR method shows a certain ordinal distribution in the feature space, while POE learns a relatively regular ordinal distribution. Although their learning strategies can model the ordinal relationship in the latent space, they cannot guarantee the global ordinal relationship and lack within-class compactness ( Figure 3 at the red arrow).

[0077] Using our PnP-FOR, all results can be adjusted to a space close to one-dimensional, which is consistent with the true ordinal label space, which also verifies the effectiveness of our motivation. In addition, PnP-FOR not only preserves the local ordinal relationship but also the global distribution. On the other hand, PnP-FOR encourages within-class compactness; for example, POE maintains the global ordinal distribution. However, POE+PnP-FOR further compresses the data points while maintaining a better ordinal distribution.

[0078] Experimental Example 3: Influence of Different Sampling Strategies on PnP-FOR

[0079] We studied two sampling strategies for batch training using PnP-FOR: 1) Random sampling, and 2) Stratified sampling, i.e., all samples in a mini-batch have different labels. Figure 4 It shows that the results of stratified sampling are not as good as those of random sampling.

[0080] Stratified sampling will cause SORD to slightly overfit after 25 epochs. This can be explained from two aspects. First, as shown in Equation (6), stratified sampling will pay more attention to the distance between adjacent ordered labels, because labels at a long distance hardly affect the probability magnitude, while random sampling helps with the compactness between classes. Second, the probability in Equation (6) approximates the probability of the batch samples. Therefore, it is difficult for stratified sampling to learn the true distribution of the dataset.

[0081] References:

[0082] [1] Haiping Zhu, Hongming Shan, Yuheng Zhang, Lingfu Che, Xiaoyang Xu, Junping Zhang, Jianbo Shi, and Fei-Yue Wang. Convolutional ordinal regression forest for image ordinal estimation. IEEE Transactions on Neural Networks and Learning Systems, 2021.

[0083] [2] Wei Shen, Yilu Guo, Yan Wang, Kai Zhao, Bo Wang, and Alan L Yuille. Deep regression forests for age estimation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 2304-2313, 2018.

[0084] [3] Wei Shen, Kai Zhao, Yilu Guo, and Alan Yuille. Label distribution learning forests. In Advances in Neural Information Processing Systems, 2017.

[0085] [4] Hongyu Pan, Hu Han, Shiguang Shan, and Xilin Chen. Mean-variance loss for deep age estimation from a face. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 5285–5294, 2018.

[0086] [5] Wanhua Li, Xiaoke Huang, Jiwen Lu, Jian-jiang Feng, and Jie Zhou. Learning probabilistic ordinal embeddings for uncertainty-aware regression. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 13896-13905, 2021.

[0087] [6] Raul Diaz and Amit Marathe. Soft labels for ordinal regression. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 4738–4747, 2019.

[0088] [7] Zhenxing Niu, Mo Zhou, Le Wang, Xinbo Gao, and Gang Hua. Ordinal regression with multiple output CNN for age estimation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 4920–4928, 2016.

[0089] [8] Xiaofeng Liu, Yang Zou, Yuhang Song, Chao Yang, Jane You, and BV K Vijaya Kumar. Ordinal regression with neuron stick-breaking for medical diagnosis. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0–0, 2018.

[0090] [9] Christopher Beckham and Christopher Pal. Unimodal probability distributions for deep ordinal classification. In International Conference on Machine Learning, pages 411–419, 2017.

[0091]

[10] Víctor Manuel Vargas, Pedro Antonio Gutiérrez, and César Hervás-Martínez. Unimodal regularisation based on beta distribution for deep ordinal regression. Pattern Recognition, 122:108310, 2022.

[0092]

[11] Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9(11), 2008.

[0093]

[12] K.Ricanek and T.Tesafaye. MORPH: a longitudinal image database of normal adult age-progression. In 7th International Conference on Automatic Face and Gesture Recognition (FGR06), pages 341–345, 2006.

[0094]

[13] Gabriel Panis, Andreas Lanitis, Nicolas Tsapatsoulis, and Timothy F. Cootes. Overview of research on facial ageing using the FG-NET ageing database. IET Biom., 5(2):37–46, 2016.

[0095]

[14] Botong Wu, Xinwei Sun, Lingjing Hu, and Yizhou Wang. Learning with unsure data for medical image diagnosis. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 10590–10599, 2019.

[0096]

[15] Walid Al-Dhabyani, Mohammed Gomaa, Hussien Khaled, and Aly Fahmy. Dataset of breast ultrasound images. Data in Brief, 28:104863, 2020.

[0097]

[16] Frank Palermo, James Hays, and Alexei A. Efros. Dating historical color images. In Computer Vision - ECCV European Conference on Computer Vision, Florence, Italy, October 7 - 13, 2012, Proceedings, Part VI, volume 7577, pages 499–512, 2012.

[0098]

[17] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large - scale image recognition. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7 - 9, 2015, 2015.

[0099]

[18] Botong Wu, Xinwei Sun, Lingjing Hu, and Yizhou Wang. Learning with unsure data for medical image diagnosis. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 10590–10599, 2019.

[0100]

[19] Frank Palermo, James Hays, and Alexei A. Efros. Dating historical color images. In Computer Vision - ECCV European Conference on Computer Vision, Florence, Italy, October 7 - 13, 2012, Proceedings, Part VI, volume 7577, pages 499–512, 2012.

[0101]

[20] Yanzhu Liu,Adams Wai-Kin Kong,and Chi Keong Goh.A constraineddeep neural network for ordinal regression.In 2018 IEEE Conference on Com-puter Vision and Pattern Recognition,CVPR 2018,Salt Lake City,UT,USA,June 18-22,2018,pages 831–839,2018.

[0102]

[21] Yanzhu Liu,Fan Wang,and Adams Wai-Kin Kong.Probabilistic deepordinal regression based on Gaussian processes.In 2019 IEEE / CVF InternationalConference on Computer Vision,ICCV 2019,Seoul,Ko-rea(South),October 27-November 2,2019,pages 5300–5308,2019.

[0103]

[22] Nitish Srivastava,Geoffrey Hinton,Alex Krizhevsky,Ilya Sutskever,and Ruslan Salakhutdi-nov.Dropout:a simple way to prevent neural networksfrom overfitting.Journal of Machine Learning Research,15(1):1929–1958,2014.

[0104]

[23] Geoffrey Hinton and Sam T Roweis.Stochastic neighbor embedding.InNIPS,vol-ume 15,pages 833–840.Citeseer,2002。

Claims

1. A generative feature order regularization method for a deep ordered regression model, which is applicable to three types of images, namely face age estimation, medical image classification, and historical image age classification, and is characterized in that The specific steps are as follows: (1) Select an ordered regression model for discriminant analysis; The deep ordered regression method is based on the following discriminant model: p(Y|X) = g(f θ (X)); (1) Meanwhile, construct a Markov chain: X→Z=f θ (X)→Y; (2) where both X and Y are random variables, representing a sample point in the input space and the class label corresponding to this sample point respectively, Z is a latent variable, and f θ (·) is a feature extraction network, specifically a convolutional neural network with parameters θ, and g(·) is a label mapping function; specifically, there are N independent samples {(x1, y1), (x2, y2),..., (x N , y N )}, then the labels satisfy: y ∈ {r1, r2,..., r K}, and r1 ≤ r2 ≤ … ≤ r K ; It can be obtained that the objective of the ordered regression model is to learn a mapping from X to Y, where K represents the number of ordered labels; (2) Model ordered regression from the perspective of a generative model; The feature representation learned by the discriminant ordered regression model in step (1) has many features unrelated to the ordered information and cannot guarantee the ordered relationship of features in the latent space; Regarding X as being generated by the label Y, then this generative model is expressed as: p(X, Y) = p(Y)p(X|Y), (3) The conditional probability p(X|Y) indicates that the input variable X in the high-dimensional space and the ordered label Y have the same intrinsic dimension in the latent space; Then making the feature vector Z in the latent space have the same dimension as Y will simplify the learning objective; Regarding Y as the low-dimensional representation of X embedded in a high-dimensional manifold, then Y is the latent variable of X with an intrinsic dimension of 1; And the intrinsic dimension of the latent variable Z is also 1, thus obtaining the following Markov chain: Y → Z → X, (4) Use the information bottleneck criterion to describe the relevant information between X and Y, that is, mutual information; Y implicitly determines the features in X that are related or unrelated to the ordered information, then the following Markov chain is obtained: Y → X → Z, (5) For an X, the optimal feature representation Z should capture the features most relevant to the category Y and make the prediction of X unaffected by irrelevant information; Therefore, maximize the mutual information between Y and Z to maintain the ordered information in Z; (3) Construct the probability distribution of the ordered labels; Reduce the dimension of the feature representation of the samples to align it with the one-dimensional distribution of the ordered labels; First, map the ordered labels to a probability distribution: Here, P ii = 0, the ordered label is converted into a probability representing the distance between any two ordered labels, where y i , y j , y k , y l represent the ordered labels corresponding to different samples in a batch, respectively; (4) Construct the probability distribution of the feature representation; Similar to the label probability distribution, the probability distribution of the feature representation is: Here, Q ii = 0, z = f θ (x), where z i , z j , z k , z l respectively represent the feature representations of different samples in a batch; (5) Ordered feature regularization; So far, the probability distribution representations of the sample feature representation and the ordered labels are obtained; Use the KL divergence to measure the distance between the probability distribution of the feature representation and the probability distribution of the labels: Formula (8) aligns the feature distribution and the label distribution in the latent space; Since the label space is globally ordered, the batch-training-based method can learn the global ordered information; (6) Construct the loss function; The loss function used is: Among them, λ is a hyperparameter; represents the loss function involved in the existing deep ordinal regression method.

2. The generative feature order regularization method for a depth-ordered regression model according to claim 1, wherein The feature extraction network uses VGG-16 as the backbone network.