Model training method based on diversified context data generation and classifier correction

By generating a diversified context data set and performing classifier correction, the overfitting and information distortion problems in long-tail distribution data processing are solved, and the classification performance of the model on long-tail distribution data is improved.

CN120014405APending Publication Date: 2025-05-16CHINA UNICOM (FUJIAN) IND INTERNET CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411981382.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is prone to overfitting and information distortion problems when processing long-tail distribution data, and decoupled learning may lead to poor classifier training results when training a classifier.

Method used

Through a model training method based on diversified context data generation and classifier correction, a diversified context data set is generated and a category balanced data set is constructed through a random selection strategy, which is used to train the initial model and optimize the model.

Benefits of technology

The feature expression ability and classification performance of the model on the long-tail distribution data is improved, the risk of overfitting and information distortion is reduced, and the classification accuracy of a few categories is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014405A_ABST
    Figure CN120014405A_ABST
Patent Text Reader

Abstract

The invention provides a model training method based on diversified context data generation and classifier correction. The method comprises the following steps: S1, generating a diversified context data set through context information fine control and a diversity generation strategy based on an original long-tail distribution data set; s2, generating a category balance data set through a random selection strategy based on the original long-tail distribution data set; s3, inputting the diversified context data set into the initial model for model training to obtain a trained feature extractor and classifier; s4, inputting the category balance data set into the optimization model to perform correction training of a classifier, and obtaining a final optimization model for data category prediction; the optimization model shares a feature extractor of the initial model, and a classifier trained by the initial model is used as an initial classifier of the optimization model for correction training. According to the method, the feature expression ability of the final optimization model on the tail class sample is improved, and long-tail distribution data can be better classified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification, and in particular to a model training method based on diversified context data generation and classifier correction. Background Art

[0002] LT (Long Tail Distribution) refers to the situation in which the number of samples in the minority class (tail class) is much larger than the number of samples in the majority class (head class) in a data set. This data distribution presents a long-tail shape, that is, the majority class (Majority) is a few classes with a high frequency in the data set, while the minority class (Minority) is a large number of classes with a low frequency.

[0003] CMO (Context-rich minority oversampling) refers to the use of contextual information of samples in the oversampling process to better generate new minority samples. By introducing contextual information, the newly generated samples can better maintain the characteristics and distribution of the original samples, avoiding the problem of sample overfitting caused by simple copy and paste. The advantage of this method is that it can more effectively utilize the information in the data set, improve the model's learning ability for minority samples, and thus improve the model's performance in the case of class imbalance.

[0004] Decoupled learning has been widely used in fields such as deep learning, and helps optimize the training process and performance of complex models. It usually decouples the two processes of representation learning and classifier training, which can effectively solve the recognition problem under long-tail data distribution.

[0005] However, the CMO algorithm has the following disadvantages: (1) Overfitting risk: When the context information is too rich or the generated new samples are too dependent on the original data, there is a risk of overfitting. This may cause the model to perform well on the training set, but the generalization ability on unknown data is reduced; (2) Information distortion: When extracting and representing context information, information distortion may occur, resulting in the generated new samples being unrealistic or unrepresentative; this may affect the model's learning effect on minority class samples.

[0006] When retraining a classifier, decoupled learning may lead to poor training results when the number of training samples is small or the data is still unbalanced. In this case, due to the scarcity of training samples or uneven distribution of categories, the classifier may not be able to fully learn the feature differences between categories, thus affecting the final classification performance.

[0007] In view of this, the present invention proposes a model training method based on diversified context data generation and classifier correction. Summary of the invention

[0008] The purpose of the present invention is to propose a model training method based on diversified context data generation and classifier correction. The final optimized model obtained has improved feature expression ability on tail class samples and can better classify long-tail distribution data.

[0009] To achieve the above object, the technical solution of the present invention is: a model training method based on diversified context data generation and classifier correction, which specifically includes the following steps:

[0010] S1. Generate a diversified context dataset based on the original long-tail distribution dataset through fine control of context information and diversity generation strategy;

[0011] S2, generating a class-balanced dataset based on the original long-tail distribution dataset through a random selection strategy;

[0012] S3, input the diversified context data set into the initial model for model training to obtain a trained feature extractor and classifier;

[0013] S4. Input the category-balanced data set into the optimization model for correction training of the classifier to obtain the final optimization model for data category prediction; the optimization model shares the feature extractor of the initial model, and uses the classifier trained by the initial model as the initial classifier of the optimization model for correction training.

[0014] Preferably, the S1 specifically comprises the following steps:

[0015] S11. Get the original long-tail distribution data set y i ∈{1,2,...,C}, where x i is the i-th image sample, y i is the label information corresponding to the i-th image sample, N and C represent the total number of samples and the total number of categories in the dataset respectively;

[0016] S12, use the visual language model to transform each image sample x i Converted into text features x' i ;

[0017] S13. Use the self-attention mechanism to calculate the correlation between each text feature and other text features in the data set to obtain the attention weight;

[0018] S14, each text feature uses the calculated attention weight as a weighting coefficient to perform weighted summation with other text features of the data set to obtain a feature sample with context information;

[0019] S15: Integrate the feature samples with context information and the corresponding label information of all image samples to obtain a diversified context dataset D dcdg .

[0020] Preferably, the visual language model adopts the CLIP model.

[0021] Preferably, S2 is specifically: selecting m samples from each category of the original long-tail distribution data set through a random selection strategy as the category-balanced data set D bal , v c,i Express the i-th sample of the c-th category, including the image sample and the corresponding label information.

[0022] Preferably, the random selection strategy is used to select m samples from each category of the original long-tail distribution data set, specifically, according to the number of samples n in the cth category c Do the following:

[0023] If m≥n c , directly randomly select m samples from the cth category samples;

[0024] If m <n c , copy all samples of the cth category and randomly select them cyclically until the total number of copied and randomly selected samples reaches m, and obtain m samples of the cth category. For the repeated samples obtained by random selection, data enhancement is used to replace the repeated samples.

[0025] Preferably, the repeated sample replacement by data enhancement for the randomly selected repeated samples is specifically: performing data enhancement operations including image cropping, color jittering or noise addition on the image samples in the repeated samples, so that the repeated image samples are distinguishable from the original image samples.

[0026] Preferably, S3 is specifically:

[0027] Diversify the context dataset D dcdg The feature samples with context information are input into the initial model. The feature samples are passed through the feature extractor of the initial model to obtain the projection head. After being input into the classifier, the predicted classification head g is obtained. dcdg ;

[0028] Classification Header dcdg The cross entropy loss L is formed with the label information of the diverse context dataset ceApply constraints to the initial model;

[0029] At the end of the training, a feature extractor and a classifier are obtained that represent the fully trained learning. The initial model parameters after training are denoted as w = {θ, φ}, where θ and φ represent the feature extractor and the classifier, respectively.

[0030] Preferably, the network structure of the initial model is specifically as follows: ResNet-32 is used as the backbone network, and a batch normalization layer and a fully connected classification are introduced as the network structure of the initial model.

[0031] Preferably, the network structure of the initial model is specifically as follows: ResNet-10 is used as the backbone network, and a batch normalization layer and a fully connected classification are introduced as the network structure of the initial model.

[0032] Preferably, S4 is specifically:

[0033] The class-balanced dataset D bal The image sample is input into the optimization model, and the corresponding projection head is obtained through the feature extractor of the optimization model. The projection head is input into the classifier of the optimization model to obtain the predicted classification head g bal ;

[0034] Classification Header bal The cross entropy loss L is formed with the label information of the class-balanced dataset ce Apply constraints to the optimization model;

[0035] The feature extractor is fixed and multiple rounds of classifier correction training are performed to obtain the final optimized model. The parameters of the final optimized model are represented by w′={θ,φ′}, where φ′ represents the classifier after correction training.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] The present invention proposes a fine control and diversity generation strategy for context information, which increases the attention to long-tail categories by increasing the weight of the attention mechanism, ensuring that the model reasonably pays attention to categories with low frequency in the training data, and that the richness of information does not lead to overfitting, thereby increasing the diversity of generated samples and reducing the model's dependence on the original data. The diversified context data set constructed by this strategy can be used as a subsequent training set to alleviate the imbalance problem of the original long-tail distribution data to a certain extent, help improve the performance of the model on long-tail distribution data, and improve the classification accuracy of minority categories.

[0038] The present invention can help optimize the decision boundary of the classifier by using a balanced data set to perform classifier correction, thereby effectively alleviating the classifier bias problem and improving the classification effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic diagram of the overall framework of the model training method of the present invention;

[0040] Figure 2 This is a comparative experimental effect diagram of training with and without classifier correction based on the imbalance factor ρ=100 on the CIFAR-10 dataset in one embodiment of the present invention;

[0041] Figure 3 The figure is a comparative experimental effect diagram of the training with and without classifier correction based on the imbalance factor ρ=100 on the CIFAR-100 dataset in one embodiment of the present invention. DETAILED DESCRIPTION

[0042] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.

[0043] The present invention proposes a model training method based on diverse contextual data generation and classifier calibration (DiverseContextual Data Generation and Classifier Calibration Method, DCGC), the overall framework is as follows Figure 1 As shown in the figure, it mainly includes two parts: the dataset preparation stage and the student model decoupled learning and training stage. Specifically, (1) in the dataset preparation stage, through fine control of context information and diversity generation strategy, self-attention is used to adjust the sample weights of long-tail categories, and similarity measurement is introduced as authenticity constraint to construct a diversified context dataset, and a category-balanced dataset is constructed through random selection strategy; (2) in the model decoupled learning and training stage, firstly, the model is updated and trained based on the commonly used cross entropy loss using the diversified context dataset training to obtain the initial model; secondly, the feature extractor of the initial model is fixed, and the classifier is corrected and trained using the constructed category-balanced dataset to further improve the classification accuracy of the model on the tail category and obtain the final model.

[0044] The detailed steps are as follows:

[0045] 1. Long-tail distribution data set y i ∈{1,2,...,C}, where N and C represent the total number of samples and the total number of categories in the dataset, respectively. i represents the i-th training sample, y i represents the true label corresponding to the i-th training sample. Due to the uneven distribution of the data set, the number of training samples in each category is highly unbalanced. Assuming that the number of samples corresponding to the c-th category is n c , then arrange the categories in descending order of their sample size to get n min<... <n c <... <n max .

[0046] 2. Dataset preparation phase - Diversified context dataset D dcdg :(1) Data preparation: Prepare the original dataset Including image samples and corresponding label information. (2) Text feature generation: Use existing visual language models (such as CLIP) to convert each image into a text feature x' i . These text features can be used as sequence representations, where each text feature corresponds to each image sample in the original dataset, and the label is consistent with the label information corresponding to the original image sample. (3) Self-attention calculation: For the generated text feature sequence, the self-attention mechanism is used to calculate the attention weight Q between different samples. In this step, the model will learn the dependencies between different positions in the text feature sequence and dynamically assign attention weights accordingly. (4) Context information integration: Based on the calculated attention weight Q, it is used as a weight coefficient to perform weighted summation of each text feature and other text features to obtain the context information of each sample. This step can obtain feature samples with context information. (5) Dataset generation: The generated feature samples with context information and the corresponding label information are integrated into a new dataset - a diversified context dataset D dcdg , each sample includes image features, context information and corresponding labels.

[0047] 3. Dataset preparation stage - category balanced dataset D bal :The balanced data set is selected from the original long-tail distribution data through a random selection strategy, which can be expressed as Where m controls the number of samples selected for each category. In the selection process, according to m and the number of class samples n of the cth category c When m≥n c , directly randomly select m samples from the training samples of this category. When m <n c When , copy all samples of this category and randomly select them in a loop until the number of samples of this category reaches m. In order to alleviate the problems of overfitting caused by repeated use of the same samples for training, some additional data enhancement methods are used to replace repeated samples. Image cropping can be used to retain key information, such as cropping different parts of the object image; or color jittering to slightly change the image color; or adding noise, such as Gaussian noise, to make the image slightly change. .

[0048] 4. Model decoupling learning and training phase - initial model training: First, ResNet-32 / ResNet-10 (the former is used for the CIFAR dataset, and the latter is used for the ImageNet ILSVRC 2012 dataset) is used as the backbone network, and batch normalization layers and fully connected classification are introduced as the final network structure. The images of the diverse context dataset are input into the initial model network. The images can be projected by the feature extractor of the network, and then further input into the classifier to obtain the predicted classification head g. dcdg , the classification head and the real label y of the diverse context dataset constitute a cross entropy loss L ce After the training is completed, the feature extractor and classifier that represent the learning training are finally obtained. The initial model parameters are expressed as θ and They represent feature extractor and classifier respectively.

[0049] 5. Model decoupling learning training phase - optimization model training: Input the class-balanced dataset images into the final model network. The model shares the feature extractor of the initial model. The image passes through the feature extractor to obtain the projection head corresponding to the class-balanced dataset. At the same time, download the classifier trained by the initial model as the initial classifier of the optimization model. The projection head is further input into the classifier to obtain the predicted classification head g bal , the classification head and the true label of the category-balanced dataset also constitute the cross entropy loss L ce Constrain the model. After multiple rounds of classifier correction training, the optimized model is finally obtained, and the parameters are expressed as

[0050] 6. After the above training, the optimized model has improved its ability to express features in the tail class samples, and can better classify the long-tail distribution data. In the test phase, the optimized model is used to predict the category of the test data set and calculate the classification of the samples;

[0051] 7. Calculate Top-K (k=1), the classification accuracy of each category and the overall mean average precision (mAP) based on the classification situation and classification evaluation indicators.

[0052] The following two sets of specific simulation experiments are provided to verify the effectiveness of this scheme:

[0053] In experiment 1, the present invention is used to perform image classification on two datasets: CIFAR-10 / CIFAR-100.

[0054] In order to verify the effectiveness of this algorithm, comparative experiments are conducted on the CIFAR-10 / CIFAR-100 test set. Tables 1, 2, and 3 show the experimental results. Among them, CE means that only the cross entropy loss L is used on the original long-tail distribution data set.ce , CMO means using only cross entropy loss L on diverse context datasets CE , SECC represents the method of this solution. CIFAR-10-Top-1 and CIFAR-100-Top-1 represent the average accuracy of two CIFAR datasets when the imbalance factors are 10, 50, and 100, respectively. The experimental results show that the method proposed in the present invention has a significant performance improvement on the classification task of long-tail distribution problems, which verifies the effectiveness of the method of the present invention.

[0055] Table 1. Comparative experiments of the present invention on the CIFAR-10 / CIFAR-100 test set with an imbalance factor of 10

[0056] Method CIFAR-10-Top-1 CIFAR-100-Top-1 CE 86.39 55.71 CMO - 59.50 SECC 89.69 62.20

[0057] Table 2. Comparative experiments of the present invention on the CIFAR-10 / CIFAR-100 test set with an imbalance factor of 50

[0058] Method CIFAR-10-Top-1 CIFAR-100-Top-1 CE 74.81 43.75 CMO - 48.30 SECC 84.94 51.45

[0059] Table 3. Comparative experiments of the present invention on the CIFAR-10 / CIFAR-100 test set with an imbalance factor of 100

[0060] Method CIFAR-10-Top-1 CIFAR-100-Top-1 CE 70.36 38.32 CMO - 43.90 SECC 82.16 47.41

[0061] Experiment 2: Using the present invention to perform image classification on the ImageNet2012-LT dataset.

[0062] In order to verify the effectiveness of the algorithm, the algorithm was tested on the ImageNet2012-LT dataset. Table 4 shows the experimental results. Among them, the CE, CMO and SECC representation methods are the same as those in Experiment 1. Decouple-LWS represents the decoupled representation and classifier learning method. Only the cross entropy loss L is used on the original long-tail distribution data set. ce From the results, it can be found that the method based on diverse context data generation and classifier correction proposed in the present invention also has excellent performance improvement on the ImageNet2012-LT dataset.

[0063] Table 4. Comparative experiments of the present invention on the ImageNet2012-LT test set

[0064] Method ImageNet2012-LT CE 35.27 CMO 41.35 Decouple-LWS 41.40 SECC 42.98

[0065] Combining Experiments 1 and 2, the present invention has significant performance advantages on the three existing long-tail distribution data sets, surpassing the highest level in the current academic field, verifying that the method proposed in the present invention effectively improves the feature expression ability of tail class samples and successfully selectively distills the effective knowledge of the teacher model.

[0066] In addition, this embodiment also conducts comparative experiments with and without classifier correction training based on the imbalance factor ρ=100 on the CIFAR-10 and CIFAR-100 datasets. Figure 2 , Figure 3 As shown in the figure, as the number of training rounds increases, the classification performance of the model can be improved by using the class-balanced data set to calibrate the classifier. At the same time, after the classifier is calibrated, the convergence speed of the model can be further accelerated. In summary, classifier calibration can not only further improve the performance of the model, but also increase the speed of model convergence.

[0067] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.

Claims

1. A model training method based on diverse context data generation and classifier calibration, characterized in that: The specific steps include: S1. Generate a diversified context dataset based on the original long-tail distribution dataset through fine control of context information and diversity generation strategy; S2, generating a class-balanced dataset based on the original long-tail distribution dataset through a random selection strategy; S3, input the diversified context data set into the initial model for model training to obtain a trained feature extractor and classifier; S4. Input the category-balanced data set into the optimization model for correction training of the classifier to obtain the final optimization model for data category prediction; the optimization model shares the feature extractor of the initial model, and uses the classifier trained by the initial model as the initial classifier of the optimization model for correction training.

2. The model training method based on diversified context data generation and classifier calibration according to claim 1, characterized in that: The S1 specifically includes the following steps: S11. Get the original long-tail distribution data set y i ∈{1,2,...,C}, where x i is the i-th image sample, y i is the label information corresponding to the i-th image sample, N and C represent the total number of samples and the total number of categories in the dataset respectively; S12, use the visual language model to transform each image sample x i Converted into text features x' i ; S13. Use the self-attention mechanism to calculate the correlation between each text feature and other text features in the data set to obtain the attention weight; S14, each text feature uses the calculated attention weight as a weighting coefficient to perform weighted summation with other text features of the data set to obtain a feature sample with context information; S15: Integrate the feature samples with context information and the corresponding label information of all image samples to obtain a diversified context dataset D dcdg .

3. The model training method based on diversified context data generation and classifier calibration according to claim 2, characterized in that: The visual language model adopts the CLIP model.

4. The model training method based on diversified context data generation and classifier calibration according to claim 1, characterized in that: Specifically, S2 is: m samples are selected from each category of the original long-tail distribution data set through a random selection strategy as the category-balanced data set D bal , v c,i Express the i-th sample of the c-th category, including the image sample and the corresponding label information.

5. The model training method based on diversified context data generation and classifier calibration according to claim 4, characterized in that: The random selection strategy is used to select m samples from each category of the original long-tail distribution data set. Specifically, according to the number of samples n in the cth category, c Do the following: If m≥n c , directly randomly select m samples from the cth category samples; If m <n c , copy all samples of the cth category and randomly select them cyclically until the total number of copied and randomly selected samples reaches m, and obtain m samples of the cth category. For the repeated samples obtained by random selection, data enhancement is used to replace the repeated samples.

6. The model training method based on diversified context data generation and classifier calibration according to claim 5, characterized in that: The method of replacing the randomly selected repeated samples by data enhancement specifically includes: performing data enhancement operations including image cropping, color jittering or noise addition on the image samples in the repeated samples, so that the repeated image samples are distinguishable from the original image samples.

7. The model training method based on diversified context data generation and classifier calibration according to claim 2, characterized in that: The S3 is specifically: Diversify the context dataset D dcdg The feature samples with context information are input into the initial model. The feature samples are passed through the feature extractor of the initial model to obtain the projection head. After being input into the classifier, the predicted classification head g is obtained. dcdg ; Classification Header dcdg The cross entropy loss L is formed with the label information of the diverse context dataset ce Apply constraints to the initial model; At the end of the training, a feature extractor and a classifier are obtained that represent the fully trained learning. The initial model parameters after training are denoted as w = {θ, φ}, where θ and φ represent the feature extractor and the classifier, respectively.

8. The model training method based on diversified context data generation and classifier calibration according to claim 7, characterized in that: The network structure of the initial model is specifically as follows: ResNet-32 is used as the backbone network, and a batch normalization layer and a fully connected classification are introduced as the network structure of the initial model.

9. The model training method based on diversified context data generation and classifier calibration according to claim 7, characterized in that: The network structure of the initial model is specifically as follows: ResNet-10 is used as the backbone network, and a batch normalization layer and a fully connected classification are introduced as the network structure of the initial model.

10. The model training method based on diversified context data generation and classifier calibration according to claim 4, characterized in that: The S4 is specifically: The class-balanced dataset D bal The image sample is input into the optimization model, and the corresponding projection head is obtained through the feature extractor of the optimization model. The projection head is input into the classifier of the optimization model to obtain the predicted classification head g bal ; Classification Header bal The cross entropy loss L is formed with the label information of the class-balanced dataset ce Apply constraints to the optimization model; The feature extractor is fixed and multiple rounds of classifier correction training are performed to obtain the final optimized model. The parameters of the final optimized model are represented by w′={θ,φ′}, where φ′ represents the classifier after correction training.