A method and medium for distilling data and a method for visual task processing
Through comparative learning technology to construct positive and negative samples in data distillation and supervised training, the problem that the existing technology is difficult to use in data sets other than classified data sets is solved, and efficient distillation and multi-task applications for multiple label data or unlabeled data are realized.
Patent Information
- Application Number
- CN202310526044.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-05-10
AI Technical Summary
Existing data distillation techniques are difficult to use for data sets other than classification data sets and cannot be used for tasks other than classification tasks, especially in the case of unlabeled data and multiple label data.
By using contrast learning technology, data distillation is achieved by constructing positive and negative samples, using loss functions such as noise comparison estimation (NCE) for supervision and training, update the feature extraction module and data set.
Efficient distillation of multiple label data or unlabeled data is achieved. The resulting distillation data can be applied to a variety of visual tasks including classification, significantly improving the data distillation speed and reducing data redundancy.
Smart Images

Figure CN116541763B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a method for distilling data, a medium, and a visual task processing method. Background Art
[0002] The goal of the data distillation (Dataset Distillation or Dataset Condensation) task is to refine a large training dataset (or real dataset) with a quantity of N into a set of synthetic datasets with a quantity of M, where M << N, and it is expected that the same model can obtain as consistent results as possible when trained using the synthetic dataset and the large dataset, or in other words, the accuracy rates in the test set are close. The core of data distillation is to reduce the redundancy of data, synthesize multiple similar images into fewer images, and increase the density of effective information in the images.
[0003] Existing data distillation technologies are all carried out in classification tasks. Since classification tasks come with category information, the technical principles of existing data distillation technologies are basically the same. They all use classification labels on classification datasets and synthesize images of the same category into one or a cluster of synthetic images representing this category according to the category labels. It can be seen that the existing methods synthesize data using a common category label without any modification to the label. However, this also makes it the case that in datasets for other tasks and unlabeled data, such as in object detection datasets where the labels corresponding to each image are different, the label of the synthesized image must be a "comprehensive label" containing all the label information of the samples to be synthesized, and there is no good solution for how to generate this comprehensive label. It can be seen that existing distillation means are difficult to be used for datasets other than classification datasets and cannot be used for tasks other than classification tasks. Moreover, there is currently no data distillation method that can be used for other label data such as detection and segmentation labels outside of classification labels, or unlabeled data, and there is even no framework paradigm for applying distilled data to tasks other than classification.
[0004] In addition, the scale of the synthetic data of existing data distillation technologies is small, usually less than or equal to the ImageNet1k dataset, that is, 1.2 million images. Summary of the Invention
[0005] In view of some or all of the problems in the prior art, the first aspect of the present invention provides a method for distilling data, including:
[0006] Sampling a batch of samples from a first dataset and a second dataset respectively to obtain first sample data and second sample data, where the first dataset is initialized by a subset randomly sampled from the second dataset; and
[0007] Update the first data set through N iterations, where N is a natural number, and each iteration includes:
[0008] Extract the features of the first sample data and the second sample data respectively through a feature extraction module to obtain a first feature and a second feature; and
[0009] Calculate the difference between the first feature and the second feature, and use the difference as a supervision signal for backpropagation to update the first data set.
[0010] Furthermore, the method further includes: updating the feature extraction module.
[0011] Furthermore, before each update of the first data set, update the feature extraction module once.
[0012] Furthermore, one iteration of updating the feature extraction module includes:
[0013] Perform two different data augmentation processes on the second sample data to obtain positive sample pairs;
[0014] Input the positive sample pairs into the positive sample encoder and the negative sample encoder of the feature extraction module respectively to obtain a positive sample feature vector and a negative sample feature vector;
[0015] Calculate the similarity vector of the positive sample feature vector and the negative sample feature vector, and calculate the similarity matrix of the negative sample feature vector and the pre-stored negative sample feature vector;
[0016] Calculate the loss and backpropagate to update the positive sample encoder; and
[0017] Put the negative sample feature vector into the negative sample feature vector queue, and remove an equal amount of the negative sample feature vector at the end of the queue to implement queue update. At the same time, update the parameters of the negative sample encoder according to the momentum update method based on the parameters of the positive sample encoder.
[0018] Furthermore, the data augmentation process includes: translation, and / or rotation, and / or random cropping, and / or color transformation, and / or flipping.
[0019] Furthermore, the positive sample encoder includes a backbone network and a projection layer.
[0020] Furthermore, the structure of the negative sample encoder is the same as that of the positive sample encoder.
[0021] Furthermore, the loss is calculated according to the InfoNCE loss function.
[0022] Further, extracting the features of the first sample data and the second sample data respectively by the feature extraction module includes:
[0023] Performing the same data augmentation transformation on the first sample data and the second sample data to obtain the augmented first sample data and second sample data;
[0024] Extracting the feature vectors of the augmented first sample data and second sample data through the negative sample encoder of the feature extraction module.
[0025] Further, the first feature and the second feature include:
[0026] Feature maps output by one or more layers of the backbone network of the negative sample encoder; and
[0027] Feature vectors output by the projection layer of the negative sample encoder.
[0028] Further, calculating the difference between the first feature and the second feature includes:
[0029] Calculating a distance loss matrix based on the feature maps of the first feature and the second feature;
[0030] Calculating a similarity matrix based on the feature vectors of the first feature and the second feature;
[0031] Performing weighted averaging on the distance loss matrix according to the similarity matrix to obtain the final distance loss.
[0032] Based on the method for distilled data as described above, a second aspect of the present invention also provides a computer-readable storage medium for distilled data, which stores a computer program, and when the computer program runs on a processor, it executes the method for distilled data as described above.
[0033] A third aspect of the present invention also provides a visual task processing method, which executes a specified visual task through a task processing module, where the task processing module includes:
[0034] A backbone network obtained by iteratively updating the feature extraction module through a contrastive learning method using a first data set obtained by the method for distilled data as described above; and
[0035] An output head network obtained by random generation.
[0036] Further, the task processing module is also adjusted by a labeled downstream task data set.
[0037] A method for distilling data provided by the present invention realizes data distillation based on the technology of contrastive learning, and can distill various labeled data or unlabeled data. In addition, the distilled data obtained by the method can be applied to various vision tasks including classification. In other words, by applying novel artificial intelligence technologies to the field of data distillation technology, the present invention significantly improves the data distillation speed and reduces data redundancy. In addition, the present invention can also be applied to the field of image processing technology to efficiently and significantly reduce the redundancy of image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] To further clarify the above and other advantages and features of the embodiments of the present invention, a more specific description of the embodiments of the present invention will be presented with reference to the accompanying drawings. It can be understood that these drawings only depict typical embodiments of the present invention and thus will not be considered as limiting its scope. In the drawings, for clarity, the same or corresponding components will be denoted by the same or similar reference numerals.
[0039] Figure 1 A schematic flowchart of a method for distilling data showing an embodiment of the present invention;
[0040] Figure 2 A schematic diagram of the process of a method for distilling data showing an embodiment of the present invention;
[0041] Figure 3 A schematic diagram of the process of a method for processing vision tasks showing an embodiment of the present invention; and
[0042] Figure 4 A schematic diagram of the update process of the backbone network of a method for processing vision tasks showing an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In the following description, the present invention is described with reference to the embodiments. However, those skilled in the art will recognize that the embodiments can be implemented without one or more specific details or in combination with other alternative and / or additional methods or components. In other cases, well-known structures or operations are not shown or described in detail to avoid obscuring the inventive points of the present invention. Similarly, for purposes of explanation, specific numbers and configurations are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, the present invention is not limited to these specific details.
[0044] In this specification, the reference to "an embodiment" or "the embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. The phrase "in an embodiment" appearing throughout this specification does not necessarily refer to the same embodiment.
[0045] It should be noted that the embodiments of the present invention describe the method steps in a specific order. However, this is only for explaining the specific embodiment and does not limit the sequence of the steps. On the contrary, in different embodiments of the present invention, the sequence of the steps can be adjusted according to the actual requirements.
[0046] In the present invention, each functional module can be implemented by software, firmware or a combination thereof. When the module is implemented by software, the function of the module can be realized through a computer program process. For example, the module can be implemented by a code segment (such as a code segment in languages such as C, C++) stored in a storage device (such as a hard disk, memory, etc.), and when the code segment is executed by a processor, the corresponding function of the module can be realized. When the module is implemented by firmware, the function of the module can be written into a read-only memory such as an EPROM or EEPROM of the device in the form of program code, and when the program code is executed by a processor, the corresponding function of the module can be realized.
[0047] Existing distillation means are difficult to be used for other datasets except classification datasets and cannot be used for other tasks except classification tasks. In order to distill other datasets and apply the distilled data to other vision tasks, the present invention bypasses the existing supervised labels of images and directly starts from unsupervised learning, such as contrastive learning, to synthesize data. Since for unlabeled images, users expect that the features of similar images encoded by the model are closer in the high-dimensional feature space, and the features of dissimilar images are farther apart. Based on this, in contrastive learning, it is first necessary to construct positive and negative samples. For example, positive and negative samples can be obtained through data augmentation operations. Specifically, two or more images after data augmentation are regarded as positive samples, and other images and their images after data augmentation are regarded as negative samples. The positive and negative samples are respectively input into the positive and negative sample encoders to obtain corresponding feature vectors. The positive and negative sample encoders can be the same encoder or different. Then, according to the principle that the feature vectors between similar images, that is, positive samples, are as close as possible, and the feature vectors between dissimilar images, that is, between positive and negative samples, are as far apart as possible, loss functions such as Noise Contrastive Estimation (NCE) are used to supervise the model training. For example, the momentum contrastive learning framework (MoCo), SimSiam or SimCLR can be used to implement the operations described above.
[0048] The following further describes the solution of the present invention in conjunction with the accompanying drawings of the embodiments.
[0049] Figure 1 and Figure 2 respectively show the process schematic diagrams of a method for distilling data according to an embodiment of the present invention. As shown in the figure, a method for distilling data includes:
[0050] First, in step 101, sample data is obtained. A batch of samples is respectively sampled from the first data set and the second data set to obtain first sample data and second sample data, where the first data set is initialized with a subset randomly sampled from the second data set. In an embodiment of the present invention, the first data set may be referred to as synthetic image (Synthetic Data S), and the second data set may be referred to as real image (Real Data T);
[0051] Next, in step 102, feature extraction is performed. Features of the first sample data and the second sample data are respectively extracted through the same feature extraction module to obtain first features and second features. In an embodiment of the present invention, the feature extraction module is a neural network structure, which includes a positive sample encoder and a negative sample encoder, where the positive sample encoder includes a backbone network and a projection layer, and the structure of the negative sample encoder is the same as that of the positive sample encoder. The backbone network may adopt mainstream backbone networks such as alexnet, vgg, resnet, transformer, etc. To improve performance, in an embodiment of the present invention, the second data set, that is, the real image, is usually used to update the feature extraction module, or to train the neural network before feature extraction. During the update process, the negative sample encoder does not perform gradient update. Specifically, in an embodiment of the present invention, the update of the feature extraction module includes:
[0052] First, the second sample data is subjected to two different data augmentation processes to obtain positive sample pairs, where the data augmentation processes may include translation, and / or rotation, and / or random cropping, and / or color transformation, and / or flipping operations, etc.;
[0053] Next, the positive sample pairs are respectively input into the positive sample encoder and the negative sample encoder of the feature extraction module to obtain feature vectors output at the end of the model, denoted as positive sample feature vectors and negative sample feature vectors;
[0054] Next, a similarity vector of the positive sample feature vectors and the negative sample feature vectors is calculated, and a similarity matrix of the negative sample feature vectors and the negative sample feature vectors stored in a queue in advance is calculated. In an embodiment of the present invention, the initial value of the negative sample feature vectors stored in the queue in advance is a random value, which is updated during the update process of the feature extraction module;
[0055] Next, calculate the loss and backpropagate to update the positive sample encoder. In an embodiment of the present invention, the InfoNCE loss function in Momentum Contrast for Unsupervised Visual Representation Learning is used to calculate the loss. It should be understood that in other embodiments of the present invention, other loss functions such as the NCE loss function can also be used; and
[0056] Finally, put the negative sample feature vectors into the negative sample feature vector queue, and remove an equal amount of negative sample feature vectors at the end of the queue to achieve queue update. At the same time, update the parameters of the negative sample encoder according to the momentum update method based on the parameters of the positive sample encoder.
[0057] Based on this, the feature extraction of the first sample data and the second sample data includes:[[]]
[0058] Perform the same data augmentation transformation on the first sample data and the second sample data, such as translation, and / or rotation, and / or random cropping, and / or color transformation, and / or flipping operations, etc., to obtain the augmented first sample data and the second sample data; and
[0059] Extract the features of the augmented first sample data and the second sample data through the negative sample encoder of the feature extraction module. As mentioned above, the negative sample encoder includes a backbone network and a projection layer. Therefore, in an embodiment of the present invention, the first feature and the second feature include two parts: a feature map and a feature vector, where the feature map is output by one or more layers of the backbone network of the negative sample encoder, and the feature vector is output by the projection layer of the negative sample encoder;
[0060] Next, in step 103, update the first data set. Calculate the difference between the first feature and the second feature, and use the difference as a supervision signal for backpropagation to update the first data set. Usually, the update of the first data set can include N rounds of iteration until a preset condition is met, where N is a natural number, and the feature extraction module can be updated once in each iteration. The preset condition can be, for example, a specified number of iteration rounds, a data set reduction multiple, a final data volume, etc. Based on the feature extraction module structure as described above, in an embodiment of the present invention, the calculation of the difference between the first feature and the second feature includes:[[]]
[0061] First, calculate the distance between the pairwise feature maps obtained from the augmented first sample data and the second sample data in the backbone network, and then obtain a distance loss matrix. In the embodiments of the present invention, the distance function used is not limited. For example, it can be the l1 loss, the mean square error loss, the KL divergence loss, etc.;
[0062] Next, a similarity matrix is calculated based on the feature vectors in the first feature and the second feature; and
[0063] Finally, based on the similarity matrix, a weighted average is performed on the distance loss matrix to obtain the final distance loss. The first data set is updated by backpropagation according to the final distance loss.
[0064] Based on the method of distilled data as described above, Figure 3 The process schematic diagram of a visual task processing method according to an embodiment of the present invention is shown. As Figure 3 shown, a visual task processing method executes a specified visual task through a task processing module, where the task processing module includes a backbone network 301 and an output head network 302.
[0065] The backbone network 301 is a first data set obtained by using the method of distilled data as described above, and is iteratively updated for the feature extraction module through a contrastive learning method. Figure 4 The process schematic diagram of the update of the backbone network of a visual task processing method according to an embodiment of the present invention is shown. As Figure 4 shown, the update of the backbone network 301 includes:
[0066] First, samples are taken from the first data set obtained according to the foregoing method to obtain a batch of sample data, and the sample data is subjected to two different data augmentation processes to obtain a first augmented synthetic image Aug1(S) and a second augmented synthetic image Aug2(S), forming a positive sample pair, where the data augmentation process may include, for example, translation, and / or rotation, and / or random cropping, and / or color transformation, and / or flipping operations, etc.;
[0067] Next, the positive sample pair is respectively input into the positive sample encoder and the negative sample encoder of the feature extraction module obtained according to the steps as described above to obtain the feature vectors output at the end of the model, denoted as the positive sample feature vector and the negative sample feature vector;
[0068] Next, the similarity vector of the positive sample feature vector and the negative sample feature vector is calculated, and the similarity matrix of the negative sample feature vector and the negative sample feature vector stored in the queue in advance is calculated. In an embodiment of the present invention, the initial value of the negative sample feature vector stored in the queue in advance is a random value, which is updated during the update process of the feature extraction module;
[0069] Next, calculate the loss and backpropagate to update the positive sample encoder. In one embodiment of the present invention, the InfoNCE loss function in Momentum Contrast for Unsupervised Visual Representation Learning is used to calculate the loss. It should be understood that in other embodiments of the present invention, other loss functions such as the NCE loss function may also be adopted; and
[0070] Finally, put the negative sample feature vectors into the negative sample feature vector queue, and remove an equal number of negative sample feature vectors at the end of the queue to achieve queue update. At the same time, update the parameters of the negative sample encoder according to the momentum update method based on the parameters of the positive sample encoder. Thus, the update of the backbone network is completed. However, it should be understood that in other embodiments of the present invention, other unsupervised learning methods may also be used to update the backbone network.
[0071] The updated backbone network and the randomly initialized output head network form a new network. In order to enable the task processing module to better adapt to the specified task, in one embodiment of the present invention, a small amount of labeled downstream task datasets are also used to fine-tune the task processing module to obtain the final model. In one embodiment of the present invention, when fine-tuning the task processing module, only the parameters of the output head network can be fine-tuned, that is, freeze or partially freeze the parameters of the updated backbone network, regard it as a feature extractor, and use a small amount of labeled downstream task datasets to train the parameters of the output head network. It is also possible to use the parameters of the updated backbone network as initialization parameters and perform fine-tuning training together with the output head network.
[0072] Based on the method for distilling data as described above, the present invention also provides an electronic device for distilling data, which includes a memory and a processor, wherein the memory is configured to store a computer program, and the computer program executes the method for distilling data as described above when running on the processor.
[0073] The present invention also provides a computer-readable storage medium for distilling data, which stores a computer program, and the computer program executes the method for distilling data as described above when running on a processor.
[0074] A method for distilling data provided by the present invention realizes data distillation based on the technology of contrastive learning, can distill various labeled data or unlabeled data, and in addition, the distilled data obtained by the method can be applied to various visual tasks including classification.
[0075] Although the embodiments of the present invention have been described above, it should be understood that they are presented by way of example only and not as a limitation. It will be apparent to those skilled in the relevant art that various combinations, variations, and changes can be made thereto without departing from the spirit and scope of the present invention. Therefore, the breadth and scope of the present invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined only in accordance with the appended claims and their equivalents.
Claims
1. A method for distilling data, characterized in that, Including the steps: Sampling from the first data set and the second data set respectively to obtain the first sample data and the second sample data, where the first data set is randomly sampled from the second data set, and the first data set is a synthetic image, and the second data set is a real image; Updating the first data set through N iterations, where N is a natural number, and each iteration includes: Extracting the features of the first sample data and the second sample data respectively through a feature extraction module to obtain the first feature and the second feature; and Calculating the difference between the first feature and the second feature, and using the difference as a supervision signal for backpropagation to update the first data set; and Updating the feature extraction module, where one iteration of updating the feature extraction module includes the steps: Performing two different data augmentation processes on the second sample data to obtain positive sample pairs; Inputting the positive sample pairs into the positive sample encoder and the negative sample encoder of the feature extraction module respectively to obtain positive sample feature vectors and negative sample feature vectors; Calculating the similarity vector of the positive sample feature vector and the negative sample feature vector, and calculating the similarity matrix of the negative sample feature vector and the pre-stored negative sample feature vector; Calculating the loss and backpropagating to update the positive sample encoder; and Putting the negative sample feature vector into the negative sample feature vector queue, and removing an equal amount of the negative sample feature vectors at the end of the queue to implement queue update, and at the same time updating the parameters of the negative sample encoder according to the momentum update method according to the parameters of the positive sample encoder.
2. The method according to claim 1, wherein It further includes the step: before updating the first data set each time, first update the feature extraction module.
3. The method according to claim 1, wherein The data augmentation process includes: translation, and / or rotation, and / or random cropping, and / or color transformation, and / or flipping.
4. The method according to claim 1, wherein The positive sample encoder includes a backbone network and a projection layer; and The structure of the negative sample encoder is the same as that of the positive sample encoder.
5. The method according to claim 1, characterized in that, The loss is calculated according to the InfoNCE loss function.
6. The method according to claim 1, wherein Extracting the features of the first sample data and the second sample data respectively through a feature extraction module includes the steps: Performing the same data augmentation transformation on the first sample data and the second sample data to obtain the augmented first sample data and the second sample data; and Extracting the features of the augmented first sample data and the second sample data through the negative sample encoder of the feature extraction module.
7. The method according to claim 6, characterized in that, The first feature and the second feature include: Feature maps, which are output by one or more layers of the backbone network of the negative sample encoder; and Feature vectors, which are output by the projection layer of the negative sample encoder.
8. The method according to claim 7, wherein Calculating the difference between the first feature and the second feature includes: Calculating a distance loss matrix according to the feature maps of the first feature and the second feature; Calculating a similarity matrix according to the feature vectors of the first feature and the second feature; Performing weighted average on the distance loss matrix according to the similarity matrix to obtain the final distance loss.
9. A computer-readable storage medium for distilling data, characterized in that, There is a computer program stored, and when the computer program runs on a processor, it executes the method for distilling data as described in any one of claims 1 to 8.
10. A method for processing visual tasks, characterized in that, Execute a specified vision task through a task processing module, where the task processing module includes: A backbone network, which is obtained by using the first data set obtained by the method of distilling data as described in any one of claims 1 to 8, and is iteratively updated for the feature extraction module through a contrastive learning method; and An output head network, which is obtained by random generation.
11. The visual task processing method according to claim 10, wherein It further includes the step of adjusting the task processing module through a labeled downstream task data set.
Citation Information
Patent Citations
Robust data set distillation method and system
CN115761414A
Domain adaptation of ai NLP encoders with knowledge distillation
US20220318502A1