A data distillation method based on a utilization rate discrimination mechanism
By introducing a utilization rate discrimination mechanism and attention discrimination module in the data distillation method, the inefficient utilization area is screened and optimized, and the problem of inefficient utilization of non-central areas of images in the prior art is solved, and efficient utilization of limited image areas and improved data distillation effect is achieved.
Patent Information
- Application Number
- CN202211604575.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-13
AI Technical Summary
The prior art ignores the efficient use of limited image areas during data distillation, resulting in inefficient use of non-central image areas.
The data distillation method based on the utilization rate discrimination mechanism is adopted to measure the utilization of each image area, inefficient utilization areas are selected and upsampled and optimized. Combined with the attention discrimination module and the inter-class discrimination module, backpropagation is performed to update the synthetic data.
The utilization of each area of the synthetic image is improved, and the efficient utilization of limited image areas is achieved. Different categories can synthesize large, medium and small objects according to their own characteristics and arrange them arbitrarily, improving the data distillation effect.
Smart Images

Figure CN117095249B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and specifically to a data distillation method based on a utilization rate discrimination mechanism. Background Art
[0002] The goal of the data distillation (also known as dataset distillation or dataset condensation) task is to refine a large training dataset (or real dataset) into a very small set of synthetic datasets (with as few as ten or even one data per class), expecting that the same model can obtain as consistent results as possible when trained using the synthetic dataset and the large dataset (generally referring to similar accuracies in the test set).
[0003] The sizes of the images in the same dataset are the same. When the number of finally synthesized images per class is determined, the total available area for this data distillation is also determined. For example, in the CIFAR-10 dataset with 10 classes, if only 1 image is synthesized per class and the length and width of each image are both 32 (ignoring the number of channels for now), then the available area is 10 * 1 * 32 * 32 = 10240. While the original training set contains 5000 images with an original area of 5000 * 32 * 32 = 5120000. At this time, the synthetic dataset only accounts for 1 / 500 of the original training set. Therefore, how to efficiently utilize the limited image area becomes extremely important. The main focus of the existing technologies is concentrated on improving the optimization scheme design in the process of synthesizing images, while ignoring the research on efficiently utilizing the limited image area, resulting in their inefficient utilization of the non-central regions of the images (these regions are generally synthesized as the background).
[0004] Aiming at the problem of how to efficiently utilize the limited image area, the present invention proposes a data distillation method based on a utilization rate discrimination mechanism. Summary of the Invention
[0005] The purpose of the present invention is to provide a data distillation method based on a utilization rate discrimination mechanism. By measuring the utilization of each image region, the low-utilization regions are located and optimized specifically to improve the data distillation effect.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A data distillation method based on a utilization rate discrimination mechanism, the method comprising:
[0007] Step S100: First, sample a batch of samples from the synthetic dataset and the real dataset respectively;
[0008] Step S200: Then, screen out the low-utilization regions in the synthetic data and upsample them to the original image size, and combine them with the original synthetic images to form a synthetic training set;
[0009] Finally, in step S300, the synthetic training set and the real data set respectively pass through the same model to obtain their respective losses and gradients. The difference between the gradients of the two is used as a supervision signal for backpropagation to update the synthetic data.
[0010] By effectively utilizing the subgraph junction regions in Dataset Condensation via Efficient Synthetic-Data Parameterization, the utilization of each region of the synthetic image is further improved. At the same time, the same synthesis rules are not set for each category, and different categories can synthesize large, medium, and small objects according to their own characteristics and arrange them arbitrarily.
[0011] As a preferred embodiment of the present invention, the detailed steps for screening out the low-utilization regions in the synthetic data in step S200 of this method are as follows:
[0012] In step S201, after the synthetic image passes through the model forward, a feature activation map is obtained, and the original average activation value of each region is calculated.
[0013] In step S202, Gaussian noise is added to the synthetic image, and after passing through the model forward and backward again, a feature activation map is obtained, and the noise average activation value of each region is calculated.
[0014] In step S203, step S202 is repeated to obtain multiple noise average activation values.
[0015] In step S204, the regions where the mean of the multiple noise activation values is similar to the original image average activation value are determined as the low-utilization regions.
[0016] As a preferred embodiment of the present invention, the present invention further includes an attention discrimination module for screening out the low-utilization regions in the synthetic data.
[0017] As a preferred embodiment of the present invention, there are two parts to the loss in S300. One part is the loss generated by inter-class discrimination, and the other part is the cross-entropy loss between the prediction result of the model for the real data and the real label.
[0018] As a preferred embodiment of the present invention, the present invention further includes a class discrimination module for calculating the loss in step S300, which is used to reduce the cosine similarity of features of the same class and increase the cosine similarity of features between different classes, where the features of the same class and other classes are updated by the moving average algorithm.
[0019] Compared with the prior art, the beneficial effects of the present invention are:
[0020] The present invention further improves the utilization of each region of the synthetic image by effectively utilizing the sub-graph junction region in Dataset Condensation via Efficient Synthetic-Data Parameterization. At the same time, the same synthetic rules are not set for each category, and different categories can synthesize large, medium, and small objects according to their own characteristics and arrange them arbitrarily. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.
[0022] Figure 1 It is the overall flowchart of a data distillation method based on a utilization rate discrimination mechanism of the present invention;
[0023] Figure 2 It is the schematic diagram of the inter-class discrimination module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0025] Please refer to Figure 1 - Figure 2 , Figure 1 It is the overall flowchart of a data distillation method based on a utilization rate discrimination mechanism of the present invention. In the figure, the synthetic dataset is synthetic data, the real dataset is real data, the attention discrimination module is Utilization Discrimination Module.UDM, the synthetic training set is synthetic training data, and the model is network.
[0026] Embodiment 1
[0027] As Figure 1The present invention provides a data distillation method based on a utilization discrimination mechanism, and the method includes the following content: The real dataset and synthetic dataset of the present invention originate from the goal of the data distillation task, which is to refine a large training dataset (or real dataset) into a very small set of synthetic datasets (as low as ten or even one data per category), and it is expected that the same model can obtain as consistent results as possible when training with the synthetic dataset and training with the large dataset (generally referring to similar accuracies in the test set). The present invention mainly focuses on the distillation of image data, and the "data" mentioned in the text specifically refers to images.
[0028] Step S100 First, we sample a batch of samples from the synthetic data and the real data respectively.
[0029] Step S200 After that, use the Utilization Discrimination Module (UDM) to screen out the low-utilization regions in the synthetic data and upsample them to the original image size, and form a synthetic training data set with the original synthetic image.
[0030] Step S300 Finally, the synthetic training data set and the real data set pass through the same model (network) respectively to obtain their respective losses and gradients. The difference between the two gradients is used as a supervision signal for backpropagation to update the synthetic data. In this embodiment, the model generally refers to a classification model, such as convnet, alexnet, vgg, resnet, etc. In this step, the loss has two parts. One part is the loss generated by inter-class discrimination, and the other part is the cross-entropy loss between the prediction result of the model for the real data and the real label.
[0031] Embodiment 2
[0032] The present invention also includes its Utilization Discrimination Module (UDM), and its process is as follows:
[0033] Step S201 After the synthetic image passes through the model forward and backward, obtain the feature activation map and calculate the original average activation value of each region;
[0034] Step S202 Add Gaussian noise to the synthetic image, and then pass through the model forward and backward again to obtain the feature activation map and calculate the noise average activation value of each region;
[0035] Step S203 Repeat Step S202 to obtain multiple noise average activation values;
[0036] The regions where the mean of multiple noise activation values is close to the original image average activation value are the low-utilization regions.
[0037] The principle of the attention discrimination module is as follows: the activation values of low-utilization areas are very low, and in the eyes of the model, they are closer to cluttered noise than foreground areas. After adding noise, the low-utilization areas do not change much in the eyes of the model, so the areas where the difference between the noise average activation value and the original average activation value is small are more likely to be our target areas; while the model is more sensitive to changes in the foreground area (high-utilization area), so the activation value of this area changes more. Multiple repetitions are to reduce the impact of accidental and randomness brought by word operations.
[0038] Embodiment 3
[0039] The present invention also includes an inter-class discrimination module such as Figure 2 As shown, Figure 2 The schematic diagram of the inter-class discrimination module of the present invention. The module reduces the cosine similarity of features of the same category and increases the cosine similarity of features of different categories, thereby achieving better class discrimination and differentiation. The features of the same category and other categories are updated by the sliding average algorithm, which reduces the impact of noise points and discrete points and improves stability.
[0040] In summary, the utilization discrimination mechanism of the present invention measures the utilization of each image area and finds the low-utilization area for key training, thereby improving the full utilization of limited areas. On this basis, the inter-class discrimination loss is used to shorten the feature distance of data of the same category and increase the feature distance of data of different categories, thereby reducing the learning difficulty of the model and improving the performance of distilled data. That is, the effective utilization of the boundary area of the sub-images in the IDC is achieved; the utilization of each area of the synthesized image is further improved; the same synthesis rules are not set for each category, and different categories can synthesize large, medium, and small objects according to their own characteristics and arrange them arbitrarily.
[0041] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0042] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A data distillation method based on a utilization rate discrimination mechanism, characterized in that, The method includes: Step S100: First, sample a batch of samples from the synthetic dataset and the real dataset respectively; Step S200: After that, screen out the low-utilization regions in the synthetic data and upsample them to the original image size, and combine them with the original synthetic image to form a synthetic training set; Step S300: Finally, the synthetic training set and the real dataset respectively pass through the same model to obtain their respective losses and gradients. The difference between the two gradients is used as a supervision signal for backpropagation to update the synthetic data; In the method, the detailed steps of screening out the low-utilization regions in the synthetic data in step S200 are as follows: Step S201: After the synthetic data passes through the model forward and backward, obtain the feature activation map and calculate the original average activation value of each region; Step S202: Add Gaussian noise to the synthetic data, pass through the model forward and backward again to obtain the feature activation map, and calculate the noise average activation value of each region; Step S203: Repeat step S202 to obtain multiple noise average activation values; Step S204: Determine the regions where the mean of the multiple noise average activation values is similar to the original average activation value as the low-utilization regions.
2. The data distillation method based on a utilization rate discrimination mechanism according to claim 1, wherein It also includes an attention discrimination module for screening out the low-utilization regions in the synthetic data.
3. A data distillation method based on a utilization rate discrimination mechanism according to claim 2, wherein the loss in S300 has two parts. One part is the loss generated by inter-class discrimination, and the other part is the cross-entropy loss between the prediction result of the model on the real dataset and the real label.
4. A data distillation method based on a utilization rate discrimination mechanism according to claim 3, characterized in that It also includes that the loss calculation in step S300 adopts an inter-class discrimination module, which is used to reduce the cosine similarity of features in the same class and increase the cosine similarity of features between different classes, and the features of the same class and other classes are updated by the moving average algorithm.
Citation Information
Patent Citations
Image classification method, device, terminal device, and readable storage medium
CN109376786A
Ship sign image super-resolution method based on semantic information and gradient supervision
CN113935899A