Data set knowledge density evaluation method and device based on noise optimization type data set distillation

By optimizing noise in batches within the diffusion model and utilizing distribution alignment loss, maximum occupancy loss, and authenticity constraint loss, the distribution bias problem caused by random noise is addressed, thereby improving the accuracy of knowledge density assessment for datasets and the training performance of synthetic datasets.

CN120910494APending Publication Date: 2025-11-07INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510771401.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing dataset distillation methods based on diffusion models rely on random noise when synthesizing datasets, leading to distribution bias and affecting the accuracy of knowledge density assessment.

Method used

Based on the diffusion model, the denoising diffusion probability model DDPM is performed batch by batch. Random noise is optimized by distribution alignment loss, maximum occupancy loss and authenticity constraint loss. The final synthetic sample set is obtained iteratively to ensure that the training performance of the synthetic dataset is consistent with that of the original dataset.

Benefits of technology

It improves the accuracy of knowledge density assessment of datasets, eliminates distribution bias caused by randomness, is applicable to special scenarios where the original dataset is unavailable, and enhances the training performance of synthetic datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910494A_ABST
    Figure CN120910494A_ABST
Patent Text Reader

Abstract

The invention provides a data set knowledge density evaluation method and device based on noise optimization type data set distillation, and relates to the technical field of data processing.The method comprises the steps that on the basis that data set distillation is carried out based on a diffusion model, DDPM denoising is carried out batch by batch, performing random noise optimization in each step of denoising of each batch based on distribution alignment loss, maximum occupancy loss and authenticity constraint loss, and iteratively obtaining a final synthetic sample set of the last step of denoising of each batch; obtaining a distillation data set according to the final synthesis sample set of the last step of all batches of de-noised samples; the test performance of the classifier trained by the distillation data set is the same as that of the classifier trained on the original data set; and determining the knowledge density of the original data set according to the total data number of the distillation data set and the total data number of the original data set. According to the invention, the accuracy of knowledge density evaluation of the data set is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a dataset knowledge density evaluation method and device based on noise-optimized dataset distillation. BACKGROUND

[0002] Dataset distillation is to extract a large dataset into a small synthetic dataset while retaining the training effectiveness of the large dataset on downstream tasks. The dataset distillation method can be used to evaluate the knowledge density of the dataset.

[0003] When facing large-scale high-resolution datasets, a diffusion model is recently proposed to generate high-resolution and realistic synthetic datasets. The current dataset distillation method based on the diffusion model adopts denoising diffusion probabilistic model (DDPM) denoising.

[0004] Although synthetic datasets with extremely high training performance can be synthesized by adjusting the diffusion model or providing additional supervision in the denoising process, the DDPM denoising adopted is a random process, and the synthesis process highly depends on the randomly generated noise. Therefore, distribution deviation caused by random sampling of noise occurs in the synthetic dataset, which leads to the introduction of random noise in the calculation of the minimum lossless distillation ratio, and further reduces the accuracy of the knowledge density evaluation of the dataset. SUMMARY

[0005] The present application provides a dataset knowledge density evaluation method and device based on noise-optimized dataset distillation, to solve the defect that the accuracy of the knowledge density evaluation of the dataset in the prior art needs to be improved, and to realize improving the accuracy of the knowledge density evaluation of the dataset.

[0006] The present application provides a dataset knowledge density evaluation method based on noise-optimized dataset distillation, comprising: On the basis of dataset distillation based on the diffusion model, denoising diffusion probabilistic model (DDPM) denoising is performed batch by batch, and random noise optimization is performed based on distribution alignment loss, maximum occupancy loss and authenticity constraint loss in each step of each batch denoising, to iteratively obtain the final synthetic sample set of the last step of each batch denoising; According to the final synthetic sample set of the last step of all batch denoising, a distilled dataset is obtained; the test performance of the classifier trained by the distilled dataset is the same as the test performance of the classifier trained on the original dataset; According to the total number of data of the distilled dataset and the total number of data of the original dataset, the knowledge density of the original dataset is determined.

[0007] In some embodiments, the random noise is optimized based on the distribution alignment loss, the maximum occupancy loss, and the authenticity constraint loss in each step of each batch denoising, and a final synthetic sample set of a last step of each batch denoising is iteratively obtained, including: In each step of each batch denoising, a weighted sum of the distribution alignment loss of each step, the maximum occupancy loss of each step, and the authenticity constraint loss of each step is taken as a total loss of each step; Based on the total loss of each step, an optimized random noise of each step is obtained; Based on the optimized random noise of each step, a final synthetic sample set of a last step of each batch denoising is iteratively obtained.

[0008] In some embodiments, the method further includes: The following steps are performed for each step in each batch denoising process: Element-wise mean and standard deviation calculation is performed on a sampling sample set of the current step to obtain a first mean and a first standard deviation; the sampling sample set is a sample set sampled from a current step of a DDPM denoising process independent of the data set distillation process; All synthetic samples in a synthetic sample set of the current step are normalized element-wise using the first mean and the first standard deviation to obtain an element-wise normalized synthetic sample set; Element-wise mean and standard deviation calculation is performed on the element-wise normalized synthetic sample set to obtain a second mean and a second standard deviation; Based on the second mean and the second standard deviation, a first distribution alignment loss is obtained; Channel-wise mean and standard deviation calculation is performed on a sampling feature set mapped from the sampling sample set to obtain a third mean and a third standard deviation; All synthetic features in a synthetic feature set mapped from the synthetic sample set are normalized channel-wise using the third mean and the third standard deviation to obtain a channel-wise normalized synthetic feature set; Channel-wise mean and standard deviation calculation is performed on the channel-wise normalized synthetic feature set to obtain a fourth mean and a fourth standard deviation; Based on the fourth mean and the fourth standard deviation, a second distribution alignment loss is obtained; A weighted sum of the first distribution alignment loss and the second distribution alignment loss is taken as the distribution alignment loss.

[0009] In some embodiments, the method further includes: The nearest neighbor element-wise normalized synthetic sample corresponding to each element-wise normalized synthetic sample in the element-wise normalized synthetic sample set is determined; determine a first maximum occupancy loss according to each element-wise normalized synthetic sample in the element-wise normalized synthetic sample set and a nearest neighbor element-wise normalized synthetic sample corresponding to each element-wise normalized synthetic sample; determine a nearest neighbor channel-wise normalized synthetic feature corresponding to each channel-wise normalized synthetic feature in the channel-wise normalized synthetic feature set; determine a second maximum occupancy loss according to each channel-wise normalized synthetic feature in the channel-wise normalized synthetic feature set and the nearest neighbor channel-wise normalized synthetic feature corresponding to each feature; take a weighted sum of the first maximum occupancy loss and the second maximum occupancy loss as the maximum occupancy loss.

[0010] In some embodiments, the method further comprises: determine the authenticity constraint loss of each step based on an expected mode length of a standard multivariate normal distribution sample and a noise tensor of each step.

[0011] The application also provides a data set knowledge density evaluation device based on noise optimization type data set distillation, comprising: An optimization module is configured to, on the basis of data set distillation based on a diffusion model, perform denoising diffusion probability model (DDPM) denoising in batches, and perform random noise optimization based on a distribution alignment loss, a maximum occupancy loss, and an authenticity constraint loss in each step of each batch denoising, to iteratively obtain a final synthetic sample set of a last step of each batch denoising. An acquisition module is configured to obtain a distilled data set from the final synthetic sample set of the last step of all batch denoising, wherein a test performance of a classifier trained based on the distilled data set is the same as a test performance of a classifier trained based on an original data set. A determination module is configured to determine a knowledge density of the original data set based on a total number of data in the distilled data set and a total number of data in the original data set.

[0012] The application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data set knowledge density evaluation method based on noise optimization type data set distillation according to any one of the above embodiments when executing the computer program.

[0013] The application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the data set knowledge density evaluation method based on noise optimization type data set distillation according to any one of the above embodiments.

[0014] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements any of the above-mentioned data set knowledge density evaluation methods based on noise-optimized data set distillation.

[0015] The application provides a data set knowledge density evaluation method and device based on noise-optimized data set distillation. On the basis of data set distillation based on a diffusion model, denoising diffusion probability model (DDPM) denoising is performed batch by batch. In each step of each batch denoising, random noise optimization is performed based on distribution alignment loss, maximum occupancy loss, and authenticity constraint loss. The final synthetic sample set of the last step of each batch denoising is obtained iteratively. A distilled data set is obtained according to the final synthetic sample set of the last step of batch denoising. The test performance of a classifier trained by the distilled data set is the same as the test performance of a classifier trained on the original data set. The knowledge density of the original data set is determined according to the total number of data entries of the distilled data set and the total number of data entries of the original data set. The application optimizes the random noise sampled in each step of the denoising process by using distribution alignment loss, maximum occupancy loss, and authenticity constraint loss. The distribution deviation caused by randomness is solved, thereby improving the accuracy of knowledge density evaluation of the data set. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0017] Figure 1 is a flowchart of the data set knowledge density evaluation method based on noise-optimized data set distillation provided by the application.

[0018] Figure 2 is a structural schematic diagram of the data set knowledge density evaluation device based on noise-optimized data set distillation provided by the application.

[0019] Figure 3 is a structural schematic diagram of the electronic device provided by the application. DETAILED DESCRIPTION

[0020] Dataset distillation distills large datasets into small synthetic datasets while preserving the training efficacy of the large datasets on downstream tasks. On one hand, dataset distillation methods can compress large datasets into small datasets, saving training costs; on the other hand, dataset distillation methods can be used to evaluate the knowledge density of a dataset. Specifically, the knowledge density of a dataset can be estimated by calculating the minimum lossless distillation ratio of the dataset. The minimum lossless distillation ratio is the limit to which a dataset can be compressed using a specific dataset distillation method without compromising the training efficacy of the dataset. Since the size of a dataset is usually measured by the number of items per class (IPC), the distillation ratio of a synthetic dataset is the ratio between the IPC of the distilled dataset and the IPC of the original dataset.

[0021] Early methods minimize the distillation loss used in each method by directly optimizing the pixels of the synthetic dataset through pixel-by-pixel optimization. Common losses include measures of distribution difference between the synthetic dataset and the original dataset, trajectory difference or gradient difference when training a model on the synthetic dataset and the original dataset, and classification loss of a specific model trained on the synthetic dataset on the original dataset. Early methods cannot achieve satisfactory distillation results on large-scale high-resolution datasets, and are extremely costly.

[0022] Recently, some methods have proposed using diffusion models to generate high-resolution and realistic synthetic datasets. For example, the Minimax Diffusion method fine-tunes the diffusion model to achieve a trade-off between representativeness and diversity of the generated samples; the Influence Guided Diffusion (IGD) method designs an influence guiding function to provide additional influence in the denoising process of the diffusion model. In these diffusion model-based methods, the synthetic dataset is directly generated by a specific pre-trained diffusion model, which is much faster than previous methods and has excellent results on large-scale high-resolution datasets. In particular, the IGD method can achieve lossless distillation, and therefore can be used to calculate the minimum lossless distillation ratio.

[0023] Current diffusion model-based dataset distillation methods all use a random DDPM denoising process to generate synthetic datasets. Although existing methods can synthesize synthetic datasets with extremely high training performance by adjusting the diffusion model or providing additional supervision in the denoising process, the DDPM denoising used is a random process, and the synthesis process is highly dependent on randomly generated noise. Therefore, distribution bias caused by random sampling of noise will occur in the synthetic samples. This leads to suboptimal training performance of the distilled dataset on the one hand, and introduces random noise in the calculation of the minimum lossless distillation ratio, which is not conducive to evaluating the knowledge density of the dataset.

[0024] To solve the above problems, the application provides a dataset knowledge density evaluation method based on noise optimization dataset distillation, which explicitly optimizes the distribution distillation framework of noise to solve the distribution deviation caused by randomness. The framework can be combined with existing dataset distillation methods based on diffusion models to improve the performance of these methods and eliminate the problem of difficulty in calculating the minimum lossless compression ratio caused by randomness of the existing methods. At the same time, the method does not depend on the original dataset in the process of synthesizing the dataset, so it is suitable for special application scenarios where the original dataset cannot be obtained.

[0025] To make the purpose, technical solutions and advantages of the application clearer, the technical solutions in the application will be described clearly and completely below combined with the drawings in the application. Obviously, the described embodiments are part of the embodiments of the application, not all embodiments. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative labor belong to the scope of protection of the application.

[0026] Figure 1 The application provides a flowchart of the dataset knowledge density evaluation method based on noise optimization dataset distillation, as shown in Figure 1 The application provides a dataset knowledge density evaluation method based on noise optimization dataset distillation, which comprises: Step 110, on the basis of dataset distillation based on a diffusion model, denoising diffusion probability model (DDPM) denoising is performed batch by batch, and random noise optimization is performed based on distribution alignment loss, maximum occupancy loss and authenticity constraint loss in each step of each batch denoising, and the final synthesized sample set of the last step of each batch denoising is obtained iteratively.

[0027] Step 120, according to the final synthesized sample set of the last step of all batch denoising, a distilled dataset is obtained; the test performance of the classifier trained by the distilled dataset is the same as the test performance of the classifier trained on the original dataset.

[0028] Step 130, according to the total number of data of the distilled dataset and the total number of data of the original dataset, the knowledge density of the original dataset is determined.

[0029] Specifically, dataset distillation is continuously and batch by batch based on a diffusion model, that is, denoising diffusion probability model (DDPM) denoising is continuously and batch by batch.

[0030] The DDPM process is a step-by-step denoising process. In the first step, the diffusion model is based on the previous step ​ samples of predicted next step samples of mean of and standard deviation and randomly sample a noise tensor from the normal distribution corresponding to the mean and standard deviation as the sample of the current step.

[0031] To solve the distribution deviation caused by random sampling, the random noise sampled at each step of each batch DDPM denoising process is optimized explicitly. Distribution alignment loss, maximum occupancy loss, and authenticity constraint loss are proposed as the loss function for optimizing the noise at each step. Based on the distribution alignment loss, maximum occupancy loss, and authenticity constraint loss, the random noise sampled at each step of each batch DDPM denoising process is optimized.

[0032] The distribution alignment loss explicitly penalizes the difference between the sample distribution of the synthesized sample set at each step of denoising and the probability distribution of the corresponding independent denoising process at that step. The maximum occupancy loss explicitly increases the distance between each synthesized sample set to increase the diversity of the synthesized sample set. The authenticity loss ensures that the optimized noise still obeys the standard normal distribution by constraining the length of the noise at the first step.

[0033] Since there is a fixed conversion formula (DDPM denoising formula) between noise and synthesized sample set, the process of optimizing random noise can be equivalently converted into the process of optimizing synthesized sample set. The final synthesized sample set obtained at the end of each step of equivalent optimization process is the synthesized sample set obtained using the optimized random noise.

[0034] Based on the optimized random noise at each step, the final synthesized sample set at each step is obtained. The diffusion model predicts the mean and variance of the last step based on the final synthesized sample set of the last step. The final synthesized sample set of the last step of each batch denoising is obtained using the optimized random noise of the last step and the mean and variance of the last step.

[0035] After each batch synthesis, the final synthesized sample set of the last step of the batch denoising is obtained, and then all the final synthesized sample sets of the last step of the batch denoising generated so far are used to train a classifier. The classifier adopts the classifier architecture with the best test classification performance on the original data set, and then the next batch of synthesis is performed. This process is repeated until the test performance of the classifier trained in this cycle is the same as the test performance of the classifier trained on the original data set. In this way, a distilled data set is obtained, which is the set of all the final synthesized sample sets of the last step of the batch denoising generated before the cycle terminates. At this time, the total number of data in the distilled data set and the total number of data in the original data set are determined.​

[0036] Divide the total number of data items of the distilled data set by the total number of data items of the original data set to obtain a minimum lossless distillation ratio. The minimum lossless distillation ratio is the knowledge density of the original data set. The expression of the minimum lossless distillation ratio is as follows: In the formula, denotes the minimum lossless distillation ratio, denotes the total number of data items of the distilled data set, denotes the total number of data items of the original data set.

[0037] The data set knowledge density evaluation method based on noise optimization type data set distillation provided by the present application, on the basis of data set distillation based on a diffusion model, denoises in batches by using a denoising diffusion probability model (DDPM), and in each step of each batch denoising, optimizes random noise based on a distribution alignment loss, a maximum occupancy loss, and a reality constraint loss, and iteratively obtains a final synthesized sample set of the last step of each batch denoising; obtains a distilled data set according to the final synthesized sample set of the last step of each batch denoising; the test performance of a classifier trained by the distilled data set is the same as the test performance of a classifier trained on the original data set; and the knowledge density of the original data set is determined according to the total number of data items of the distilled data set and the total number of data items of the original data set. The present application optimizes the random noise sampled in each step of the denoising process by using the distribution alignment loss, the maximum occupancy loss, and the reality constraint loss, solves the distribution deviation caused by randomness, and thus improves the accuracy of the knowledge density evaluation of the data set.

[0038] In some embodiments, the data set knowledge density evaluation method based on noise optimization type data set distillation provided by the present application further comprises: The following steps are performed for each step in each batch denoising process: Element-wise mean and standard deviation calculation is performed on the sampling sample set of the current step to obtain a first mean and a first standard deviation; the sampling sample set is a sample set sampled from the current step of a DDPM denoising process independent of the data set distillation process; All synthesized samples in the synthesized sample set of the current step are normalized element-wise using the first mean and the first standard deviation to obtain an element-wise normalized synthesized sample set; Element-wise mean and standard deviation calculation is performed on the element-wise normalized synthesized sample set to obtain a second mean and a second standard deviation; Based on the second mean and the second standard deviation, a first distribution alignment loss is obtained; Channel-wise mean and standard deviation calculation is performed on the sampling feature set mapped from the sampling sample set to obtain a third mean and a third standard deviation; All synthetic features in the synthetic feature set mapped from the synthetic sample set are normalized channel by channel using the third mean and the third standard deviation to obtain the channel-normalized synthetic feature set. The mean and standard deviation of the channel-by-channel normalized synthetic feature set are calculated to obtain the fourth mean and the fourth standard deviation. Based on the fourth mean and the fourth standard deviation, the second distribution alignment loss is obtained; The weighted sum of the first distribution alignment loss and the second distribution alignment loss is used as the distribution alignment loss.

[0039] Specifically, the following section will focus on noise reduction... This example illustrates the steps; other steps are similar and will not be repeated here.

[0040] From a standalone DDPM denoising process Step-by-step collection Each sample constitutes a sample set. A standalone DDPM denoising process refers to a standard DDPM denoising process independent of dataset distillation. (Sample set) The expression is as follows: In the formula, Indicates the first Step-by-step sample collection Indicates the collection of sample sets The first in One sample, sample Shape writing , This indicates the channel dimension size of the sample. Indicates the height of the sample. Indicates the width of the sample. The range of values ​​for is [1, ...]. ], Indicates the collection of sample sets The number of samples in the sample.

[0041] Distribution Alignment Loss It consists of two parts: the first distribution alignment loss Second distribution alignment loss First distribution alignment loss It is an element-wise distribution alignment loss, and a second distribution alignment loss. It is a channel-by-channel distributed alignment loss.

[0042] To calculate the alignment loss of the first distribution First, for the sample set Calculate the mean and standard deviation for each element to obtain the first mean. and the first standard deviation Both of them have the following shapes .

[0043] Secondly, synthesize the sample set All synthetic samples were taken with the first mean. and the first standard deviation Element-wise normalization is performed to obtain the element-wise normalized synthetic sample set. Element-wise normalized synthetic sample set The expression is as follows: In the formula, This represents a composite sample set normalized element by element. The first element in the element-normalized synthetic sample set represents the first element. Each element-wise normalized composite sample This represents the number of samples in the element-wise normalized synthetic sample set. Indicates the first sample in the synthetic sample set One synthetic sample, The range of values ​​for is [1, ...]. ], This represents the first mean. This represents the first standard deviation.

[0044] Then, the element-wise normalized synthetic sample set The second mean is obtained by calculating the mean and standard deviation for each element. Second standard deviation Both of them have the following shapes .

[0045] Finally, based on the second mean Second standard deviation The first distribution alignment loss is obtained. First distribution alignment loss The expression is as follows: In the formula, This represents the alignment loss of the first distribution. Indicates the second mean The ( ) )element, Indicates the second standard deviation The ( ) )element, This represents the sum of the absolute value and the square. , Indicates operands, The range of values ​​for is [1, ...]. ], The range of values ​​for is [1, ...]. ], The range of values ​​for is [1, ...]. ].

[0046] To calculate the second distribution alignment loss First, a mapper is used to sample the sample set. Mapped to sampled feature set Sampling feature set The shape of the feature tensor in is ,satisfy The mapper consists of one 3x3Stride1 convolution layer activated by LeakyReLU and three 4x4Stride2 convolution layers also activated by LeakyReLU, in sequence.

[0047] For the sampled feature set The mean and standard deviation of each channel are calculated to obtain the third mean. and the third standard deviation Both of them have the following shapes .

[0048] In the formula, Indicates the third mean The ( ) )element, Represents the sampling feature set The Middle The ( )th feature )element, Indicates the first Step-by-step sampling sample set The number of samples in Indicates the third standard deviation The ( ) )element, The range of values ​​for is [1, ...]. ], The range of values ​​for is [1, ...]. ], The range of values ​​for is [1, ...]. ].

[0049] Secondly, the above mapper is used to synthesize the sample set. Mapping to the synthetic feature set All the synthetic features in the synthetic feature set are replaced by the third mean and the third standard deviation . The third mean and the third standard deviation are normalized channel by channel to obtain the channel-wise normalized synthetic feature set . The expression of the channel-wise normalized synthetic feature set wherein denotes the channel-wise normalized synthetic feature set, denotes the th channel-wise normalized synthetic feature in the channel-wise normalized synthetic feature set , denotes the number of synthetic features in the channel-wise normalized synthetic feature set , denotes the th synthetic feature in the synthetic feature set , denotes the third mean, denotes the third standard deviation.

[0050] Then, the channel-wise mean and standard deviation of the channel-wise normalized synthetic feature set are calculated to obtain the fourth mean and the fourth standard deviation .

[0051] wherein denotes the th element in the fourth mean , denotes the th element in the th channel-wise normalized synthetic feature in the channel-wise normalized synthetic feature set , denotes the number of samples in the synthetic sample set , denotes the th element in the fourth standard deviation .

[0052] Finally, based on the fourth mean and the fourth standard deviation , the second distribution alignment loss is obtained. The second distribution alignment loss The expression is as follows: In the formula, represents the second distribution alignment loss, represents the fourth mean element in the fourth mean, represents the fourth standard deviation element in the fourth standard deviation, represents the sum of absolute values and squares.

[0053] After obtaining the first distribution alignment loss and the second distribution alignment loss , the sum of the first distribution alignment loss and the second distribution alignment loss is weighted as the distribution alignment loss. The goal of the distribution alignment loss is to reduce the distribution difference between the sampled sample set and the synthesized sample set, and achieve distribution alignment.

[0054] In some embodiments, the dataset knowledge density evaluation method for noise-based optimized dataset distillation provided by the present application further comprises: determining the nearest neighbor element-wise normalized synthesized sample corresponding to each element-wise normalized synthesized sample in the element-wise normalized synthesized sample set; obtaining a first maximum occupancy loss according to each element-wise normalized synthesized sample in the element-wise normalized synthesized sample set and the nearest neighbor element-wise normalized synthesized sample corresponding to each element-wise normalized synthesized sample; determining the nearest neighbor element-wise normalized synthesized feature corresponding to each element-wise normalized synthesized feature in the element-wise normalized synthesized feature set; obtaining a second maximum occupancy loss according to each element-wise normalized synthesized feature in the element-wise normalized synthesized feature set and the nearest neighbor element-wise normalized synthesized feature corresponding to each feature; weighting the sum of the first maximum occupancy loss and the second maximum occupancy loss as the maximum occupancy loss.

[0055] Specifically, the maximum occupancy loss is composed of two parts: the first maximum occupancy loss and the second maximum occupancy loss .

[0056] For any sample set and a sample in the sample set, the nearest neighbor sample corresponding to the sample is determined by the cosine similarity.

[0057] ​​ wherein, denotes the nearest neighbor sample of sample , denotes any sample in the sample set, denotes the samples in the sample set except .

[0058] To calculate the first maximum occupancy loss , the nearest neighbor element-wise normalized synthetic sample corresponding to each element-wise normalized synthetic sample in the element-wise normalized synthetic sample set is determined, cosine similarity calculation is performed on each element-wise normalized synthetic sample and the nearest neighbor element-wise normalized synthetic sample corresponding to each element-wise normalized synthetic sample, and all cosine similarities are accumulated to obtain the first maximum occupancy loss . The expression of the first maximum occupancy loss is as follows: wherein, denotes the first maximum occupancy loss, denotes the element-wise normalized synthetic sample set, denotes the element-wise normalized synthetic sample, denotes the nearest neighbor element-wise normalized synthetic sample corresponding to the element-wise normalized synthetic sample .

[0059] To calculate the second maximum occupancy loss , the nearest neighbor channel-wise normalized synthetic feature corresponding to each channel-wise normalized synthetic feature in the channel-wise normalized synthetic feature set is determined, cosine similarity calculation is performed on each channel-wise normalized synthetic feature and the nearest neighbor channel-wise normalized synthetic feature corresponding to each channel-wise normalized synthetic feature, and all cosine similarities are accumulated to obtain the second maximum occupancy loss . The expression of the second maximum occupancy loss is as follows: wherein, denotes the second maximum occupancy loss, denotes the channel-wise normalized synthetic feature set, denotes the channel-wise normalized synthetic feature, denotes the nearest neighbor channel-wise normalized synthetic feature corresponding to the channel-wise normalized synthetic feature .

[0060] After obtaining the first maximum occupancy loss and the second maximum occupancy loss After that, the weighted sum of the first maximum occupancy loss and the second maximum occupancy loss is taken as the maximum occupancy loss.

[0061] In some embodiments, the dataset knowledge density evaluation method based on noise optimization dataset distillation provided by the present application further comprises: Based on the expected norm of the standard multivariate normal distribution sample and the noise tensor of each step, the authenticity constraint loss is obtained.

[0062] Specifically, since the norm of each sample in a d-dimensional standard multivariate normal distribution is close to Therefore, based on the expected norm of the standard multivariate normal distribution sample and the noise tensor of each step, the authenticity constraint loss of each step is obtained. The expression of the authenticity constraint loss is as follows: In the formula, represents the authenticity constraint loss, represents the noise tensor of the t-th step, is the random noise to be optimized, represents the expected norm of the standard multivariate normal distribution sample.

[0063] In some embodiments, the random noise is optimized based on the distribution alignment loss, the maximum occupancy loss, and the authenticity constraint loss in each step of each batch denoising, and the final synthetic sample set of the last step of each batch denoising is iteratively obtained, comprising: In each step of each batch denoising, the weighted sum of the distribution alignment loss of each step, the maximum occupancy loss of each step, and the authenticity constraint loss of each step is taken as the total loss of each step; Based on the total loss of each step, the optimized random noise of each step is obtained; Based on the optimized random noise of each step, the final synthetic sample set of the last step of each batch denoising is iteratively obtained.

[0064] Specifically, for each step in the DDPM denoising process of each batch, the weighted sum of the distribution alignment loss of each step, the maximum occupancy loss of each step, and the authenticity constraint loss of each step is taken as the total loss of each step. It should be noted that the weight of the authenticity constraint loss is artificially set to 1, and the weights of the other two losses are determined through parameter tuning experiments (synthetic samples are generated using different weights, and the weight corresponding to the synthetic sample with the best training performance is selected).

[0065] The total loss is the objective function used to optimize random noise, and the optimization direction for random noise is to reduce the total loss. This refers to the noise tensor used in this step. A 200-step optimization was performed using a stochastic gradient descent (SGD) optimizer with a learning rate of 0.1 and a momentum of 0.9. The final synthetic sample set for each step was then calculated using the optimized random noise.

[0066] The diffusion model predicts the mean and variance of the current step based on the final synthetic sample set of the previous step. It then uses the optimized random noise of the current step, along with the mean and variance of the current step, to calculate the final synthetic sample set of the current step based on the DDPM denoising formula.

[0067] In the formula, express The final synthetic sample set, The diffusion model is based on The final synthetic sample set predicted The mean and standard deviation, express The final random noise obtained after optimization.

[0068] The process iterates sequentially to the final denoising step. The diffusion model predicts the mean and variance of the final step based on the final synthetic sample set of the previous step. Using the optimized random noise of the final step and the mean and variance of the final step, the final synthetic sample set of the final step of each batch of denoising is obtained.

[0069] The Noise-Optimized Distribution Distillation (NODD) algorithm proposed in this invention can be combined with other similar algorithms to obtain synthetic datasets with stronger training performance.

[0070] The NODD method proposed in this invention does not involve the original dataset during dataset synthesis, thus enabling dataset distillation even when the original dataset is unavailable, whereas other existing methods require the original dataset for synthesis. Compared to other methods, this invention offers higher performance and better generalizability. Furthermore, compared to other diffusion-based methods, this invention exhibits higher synthesis efficiency. Finally, when used to estimate the knowledge density of a dataset, this invention eliminates the randomness of the minimum lossless distillation ratio calculation and removes the influence of noise, thereby enabling a more accurate estimation of the knowledge density of the original dataset.

[0071] The present application mainly focuses on how to weaken the deviation between the synthetic dataset generated by the denoising process of the diffusion model and the denoising probability distribution of the diffusion model by eliminating the random sampling bias. As an extension, the present solution can be further enhanced by introducing the distribution information of the original dataset, changing the target distribution of distribution alignment from the denoising probability distribution of the diffusion model to the sample distribution of the original dataset. In addition, the distribution alignment capability can be further improved by introducing an explicit distribution alignment loss function such as Maximum Mean Discrepancy (MMD), so as to obtain better distillation performance.

[0072] It should be noted that the synthetic sample set in the present application is in the denoising process, the synthetic dataset is used to train the classifier, and the final synthetic sample set of the last step of each batch denoising is used to train the classifier, so the final synthetic sample set of the last step of each batch denoising is also the synthetic dataset.

[0073] The data set knowledge density evaluation device based on noise optimization type data set distillation provided by the present application is described below, and the data set knowledge density evaluation device based on noise optimization type data set distillation described below can be correspondingly referred to the data set knowledge density evaluation method based on noise optimization type data set distillation described above.

[0074] Figure 2 The structure diagram of the data set knowledge density evaluation device based on noise optimization type data set distillation provided by the present application is shown in Figure 2 The data set knowledge density evaluation device based on noise optimization type data set distillation provided by the present application includes: The optimization module 210 is used for denoising the diffusion probability model DDPM in batches based on data set distillation based on the diffusion model, and iteratively obtains the final synthetic sample set of the last step of each batch denoising based on distribution alignment loss, maximum occupancy loss and authenticity constraint loss in each step of each batch denoising. The acquisition module 220 is used for obtaining the distillation dataset according to the final synthetic sample set of the last step of all batch denoising; and the test performance of the classifier trained by the distillation dataset is the same as the test performance of the classifier trained on the original dataset. The determination module 230 is used for determining the knowledge density of the original dataset according to the total data quantity of the distillation dataset and the total data quantity of the original dataset.

[0075] In some embodiments, the optimization module 220 is specifically used for: aligning the distribution loss of each step, the maximum occupancy loss of each step, and the authenticity constraint loss of each step as the total loss of each step in each batch denoising step; obtaining the optimized random noise of each step based on the total loss of each step; iteratively obtaining the final synthesized sample set of the last step of each batch denoising based on the optimized random noise of each step.

[0076] In some embodiments, the apparatus further comprises a distribution alignment loss obtaining module for: performing the following steps for each step in each batch denoising process: performing element-wise mean and standard deviation calculation on the sampling sample set of the current step to obtain the first mean and the first standard deviation; the sampling sample set is a sample set sampled from the current step of the DDPM denoising process independent of the data set distillation process; performing element-wise normalization on all synthesized samples in the synthesized sample set of the current step using the first mean and the first standard deviation to obtain an element-wise normalized synthesized sample set; performing element-wise mean and standard deviation calculation on the element-wise normalized synthesized sample set to obtain the second mean and the second standard deviation; obtaining the first distribution alignment loss based on the second mean and the second standard deviation; performing channel-wise mean and standard deviation calculation on the sampling feature set mapped from the sampling sample set to obtain the third mean and the third standard deviation; performing channel-wise normalization on all synthesized features in the synthesized feature set mapped from the synthesized sample set using the third mean and the third standard deviation to obtain a channel-wise normalized synthesized feature set; performing channel-wise mean and standard deviation calculation on the channel-wise normalized synthesized feature set to obtain the fourth mean and the fourth standard deviation; obtaining the second distribution alignment loss based on the fourth mean and the fourth standard deviation; obtaining the weighted sum of the first distribution alignment loss and the second distribution alignment loss as the distribution alignment loss.

[0077] In some embodiments, the apparatus further comprises a maximum occupancy loss obtaining module for: determining the nearest neighbor element-wise normalized synthesized sample corresponding to each element-wise normalized synthesized sample in the element-wise normalized synthesized sample set; obtaining the first maximum occupancy loss according to each element-wise normalized synthesized sample in the element-wise normalized synthesized sample set and the nearest neighbor element-wise normalized synthesized sample corresponding to each element-wise normalized synthesized sample. determining a nearest neighbor per-channel normalized synthesized feature corresponding to each per-channel normalized synthesized feature in the per-channel normalized synthesized feature set; obtaining a second maximum occupancy loss according to each per-channel normalized synthesized feature in the per-channel normalized synthesized feature set and the nearest neighbor per-channel normalized synthesized feature corresponding to each per-channel normalized synthesized feature; taking a weighted sum of the first maximum occupancy loss and the second maximum occupancy loss as the maximum occupancy loss.

[0078] In some embodiments, the apparatus further comprises a reality constraint loss obtaining module configured to: obtain the reality constraint loss of each step based on an expected length of a standard multivariate normal distribution sample and a noise tensor of each step.

[0079] It should be noted that the above-described data set knowledge density evaluation apparatus for noise-optimized data set distillation provided by the present application can realize all the method steps realized by the method embodiment described above and achieve the same technical effects, and thus the same parts and beneficial effects in the method embodiment will not be described in detail herein.

[0080] Figure 3 is a structural schematic diagram of an electronic device provided by the present application, as Figure 3 shown, the electronic device can include a processor 310, a communications interface 320, a memory 330 and a communications bus 340, wherein the processor 310, the communications interface 320 and the memory 330 complete mutual communication through the communications bus 340. The processor 310 can invoke a logical instruction in the memory 330 to execute a data set knowledge density evaluation method based on noise-optimized data set distillation, which includes: on the basis of data set distillation based on a diffusion model, performing denoising diffusion probability model DDPM denoising in batches, in each step of each batch denoising, based on distribution alignment loss, maximum occupancy loss and reality constraint loss, performing random noise optimization, iteratively obtaining a final synthesized sample set of the last step of each batch denoising; obtaining a distilled data set according to the final synthesized sample set of the last step of all batch denoising; the test performance of a classifier trained by the distilled data set is the same as the test performance of a classifier trained on an original data set; determining the knowledge density of the original data set according to the total number of data in the distilled data set and the total number of data in the original data set.

[0081] In addition, the logical instructions in the memory 330 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0082] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor, so that the computer can execute the data set knowledge density evaluation method based on noise optimization data set distillation provided by the above-mentioned method, the method comprises: on the basis of data set distillation based on a diffusion model, denoising diffusion probability model DDPM denoising is performed batch by batch, random noise optimization is performed based on distribution alignment loss, maximum occupation loss and authenticity constraint loss in each step of each batch denoising, and finally the final synthetic sample set of the last step of each batch denoising is obtained by iteration; according to the final synthetic sample set of the last step of all batch denoising, a distilled data set is obtained; the test performance of the classifier trained by the distilled data set is the same as the test performance of the classifier trained on the original data set; according to the total number of data of the distilled data set and the total number of data of the original data set, the knowledge density of the original data set is determined.

[0083] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements a method for evaluating knowledge density of a data set based on noise-optimized data set distillation provided by the above method, the method comprising: on the basis of data set distillation based on a diffusion model, performing denoising diffusion probability model (DDPM) denoising in batches, performing random noise optimization based on distribution alignment loss, maximum occupancy loss and authenticity constraint loss in each step of each batch denoising, and iteratively obtaining a final synthetic sample set of the last step of each batch denoising; obtaining a distilled data set according to the final synthetic sample set of the last step of all batch denoising; the test performance of a classifier trained by the distilled data set is the same as the test performance of a classifier trained on the original data set; and determining the knowledge density of the original data set according to the total number of data in the distilled data set and the total number of data in the original data set.

[0084] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0085] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary general hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0086] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A dataset knowledge density evaluation method based on noise optimization dataset distillation, characterized in that, The method comprises the following steps: On the basis of data set distillation based on a diffusion model, denoising diffusion probability model (DDPM) denoising is performed batch by batch, and random noise is optimized based on a distribution alignment loss, a maximum occupancy loss and a reality constraint loss in each step of batch denoising, so that the final synthetic sample set of the last step of each batch denoising is obtained iteratively; A distillation data set is obtained according to the final synthetic sample set of the last step of all batch denoising; the test performance of a classifier trained by using the distillation data set is the same as the test performance of a classifier trained by using the original data set; The knowledge density of the original data set is determined according to the total number of data in the distillation data set and the total number of data in the original data set.

2. The method of claim 1, wherein, The method further comprises the following steps: In each step of batch denoising, a weighted sum of the distribution alignment loss of each step, the maximum occupancy loss of each step and the reality constraint loss of each step is taken as a total loss of each step; Based on the total loss of each step, the random noise optimized in each step is obtained; Based on the random noise optimized in each step, the final synthetic sample set of the last step of each batch denoising is obtained iteratively.

3. The method according to claim 1 or 2, wherein, The method further comprises the following steps: The following steps are performed for each step in each batch denoising process: Element-wise mean and standard deviation calculation is performed on a sampling sample set of the current step to obtain a first mean and a first standard deviation; the sampling sample set is a sample set sampled from the current step of a DDPM denoising process independent of the data set distillation process; All synthetic samples in a synthetic sample set of the current step are normalized element by element using the first mean and the first standard deviation to obtain an element-wise normalized synthetic sample set; Element-wise mean and standard deviation calculation is performed on the element-wise normalized synthetic sample set to obtain a second mean and a second standard deviation; Based on the second mean and the second standard deviation, a first distribution alignment loss is obtained; Channel-wise mean and standard deviation calculation is performed on a sampling feature set mapped from the sampling sample set to obtain a third mean and a third standard deviation; All synthetic features in a synthetic feature set mapped from the synthetic sample set are normalized channel by channel using the third mean and the third standard deviation to obtain a channel-wise normalized synthetic feature set; Channel-wise mean and standard deviation calculation is performed on the channel-wise normalized synthetic feature set to obtain a fourth mean and a fourth standard deviation; Based on the fourth mean and the fourth standard deviation, a second distribution alignment loss is obtained; A weighted sum of the first distribution alignment loss and the second distribution alignment loss is taken as the distribution alignment loss.

4. The noise-optimized dataset distillation-based dataset knowledge density evaluation method according to claim 3, characterized in that, The method further comprises the following steps: The nearest neighbor element-wise normalized synthetic sample corresponding to each element-wise normalized synthetic sample in the element-wise normalized synthetic sample set is determined. According to each element-wise normalized synthetic sample in the element-wise normalized synthetic sample set and the nearest neighbor element-wise normalized synthetic sample corresponding to each element-wise normalized synthetic sample, a first maximum occupancy loss is obtained; Determine the nearest neighbor channel-wise normalized synthetic feature corresponding to each channel-wise normalized synthetic feature in the channel-wise normalized synthetic feature set; According to each channel-wise normalized synthetic feature in the channel-wise normalized synthetic feature set and the nearest neighbor channel-wise normalized synthetic feature corresponding to each feature, a second maximum occupancy loss is obtained; The weighted sum of the first maximum occupancy loss and the second maximum occupancy loss is taken as the maximum occupancy loss.

5. The noise-optimized dataset distillation-based dataset knowledge density evaluation method according to claim 4, characterized in that, The method further comprises: Based on the expected mode length of the standard multivariate normal distribution sample and the noise tensor of each step, the authenticity constraint loss of each step is obtained.

6. A data set knowledge density evaluation apparatus based on noise-optimized data set distillation, characterized by It comprises: An optimization module is used to perform denoising diffusion probability model (DDPM) denoising in batches based on dataset distillation based on diffusion model, and to perform random noise optimization based on distribution alignment loss, maximum occupancy loss and authenticity constraint loss in each step of each batch denoising, and to iteratively obtain the final synthetic sample set of the last step of each batch denoising; An acquisition module is used to obtain a distillation dataset according to the final synthetic sample set of the last step of all batch denoising; the test performance of the classifier trained by the distillation dataset is the same as the test performance of the classifier trained on the original dataset; A determination module is used to determine the knowledge density of the original dataset according to the total number of data in the distillation dataset and the total number of data in the original dataset.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to realize the data set knowledge density evaluation method based on noise optimization type data set distillation according to any one of claims 1-5.

8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the data set knowledge density evaluation method based on noise optimization type data set distillation according to any one of claims 1-5.

9. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the data set knowledge density evaluation method based on noise optimization type data set distillation according to any one of claims 1-5.