Data set distillation method based on attribute adaptive distribution matching

By constructing a parameter hybrid model of random initialization and pre-trained expert models, using the distribution map prompter to adaptively match the data set attributes, the problem of high computing resources and suboptimal performance in the existing technology is solved, and efficient data set distillation effect is achieved.

CN120372279APending Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510385276.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing dataset distillation method based on distribution matching ignores dataset properties, resulting in high computing resource consumption and suboptimal performance.

Method used

By constructing a parameter hybrid model of random initialization model and pre-trained expert model, a distribution map prompt is used to predict hyperparameters, a mixed distribution mapper is generated for data distillation, and the data set attributes are adaptively matched.

Benefits of technology

Adaptive feature extraction based on distillation tasks is realized according to different data sets, improving the performance of synthetic data sets and reducing computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372279A_ABST
    Figure CN120372279A_ABST
Patent Text Reader

Abstract

The invention discloses a data set distillation method based on attribute adaptive distribution matching. The method comprises the following steps: acquiring a to-be-distilled data set; constructing a corresponding model parameter pair and a parameter hybrid model based on the random initialization model and the pre-trained expert model; predicting hyper-parameters corresponding to the attributes according to the attributes of the to-be-distilled data set by using a pre-constructed distribution mapping prompter; generating a hybrid distribution mapper by using the hyper-parameter control parameter hybrid model; and performing data distillation on the to-be-distilled data set by using a mixed distribution mapper to obtain a data set with attribute adaptive distribution matching, and since the mixed distribution mapper is obtained after generating a distribution mapping prompter based on attribute adaptive, obtaining an optimal mixed distribution mapper for different data set distillation tasks to perform feature extraction. And obtaining a data set with optimal performance and adaptive attribute distribution matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more particularly to a dataset distillation method based on attribute adaptive distribution matching. Background Art

[0002] The growth of data has brought significant challenges to model training. To address this issue, various methods for improving data processing efficiency have been proposed, and among them, dataset distillation, as an emerging technical method, has attracted the attention of many scholars. Dataset distillation aims to efficiently compress a large dataset into a smaller and more representative subset. The dataset distillation method based on bilevel optimization was first proposed. However, this method consumes a large amount of computing resources. Therefore, many different dataset distillation methods have been proposed, and among them, the dataset distillation method based on distribution matching has been widely studied. The method based on distribution matching uses a mapper to extract features from the training set and the synthetic set, and then aligns the statistics of their features in the distribution space. In the prior art, it is proposed to use the intermediate model during the training process as the mapper to extract features. However, whether using a randomly initialized model or an intermediate model as the feature extractor, the existing dataset distillation methods based on distribution matching ignore the dataset attributes (such as: training set attributes (number of images, image channels, etc.) and synthetic dataset attributes (number of images per class, etc.). Summary of the Invention

[0003] This application aims to at least solve one of the technical problems existing in the prior art. For this purpose, in the first aspect of this application, a dataset distillation method based on attribute adaptive distribution matching is proposed, including: obtaining the dataset to be distilled; constructing corresponding model parameter pairs and a parameter mixture model based on a randomly initialized model and a pre-trained expert model; using a pre-constructed distribution mapping prompt to predict the hyperparameters corresponding to the attributes according to the attributes of the dataset to be distilled; using the hyperparameters to control the parameter mixture model to generate a mixed distribution mapper, where the parameters of the mixed distribution mapper are determined based on the model parameter pairs adjusted by the hyperparameters;

[0004] Using the mixed distribution mapper to perform data distillation on the dataset to be distilled to obtain a dataset with attribute adaptive distribution matching.

[0005] Optionally, the distribution mapping prompt includes: a fully connected layer, with a total of four layers, where the input of the fully connected layer is the attribute of the dataset, and the output includes the hyperparameters of the mixed distribution mapper.

[0006] Optionally, the attributes of the dataset include the union of the attributes of the training dataset and the synthetic dataset; wherein, the attributes of the training dataset include: image resolution, number of image channels, number of classes, and number of images in each class; the attributes of the synthetic dataset include: number of images, scaling factor, relative information amount of the randomly initialized model, and relative information amount of the pre-trained expert model.

[0007] Optionally, before predicting the attributes of the dataset to be distilled using the pre-constructed distribution mapping prompt, the method further includes: constructing a meta-training set; training the pre-constructed distribution mapping prompt in a manner of inner loop and outer loop, wherein, in the inner loop, by sampling the meta-training set and calculating the sum of the intra-class loss and the inter-class loss, using the sum of the intra-class loss and the inter-class loss as the objective function, performing the dataset distillation task, and obtaining the corresponding optimal hyperparameters; in the outer loop, using the attributes of the meta-training set and the optimal hyperparameters to train the distribution mapping prompt, and obtaining the distribution mapping prompt with optimized parameters.

[0008] Optionally, use a hybrid distribution mapper to extract features from the training dataset and the synthetic dataset obtained by performing the dataset distillation task, and calculate the intra-class loss; the intra-class loss function used to calculate the intra-class loss is:

[0009]

[0010] wherein, T and S respectively represent the training dataset and the synthetic dataset, A(·, ω) represents image enhancement processing, ||·|| 2 is the L2 norm, x i represents the image i in the training dataset, represents the hybrid distribution mapper, s j represents the image j in the synthetic dataset, |·| represents the L1 norm.

[0011] Optionally, use the pre-trained expert model to extract features from the training set and the synthetic dataset obtained by performing the dataset distillation task, and calculate the inter-class loss; the inter-class loss function used to calculate the inter-class loss is:

[0012]

[0013] wherein, S represents the synthetic dataset, L CE represents the cross-entropy loss function, represents the pre-trained expert model, x j represents an image in the synthetic dataset, y j is the label of.

[0014] Optionally, the loss function used to train the distribution mapping prompt using the attributes of the training set and the optimal hyperparameters is:

[0015]

[0016] Among them, A k represents different attributes of the input meta-training tasks, and η k represents the optimized hyperparameters obtained from the k-th meta-training task, represents the distribution mapping promptor.

[0017] Optionally, constructing the corresponding model parameter pairs and parameter mixture models based on the randomly initialized model and the pre-trained expert model includes: obtaining the randomly initialized model; processing the randomly initialized model using a preset loss function to obtain the corresponding pre-trained expert model; constructing a parameter mixture model based on the randomly initialized model and the corresponding pre-trained expert model, and obtaining the corresponding mixture distribution mapper based on the parameter mixture model.

[0018] Optionally, the parameter expression of the mixture distribution mapper is:

[0019]

[0020] Among them, represents the j-th parameter in the i-th mixture distribution mapper, and represent the j-th parameters of the i-th randomly initialized model and the i-th pre-trained model respectively, and I η (λ) represents the indicator function, λ ∼ U(0,1). When λ > η, the parameters of the mixture distribution mapper are obtained from the randomly initialized model. When λ ≤ η, the parameters of the mixture distribution mapper are selected from the expert model.

[0021] Optionally, after using the mixture distribution mapper to perform data distillation on the dataset to be distilled to obtain a dataset with attribute-adaptive distribution matching, the method further includes: the distribution mapping promptor determines the hyperparameters according to the attributes of the synthetic dataset; controls the parameter mixture model to generate a mixture distribution mapper according to the hyperparameters; uses the mixture distribution mapper to extract features from the attributes of the synthetic dataset and calculates the intra-class loss; uses the pre-trained expert model to extract features from the attributes of the synthetic dataset and calculates the inter-class loss; takes the sum of the intra-class loss and the inter-class loss as the objective function to minimize the distribution difference between the synthetic dataset and the real data in the feature space.

[0022] The embodiments of the present application provide a dataset distillation method based on attribute adaptive distribution matching. Compared with the prior art, the beneficial effects are as follows: obtaining a dataset to be distilled; constructing corresponding model parameter pairs and a parameter mixture model based on a randomly initialized model and a pre-trained expert model; using a pre-constructed distribution mapping prompt to predict hyperparameters corresponding to attributes according to the attributes of the dataset to be distilled; using the hyperparameters to control the parameter mixture model to generate a mixed distribution mapper; using the mixed distribution mapper to perform data distillation on the dataset to be distilled to obtain a dataset with attribute adaptive distribution matching. In the dataset distillation stage of the present application, the distribution mapping prompt adaptively predicts the attributes of a specific dataset distillation task, and controls the parameter mixture model to generate a mixed distribution mapper through the hyperparameters corresponding to the attributes, and uses the mixed distribution mapper to perform feature extraction to obtain synthetic data with optimal performance. Description of the Drawings

[0023] To more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] Figure 1 It is a flowchart of a dataset distillation method based on attribute adaptive distribution matching provided by the embodiments of the present application. Detailed Embodiments

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0026] This specification provides the method operation steps as described in the embodiments or the flowchart, but may include more or fewer operation steps based on routine or non-creative labor. When actually executed in a system or server product, it can be executed in the order shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment).

[0027] Using a general mapper for dataset distillation tasks with different attributes results in suboptimal performance of the synthetic dataset. For example, when the number of images per class in the synthetic dataset is small, the synthetic images need to contain more general (uncertain) information, while using a pre-trained model will make the synthetic images contain more semantic information, leading to suboptimal performance. On the contrary, when the number of images per class in the synthetic dataset is large, the synthetic images need to contain more semantic (certain) information, and using a randomly initialized model cannot meet this requirement.

[0028] This application provides a dataset distillation method based on attribute-adaptive distribution matching. The attribute-adaptive distribution matching method is different from other distribution matching methods that use a general mapper. The attribute-adaptive way is to adaptively obtain a mapper for feature extraction according to the attributes of different dataset distillation tasks. First, a model pool of pairs (randomly initialized model - pre-trained model) is constructed using the original dataset. Then, a mixed distribution mapper is obtained through a parameter mixing model controlled by hyperparameters. Secondly, a distribution mapping prompt is constructed to adaptively output the hyperparameters of the parameter mixing model according to the dataset distillation tasks with different attributes. Then, a meta-learning task is constructed to train the hyperparameter predictor. Finally, in the dataset distillation stage, the distribution mapping prompt adaptively predicts the attributes of a specific dataset distillation task, and a mixed distribution mapper is generated through the parameter mixing model for feature extraction.

[0029] Specifically, referring to Figure 1 , this application proposes a dataset distillation method based on attribute-adaptive distribution matching. This method can be executed by a processor on the server side or the client side. The dataset distillation method based on attribute-adaptive distribution matching may include:

[0030] S10. Obtain the dataset to be distilled;

[0031] S20. Construct corresponding model parameter pairs and a parameter mixing model based on a randomly initialized model and a pre-trained expert model;

[0032] S30. Use hyperparameters to control the parameter mixing model to generate a mixed distribution mapper, where the parameters of the mixed distribution mapper are determined based on the model parameter pairs adjusted by the hyperparameters;

[0033] S40. Use hyperparameters to control the parameter mixing model to generate a mixed distribution mapper;

[0034] S50. Use the mixed distribution mapper to perform data distillation on the dataset to be distilled, and obtain a dataset with attribute-adaptive distribution matching.

[0035] In the above steps, by obtaining the dataset to be distilled; constructing corresponding model parameter pairs and parameter mixture models based on the randomly initialized model and the pre-trained expert model. Using the pre-constructed distribution mapping prompt to predict the hyperparameters corresponding to the attributes according to the attributes of the dataset to be distilled, the hyperparameters of the parameter mixture model are adaptively output according to the distillation tasks of datasets with different attributes, so as to use the hyperparameters to control the parameter mixture model to generate a mixture distribution mapper. The mixture distribution mapper is the best mixture distribution mapper, and using the best mixture distribution mapper for feature extraction can obtain the synthetic dataset with the optimal performance.

[0036] In an embodiment of the present application, the constructing a parameter mixture model based on the randomly initialized model and the pre-trained expert model may include:

[0037] Obtain a randomly initialized model;

[0038] Process the randomly initialized model using a preset loss function to obtain the corresponding pre-trained expert model;

[0039] Construct a parameter mixture model based on the randomly initialized model and the corresponding pre-trained expert model.

[0040] Among them, for dataset distillation based on distribution matching, a mixture distribution mapper is used to extract features from the training dataset and the synthetic dataset. The randomly initialized model for distribution mapping can provide more general information, and the pre-trained expert model can provide more definite semantic information, and the expert model parameters obtained after training different randomly initialized models vary greatly. Therefore, in order to obtain the mixture distribution mapper, the present application first constructs a paired model pool. For one randomly initialized model θ I,i , a corresponding pre-trained expert model θ E,i is obtained according to the preset formula (1) for training.

[0041]

[0042] Where f θ is a model with parameter θ, x i ∈ T, T is the training dataset, including m images. Then repeat the same training process K times to construct a model pool containing K pairs (randomly initialized model - expert model).

[0043] For a given training dataset, train the model based on the above cross-entropy loss function to construct a paired model (randomly initialized model θ I,i - pre-trained model θ E,i ) library P = {(θ I,i , θ E,i ), i ∈ {1, N}.

[0044] In an embodiment of the present application, the expression of the hybrid distribution mapper parameters is as follows:

[0045]

[0046] Wherein, represents the j-th parameter in the i-th hybrid distribution mapper, and respectively represent the j-th parameters of the i-th randomly initialized model and the i-th pre-trained model, I η (λ) represents the indicator function, λ~U(0,1). When λ>η, the parameters of the hybrid distribution mapper are obtained from the randomly initialized model. When λ≤η, the parameters of the hybrid distribution mapper are selected from the expert model.

[0047] In an embodiment of the present application, the distribution mapping prompt can include:

[0048] A fully connected layer, wherein the input of the fully connected layer is the attribute of the data set, and the output includes the hyperparameters of the hybrid distribution mapper.

[0049] It should be noted that there are four fully connected layers in total to improve the accuracy of feature extraction.

[0050] In this step, the processor adaptively predicts prompts for the data set distillation task of different attributes by constructing a distribution mapping prompt. The prediction prompt is used to prompt the parameter mixing model to generate a hybrid distribution mapper. The input of the distribution mapping prompt is the attribute of the data set distillation task, and the output is the prompt for guiding the parameter mixing model, that is, the value of the hyperparameter.

[0051] In an embodiment of the present application, the attributes of the data set may include the union of the attributes of the training data set and the synthetic data set;

[0052] Among them, the attributes of the training data set include: image resolution, number of image channels, number of classes, and number of images of each class; the attributes of the synthetic data set include: number of images, scaling factor, relative information amount of the randomly initialized model, and relative information amount of the pre-trained expert model.

[0053] Specifically, the present application defines the attribute A k of the data set distillation task as the attribute A k1 of the training data set k2 and the attribute A k1 of the synthetic data set. Exemplarily, the attribute A k1 of the training data set C includes the image resolution R, the number of image channels C, the number of classes N IPC and the number of images of each class N k2Include the IPC of various image quantities, the scaling factor S, and the relative information quantity I I and I E . Among them, the relative information quantity I I and I E respectively represent the information quantity of the training data set relative to the randomly initialized model and the pre-trained model. Obtained according to formula (3) and formula (4).

[0054]

[0055] where p j (x i ) is the predicted probability of the randomly initialized model or the pre-trained parameter hybrid model for the image x i , and m represents the number of images in the training data set.

[0056] In an embodiment of the present application, step S30 may include the following execution process:

[0057] S301. Construct a meta-training set;

[0058] It should be particularly noted that the meta-training set corresponds to meta-learning. The meta-training set contains K meta-tasks, and the meta-task is the data set distillation task. The data of the data set distillation task consists of the training data set and the synthetic data set.

[0059] S302. Train the pre-constructed distribution mapping prompt using the inner loop and the outer loop. Among them, in the inner loop, by sampling the meta-training set and calculating the sum of the intra-class loss and the inter-class loss, using the sum of the intra-class loss and the inter-class loss as the objective function, execute the data set distillation task to obtain the corresponding optimal hyperparameters;

[0060] S303. In the outer loop, use the attributes of the meta-training set and the optimal hyperparameters to train the distribution mapping prompt to obtain the distribution mapping prompt with optimized parameters.

[0061] Based on the distribution mapping prompt constructed in step S30, it can adaptively guide the parameter hybrid model to generate a hybrid distribution mapper. To train this distribution mapping prompt, the processor executes a distribution mapping prompt learning method based on meta-learning. The processor first constructs a meta-training set D of size K k train . Each meta-task in the meta-training set is a data set distillation task with different attribute settings. In the inner loop, by setting the value of the prompt η, different hybrid distribution mappers are obtained. The value of the hyperparameter η is a floating point number after the decimal point between 0 and 1. The processor obtains synthetic data sets with different performances by traversing the values of the hyperparameter η, and outputs the prompt η corresponding to the synthetic data set with the best performance.

[0062] For the optimization method of dataset distillation in the inner loop, this application uses intra-class and inter-class losses. To map the synthetic dataset and the training dataset to a suitable distribution space for alignment, the intra-class loss uses a mixture distribution mapper for feature extraction and optimizes the above formula (5). The task of the inner loop is to determine different mixture distribution mappers according to different preset hyperparameters, use this mixture distribution mapper for dataset distillation, and obtain the optimal mixture distribution mapper and hyperparameters under this attribute by comparing the performance of the distilled dataset.

[0063] Specifically, use the mixture distribution mapper to extract features from the training dataset and the synthetic dataset obtained by performing the dataset distillation task; the intra-class loss function used to calculate the intra-class loss is:

[0064]

[0065] where T and S represent the training dataset and the synthetic dataset respectively, A(·, ω) represents image enhancement processing, ||·|| 2 is the L2 norm, xi represents the image i in the training dataset, represents the mixture distribution mapper, s j represents the image j in the synthetic dataset, and |·| represents the L1 norm.

[0066] For the inter-class loss, it should provide a more accurate classification boundary. Therefore, a pre-trained expert model is used to extract features from the training dataset and the synthetic dataset obtained by performing the dataset distillation task, and the inter-class loss is calculated. The inter-class loss function is formula (7).

[0067]

[0068] where S represents the synthetic dataset, L CE represents the cross-entropy loss function, represents the pre-trained expert model, x j represents an image in the synthetic dataset, and y j is the label of.

[0069] For the outer loop, the outer loop optimizes the distribution mapping prompt according to formula (7), trains the mixture distribution mapper by sampling the meta-training set, and finally obtains a distribution mapping prompt with optimized parameters. This prompt can adaptively generate prompts according to the attributes of different dataset distillation tasks, and guide the parameter mixture model to generate the best mixture distribution mapper.

[0070] In an embodiment of this application, the loss function used to train the distribution mapping prompt with the attributes of the training dataset and the optimal hyperparameters is:

[0071]

[0072] Among them, A k is the different attributes of the input meta-training tasks, and η k represents the optimized hyperparameters obtained from the k-th meta-training task, represents the distribution mapping promptor.

[0073] It should be noted that for a dataset distillation task, its training dataset is and its corresponding attribute is A k . The attribute is input into the mixed distribution mapper to obtain the prompt for this task, and then the prompt is input into the parameter mixing model to obtain the mixed distribution mapper. Finally, the mixed distribution mapper and its corresponding pre-trained expert model are used to optimize the synthetic dataset according to formula (8). Finally, the synthetic dataset

[0074] L total = L inter + L inner (8)

[0075] In an embodiment of the present application, after using the mixed distribution mapper to perform data distillation on the dataset to be distilled to obtain a dataset with attribute adaptive distribution matching, the method further includes:

[0076] The distribution mapping promptor determines the hyperparameters according to the attributes of the synthetic dataset;

[0077] Control the parameter mixing model to generate a mixed distribution mapper according to the hyperparameters;

[0078] Use the mixed distribution mapper to extract features from the attributes of the synthetic dataset and calculate the intra-class loss;

[0079] Use the pre-trained expert model to extract features from the attributes of the synthetic dataset and calculate the inter-class loss;

[0080] Based on the intra-class loss and the inter-class loss as the objective function, train the mixed distribution mapper to minimize the distribution difference between the synthetic dataset and the real data in the feature space.

[0081] For the synthetic dataset obtained in the dataset distillation stage, the present application uses the classification accuracy provided by formula (9) to verify its performance.

[0082] The cross-entropy loss function is used when training the classification network with the synthetic dataset:

[0083]

[0084] where si represents the image in the synthetic dataset and yi is the corresponding label, Represents a classification network trained using a synthetic dataset.

[0085] When validating the synthetic dataset obtained in the dataset distillation stage, the formula η k = argmax η P k,η is used for validation, and if the validation meets the conditions, the hyperparameter η is retained.

[0086] where P k,η is the k-th meta-task, and the test performance of the synthetic dataset when the hyperparameter is η.

[0087] The attribute-adaptive hybrid distribution mapper maps the training dataset and the synthetic dataset to be optimized into an optimal distribution space, which is both semantic and general. Performing distribution matching in this space can effectively improve the test performance of the synthetic dataset. Finally, accuracies of 49.0 (IPC = 1), 72.1 (IPC = 10), and 78.6 (IPC = 50) are obtained on the Cifar10 dataset distillation task, and accuracies of 29.5 (IPC = 1), 51.5 (IPC = 10), and 56.0 (IPC = 50) are obtained on the Cifar100 dataset distillation task. On subsets of the higher-resolution (128*128) dataset ImageNet1K, accuracies of 63.7 (IPC = 1) and 82.8 (IPC = 10) are obtained in ImageNettle, and accuracies of 31.9 (IPC = 1) and 58.4 (IPC = 10) are obtained in ImageWoof. Improvements to varying degrees are achieved compared to the state-of-the-art algorithms.

[0088] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or also elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0089] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0090] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A dataset distillation method based on attribute adaptive distribution matching, characterized in that Including: Obtain the dataset to be distilled; Construct corresponding model parameter pairs and parameter mixture models based on the randomly initialized model and the pre-trained expert model; Use the pre-constructed distribution mapping prompt to predict the hyperparameters corresponding to the attributes according to the attributes of the dataset to be distilled; Use the hyperparameters to control the parameter mixture model to generate a mixture distribution mapper, where the parameters of the mixture distribution mapper are determined based on the model parameter pairs adjusted by the hyperparameters; Use the mixture distribution mapper to perform data distillation on the dataset to be distilled to obtain a dataset with attribute-adaptive distribution matching.

2. The dataset distillation method based on attribute adaptive distribution matching according to claim 1, wherein The distribution mapping prompt includes: A fully connected layer, where the input of the fully connected layer is the attribute of the dataset, and the output includes the hyperparameters of the mixture distribution mapper.

3. The dataset distillation method based on attribute adaptive distribution matching according to claim 2, characterized in that, The attributes of the dataset include the union of the attributes of the training dataset and the synthetic dataset; Among them, the attributes of the training dataset include: Image resolution, number of image channels, number of classes, and number of images in each class; The attributes of the synthetic dataset include: Number of images, scaling factor, relative information amount of the randomly initialized model, and relative information amount of the pre-trained expert model.

4. The dataset distillation method based on attribute adaptive distribution matching according to claim 1, wherein, Before using the pre-constructed distribution mapping prompt to predict the attributes of the dataset to be distilled, the method further includes: Construct a meta-training set; Train the pre-constructed distribution mapping prompt in the way of inner loop and outer loop; Among them, in the inner loop, by sampling the meta-training set and calculating the sum of the intra-class loss and the inter-class loss, using the sum of the intra-class loss and the inter-class loss as the objective function, perform the dataset distillation task to obtain the corresponding optimal hyperparameters; In the outer loop, use the attributes of the meta-training set and the optimal hyperparameters to train the distribution mapping prompt to obtain the distribution mapping prompt with optimized parameters.

5. The method for dataset distillation based on attribute adaptive distribution matching according to claim 4, wherein Use the mixture distribution mapper to extract features from the training dataset and the synthetic dataset obtained by performing the dataset distillation task, and calculate the intra-class loss; The intra-class loss function used to calculate the intra-class loss is: Among them, T and S represent the training dataset and the synthetic dataset respectively, represents image enhancement processing, is the L2 norm, x i represents the image i in the training dataset, represents the mixture distribution mapper, s j represents the image j in the synthetic dataset, and |·| represents the L1 norm.

6. The method for dataset distillation based on attribute adaptive distribution matching according to claim 4, wherein Use the pre-trained expert model to extract features from the training set and the synthetic dataset obtained by performing the dataset distillation task, and calculate the inter-class loss; The inter-class loss function used to calculate the inter-class loss is: Among them, S represents the synthetic dataset, and L CE represents the cross-entropy loss function, represents the pre-trained expert model, and x j represents an image in the synthetic dataset, and y j is the label of 7. The dataset distillation method based on attribute adaptive distribution matching according to claim 4, wherein The loss function used to train the distribution mapping prompt using the attributes of the meta-training set and the optimal hyperparameters is: Among them, A k is the different attributes of the input meta-training task, and η k represents the optimized hyperparameters obtained from the k-th meta-training task, represents the distribution mapping promptor.

8. The dataset distillation method based on attribute adaptive distribution matching according to claim 1, characterized in that The constructing of the corresponding model parameter pairs and parameter mixture models based on the randomly initialized model and the pre-trained expert model includes: Obtain the randomly initialized model; Process the randomly initialized model using a preset loss function to obtain the corresponding pre-trained expert model; Construct a parameter mixture model and the corresponding model parameter pairs based on the randomly initialized model and the corresponding pre-trained expert model.

9. The dataset distillation method based on attribute adaptive distribution matching according to claim 1, wherein, The parameter expression of the mixture distribution mapper is: Among them, represents the j-th parameter in the i-th mixture distribution mapper, and represent the j-th parameters of the i-th randomly initialized model and the i-th pre-trained model respectively, η (λ) represents the indicator function, λ ∼ U(0, 1). When λ > η, the parameters of the mixture distribution mapper are obtained from the randomly initialized model. When λ ≤ η, the parameters of the mixture distribution mapper are selected from the expert model, where η represents the hyperparameter.

10. The method for dataset distillation based on attribute adaptive distribution matching according to claim 1, wherein After using the mixture distribution mapper to perform data distillation on the dataset to be distilled to obtain a dataset with attribute-adaptive distribution matching, the method further includes: The distribution mapping prompt determines the hyperparameters according to the attributes of the synthetic dataset; Control the parameter mixture model to generate a mixture distribution mapper according to the hyperparameters; Use the mixture distribution mapper to extract features from the attributes of the synthetic dataset and calculate the intra-class loss; Extract features of the attributes of the synthetic dataset using a pre-trained expert model, and calculate the inter-class loss; Use the sum of the intra-class loss and the inter-class loss as the objective function to train the synthetic dataset, minimizing the distribution difference between the synthetic dataset and the real data in the feature space.