Simulation data generation method, system and equipment based on diffusion model and medium
Through diffusion model and binary convolutional neural network, real historical sample data is processed to generate high-quality simulation data that is highly similar to the real data, solving the problems of unstable simulation data generation and incomplete distribution law simulation in the existing technology.
Patent Information
- Application Number
- CN202510350701.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing simulation data generation methods are difficult to simulate the overall distribution law of real historical sample data. The generated data is far from the real data, and the generative adversarial network is unstable during the training process, making it difficult to play a role in practical applications.
The diffusion model is used to divide and format the real historical sample data, and combine it with a binary convolutional neural network to directly fit the overall joint probability distribution of the data to generate high-quality simulation data.
The generated simulation data is highly similar to the real historical sample data, with high clarity and fewer deviations, which solves the problem of data acquisition and improves the stability and consistency of the simulation data.
Smart Images

Figure CN120278012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and particularly relates to a simulation data generation method, system, device and medium based on a diffusion model. Background Art
[0002] Simulation data refers to data generated by a computer rather than captured from real activities. This is a type of data created algorithmically that replicates the statistical part of real data. With the decline in storage costs, the popularity of algorithms such as generative adversarial networks, and the development in the field of computing power, the demand for simulation data has become increasingly evident.
[0003] Simulation data is widely used in multiple fields, including providing data for the verification and testing of new products and models, and clinical scientific experiments; it can also be used in agile development and operation to accelerate the cycle of testing and quality assurance; financial institutions can use simulation data to test and train fraud detection systems, effectively protecting sensitive information from being leaked and preventing losses caused by data breaches.
[0004] In the prior art, generative adversarial networks (GANs) are commonly used to generate simulation data. However, during the training process of generative adversarial networks, mode collapse is likely to occur, that is, it is difficult for the adversarial process between the generator and the discriminator to reach the Nash equilibrium, and the generated data may deviate from the real distribution. Especially when dealing with multi-modal data, the distribution deviation will increase significantly, resulting in a large gap between the generated simulation data and the real historical sample data, and it cannot be put into practical applications. In addition to generative adversarial networks, other simulation data generation methods are difficult to simulate the overall distribution law of real historical sample data and can only simulate it feature by feature, thus losing the internal connection between the various features of the data, resulting in serious limitations in the application of existing simulation data and failing to play a safe and efficient role in model training.
[0005] The patent application document with the publication number CN117574263A discloses a sample data generation method, device, equipment and readable storage medium. This method constructs a multivariate distribution sampler for real data through the Copula function, and makes the generated data as close to the real data as possible by learning to find the neighborhood of similar data. The generated data can better balance the two extreme situations of overfitting data and random noise data, effectively improving the quality of the data; however, this method can only simulate the distribution law of data feature by feature and attribute, and cannot capture the internal connection between multi-dimensional data, that is, it cannot simulate the distribution law of data as a whole.
[0006] The patent application document with the publication number CN118779691A discloses a sample generation method, a model training method, a sorting method, a device and a device based on a large model. The method determines a plurality of metrics from a plurality of initial metrics included in a metric database in response to a sample generation request, where the sample generation request includes an example sample; uses the question example included in the example sample as the basic corpus, and generates a plurality of alternative questions based on the plurality of metrics; recalls a plurality of alternative metrics corresponding to each alternative question from the plurality of initial metrics; and uses the plurality of example samples as the basic corpus, and generates a plurality of target samples based on the plurality of alternative questions and the plurality of alternative metrics corresponding to each alternative question. However, this method only extracts the features of two types of data respectively and performs confidence matching, and does not deeply mine the distribution law of the data. The unsupervised data it relies on also needs to be provided by a real physical scenario, which greatly limits the large-scale and low-cost acquisition of target data training samples, and there are also serious problems of information security leakage. Summary of the Invention
[0007] In order to overcome the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide a simulation data generation method, system, device and medium based on a diffusion model. The method divides real historical sample data, and processes and transforms the format of the divided real historical sample data to facilitate input into subsequent models; uses the score-based generative diffusion model SMLD to directly fit the overall joint probability distribution of the data, and then combines it with a trained binary classification convolutional neural network model to improve the consistency of the distribution laws of the simulation data and the real historical data, thereby generating high-quality simulation data. The simulation data generated by the present invention is highly similar to the real historical sample data, has high clarity and less deviation, and solves the problem of data acquisition faced in the fields of artificial intelligence and machine learning.
[0008] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0009] A simulation data generation method based on a diffusion model, comprising the following steps:
[0010] Step 1: Divide the real historical sample data, process and transform the format of the divided real historical sample data to obtain a second data Y;
[0011] Step 2: Generate a first simulation data F1Y of the same order of magnitude as the second data Y in Step 1; after respectively identifying the first simulation data F1Y and the second data Y, train a binary classification convolutional neural network by constructing a binary classification data set;
[0012] Step 3: Generate second simulation data F2Y with the same order of magnitude as the second data Y in Step 1; Use the trained binary convolutional neural network in Step 2 to screen out the third simulation data SY that meets the threshold and the required quantity from the second simulation data F2Y, and convert it into fourth simulation data SX.
[0013] Further, the specific steps of Step 1 are as follows:
[0014] Step 1.1: Divide the real historical sample data into ordered data and unordered data according to business meaning, and perform one-hot encoding on the unordered data;
[0015] Step 1.2: Concatenate the ordered data and the encoded unordered data in Step 1.1 column by column to generate the first data X;
[0016] Step 1.3: Convert each column in the first data X in Step 1.2 into the second data Y through Formula 1; The formula 1 is:
[0017]
[0018] Where: X is the column to be processed in the first data; X min is the minimum value of this column to be processed in the first data; X max is the maximum value of this column to be processed in the first data; eps = 10 -10 ; Y is the second data.
[0019] Further, the specific steps of Step 2 are as follows:
[0020] Step 2.1: Use the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to generate first simulation data F1Y with the same order of magnitude as the second data Y in Step 1;
[0021] Step 2.2: Mark the data in the second data Y in Step 1 as true, that is, 1, mark the data in the first simulation data F1Y in Step 2.1 as false, that is, 0, and construct a binary classification dataset from the marked data in the second data Y and the marked data in the first simulation data F1Y;
[0022] Step 2.3: Use the binary classification dataset in Step 2.2 to train a binary convolutional neural network.
[0023] Further, the specific steps of Step 3 are as follows:
[0024] Step 3.1: Use the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to generate the second simulation data F2Y with the same order of magnitude as the second data Y in Step 1;
[0025] Step 3.2: Use the trained binary classification convolutional neural network in Step 2 to score and sort the second simulation data F2Y in Step 3.1, and select the second simulation data F2Y whose scores meet the threshold;
[0026] Step 3.3: Iterate the operation in Step 3.2 until the sample quantity reaches the required amount and then stop, and record all the second simulation data F2Y whose scores meet the threshold as the third simulation data SY;
[0027] Step 3.4: Convert each column in the third simulation data SY in Step 3.3 into the fourth simulation data SX through Formula 2; the Formula 2 is:
[0028]
[0029] where: Y represents the column to be processed in the third simulation data; X min represents the minimum value of this column to be processed in the first data; X max represents the maximum value of this column to be processed in the first data; e represents the natural constant.
[0030] Further, the threshold in Step 3.2 is set to the top 1 - 5% of the score ranking of the second simulation data F2Y.
[0031] Further, Step 3.4 further includes: for the columns of the one-hot encoding corresponding to the unordered class data, those greater than 0.5 are assigned 1, those less than 0.5 are assigned 0, and the simulation data equal to 0.5 are deleted; then according to the one-hot encoding rule, the columns of the one-hot encoding corresponding to the unordered class data are converted into the form corresponding to the real historical sample data.
[0032] A simulation data generation system based on a diffusion model, comprising:
[0033] A data processing module: divide the real historical sample data, process and convert the format of the divided real historical sample data to obtain the second data Y;
[0034] An identification module: generate the first simulation data F1Y with the same order of magnitude as the second data Y; after respectively identifying the first simulation data F1Y and the second data Y, train a binary classification convolutional neural network by constructing a binary classification data set;
[0035] Simulation data generation module: generates second simulation data F2Y of the same order of magnitude as the second data Y; uses a trained binary classification convolutional neural network to filter out third simulation data SY that meets the threshold and the required quantity from the second simulation data F2Y, and converts it into fourth simulation data SX.
[0036] A device for generating simulation data based on a diffusion model, comprising:
[0037] Memory: used for storing a computer program to implement the above-mentioned simulation data generation method based on the diffusion model;
[0038] Processor: used to implement the above-mentioned method for generating simulation data based on the diffusion model when executing the computer program.
[0039] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a method for generating simulation data based on a diffusion model are implemented.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] 1. The present invention uses the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to directly fit the overall joint probability distribution of the data, and obtains simulation data that is closer to the distribution law of the real historical sample data, so that the simulation data shows higher clarity and less deviation; the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) is more stable than the existing generative adversarial network, does not require adversarial training, and reduces the instability of generated simulation data.
[0042] 2. Based on simulation data and real historical data, the present invention constructs a binary classification convolutional neural network, which can screen out simulation data that meets the authenticity requirements, effectively improve the realism of simulation data, and improve the consistency of the distribution laws of simulation data and real historical data.
[0043] In summary, the present invention divides the real historical sample data, processes and transforms the format of the divided real historical sample data to facilitate input into subsequent models; uses the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to directly fit the overall joint probability distribution of the data, and then combines the trained binary classification convolutional neural network model to improve the consistency of the distribution laws of the simulation data and the real historical data, thereby generating high-quality simulation data; the simulation data generated by the present invention is highly similar to the real historical sample data, has high clarity and few deviations, and solves the problem of data acquisition faced in the fields of artificial intelligence and machine learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the simulation data generation method based on the diffusion model of the present invention.
[0045] Figure 2 It is the comparison effect between the handwritten digit image set and the simulation data generated by the method of the present invention Figure 1 ; wherein, Figure 2 (a) in is the handwritten digit image of 0; Figure 2 (b) in is the effect diagram of the simulation data of 0.
[0046] Figure 3 It is the comparison effect between the handwritten digit image set and the simulation data generated by the method of the present invention Figure 2 ; wherein, Figure 3 (a) in is the handwritten digit image of 1; Figure 3 (b) in is the effect diagram of the simulation data of 1.
[0047] Figure 4 It is the comparison effect between the handwritten digit image set and the simulation data generated by the method of the present invention Figure 3 ; wherein, Figure 4 (a) in is the handwritten digit image of 2; Figure 4 (b) in is the effect diagram of the simulation data of 2.
[0048] Figure 5 It is the comparison effect between the handwritten digit image set and the simulation data generated by the method of the present invention Figure 4 ; wherein, Figure 5 (a) in is the handwritten digit image of 3; Figure 5 (b) in is the effect diagram of the simulation data of 3.
[0049] Figure 6 It is the comparison effect between the handwritten digit image set and the simulation data generated by the method of the present invention Figure 5 ; wherein, Figure 6In (a) is a handwritten digit image of 4; Figure 6 In (b) is the simulation data effect diagram of 4.
[0050] Figure 7 Is the comparison effect of the handwritten digit image set and the simulation data generated by the method of the present invention Figure 6 ; wherein, Figure 7 In (a) is a handwritten digit image of 5; Figure 7 In (b) is the simulation data effect diagram of 5.
[0051] Figure 8 Is the comparison effect of the handwritten digit image set and the simulation data generated by the method of the present invention Figure 7 ; wherein, Figure 8 In (a) is a handwritten digit image of 6; Figure 8 In (b) is the simulation data effect diagram of 6.
[0052] Figure 9 Is the comparison effect of the handwritten digit image set and the simulation data generated by the method of the present invention Figure 8 ; wherein, Figure 9 In (a) is a handwritten digit image of 7; Figure 9 In (b) is the simulation data effect diagram of 7.
[0053] Figure 10 Is the comparison effect of the handwritten digit image set and the simulation data generated by the method of the present invention Figure 9 ; wherein, Figure 10 In (a) is a handwritten digit image of 8; Figure 10 In (b) is the simulation data effect diagram of 8.
[0054] Figure 11 Is the comparison effect of the handwritten digit image set and the simulation data generated by the method of the present invention Figure 10 ; wherein, Figure 11 In (a) is a handwritten digit image of 9; Figure 11 In (b) is the simulation data effect diagram of 9. Specific embodiments
[0055] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0056] Real historical sample data is attached to real existing objects. When the research object has a complex structure and is not easy to operate, collecting real historical sample data requires a large amount of manpower and material resources; in addition, issues such as data compliance and data security involved in real historical sample data make it difficult to achieve interaction and sharing between different enterprises; while simulation data is no longer restricted by the real physical world and is no longer limited by time, and only a small amount of real data is needed to simulate the probability distribution law of real data.
[0057] See Figure 1 , a simulation data generation method based on a diffusion model, comprising the following steps:
[0058] Step 1: Divide the real historical sample data, process and transform the format of the divided real historical sample data to obtain the second data Y;
[0059] Further, the specific content of the said Step 1 is:
[0060] Step 1.1: Divide the real historical sample data into ordered data and unordered data according to the business meaning, and perform one-hot encoding on the unordered data; among them, the ordered data refers to the data with a certain order relationship between data elements, such as the age of a user, the number of employees in an enterprise, etc., and there is a size difference between these data; the unordered data refers to the data with no clear order relationship between data elements, such as a person's nationality, weather conditions, etc., and it is impossible to compare the sizes between these data; therefore, corresponding one-hot encoding is required for the unordered data.
[0061] In this embodiment, the data in the financial field is often used as the real historical sample data. See Table 1:
[0062] Table 1 Example of real historical sample data
[0063]
[0064] Based on Table 1, Table 2 is obtained. Table 2 contains the features of three dimensions: gender, household registration authentication, and current status. The following demonstrates the one-hot encoding process of unordered data through Table 2:
[0065] Table 2: Summary of borrowing information
[0066] Borrowing ID Gender Household Registration Authentication Current Status Sample 1 Male Failed Authentication Not Repaid Sample 2 Female Failed Authentication Repaid Sample 3 Female Successful Authentication Not Repaid Sample 4 Male Successful Authentication Repaid
[0067] Convert Table 2 into sample-feature data represented numerically, as shown in Table 3 below:
[0068] Table 3 Sample-feature data table
[0069] Serial Number Gender Household Registration Authentication Current Status Sample 1 0 0 0 Sample 2 1 0 1 Sample 3 1 1 0 Sample 4 0 1 1
[0070] For features with two states of 0 and 1, such as gender, household registration authentication, and current status in Table 3, 2 state bits can be used to represent them. 0 is represented as [1,0], and 1 is represented as [0,1]. Thus, Table 3 can be converted into Table 4 below:
[0071] Table 4: Sample-feature one-hot encoding table
[0072] Serial Number Gender Department Major Sample 1 10 10 10 Sample 2 01 10 01 Sample 3 01 01 10 Sample 4 10 01 01
[0073] The feature vectors of the four encoded samples can be visually obtained from Table 4:
[0074] Sample 1: [1, 0, 1, 0, 1, 0]; Sample 2: [0, 1, 1, 0, 0, 1]; Sample 3: [0, 1, 0, 1, 1, 0]; Sample 4: [1, 0, 0, 1, 0, 1].
[0075] Step 1.2: Concatenate the ordered class data and the encoded unordered class data in step 1.1 column by column to generate the first data X;
[0076] Step 1.3: Convert each column in the first data X in step 1.2 into the second data Y through Formula 1; the Formula 1 is:
[0077]
[0078] where: X is the column to be processed in the first data; X min is the minimum value of this column to be processed in the first data; X max is the maximum value of this column to be processed in the first data; eps = 10 -10 ; Y is the second data.
[0079] Through the above Step 1, the real historical sample data is partitioned and formatted to obtain a new historical sample data, i.e., the second data Y, which is convenient for subsequent input into the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics).
[0080] Step 2: Generate the first simulated data F1Y with the same order of magnitude as the second data Y in Step 1; after respectively identifying the first simulated data F1Y and the second data Y, construct a binary classification dataset to train a binary classification convolutional neural network;
[0081] Furthermore, the specific steps of Step 2 are as follows:
[0082] Step 2.1: Use the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to generate the first simulated data F1Y with the same order of magnitude as the second data Y in Step 1; the score-based generative diffusion model SMLD in this embodiment is mainly used to generate simulated data, and a U-Net model is used to predict noise in the noise prediction part. The network depth is 32 layers, and a convolutional kernel of size 3 is used throughout the convolutional layer;
[0083] The score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) specifically includes:
[0084] A variational autoencoder component, including an encoder and a decoder; the encoder compresses an image of 28×28 pixels into a 2×2 model in the latent space; the decoder restores the model from the latent space to a full-size 28×28 pixel image;
[0085] A forward diffusion component that gradually adds Gaussian noise to the image until only random noise remains;
[0086] Reverse diffusion, gradually denoising from the random noise until simulated data is generated;
[0087] In the variational autoencoder component, the encoder consists of 4 convolutional layers (Conv2d) with a convolutional kernel size of 3×3. The first convolutional layer has a stride of 1, and the last three have a stride of 2; a group normalization layer (GroupNormalization) is connected after each convolutional layer; in the variational autoencoder component, the decoder consists of 4 transposed convolutional layers (ConvTranspose2d) with a convolutional kernel size of 3×3. The stride of the first three transposed convolutional layers is 2, and the stride of the last convolutional layer is 1; a group normalization layer (Group Normalization) is connected after each convolutional layer.
[0088] The present invention uses the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to directly fit the overall joint probability distribution of the data, obtaining simulated data that more closely approximates the distribution law of real historical sample data, thereby making the simulated data exhibit higher clarity and less deviation; the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) is more stable than existing generative adversarial networks, does not require adversarial training, and reduces the instability of generating simulated data.
[0089] Step 2.2: Mark the data in the second data Y in Step 1 as true, i.e., 1, and mark the data in the first simulated data F1Y in Step 2.1 as false, i.e., 0, and construct the marked data in the second data Y and the marked data in the first simulated data F1Y into a binary classification dataset;
[0090] Step 2.3: Use the binary classification dataset in Step 2.2 to train a binary classification convolutional neural network; in this embodiment, the binary classification convolutional neural network is designed in the residual network (ResNet) mode, with 16 layers, and a convolutional kernel of size 3 is used in the convolutional layer, and a residual (resnet) is connected between every two layers;
[0091] The binary classification convolutional neural network includes a single convolutional layer, three residual blocks, and a fully connected layer; the input image of the single convolutional layer is single-channel, and the output is 64 channels; each residual block includes two sub-modules. The first sub-module consists of a convolutional layer (Conv2d), a normalization layer (BatchNorm2d), a rectified linear unit (ReLU), followed by another convolutional layer (Conv2d) and a normalization layer (BatchNorm2d), and finally, after connecting the initial value through a residual connection, it is connected to a rectified linear unit (ReLU); the sizes of both convolutional layers are 3×3, the stride is 2, and other parameters are default values; the second sub-module consists of a convolutional layer (Conv2d), a normalization layer (BatchNorm2d), a rectified linear unit (ReLU), followed by another convolutional layer (Conv2d) and a normalization layer (BatchNorm2d), and finally, after connecting the initial value through a residual connection, it is connected to a rectified linear unit (ReLU); the sizes of both convolutional layers are 3×3, the stride is 1, and other parameters are default values; the input of the first residual block is 64 channels, and the output is 64 channels; the input of the second residual block is 64 channels, and the output is 128 channels; the input of the third residual block is 128 channels, and the output is 256 channels; the fully connected layer includes a linear connection layer and an activation layer (Sigmoid), and the output result is converted into a number between 0 and 1, representing the probability that the input data is the real historical sample data.
[0092] The binary classification convolutional neural network is mainly used to select the part that meets the authenticity requirements from the simulation data, ensure that the simulation data meets the required number of samples, effectively improve the fidelity of the simulation data, and improve the consistency of the distribution laws of the simulation data and the real historical data.
[0093] Step 3: Generate second simulation data F2Y with the same order of magnitude as the second data Y in Step 1; use the trained binary classification convolutional neural network in Step 2 to screen out the third simulation data SY that meets the threshold and the required number from the second simulation data F2Y, and convert it into fourth simulation data SX;
[0094] Furthermore, the specific content of Step 3 is as follows:
[0095] Step 3.1: Use the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to generate second simulation data F2Y with the same order of magnitude as the second data Y in Step 1;
[0096] Step 3.2: Use the binary classification convolutional neural network trained in Step 2 to score and sort the second simulation data F2Y in Step 3.1, and select the second simulation data F2Y whose scores meet the threshold;
[0097] In Step 3.2, the threshold is set to the top 1-5% of the score ranking of the second simulation data F2Y. In this embodiment, the threshold is set to the top 5% of the score ranking of the second simulation data F2Y; by setting the threshold within this range, the finally generated fourth simulation data SX has a similar data distribution law to the real historical sample data;
[0098] Step 3.3: Iterate the operation in Step 3.2 until the sample quantity reaches the required quantity and stop, and record all the second simulation data F2Y whose scores meet the threshold as the third simulation data SY;
[0099] Step 3.4: Convert each column in the third simulation data SY in Step 3.3 into the fourth simulation data SX through Formula 2; the Formula 2 is:
[0100]
[0101] Where: Y represents the column to be processed in the third simulation data; X min represents the minimum value of this column to be processed in the first data; X max represents the maximum value of this column to be processed in the first data; e represents the natural constant.
[0102] Further, Step 3.4 further includes: for the columns of one-hot encoding corresponding to unordered data, assign a value of 1 if it is greater than 0.5, assign a value of 0 if it is less than 0.5, and delete the simulation data equal to 0.5; then, according to the one-hot encoding rule, convert the columns of one-hot encoding corresponding to unordered data into the form corresponding to the real historical sample data.
[0103] In this embodiment, since 0.5 is in the middle of 0 and 1 and it is impossible to determine whether to assign a value of 0 or 1, so choose to delete the simulation data of 0.5 to avoid affecting the overall distribution of the simulation data and improve the quality and accuracy of the simulation data; in addition, the following specific example is used to illustrate how to convert the columns of one-hot encoding corresponding to unordered data into the form corresponding to the real historical sample data:
[0104] Based on Table 2, if the assigned value here is [1,0,1,0,1,0], according to the one-hot encoding rule, the first two [1,0] represent male gender, the middle [1,0] represents unsuccessful authentication, and the last [1,0] represents outstanding balance, then this operation can convert [1,0,1,0,1,0] into [male, unsuccessful authentication, outstanding balance].
[0105] A simulation data generation system based on a diffusion model, comprising:
[0106] A data processing module: divides real historical sample data, processes and transforms the format of the divided real historical sample data to obtain a second data Y; this module is mainly used to process the data format to facilitate subsequent input into the model;
[0107] A discrimination module: generates a first simulation data F1Y of the same order of magnitude as the second data Y; after respectively identifying the first simulation data F1Y and the second data Y, trains a binary classification convolutional neural network by constructing a binary classification data set; this module trains a binary classification convolutional neural network for discriminating the authenticity of simulation data, and is subsequently used to select the simulation data closest to the real data;
[0108] A simulation data generation module: generates a second simulation data F2Y of the same order of magnitude as the second data Y; uses the trained binary classification convolutional neural network to screen out a third simulation data SY that meets the threshold and the required quantity from the second simulation data F2Y, and transforms it into a fourth simulation data SX; this module selects the simulation data that meets the authenticity requirements from the simulation data through the binary classification convolutional neural network until the simulation data meets the required sample size.
[0109] A simulation data generation device based on a diffusion model, comprising:
[0110] A memory: used to store a computer program to implement a simulation data generation method based on a diffusion model as described above;
[0111] A processor: used to implement a simulation data generation method based on a diffusion model as described above when executing the computer program.
[0112] A computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a simulation data generation method based on a diffusion model as described above are implemented.
[0113] The application effect of the present invention is described in detail below in combination with a simulation experiment.
[0114] The test environment is shown in Table 5. The simulation experiment is based on a handwritten digit image data set (MINIST), and is configured with CUDA (CUDA is a parallel computing platform and programming model designed and developed by NVIDIA Corporation) on the PyTorch (PyTorch is an open-source software for neural network training) framework;
[0115] Table 5 Simulation Experiment Configuration
[0116]
[0117]
[0118] As shown in Figure 2-11 (a) therein, ten handwritten digit images from 0 to 9 are selected and processed by the method of the present invention to obtain a simulated data image as shown in Figure 2-11 (b) therein. It can be seen by comparison that the original handwritten digits and the simulated data generated based on the method of the present invention are highly similar, with high clarity and less deviation.
[0119] In order to verify whether the generated simulated data has a similar data distribution law as the original handwritten digits, the following operations are carried out:
[0120] Mix the Figure 2-11 handwritten digit image samples in (a) therein, Figure 2-11 mix the simulated data samples in (b) therein, and then input them into a convolutional neural network model for training respectively to obtain the corresponding first model Model X and the second model Model SX ;
[0121] Calculate the metrics of the first model Model X and the second model Model SX respectively, and obtain the area under the receiver operating characteristic curve (AUC) of the two models as shown in Table 6;
[0122] Table 6 Comparison of the metrics of the first model Model X and the second model Model SX
[0123]
[0124]
[0125] Generally speaking, the smaller the difference in the metric results between the first model Model X and the second model Model SX , the better the consistency of the data distribution law between the generated simulated data and the original handwritten digits; the differences in the 10 groups of metric comparisons in Table 6 are all less than 2%, so the simulated data generated based on the method of the present invention has a similar data distribution law as the original handwritten digits;
[0126] If the difference in the metric results between the first model Model X and the second model Model SX is greater than 2%, it is necessary to return to step 3.2, appropriately reduce the threshold on the basis of the top 5%, and control it within the range of 1%-5%, so that the first model Model Xand the second model Model SX The difference in the index results of is reduced, and then simulation data with a data distribution law similar to that of the original handwritten digits is obtained.
[0127] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, and improvement made within the technical scope disclosed by the present invention and within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.
Claims
1. A simulation data generation method based on a diffusion model, characterized in that: It includes the following steps: Step 1: Divide the real historical sample data, process and transform the format of the divided real historical sample data to obtain the second data Y; Step 2: Generate the first simulation data F1Y with the same order of magnitude as the second data Y in Step 1; after respectively identifying the first simulation data F1Y and the second data Y, train a binary classification convolutional neural network by constructing a binary classification dataset; Step 3: Generate the second simulation data F2Y with the same order of magnitude as the second data Y in Step 1; use the trained binary classification convolutional neural network in Step 2 to screen out the third simulation data SY that meets the threshold and the required quantity from the second simulation data F2Y, and transform it into the fourth simulation data SX.
2. The simulation data generation method based on a diffusion model according to claim 1, wherein: The specific content of Step 1 is as follows: Step 1.1: Divide the real historical sample data into ordered data and unordered data according to the business meaning, and perform one-hot encoding on the unordered data; Step 1.2: Horizontally concatenate the ordered data and the encoded unordered data in Step 1.1 to generate the first data X; Step 1.3: Convert each column in the first data X in Step 1.2 into the second data Y through Formula 1; the Formula 1 is: Where: X is the column to be processed in the first data; X min is the minimum value of this column to be processed in the first data; X max is the maximum value of this column to be processed in the first data; eps = 10 -10 ; Y is the second data.
3. The simulation data generation method based on a diffusion model according to claim 1, wherein: The specific content of Step 2 is as follows: Step 2.1: Use the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to generate the first simulation data F1Y with the same order of magnitude as the second data Y in Step 1; Step 2.2: Mark the data in the second data Y in Step 1 as true, i.e., 1, mark the data in the first simulation data F1Y in Step 2.1 as false, i.e., 0, and construct the marked data in the second data Y and the marked data in the first simulation data F1Y into a binary classification dataset; Step 2.3: Use the binary classification dataset in Step 2.2 to train a binary classification convolutional neural network.
4. A method for generating simulation data based on a diffusion model according to claim 1, characterized in that: The specific content of Step 3 is as follows: Step 3.1: Use the score-based generative diffusion model SMLD (Score Matching with Langevin Dynamics) to generate the second simulation data F2Y with the same order of magnitude as the second data Y in Step 1; Step 3.2: Use the trained binary classification convolutional neural network in Step 2 to score and sort the second simulation data F2Y in Step 3.1, and select the second simulation data F2Y whose scores meet the threshold; Step 3.3: Iterate the operation in Step 3.2 until the sample quantity reaches the required quantity and then stop, and record all the second simulation data F2Y whose scores meet the threshold as the third simulation data SY; Step 3.4: Convert each column in the third simulation data SY in Step 3.3 into the fourth simulation data SX through Formula 2; the Formula 2 is: Where: Y represents the column to be processed in the third simulation data; X min represents the minimum value of this column to be processed in the first data; X max represents the maximum value of this column to be processed in the first data; e represents the natural constant.
5. The simulation data generation method based on a diffusion model according to claim 4, characterized in that: In Step 3.2, the threshold is set to the top 1-5% of the score ranking of the second simulation data F2Y.
6. The simulation data generation method based on a diffusion model according to claim 4, wherein: Step 3.4 further includes: for the columns of the one-hot encoding corresponding to unordered data, values greater than 0.5 are assigned 1, values less than 0.5 are assigned 0, and simulation data equal to 0.5 are deleted; then, according to the one-hot encoding rule, the columns of the one-hot encoding corresponding to unordered data are converted into the form corresponding to the real historical sample data.
7. A simulation data generation system based on a diffusion model, characterized in that: It includes: Data processing module: dividing the real historical sample data, processing and transforming the format of the divided real historical sample data to obtain the second data Y; Identification module: generating the first simulation data F1Y with the same order of magnitude as the second data Y; after respectively identifying the first simulation data F1Y and the second data Y, training a binary classification convolutional neural network by constructing a binary classification data set; Simulation data generation module: generating the second simulation data F2Y with the same order of magnitude as the second data Y; using the trained binary classification convolutional neural network to screen out the third simulation data SY that meets the threshold and the required quantity from the second simulation data F2Y, and converting it into the fourth simulation data SX.
8. A simulation data generation device based on a diffusion model, characterized in that: It includes: Memory: used to store a computer program to implement a simulation data generation method according to any one of claims 1-6; Processor: used to implement a simulation data generation method according to any one of claims 1-6 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a simulation data generation method according to any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Sample data generation method, apparatus and device, and readable storage medium
CN117574263A
Sample generation method, model training method, sorting method, device and equipment based on large model
CN118779691A