Simulation Service Platform Data Similarity Evaluation Method, Device, and Electronic Device

By generating synthetic data and calculating compression distances, the accuracy and intuitiveness of data generation and evaluation are solved, and the data analysis and processing capabilities of the simulation service platform are improved.

CN119226814BActive Publication Date: 2025-07-08BEIJING FUZHI GONGCHUANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411370714.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-07-08
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

In the industrial design and virtual simulation service platform, the data generation quality is not high, the discriminant model accuracy is insufficient, and the similarity evaluation method lacks intuitiveness, resulting in poor data generation and evaluation results.

Method used

By generating adversarial networks to generate synthetic data that is highly similar to the real data distribution, and calculate the compressed distance between the real data and the synthetic data, evaluate its similarity, and use neural network models to support data augmentation and simulation experiments.

Benefits of technology

It improves the accuracy of data generation and the intuitiveness of similarity evaluation, provides effective means to analyze and process simulated data, and is suitable for any one-dimensional or multi-dimensional continuous data set, and is used in fields such as data compression effect, data denoising and signal recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119226814B_ABST
    Figure CN119226814B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, and electronic device for evaluating data similarity in a simulation service platform, and relates to the technical field of simulation data processing. The method includes: obtaining real data, generating synthetic data based on the real data through a generative adversarial network, calculating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data, determining the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate, and evaluating the similarity between the real data and the synthetic data based on the compression distance. The solution provided by the present application can generate synthetic data highly similar to the distribution of real data through a generative adversarial network, provide strong support for data augmentation and simulation experiments by combining neural network models, and evaluate the similarity by calculating the compression distance value between the real data and the synthetic data, providing an effective means for simulation data analysis and processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of simulation data processing, and particularly to a method, device, and electronic device for evaluating data similarity on a simulation service platform. Background Art

[0002] In the context of the data-driven era, the data generation, discrimination, and similarity evaluation technologies of industrial design and virtual simulation service platforms are crucial for improving product design quality, optimizing simulation experiment effects, and accelerating the product development cycle.

[0003] With the continuous progress of technology, significant advancements have been made in data generation, discrimination, and similarity evaluation on industrial design and virtual simulation service platforms, and model-based data generation methods have gradually emerged. These methods capture the distribution characteristics in data by training specific machine learning models, thereby generating synthetic data that is closer to the real data distribution. However, the accuracy of these methods is still limited by the complexity of the models and the diversity of training data, and they still face problems such as low data generation quality, insufficient accuracy of discrimination models, and lack of intuitiveness in similarity evaluation methods. Summary of the Invention

[0004] To solve or partially solve the problems existing in the related technologies, the present application provides a method, device, and electronic device for evaluating data similarity on a simulation service platform, which can generate synthetic data highly similar to the real data distribution through a generative adversarial network, provide strong support for data augmentation and simulation experiments by combining neural network models, and evaluate similarity by calculating the compression distance value between the real data and the synthetic data, providing an effective means for data analysis and processing.

[0005] The first aspect of the present application provides a method for evaluating data similarity on a simulation service platform, including:

[0006] Obtain real data, and generate synthetic data based on the real data through a generative adversarial network;

[0007] Calculate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data;

[0008] Determine the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate;

[0009] Evaluate the similarity between the real data and the synthetic data based on the compression distance.

[0010] Preferably, the generating synthetic data based on the real data through a generative adversarial network includes:

[0011] Generate noise according to the data characteristics of the real data;

[0012] Generate the synthetic data based on the noise and the real data through the generative adversarial network.

[0013] Preferably, calculating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data includes:

[0014] Calculate the variance of the noise and calculate the compression rate according to the variance;

[0015] Subtract the noise from the synthetic data to generate simulated real data;

[0016] Calculate the decompression rate according to the simulated real data and the synthetic data.

[0017] Preferably, calculating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data further includes:

[0018] Calculate the feature matrices of the real data and the synthetic data respectively;

[0019] Calculate the eigenvalues and eigenvectors respectively based on their respective feature matrices;

[0020] Determine the number of principal components to be retained for the real data and the synthetic data according to their respective eigenvalues;

[0021] Input the real data and the synthetic data respectively combined with the number of principal components into a preset compression rate function to generate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data respectively.

[0022] Preferably, inputting the real data matrix and the synthetic data matrix respectively combined with the number of principal components into a preset compression rate function to generate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data respectively includes:

[0023] Determine a number of eigenvalues equal to the number of principal components from the eigenvalues, and the target eigenvectors corresponding to the number of eigenvalues equal to the number of principal components, and determine the target eigenvectors as the principal components;

[0024] Project the real data matrix and the synthetic data matrix onto their respective principal components respectively to generate the real data after dimensionality reduction and the synthetic data after dimensionality reduction;

[0025] Calculate the compression ratio of the real data and the compression ratio of the synthetic data respectively through the preset compression ratio function according to the number of the respective eigenvalue and the number of the principal components;

[0026] Determine the compression ratio of the synthetic data as the compression rate, and determine the compression ratio of the real data as the decompression rate.

[0027] Preferably, the determining the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate includes:

[0028] Compare the magnitudes of the compression rate and the decompression rate;

[0029] Determine the larger one of the compression rate and the decompression rate as the compression distance.

[0030] Preferably, the generative adversarial network includes a generator network and a discriminator network. Both the generator network and the discriminator network are multi-layer fully connected neural networks. The generative adversarial network is obtained by the following method:

[0031] Input a first training sample into an initial generative adversarial network;

[0032] The initial generator network generates a first synthetic sample of the first training sample;

[0033] Train the initial generative adversarial network by using the first training sample and the first synthetic sample to obtain a preliminary generative adversarial network;

[0034] Input a second training sample into the preliminary generative adversarial network;

[0035] The preliminary generator network generates a second synthetic sample of the second training sample;

[0036] Train the preliminary discriminator network by using the second training sample and the second synthetic sample to generate a trained generative adversarial network.

[0037] Preferably, the training the initial generative adversarial network by using the first training sample and the first synthetic sample to obtain a preliminary generative adversarial network includes:

[0038] Input the first training sample and the first synthetic sample into the initial discriminator network respectively, calculate a first loss by using a cross-entropy loss function, and adjust the initial generative adversarial network according to the first loss until a first preset condition is satisfied to obtain the preliminary generative adversarial network;

[0039] The training the preliminary discriminator network by using the second training sample and the second synthetic sample to generate a trained generative adversarial network includes:

[0040] Add labels to the second training sample and the second synthetic sample respectively to construct a training data set pair;

[0041] Use the training data set pair to train the preliminary discrimination network until the preliminary discrimination network meets the second preset condition to generate the generative adversarial network.

[0042] A second aspect of the present application provides a simulation service platform data similarity evaluation device, including:

[0043] An acquisition module, configured to acquire real data and generate synthetic data based on the real data through a generative adversarial network;

[0044] A calculation module, configured to calculate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data;

[0045] A determination module, configured to determine the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate;

[0046] An evaluation module, configured to evaluate the similarity between the real data and the synthetic data based on the compression distance.

[0047] A third aspect of the present application provides an electronic device, including:

[0048] A processor; and

[0049] A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is made to execute the method as described above.

[0050] The technical solution provided by the present application may include the following beneficial effects: An embodiment of the present application discloses a method for evaluating the similarity of simulation service platform data, including acquiring real data, generating synthetic data based on the real data through a generative adversarial network, calculating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data, determining the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate, and evaluating the similarity between the real data and the synthetic data based on the compression distance. It can generate synthetic data highly similar to the real data distribution through a generative adversarial network, provide strong support for data enhancement and simulation experiments by combining neural network models, and evaluate the similarity by calculating the compression distance value between the real data and the synthetic data, providing an effective means for simulation data analysis and processing.

[0051] The technical solution of the present application can also: directly reflect the information loss degree of data during the compression and decompression processes through the compression distance, making the evaluation result more intuitive and understandable. It can be applied to any one-dimensional or multi-dimensional continuous data set, and is not limited to a specific compression algorithm or denoising technique. It has broad application prospects in the fields of data compression effect evaluation, data denoising, signal recovery, etc., and can provide valuable information support for users.

[0052] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] By describing the exemplary embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. Among them, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.

[0054] Figure 1 is a schematic flowchart of a method for evaluating data similarity of a simulation service platform shown in an embodiment of the present application;

[0055] Figure 2 is another schematic flowchart of a method for evaluating data similarity of a simulation service platform shown in an embodiment of the present application;

[0056] Figure 3 is another schematic flowchart of a method for evaluating data similarity of a simulation service platform shown in an embodiment of the present application;

[0057] Figure 4 is a schematic flowchart of a method for generating a generative adversarial network shown in an embodiment of the present application;

[0058] Figure 5 is a schematic structural diagram of a device for evaluating data similarity of a simulation service platform shown in an embodiment of the present application;

[0059] Figure 6 is a schematic structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The embodiments of the present application will be described in more detail below with reference to the drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0061] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "said", and "the" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0062] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, the meaning of "a plurality" is two or more unless otherwise specifically defined.

[0063] With the continuous progress of technology, the industrial design and virtual simulation service platform has made remarkable progress in data generation, discrimination, and similarity evaluation, and the model-based data generation method has gradually emerged. These methods capture the distribution characteristics in the data by training specific machine learning models, thereby generating synthetic data that is closer to the real data distribution. However, the accuracy of these methods is still limited by the complexity of the model and the diversity of the training data, and still faces problems such as low data generation quality, insufficient accuracy of the discrimination model, and lack of intuitiveness in the similarity evaluation method.

[0064] In view of the above problems, the embodiments of this application provide a method for evaluating the similarity of data on a simulation service platform, which can generate synthetic data highly similar to the real data distribution through a generative adversarial network, provide strong support for data enhancement and simulation experiments, and evaluate the similarity by calculating the compression distance value between the real data and the synthetic data, providing an effective means for data analysis and processing.

[0065] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings.

[0066] Figure 1 It is a schematic flowchart of the method for evaluating the similarity of data on a simulation service platform shown in the embodiments of this application.

[0067] See Figure 1 , the method includes:

[0068] Step 101, obtain real data and generate synthetic data based on the real data through a generative adversarial network.

[0069] A generative adversarial network (GAN, Generative Adversarial Networks) contains a generative network and a discriminative network. Among them, the generative network is responsible for capturing the distribution of sample data, and the discriminative network is generally a binary classifier that discriminates whether the input is real data or generated samples.

[0070] In the embodiments of the present application, the real data X can be a vector containing n m-dimensions, representing the observed values of n samples on m features, or an original continuous data set. The real data X can be any form of data, such as image, text, or time series data. The generative adversarial network can generate synthetic data Y that is as close as possible to the real data according to the real data X.

[0071] Step 102, calculate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data.

[0072] After generating the synthetic data Y, calculate the compression rate from the real data X to the synthetic data Y and the decompression rate from the synthetic data Y to the real data X respectively.

[0073] Step 103, determine the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate.

[0074] According to the compression rate and the decompression rate, the compression distance between the real data X and the synthetic data Y can be determined, and the compression rate from the real data X to the synthetic data Y, the decompression rate from the synthetic data Y to the real data X, and the compression distance between the real data X and the synthetic data Y can also be output for users to reference and analyze.

[0075] Step 104, evaluate the similarity between the real data and the synthetic data based on the compression distance.

[0076] Evaluate the similarity between the real data X and the synthetic data Y according to the compression distance. The greater the compression distance, the greater the difference in information loss between the real data X and the synthetic data Y during the compression process, thus reflecting the lower similarity or correlation between them.

[0077] An embodiment of the present application discloses a method for evaluating data similarity of a simulation service platform, including obtaining real data, generating synthetic data based on the real data through a generative adversarial network, calculating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data, determining the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate, and evaluating the similarity between the real data and the synthetic data based on the compression distance. It can generate synthetic data highly similar to the distribution of real data through a generative adversarial network, provide strong support for data augmentation and simulation experiments by combining neural network models, and evaluate the similarity by calculating the compression distance value between the real data and the synthetic data, providing an effective means for simulation data analysis and processing.

[0078] Figure 2 It is another flowchart of the method for evaluating data similarity of the simulation service platform shown in the embodiment of the present application.

[0079] See Figure 2 , the method includes:

[0080] Step 201, obtain real data and generate noise according to the data characteristics of the real data.

[0081] The embodiment of the present application indirectly reflects the similarity or difference between data by simulating the compression and decompression processes of real data affected by noise, providing strong support for tasks such as evaluating data compression effects, data denoising, and signal recovery.

[0082] The real data X can be an original continuous data set, which can be any one-dimensional or multi-dimensional continuous data point set such as time series, image pixel data, sensor readings, etc. According to the application scenario and the data characteristics of the real data, a reasonable noise variance Y_noise_variance is set, and a random number generator (such as a Gaussian noise generator) is used to generate noise N of the same size as the real data X (i.e., the same number of data points and the same dimension for each data point) according to the given noise variance Y_noise_variance.

[0083] Step 202, generate synthetic data based on the noise and the real data through a generative adversarial network.

[0084] Using the generated noise N, the generative adversarial network adds the generated noise N to each corresponding data point of the real data X to obtain synthetic data Y on the basis of the real data X, that is, synthetic data Y = X + N. This step simulates the noise pollution that data may suffer during actual transmission, storage, or processing.

[0085] Step 203, calculate the variance of the noise and calculate the compression rate according to the variance.

[0086] Calculate the variance of the noise N in the synthetic data Y, denoted as Var(N). Take the logarithm (base 10) of the variance Var(N) of the noise N as an approximation of the compression rate R_X_given_Y from the real data X to the synthetic data Y, that is . The compression rate can be regarded as a quantitative indicator of the degree of information loss. The larger the noise variance, the more information is lost and the higher the compression rate.

[0087] Step 204: Subtract the noise from the synthetic data to generate simulated real data.

[0088] Since the added noise N is known, the simulated real data X_hat can be directly generated by subtracting the noise N from the synthetic data Y, that is, X_hat = Y - N. This step simulates the complete decompression process under ideal conditions, that is, completely removing the noise and restoring the original data.

[0089] Step 205: Calculate the decompression rate based on the simulated real data and the synthetic data.

[0090] Based on the simulated real data X_hat and the synthetic data Y, the decompression rate R_Y_given_X can be calculated. In the embodiments of the present application, for the sake of simplifying the calculation process, the decompression rate R_Y_given_X is directly set to be the same as the compression rate R_X_given_Y, that is, R_Y_given_X = R_X_given_Y. However, in practical applications, if the efficiency of the decompression algorithm or the quality of data recovery has a significant impact on the results, a more complex model can be introduced to calculate the decompression rate. The embodiments of the present application do not limit the calculation method of the decompression rate.

[0091] Step 206: Compare the magnitudes of the compression rate and the decompression rate.

[0092] After obtaining the decompression rate R_Y_given_X, the magnitudes of the compression rate R_X_given_Y and the decompression rate R_Y_given_X can be compared.

[0093] Step 207: Determine the larger of the compression rate and the decompression rate as the compression distance.

[0094] Take the larger value of the compression rate R_X_given_Y and the decompression rate R_Y_given_X as the overlapping compression distance CD_lsy between continuous data, that is, CD_lsy = max(R_X_given_Y, R_Y_given_X ).

[0095] Step 208: Evaluate the similarity between the real data and the synthetic data based on the compression distance.

[0096] For ease of understanding, the compression distance CD_lsy can be converted into a more understandable percentage or ratio form. When evaluating similarity, a preset threshold can be set for the compression distance CD_lsy, and the compression distance CD_lsy is compared with the preset threshold to evaluate whether the similarity or difference between the real data and the synthetic data reaches a predetermined standard. In one example, the calculated compression distance CD_lsy can be output to the user and presented in the form of a Graphical User Interface (GUI), a command-line interface (CLI), or a data report, etc. The user can judge the degree of similarity or difference between the real data X and the synthetic data Y according to the value of the compression distance CD_lsy, and then make corresponding data analysis or decision support.

[0097] In an alternative embodiment of the present application, if the real data X and the synthetic data Y have complete information overlap, for this special case, a random process M is introduced to simulate the complete overlap between the data, and based on this process, the calculation of the bidirectional compression rate between the real data X and the synthetic data Y is realized, and then an accurate compression distance is obtained. The random process M can be Gaussian white noise, Brownian motion or any other independent random process that does not depend on the real data X, ensuring that the generation of M can cover all possible variation ranges of X while maintaining its independence. The specific process is as follows: First, a complete overlap model under continuous data is constructed. By superimposing an independent random process M on the data set X (real data X), the synthetic data Y (Y = X + M) is generated through a generative adversarial network. The parameters of M, such as the mean, variance or correlation function, can be adjusted according to actual requirements and application scenarios. The synthetic data Y (i.e., X + M) is compressed using a preset compressor to obtain the compressed data representation S_(X + M). The preset compressor can effectively process and compress data containing random noise. The length or number of bits of S_(X + M) is defined as the compression distance R_(Y|X)(D) required to successfully decompress from the real data X to the synthetic data Y. This process simulates the information transfer and compression process from the real data X to the synthetic data Y. The synthetic data Y is compressed using the preset compressor to obtain the compressed data S_Y. The synthetic data Y is decompressed to obtain the estimated value Ŷ of the synthetic data Y. The real data X (or an approximation of X) is recovered from the estimated value Ŷ through filtering, denoising algorithms or model-based methods. The length or number of bits of S_Y is defined as the compression distance R_(X|Y)(D) from the synthetic data Y to the real data X. Compare R_(Y|X)(D) and R_(X|Y)(D), and determine the larger of the two as the compression distance CD_lsy(X,Y). The compression distance is output as an index for evaluating the similarity or compression characteristics between the real data X and the synthetic data Y, and can be used in multiple fields such as the performance evaluation of data compression algorithms, data redundancy removal, and feature extraction.

[0098] The embodiments of the present application disclose a method for evaluating the similarity of data in a simulation service platform, including obtaining real data, generating noise according to the data characteristics of the real data, generating synthetic data through a generative adversarial network based on the noise and the real data, calculating the variance of the noise, calculating the compression rate according to the variance, subtracting the noise from the synthetic data to generate simulated real data, calculating the decompression rate according to the simulated real data and the synthetic data, comparing the magnitudes of the compression rate and the decompression rate, determining the larger of the compression rate and the decompression rate as the compression distance, and evaluating the similarity between the real data and the synthetic data based on the compression distance. It can directly reflect the degree of information loss of data during the compression and decompression processes through the compression distance, making the evaluation result more intuitive and understandable. It can be applied to any one-dimensional or multi-dimensional continuous data set, and is not limited to a specific compression algorithm or denoising technique. It has a wide range of application prospects in the fields of data compression effect evaluation, data denoising, signal recovery, etc., and can provide valuable information support for users.

[0099] Figure 3 It is another schematic flowchart of the method for evaluating the similarity of data in the simulation service platform shown in the embodiments of the present application.

[0100] See Figure 3 , the method includes:

[0101] Step 301, obtain real data and generate synthetic data through a generative adversarial network based on the real data.

[0102] In the embodiments of the present application, a compression rate simulation method of principal component analysis (PCA, Principal Components Analysis) is adopted. Using this method, the compression rate is calculated to evaluate the similarity or correlation between two data sets. By simulating the data compression process, the amount of information loss of the data before and after compression is calculated, so as to indirectly reflect the dependence and similarity between the data. PCA is a commonly used data dimensionality reduction technology. Through linear transformation, the original data is mapped into a new coordinate system, so that the variance of the transformed data is the largest on the first coordinate (i.e., the first principal component), the second largest on the second coordinate, and so on. In the embodiments of the present application, PCA is used to simulate the data compression process.

[0103] In the embodiments of the present application, it is assumed that the real data X is stored in matrix form, where each real data X contains n m-dimensional vectors, representing the observed values of n samples on m features. The generative adversarial network can generate synthetic data Y that is as close as possible to the real data according to the real data X. The synthetic data Y also contains n m-dimensional vectors. If the real data X and the synthetic data Y are not in matrix form, they can be converted into matrix form.

[0104] Step 302: Calculate the feature matrices of the real data and the synthetic data respectively.

[0105] Calculate the feature matrices of the real data and the synthetic data respectively. The feature matrix can be a covariance matrix or a correlation matrix.

[0106] Step 303: Calculate the eigenvalues and eigenvectors respectively based on their respective feature matrices.

[0107] Calculate their respective eigenvalues and eigenvectors respectively based on their respective feature matrices. The eigenvectors correspond to the main change directions of the data, i.e., the principal components. At the same time, record the number m of eigenvalues.

[0108] Step 304: Determine the number of principal components to be retained for the real data and the synthetic data according to their respective eigenvalues.

[0109] According to the magnitudes of the eigenvalues, determine the number numComponents of principal components to be retained for the real data X and the synthetic data Y, and select the eigenvectors corresponding to the first numComponents largest eigenvalues as the principal components.

[0110] Step 305: Input the real data and the synthetic data respectively together with the number of principal components into a preset compression rate function to generate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data respectively.

[0111] Define a preset compression rate function simulatedCompressionRate(data, numComponents), which receives the real data X, the synthetic data Y and the number numComponents of principal components to be retained for each as inputs respectively. In one example, if the PCA transformation has not been performed, inside the preset compression rate function, first perform the PCA transformation on the real data X or the synthetic data Y, and then generate the compression rate from the real data X to the synthetic data Y and the decompression rate from the synthetic data Y to the real data X respectively.

[0112] In an optional embodiment of the present application, Step 305 includes:

[0113] Sub-step 3051: Determine the number of eigenvalues equal to the number of principal components from the eigenvalues, and the target eigenvectors corresponding to the number of eigenvalues equal to the number of principal components, and determine the target eigenvectors as the principal components.

[0114] According to the magnitudes of the eigenvalues, determine the number of principal components numComponents to be retained for the real data X and the synthetic data Y. Select the eigenvectors corresponding to the top numComponents largest eigenvalues as the target eigenvectors, and determine these target eigenvectors as the principal components. Then retain the first numComponents principal components, which will be used to construct a new coordinate system.

[0115] Sub-step 3052: Project the real data and the synthetic data onto their respective principal components, respectively generating the reduced-dimensional real data and the reduced-dimensional synthetic data.

[0116] The real data X and the synthetic data Y will be projected along the principal components. Project the real data X and the synthetic data Y onto their respective principal components to obtain the reduced-dimensional data sets X_reduced and Y_reduced.

[0117] Sub-step 3053: Calculate the compression rate of the real data and the compression rate of the synthetic data respectively through a preset compression rate function according to the number of their respective eigenvalues and the number of principal components.

[0118] The compression rate L(X) of the real data and the compression rate L(Y) of the synthetic data can be calculated as log2(number of eigenvalues m / number of retained principal components numComponents).

[0119] Sub-step 3054: Determine the compression rate of the synthetic data as the compression speed, and determine the compression rate of the real data as the decompression speed.

[0120] To evaluate the similarity or correlation between the real data X and the synthetic data Y, assume that the compression rate L(Y|X) from the real data X to the synthetic data Y can be approximated by projecting Y onto the principal components of X and calculating the compression rate. Similarly, the compression rate L(X|Y) from the synthetic data Y to the real data X can also be approximated by projecting X onto the principal components of Y and calculating the compression rate. In one example, to improve the calculation speed, a simplified approximation method can be adopted, that is, use the principal components of the real data X and the synthetic data Y themselves to calculate the compression rate and the decompression rate respectively, that is, the compression rate L(Y|X) from the real data X to the synthetic data Y = L(Y), and the decompression rate L(X|Y) from the synthetic data Y to the real data X = L(X).

[0121] Step 306: Compare the magnitudes of the compression rate and the decompression rate.

[0122] Compare the compression rate L(Y) and the decompression rate L(X) to determine the larger one of the two.

[0123] Step 307, determine the larger value between the compression rate and the decompression rate as the compression distance.

[0124] Take the larger value between L(Y) and L(X) as the compression distance CD(X, Y) between the real data X and the synthetic data Y. This choice is based on the assumption that if the difference in information loss between two data sets during compression is large, then their similarity or correlation may be low. The calculated compression rate from the real data X to the synthetic data Y (approximately L(Y)), the decompression rate from the synthetic data Y to the real data X (approximately L(X)), and the compression distance CD(X, Y) between the real data X and the synthetic data Y can also be presented to the user in the form of a table or a graph.

[0125] Step 308, evaluate the similarity between the real data and the synthetic data based on the compression distance.

[0126] According to the value of the compression distance CD(X, Y), the similarity or correlation between the real data X and the synthetic data Y can be evaluated. For example, if CD(X, Y) is small, it indicates that the difference in information loss between the real data X and the synthetic data Y during compression is not large, and they may have a high similarity or correlation; on the contrary, if CD(X, Y) is large, it indicates that the similarity or correlation between the real data X and the synthetic data Y may be low.

[0127] In an alternative embodiment of the present application, if the real data X and the synthetic data Y have complete information overlap, for this special case, a random process M is introduced to simulate the complete overlap between the data, and based on this process, the calculation of the bidirectional compression rate between the real data X and the synthetic data Y is realized, and then an accurate compression distance is obtained. The random process M can be Gaussian white noise, Brownian motion, or any other independent random process that does not depend on the real data X, ensuring that the generation of M can cover all possible variation ranges of X while maintaining its independence. The specific process is as follows: First, a complete overlap model under continuous data is constructed. By superimposing an independent random process M on the dataset X (real data X), the synthetic data Y (Y = X + M) is generated through a generative adversarial network. The parameters of M, such as the mean, variance, or correlation function, can be adjusted according to actual requirements and application scenarios. The synthetic data Y (i.e., X + M) is compressed using a preset compressor to obtain the compressed data representation S_(X + M). The preset compressor can effectively process and compress data containing random noise. The length or number of bits of S_(X + M) is defined as the compression distance R_(Y|X)(D) required to successfully decompress from the real data X to the synthetic data Y. This process simulates the information transfer and compression process from the real data X to the synthetic data Y. The synthetic data Y is compressed using the preset compressor to obtain the compressed data S_Y. The synthetic data Y is decompressed to obtain an estimated value Ŷ of the synthetic data Y. The real data X (or an approximation of X) is recovered from the estimated value Ŷ through filtering, denoising algorithms, or model-based methods. The length or number of bits of S_Y is defined as the compression distance R_(X|Y)(D) from Y to X. R_(Y|X)(D) and R_(X|Y)(D) are compared, and the larger of the two is determined as the compression distance CD_lsy(X,Y). The compression distance is output as an index for evaluating the similarity or compression characteristics between the real data X and the synthetic data Y, and can be used for performance evaluation of data compression algorithms, data redundancy removal, feature extraction, and other fields.

[0128] An embodiment of the present application discloses a method for evaluating data similarity of a simulation service platform, which includes obtaining real data, generating synthetic data based on the real data through a generative adversarial network, respectively calculating the feature matrices of the real data and the synthetic data, calculating the eigenvalues and eigenvectors based on their respective feature matrices, determining the number of principal components to be retained for the real data and the synthetic data according to their respective eigenvalues, inputting the real data and the synthetic data respectively combined with the number of principal components into a preset compression rate function, generating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data respectively, comparing the magnitudes of the compression rate and the decompression rate, and determining the larger one of the compression rate and the decompression rate as the compression distance, so that users can adjust the number of principal components to be retained according to actual needs to balance the compression effect and the degree of information retention. It can be applied to various types of data sets, including but not limited to images, texts, time series, etc., has wide applicability, and can play an important role in fields such as data compression, data redundancy removal, feature extraction, and anomaly detection, providing strong support for research and applications in the fields of data science and machine learning.

[0129] Figure 4 It is a schematic flowchart of a method for generating a generative adversarial network shown in an embodiment of the present application.

[0130] See Figure 4 , the generative adversarial network includes a generative network and a discriminative network, and both the generative network and the discriminative network are multi-layer fully connected neural networks. The method includes:

[0131] Step 401, input the first training sample into the initial generative adversarial network.

[0132] The generative adversarial network includes a generator network (Generator Network, G) and a discriminator network (Discriminator Network, D). The generator network is responsible for receiving random noise or real data as input and generating synthetic data that is as close as possible to the real data distribution; the discriminator network is responsible for distinguishing whether the input data is real or generated by the generator network. In the embodiments of the present application, both the generator network and the discriminator network may include multiple fully connected layers, ReLU (Rectified Linear Unit) activation layers, and a specific output layer (such as a regression layer or a sigmoid layer). The generator network is a multi-layer fully connected neural network. The input layer receives a random noise vector of a fixed length (for example, a vector of length 100), followed by several hidden layers, each hidden layer containing a certain number of neurons (such as 256 neurons per layer), and using ReLU as the activation function. The number of neurons in the output layer matches the dimension of the input data. For example, for image data, the number of neurons in the output layer should be equal to the number of pixels in the image. The discriminator network is also a multi-layer fully connected neural network. The input layer receives data of the same dimension as the real data (including the data output by the generator network and the real data), followed by several hidden layers (each layer containing, for example, 512 neurons and using ReLU activation), and finally an output layer containing one neuron, using the sigmoid function to output the judgment result (0 represents synthetic data, and 1 represents real data). Before training, the weights and biases of the generator network and the discriminator network can be randomly initialized, and then the first training sample is input into the initial generator network of the initial generative adversarial network.

[0133] Step 402, the initial generator network generates a first synthetic sample of the first training sample.

[0134] The initial generator network generates a corresponding first synthetic sample according to the first training sample and random noise.

[0135] Step 403, input the first training sample and the first synthetic sample into the initial generative adversarial network respectively, calculate the first loss using the cross-entropy loss function, and adjust the initial generative adversarial network according to the first loss until the first preset condition is met, obtaining a preliminary generative adversarial network.

[0136] Input the first training sample and the first synthetic sample into the initial generative adversarial network respectively. The initial generative adversarial network calculates the first loss according to the cross-entropy loss function, and uses the backpropagation algorithm and an appropriate optimizer (such as the Adam optimizer) to update the parameters of the initial generator network and the initial discriminator network until the first preset condition is met, obtaining a preliminary generative adversarial network. The first preset condition may be reaching a preset maximum number of iterations or meeting other stopping conditions. The embodiments of the present application do not limit the first preset condition.

[0137] Step 404: Input the second training sample into the preliminary generative adversarial network.

[0138] After obtaining the preliminary generative adversarial network, further training can be carried out by inputting the second training sample into the preliminary generative adversarial network.

[0139] Step 405: The preliminary generative network generates a second synthetic sample of the second training sample.

[0140] The preliminary generative network of the preliminary generative adversarial network generates a corresponding second synthetic sample according to the second training sample.

[0141] Step 406: Add labels to the second training sample and the second synthetic sample respectively to construct a training data set pair.

[0142] Add labels to the second training sample and the second synthetic sample respectively. For example, add the label "1" to the second training sample and the label "0" to the second synthetic sample. For any second training sample, construct a training data pair (X_train, Y_train) according to the second training sample and the second synthetic sample corresponding to the second training sample, and traverse each second training sample to generate a training data set pair.

[0143] Step 407: Use the training data set pair to train the preliminary discriminant network until the preliminary discriminant network meets the second preset condition to generate a generative adversarial network.

[0144] Use the training data set pair to train the preliminary discriminant network to improve its ability to distinguish between real data and synthetic data. Repeat the training process until the preliminary discriminant network meets the second preset condition to generate the final generative adversarial network. The second preset condition can be that the performance of the discriminant network reaches stability, reaches a preset training goal, the similarity between synthetic data and real data reaches a preset condition, etc. The embodiments of the present application do not limit the second preset condition.

[0145] In an alternative embodiment of the present application, the method further includes:

[0146] Step S11: Obtain real data and generate synthetic data based on the real data through a generative adversarial network.

[0147] After generating the final generative adversarial network, real data X can be obtained and input into the generative adversarial network to generate synthetic data Y.

[0148] Step S12: Calculate the compression distance between the real data and the synthetic data through a preset compression distance function.

[0149] Select a suitable preset compression distance function, such as Kolmogorov complexity or an approximate compression distance based on a specific compression algorithm (such as the compression rate difference based on algorithms like ZIP or LZ77). Compress the real data and the synthetic data respectively through the preset compression distance function, and record the sizes after compression for each. Calculate the difference (or ratio) of the sizes of the two sets of data after compression as the compression distance between them.

[0150] Step S13, evaluate the similarity between the real data and the synthetic data based on the compression distance.

[0151] Output the calculated compression distance value as the evaluation result of the similarity between the synthetic data Y and the real data X. According to actual requirements, a threshold can be set to determine whether the quality of the synthetic data meets the requirements. It can be understood that if the synthetic data does not meet the requirements, the generative adversarial network can continue to be trained or the parameters of the generative adversarial network can be changed until the quality of the synthetic data meets the requirements.

[0152] The embodiment of the present application discloses a method for generating a generative adversarial network. The generative adversarial network includes a generator network and a discriminator network, and both the generator network and the discriminator network are multi-layer fully connected neural networks. The method includes: inputting the first training sample into the initial generative adversarial network, the initial generator network generating the first synthetic sample of the first training sample, inputting the first training sample and the first synthetic sample into the initial discriminator network respectively, calculating the first loss using the cross-entropy loss function, and adjusting the initial generative adversarial network according to the first loss until the first preset condition is met to obtain a preliminary generative adversarial network. Input the second training sample into the preliminary generative adversarial network, the preliminary generator network generating the second synthetic sample of the second training sample, adding labels to the second training sample and the second synthetic sample respectively to construct a training data set pair, and training the preliminary discriminator network using the training data set pair until the preliminary discriminator network meets the second preset condition to generate a generative adversarial network. It can generate synthetic data highly similar to the real data distribution based on the trained generative adversarial network, providing strong support for data augmentation and simulation experiments. Using the real data and the synthetic data for training the discriminator network can significantly improve the discrimination ability of the discriminator model for the real data and the synthetic data. This method can be flexibly applied to data generation and discrimination tasks in different fields, and the network structure and training parameters can be adjusted according to actual requirements to adapt to different application scenarios.

[0153] Corresponding to the foregoing embodiment of the application function implementation method, the present application also provides a device for evaluating data similarity of a simulation service platform, an electronic device, and corresponding embodiments.

[0154] Figure 5 It is a schematic structural diagram of a device for evaluating data similarity of a simulation service platform shown in the embodiment of the present application.

[0155] See Figure 5 , the data similarity evaluation device 500 includes:

[0156] An acquisition module 501, configured to acquire real data and generate synthetic data based on the real data through a generative adversarial network;

[0157] A calculation module 502, configured to calculate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data;

[0158] A determination module 503, configured to determine the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate;

[0159] An evaluation module 504, configured to evaluate the similarity between the real data and the synthetic data based on the compression distance.

[0160] In an alternative embodiment of the present application, the acquisition module 501 includes:

[0161] A noise sub-module, configured to generate noise according to the data characteristics of the real data;

[0162] A synthetic data sub-module, configured to generate synthetic data based on the noise and the real data through a generative adversarial network.

[0163] In an alternative embodiment of the present application, the calculation module 502 includes:

[0164] A variance sub-module, configured to calculate the variance of the noise and calculate the compression rate according to the variance;

[0165] A simulation sub-module, configured to subtract the noise from the synthetic data to generate simulated real data;

[0166] A decompression sub-module, configured to calculate the decompression rate according to the simulated real data and the synthetic data.

[0167] In an alternative embodiment of the present application, the calculation module 502 further includes:

[0168] A matrix sub-module, configured to calculate the feature matrices of the real data and the synthetic data respectively;

[0169] A feature sub-module, configured to calculate eigenvalues and eigenvectors respectively based on the respective feature matrices;

[0170] A principal component sub-module, configured to determine the number of principal components to be retained for the real data and the synthetic data according to the respective eigenvalues;

[0171] A calculation sub-module, configured to separately input real data and synthetic data together with the number of principal components into a preset compression rate function, and respectively generate a compression rate from the real data to the synthetic data and a decompression rate from the synthetic data to the real data.

[0172] In an optional embodiment of the present application, the calculation sub-module includes:

[0173] A principal component unit, configured to determine a number of principal component eigenvalues from the eigenvalues, and the target eigenvectors corresponding to the number of principal component eigenvalues, and determine the target eigenvectors as the principal components;

[0174] A projection unit, configured to project the real data matrix and the synthetic data matrix onto their respective principal components, and respectively generate the real data after dimensionality reduction and the synthetic data after dimensionality reduction;

[0175] A calculation unit, configured to respectively calculate the compression rate of the real data and the compression rate of the synthetic data through a preset compression rate function according to the number of their respective eigenvalues and the number of principal components;

[0176] A determination unit, configured to determine the compression rate of the synthetic data as the compression rate, and determine the compression rate of the real data as the decompression rate.

[0177] In an optional embodiment of the present application, the determination module 503 includes:

[0178] A comparison sub-module, configured to compare the magnitudes of the compression rate and the decompression rate;

[0179] A compression distance sub-module, configured to determine the larger one of the compression rate and the decompression rate as the compression distance.

[0180] In an optional embodiment of the present application, the generative adversarial network includes a generator network and a discriminator network. Both the generator network and the discriminator network are multi-layer fully connected neural networks. The device further includes:

[0181] A first input module, configured to input a first training sample into the initial generative adversarial network;

[0182] A first synthesis module, configured to the initial generator network generate a first synthetic sample of the first training sample;

[0183] A first training module, configured to use the first training sample and the first synthetic sample to train the initial generative adversarial network to obtain a preliminary generative adversarial network;

[0184] A second input module, configured to input a second training sample into the preliminary generative adversarial network;

[0185] A second synthesis module, configured to the preliminary generator network generate a second synthetic sample of the second training sample;

[0186] A second training module, configured to train a preliminary discrimination network using second training samples and second synthetic samples to generate a trained generative adversarial network.

[0187] In an optional embodiment of the present application, the first training module includes:

[0188] An adjustment sub-module, configured to input the first training samples and the first synthetic samples into the initial adversarial network respectively, calculate a first loss using a cross-entropy loss function, and adjust the initial generative adversarial network according to the first loss until a first preset condition is met, obtaining a preliminary generative adversarial network;

[0189] The second training module includes:

[0190] A label sub-module, configured to add labels to the second training samples and the second synthetic samples respectively to construct a training data set pair;

[0191] A second training sub-module, configured to train the preliminary discrimination network using the training data set pair until the preliminary discrimination network meets a second preset condition, generating a generative adversarial network.

[0192] An embodiment of the present application discloses a simulation service platform data similarity evaluation device, which includes obtaining real data, generating synthetic data based on the real data through a generative adversarial network, calculating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data, determining the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate, and evaluating the similarity between the real data and the synthetic data based on the compression distance. It can generate synthetic data highly similar to the distribution of the real data through the generative adversarial network. The neural network model provides strong support for data augmentation and simulation experiments. By calculating the compression distance value between the real data and the synthetic data to evaluate the similarity, it provides an effective means for simulation data analysis and processing.

[0193] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0194] Figure 6 It is a schematic structural diagram of an electronic device shown in an embodiment of the present application.

[0195] See Figure 6 , the electronic device 600 includes a memory 610 and a processor 620.

[0196] The processor 620 can be a Central Processing Unit (CPU), or it can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0197] The memory 610 can include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 620 or other modules of the computer. The permanent storage device can be a read-write storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 6010 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 610 can include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or wired.

[0198] An executable code is stored on the memory 610, and when the executable code is processed by the processor 620, it can cause the processor 620 to execute some or all of the methods described above.

[0199] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for performing some or all of the steps in the above method of the present application.

[0200] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), the processor is caused to execute some or all of the steps of the above method according to the present application.

[0201] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.

Claims

1. A method for evaluating data similarity of a simulation service platform, characterized in that, The method includes: Obtaining real data and generating synthetic data based on the real data through a generative adversarial network; the real data includes image, text, or time series data; Calculating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data; including: calculating the feature matrices of the real data and the synthetic data respectively; calculating eigenvalues and eigenvectors based on their respective feature matrices; determining the number of principal components to be retained for the real data and the synthetic data according to their respective eigenvalues; inputting the real data and the synthetic data respectively combined with the number of principal components into a preset compression rate function to generate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data respectively; Determining the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate; Evaluating the similarity between the real data and the synthetic data based on the compression distance.

2. The method according to claim 1, wherein The generating synthetic data based on the real data through a generative adversarial network includes: Generating noise according to the data characteristics of the real data; Generating the synthetic data based on the noise and the real data through the generative adversarial network.

3. The method according to claim 2, wherein The calculating the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data further includes: Calculating the variance of the noise and calculating the compression rate according to the variance; Subtracting the noise from the synthetic data to generate simulated real data; Calculating the decompression rate according to the simulated real data and the synthetic data.

4. The method according to claim 1, wherein The inputting the real data matrix and the synthetic data matrix respectively combined with the number of principal components into a preset compression rate function to generate the compression rate from the real data to the synthetic data and the decompression rate from the synthetic data to the real data respectively includes: Determining a number of eigenvalues equal to the number of principal components from the eigenvalues and the target eigenvectors corresponding to the number of principal components, and determining the target eigenvectors as the principal components; Projecting the real data matrix and the synthetic data matrix onto their respective principal components to generate the real data after dimensionality reduction and the synthetic data after dimensionality reduction respectively; Calculating the compression rate of the real data and the compression rate of the synthetic data respectively through the preset compression rate function according to the number of their respective eigenvalues and the number of principal components; Determining the compression rate of the synthetic data as the compression rate and determining the compression rate of the real data as the decompression rate.

5. The method according to claim 1, characterized in that, The determining the compression distance between the real data and the synthetic data according to the compression rate and the decompression rate includes: Comparing the magnitudes of the compression rate and the decompression rate; Determining the larger of the compression rate and the decompression rate as the compression distance.

6. The method according to claim 1, wherein The generative adversarial network includes a generator network and a discriminator network, both the generator network and the discriminator network are multi-layer fully connected neural networks, and the generative adversarial network is obtained through the following method: Input the first training sample into the initial generative adversarial network; The initial generative network generates a first synthetic sample of the first training sample; Train the initial generative adversarial network using the first training sample and the first synthetic sample to obtain a preliminary generative adversarial network; Input the second training sample into the preliminary generative adversarial network; The preliminary generative network generates a second synthetic sample of the second training sample; Train the preliminary discriminative network using the second training sample and the second synthetic sample to generate a trained generative adversarial network.

7. The method according to claim 6, wherein The training of the initial generative adversarial network using the first training sample and the first synthetic sample to obtain a preliminary generative adversarial network includes: Input the first training sample and the first synthetic sample into the initial adversarial network respectively, calculate a first loss using the cross-entropy loss function, and adjust the initial generative adversarial network according to the first loss until a first preset condition is satisfied to obtain the preliminary generative adversarial network; The training of the preliminary discriminative network using the second training sample and the second synthetic sample to generate a trained generative adversarial network includes: Add labels to the second training sample and the second synthetic sample respectively to construct a training data set pair; Train the preliminary discriminative network using the training data set pair until the preliminary discriminative network satisfies a second preset condition to generate the generative adversarial network.

8. An apparatus for evaluating data similarity of a simulation service platform, characterized in that, The apparatus includes: An acquisition module, configured to acquire real data and generate synthetic data based on the real data through a generative adversarial network; the real data includes image, text or time series data; A calculation module, configured to calculate a compression rate from the real data to the synthetic data and a decompression rate from the synthetic data to the real data; including: a matrix sub-module, configured to calculate feature matrices of the real data and the synthetic data respectively; a feature sub-module, configured to calculate eigenvalues and eigenvectors based on their respective feature matrices; a principal component sub-module, configured to determine the number of principal components to be retained for the real data and the synthetic data according to their respective eigenvalues; a calculation sub-module, configured to input the real data and the synthetic data respectively combined with the number of principal components into a preset compression rate function to generate a compression rate from the real data to the synthetic data and a decompression rate from the synthetic data to the real data respectively; A determination module, configured to determine a compression distance between the real data and the synthetic data according to the compression rate and the decompression rate; An evaluation module, configured to evaluate the similarity between the real data and the synthetic data based on the compression distance.

9. An electronic device, characterized in that, Includes: A processor; And A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor executes the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image preprocessing method and device, defect detection method and device and computer equipment

    CN115496890A

  • Method and device for realizing data asset health degree evaluation processing based on adversarial neural network, processor and storage medium thereof

    CN117972538A