GAN-based test data automatic generation method

Through the automatic generation method of GAN-based test data, the shortcomings of the test data generation algorithm in the prior art in terms of initial information conditions and comprehensiveness of the test path collection are solved, and efficient and automated test data generation is realized, meeting the correctness and comprehensiveness of the test data.

WO2025119081A1PCT designated stage expired Publication Date: 2025-06-12CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
PCT/CN2024/135489
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-11-29
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The existing test data generation algorithm has shortcomings in the initial information conditions and comprehensiveness of the test path collection, resulting in the long execution time of the algorithm and the generated test data cannot fully cover the actual scenario.

Method used

The automatic generation method of GAN-based test data is adopted to build a generative adversarial network, and the adversarial game between the generative network and the discriminative network is used to learn the distribution of real test data, and the least squares adversarial loss value, distribution distance loss value and distribution cross-comparison loss value are used to generate test data that conforms to the real data distribution.

Benefits of technology

It significantly improves the efficiency of testing work, reduces the burden on testers, improves the automation capabilities of test tasks in the cloud platform, and ensures that the generated test data is correct, comprehensive, coherent and determinable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135489_12062025_PF_FP_ABST
    Figure CN2024135489_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is a generative adversarial network (GAN)-based test data automatic generation method, which comprises the following steps: step 1, for a required test task, acquiring training data conforming to an actual test; step 2, preprocessing the training data, wherein, in order to ensure more balanced and comprehensive learning of a GAN during a training process, triplet partitioning is used to preprocess the training data; and step 3, designing and building a GAN model, wherein the GAN comprises a generator network and a discriminator network, the generator network being used for capturing and learning the distribution of the training data, and the discriminator network being used for determining the realness of sample data, namely determining the probability of the sample data being from real training data. The present application constructs the generator network in the form of encoder-decoder, and then constructs the discriminator network in the form of processing a binary classification task, the generator network and the discriminator network jointly forming the whole architecture of the GAN.
Need to check novelty before this filing date? Find Prior Art

Description

A GAN-based automatic test data generation method

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 5, 2023, with application number 202311653401.2 and invention name “A method for automatic generation of test data based on GAN”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The technical field involved in this application particularly relates to a method for automatically generating test data based on GAN (Generative Adversarial Network). Background Art

[0004] Over the past few years, cloud computing, as a new computing paradigm, has flourished, ushering in a new wave of IT technology transformation in the distributed computing community. Today, the cloud has become a vital service in internet computing. With the continued growth and development of cloud platforms, a wide range of functional modules and software tailored to these platforms have emerged. Before release, all software used in cloud platforms undergoes comprehensive testing to verify the correctness, comprehensiveness, consistency, and predictability of its functionality. Excellent test data and use cases are crucial to ensuring test integrity. Automatically generating test data and use cases can significantly improve testing efficiency and reduce the burden on testers.

[0005] Patent CN114676042A proposes an algorithm for generating test data for the power Internet of Things that combines an improved genetic algorithm with reinforcement learning. The algorithm initializes the population chromosome by encoding historical power data, iterates the encoding parameters of the chromosome with the largest fitness at each step, and iterates the test data for training. The final test path set is then obtained by driving reinforcement learning through the Jaccard distance.

[0006] The problems with this technical solution are:

[0007] (1) The initial information condition uses the encoding classification initialization method, which deviates from the actual scenario. The multi-step update genetic population iteration is slow, which increases the time cost of algorithm execution.

[0008] (2) The optimal test path set determined by chromosome fitness and Jaccard distance is only a subset of the actual scenario test path set and cannot guarantee the comprehensiveness of the test process.

[0009] To this end, we propose a GAN-based automatic test data generation method to solve the above problems. Summary of the Invention

[0010] The purpose of this section is to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of the present application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions shall not be used to limit the scope of the present application.

[0011] In view of the above-mentioned problems existing in the existing GAN-based test data automatic generation method, this application is proposed.

[0012] Therefore, the purpose of this application is to provide a method for automatically generating test data based on GAN, which constructs a generation network in the form of a codec, and then constructs a discrimination network by processing a binary classification task. The two together constitute the overall architecture of the generative adversarial network. This application also constructs a generation adversarial loss value, a distribution distance loss value, and a distribution intersection-over-union loss value based on the least squares algorithm. The three together constrain and drive the GAN model to learn the data distribution in the real test scenario, and can automatically generate test data that conforms to the real data distribution by inputting Gaussian distributed noise. This application can maximize the efficiency of test work and reduce the burden on test personnel, and can effectively improve the automation capability of test tasks in the cloud platform.

[0013] To solve the above technical problems, this application provides the following technical solution: a method for automatically generating test data based on GAN, comprising the following steps:

[0014] Step 1: Obtain training data that meets the actual test requirements for the required test task;

[0015] Step 2: Preprocessing the training data. To make GAN learn more balanced and comprehensive during training, the training data is preprocessed using triplet partitioning.

[0016] Step 3. Design and build a GAN model. The generative adversarial network consists of a generator network and a discriminator network. The generator network is used to capture and learn the distribution of training data, and the discriminator network is used to judge the authenticity of the sample data, that is, to judge the probability that the sample data comes from the real training data.

[0017] Step 4: Calculate the adversarial loss value, assuming that the data distribution of the real sample is p data , initialized to obey Gaussian distribution p z The random noise of (z) is z, and the generation network is generated by mapping G(z; θ g ) Transform the input random noise z to generate data distribution p g , discriminator D(x;θ d) The cross entropy loss function outputs a scalar [0,1] to represent the probability that the input sample comes from the real data;

[0018] Step 5: Calculate the distribution distance loss value. The adversarial loss proposed in step 4 enables the generative network to generate test data that conforms to the data distribution of the real test scenario. However, it may cause the generated data to experience pattern collapse, that is, the generated data is relatively simple and confined to a small interval of the real distribution. To address this situation, a distribution distance loss value is proposed. By reducing the distribution distance loss, the generated data distribution is driven to continuously approach the real data distribution, making the generated data more comprehensive and coherent, ensuring that the generated data can meet the requirements of the real test task.

[0019] Step 6: Calculate the distribution intersection loss. The loss values ​​designed in steps 4 and 5 make the generated data distribution have a good real data distribution. However, when reducing the distribution distance loss value obtained in step 5 during training, it may happen that only (μ t -μ g ) 2 Or just reduce (σ t -σ g ) 2 ,In order to constrain the comprehensiveness and availability of generated data, a distribution intersection-over-union loss value is proposed;

[0020] Step 7: Calculate the total loss value of the model based on the loss function designed in steps 4, 5, and 6, reduce the loss through gradient descent, and continuously update the model parameters until the model converges and stabilizes;

[0021] Step 8: Input the noise data that conforms to the Gaussian distribution into the trained generative network. The generative network automatically generates rich test data and uses it for testing in real test tasks.

[0022] As an optional solution to the GAN-based automatic test data generation method described in this application, the specific training data in step 1 should have the properties of correctness, comprehensiveness, coherence and decidability, and include correct data, incorrect data and boundary data in the test scenario.

[0023] As an optional solution to the GAN-based test data automatic generation method described in this application, in step 2, for the correct value class, the error value class, and the boundary value class in the training data, the preprocessing adopts a combination of random sampling and resampling to perform single-case sampling from the three classes, thereby obtaining a set of comprehensive triples: Triplet i =(t i ,f i ,m i) (7)

[0024] Where: Triplet i is the i-th triplet, and the same applies to the subsequent i letters; t is the abbreviation of true, which is the correct value class; f is the abbreviation of false, which is the error value class; m is the abbreviation of margin, which is the boundary value class; the overall formula means: the i-th triplet is composed of the i-th element in the correct value class, the error value class, and the boundary value class.

[0025] As an optional solution to the GAN-based test data automatic generation method described in this application, the generation network in step 3 continuously learns and strives to generate data consistent with the real data distribution to confuse the discriminator, and the discriminator network evolves its own discrimination ability so that it can correctly discriminate the real data and discriminate the generated data as wrong; the two continuously compete against each other and finally converge to the Nash equilibrium, so that the generation network can generate data with the same distribution as the real data.

[0026] As an optional solution to the GAN-based test data automatic generation method described in this application, step 3 further includes:

[0027] Step 3.1. Design and generate a network in the form of a codec. The upsampling and downsampling layers in the network are constructed using residual blocks to reduce information loss during the encoding and decoding process and gradient vanishing and performance degradation during training.

[0028] As an optional solution of the GAN-based test data automatic generation method described in this application, wherein: in step 3.1, for the input noise z~p that meets a certain probability distribution z (z), the encoder extracts the features of the input noise through the downsampling layer and obtains its features in the latent feature space m is the feature embedding in the latent space dimension The feature embedding passes through a common residual module containing a convolutional layer and an average pooling layer, and is then input into the decoder, which performs multi-layer upsampling decoding on it to obtain the generated test sample. When the generated data needs to be rounded, a rounding function can be used to round the generated data up or down with a probability of 1 / 2. The rounding function is calculated as:

[0029] Where: Ceil function is rounding up, floor function is rounding down, and p is the probability. The meaning of this formula is: when rounding the generated data, use the rounding function to round the generated data up or down with a probability of 1 / 2.

[0030] As an optional solution to the GAN-based test data automatic generation method described in this application, step 3 further includes:

[0031] Step 3.2: The task of the discriminant network is to distinguish between generated samples and real samples. The discriminant network consists of multiple downsampling residual modules and a fully connected layer. In order to make the performance of the discriminator more stable, spectral normalization is used to impose Lipschitz constraints on the weights of each layer of the discriminator to reduce the oscillation of the loss function and make the model converge faster and more stable.

[0032] As an optional solution to the GAN-based test data automatic generation method described in this application, wherein: in step 4, the least squares adversarial loss function is used to calculate the adversarial loss value. The purpose of the generative network in the generative adversarial network is to generate data that is consistent with the real data distribution, so that the discriminant network believes that the generated data is the real data, that is, the discriminant D(G(z))→1. The adversarial loss of the generative network is calculated as:

[0033] Where c represents the discrimination value of the generated data expected by the generated network, which is 1;

[0034] The purpose of the discriminant network is to distinguish between generated data and real data, that is, to discriminate D(x)→1, D(G(z))→0. The adversarial loss of the discriminant network is calculated as:

[0035] Where a and b represent the expected discrimination values ​​of the discriminant network for the generated samples and the real samples, which are 0 and 1 respectively; is the loss of the generated network; G(z) is the generated data of the generated network based on the input noise z;

[0036] D(G(z)) is the discriminant network's discriminant value of G(z). The discriminant value of D(G(z)) can be understood as a probability. D(G(z)) = 1 means that the discriminant network believes that the probability that the generated value G(z) of the generated network is a true value is 1, that is, the generated value G(z) of the generated network is real enough and consistent with the real sample distribution; if it is 0, the probability is 0 otherwise; c represents the discriminant network's expected discriminant value of the generated data, which is 1; that is, the generated network expects the value it generates to be very real, so that the discriminant network can judge its true probability to be 1; E is the expected calculation. The meaning of this formula is that the purpose of the generated network in the generative adversarial network is to generate data that is consistent with the real data distribution. Therefore, for the generated network, it expects the discriminant network to judge its generated value to be a true value with a probability of 1; so the loss function is designed. As the loss is reduced, that is The value decreases and approaches 0, D(G(z))-c approaches 0, that is, D(G(z)) approaches c, c is 1; x is the value in the real sample, D(x) is the probability that the discriminant network discriminates the real sample value as the true value; as the loss is reduced, Approaches 0, D(x)-b approaches 0, that is, D(x) approaches b, b is 1; D(G(z))-a approaches 0, D(G(z)) approaches a, a is 0.

[0037] As an optional solution to the GAN-based test data automatic generation method described in this application, the specific method of distributing the distance loss value in step 5 includes: denoting the probability distribution of the real data as X t ~N(μ t ,σ t ), the probability distribution of generated data is X g ~N(μ g ,σ g ), where (μ,σ) are the mean and standard deviation of the distribution respectively;

[0038] For the test task scenario, the requirement is not only to expect the generated data to meet the real data distribution, but also to expect the overall distribution of the generated data to be close to the real data distribution, that is, μ g →μ t ,σ g →σ t , the distribution distance loss value is specifically calculated as: L2=(μ t -μ g ) 2 +(σ t -σ g ) 2 (11).

[0039] As an optional solution of the GAN-based test data automatic generation method described in this application, the method of distributing the intersection-over-union loss value in step 6 includes: for the probability distribution X of the real data t ~N(μ t ,σ t ), the probability distribution X of the generated data g ~N(μ g ,σ g ), and its distribution intersection loss is calculated as

[0040] The distribution intersection loss further constrains the distribution distance loss, preventing it from being constrained by a single variable, so that the generated data can perfectly fit the real data distribution.

[0041] The beneficial effects of this application are: 1. The generative network and the discriminative network continuously compete against each other and finally converge to the Nash equilibrium, ensuring the overall stability of the GAN model.

[0042] 2. Simultaneously, the designed least squares adversarial loss, distribution distance loss, and distribution intersection-over-union loss jointly constrain and drive the GAN model to learn the distribution of real test data. This allows the GAN model to continuously learn adversarially and ultimately generate test data that matches real-world test scenarios, ensuring the correctness, comprehensiveness, consistency, and decidability of the generated data use cases.

[0043] 3. GAN models can be quickly built and deployed in the environment through the Pytorch platform. The trained GAN model can quickly and automatically generate test cases in real time, which can reduce the burden on testers and improve the efficiency of testing work.

[0044] 4. The learning and generation capabilities of the GAN model can be applied to various test scenarios, can meet the requirements of different test tasks, and have good reshaping and generalization. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:

[0046] Figure 1 is a schematic diagram of the upsampling residual module of the GAN-based test data automatic generation method of this application.

[0047] Figure 2 is a schematic diagram of a common residual module of the GAN-based test data automatic generation method of this application.

[0048] Figure 3 is a schematic diagram of the downsampling residual module of the GAN-based test data automatic generation method of this application.

[0049] FIG4 is a schematic diagram of the overall architecture of the GAN model of the GAN-based test data automatic generation method of this application.

[0050] FIG5 is a schematic diagram of the real data distribution of the GAN-based test data automatic generation method of this application.

[0051] FIG6 is a schematic diagram of data distribution generated by the GAN model of the GAN-based test data automatic generation method of the present application. DETAILED DESCRIPTION

[0052] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are described in detail below in conjunction with the drawings in the specification.

[0053] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0054] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present application. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0055] Furthermore, this application is described in detail with reference to schematic diagrams. For ease of illustration, when describing the embodiments of this application, cross-sectional views of device structures may be partially enlarged and not to scale. Furthermore, these schematic diagrams are merely illustrative and should not limit the scope of protection of this application. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.

[0056] A method for automatically generating test data based on GAN, comprising the following steps:

[0057] Step 1: Obtain training data that meets the actual test requirements for the required test tasks. Specifically, the training data should be correct, comprehensive, coherent, and determinable, and include correct data, incorrect data, and boundary data in the test scenario.

[0058] Step 2: Preprocessing the training data. In order to make GAN learn more balanced and comprehensive during the training process, so as to generate test cases that meet the requirements of real test tasks, this application uses triplet partitioning to preprocess the training data. Specifically, for the true value class, false value class, and margin value class in the training data, the preprocessing adopts a combination of random sampling and resampling to perform single-case sampling from the three classes, thereby obtaining a set of comprehensive triples: i =(t i ,f i ,m i ) (13)

[0059] Where: Triplet iis the i-th triplet, and the same applies to the subsequent i letters; t is the abbreviation of true, which is the correct value class; f is the abbreviation of false, which is the error value class; m is the abbreviation of margin, which is the boundary value class; the overall formula means: the i-th triplet consists of the i-th element of the correct value class, the error value class, and the boundary value class;

[0060] Step 3. Design and build a GAN model. The generative adversarial network (GAN) usually consists of a generator network and a discriminator network. The generator network is used to capture and learn the distribution of training data, and the discriminator network is used to judge the authenticity of the sample data, that is, to judge the probability that the sample data comes from the real training data. The generator network continuously learns and strives to generate data consistent with the distribution of real data to confuse the discriminator, while the discriminator network evolves its own discrimination ability as much as possible, so that it can correctly judge the real data and judge the generated data as wrong. The two constantly compete against each other and finally converge to the Nash equilibrium, so that the generator network can generate data with the same distribution as the real data.

[0061] Step 3.1. This application designs a generative network in the form of a codec. As shown in Figures 1-4, the upper and lower sampling layers in the network are constructed in the form of residual blocks to reduce information loss in the encoding and decoding process and gradient disappearance and performance degradation in the training process. Specifically, for input noise z~p that meets a certain probability distribution, z (z), the encoder extracts the features of the input noise through the downsampling layer and obtains the input noise in the latent feature space (m is the latent space dimension) feature embedding The feature embedding passes through a common residual module containing a convolutional layer and an average pooling layer, and is then input into the decoder, which performs multi-layer upsampling decoding on it to obtain the generated test sample. When the generated data needs to be rounded, a rounding function can be used to round the generated data up (down) with a probability of 1 / 2, where the rounding function is calculated as:

[0062] Where: Ceil function is for rounding up, floor function is for rounding down, and p is the probability. The meaning of this formula is: when rounding the generated data, use the rounding function to round the generated data up or down with a probability of 1 / 2.

[0063] Step 3.2: The task of the discriminant network is to distinguish between generated samples and real samples, which is essentially a classification task. Therefore, this application designs a discriminant network in the form of a classifier. As shown in Figures 1-4, the discriminant network consists of a Togo downsampling residual module and a fully connected layer. To make the performance of the discriminator more stable, spectral normalization is used here to impose a Lipschitz constraint on the weights of each layer of the discriminator to reduce the oscillation of the loss function and make the model converge faster and more stable. The discriminator finally outputs a scalar in [0,1] representing the probability that the input sample comes from real data;

[0064] Step 4: Calculate the adversarial loss value; assume that the data distribution of the real sample is p da t a , initialized to obey Gaussian distribution p z The random noise of (z) is z, and the generation network is generated by mapping G(z; θ g ) Transform the input random noise z to generate data distribution p g , discriminator D(x;θ d ) outputs a scalar [0,1] through the cross entropy loss function to represent the probability that the input sample comes from the real data. This application uses the least squares adversarial loss function to calculate the adversarial loss value. The goal of the generative adversarial network is to generate data that is consistent with the real data distribution as much as possible, so that the discriminant network believes that the generated data is the real data, that is, the discriminant D(G(z))→1. Therefore, the adversarial loss of the generative network is calculated as:

[0065] Where c represents the discrimination value of the generated data expected by the generated network, which is 1.

[0066] The purpose of the discriminant network is to distinguish the generated data from the real data as much as possible, that is, to discriminate D(x)→1, D(G(z))→0. Therefore, the adversarial loss of the discriminant network is calculated as:

[0067] Where a and b represent the expected discrimination values ​​of the discriminant network for the generated samples and the real samples, which are 0 and 1 respectively; is the loss of the generating network; G(z) is the generated data of the generating network according to the input noise z; D(G(z)) is the discriminant network's discriminant value of G(z), and the discriminant value of D(G(z)) can be understood as a probability. D(G(z)) = 1 means that the discriminant network believes that the probability that the generated value G(z) of the generating network is a true value is 1, that is, the generated value G(z) of the generating network is real enough and consistent with the real sample distribution; if it is 0, the probability is 0 otherwise; c represents the discriminant network's expected discriminant value of the generated data, which is 1; that is, the generating network expects the value it generates to be very real, so that the discriminant network can judge its true probability to be 1; E is the expected calculation. The meaning of this formula is that the purpose of the generating network in the generative adversarial network is to generate data that is consistent with the real data distribution. Therefore, for the generating network, it expects the discriminant network to judge its generated value to be a true value with a probability of 1; so the loss function is designed. As the loss is reduced, that is The value decreases and approaches 0, D(G(z))-c approaches 0, that is, D(G(z)) approaches c, c is 1; x is the value in the real sample, D(x) is the probability that the discriminant network discriminates the real sample value as the true value; as the loss is reduced, Approaches 0, D(x)-b approaches 0, that is, D(x) approaches b, b is 1; D(G(z))-a approaches 0, D(G(z)) approaches a, a is 0;

[0068] Step 5: Calculate the distribution distance loss value. The adversarial loss proposed in step 4 enables the generative network to generate test data that conforms to the data distribution of the real test scenario. However, it may generate data with mode collapse, that is, the generated data is relatively simple and limited to a small interval of the real distribution. This obviously cannot meet the requirements of the real test task. Therefore, this application proposes a distribution distance loss value. Specifically, let the probability distribution of the real data be X t ~N(μ t ,σ t ), the probability distribution of generated data is X g ~N(μ g ,σ g ), where (μ,σ) are the mean and standard deviation of the distribution respectively; for the test task scenario, the requirement is not only to expect the generated data to meet the real data distribution, but also to expect the overall distribution of the generated data to be close to the real data distribution, that is, μ g →μt,σ g →σt. The specific calculation of the distribution distance loss value is: L2=(μ t -μ g ) 2 +(σ t -σ g ) 2 (17)

[0069] By reducing the distribution distance loss, the generated data distribution is driven to continuously approach the real data distribution, making the generated data more comprehensive and coherent, ensuring that the generated data can meet the requirements of real test tasks.

[0070] Step 6: Calculate the distribution intersection loss. The loss values ​​designed in steps 4 and 5 make the generated data distribution have a good real data distribution. However, when reducing the distribution distance loss value obtained in step 5 during training, it may happen that only (μ t -μ g ) 2 Or just reduce (σ t -σ g ) 2 In order to further constrain the comprehensiveness and availability of generated data, this application further proposes a distribution intersection loss value. Specifically, for the probability distribution X of real data t ~N(μ t ,σ t ), the probability distribution X of the generated data g ~N(μ g ,σ g ), and its distribution intersection loss is calculated as

[0071] The distribution intersection loss further constrains the distribution distance loss, preventing it from being constrained by a single variable, allowing the generated data to perfectly fit the real data distribution.

[0072] Step 7: Calculate the total loss value of the model based on the loss function designed in steps 4, 5, and 6, reduce the loss through gradient descent, and continuously update the model parameters until the model converges and stabilizes.

[0073] Step 8. Input the noise data that conforms to the Gaussian distribution into the trained generative network, and the generative network automatically generates data that conforms to the real distribution. Figures 5 and 6 show the real data distribution and the data distribution generated by the GAN model, respectively. It can be seen that the data generated by the GAN model can well fit the real data distribution.

[0074] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application, and all of these should be included in the scope of the claims of the present application.

Claims

1. A method for automatically generating test data based on GAN, characterized in that: The following steps are involved: Step 1: Obtain training data that meets the actual test requirements for the required test tasks; Step 2: Preprocessing of training data: In order to make GAN learn more balanced and comprehensive during training, the training data is preprocessed by triple division. Step 3: Design and build a GAN model. The generative adversarial network consists of a generator network and a discriminator network. The generator network is used to capture and learn the distribution of training data, and the discriminator network is used to judge the authenticity of the sample data, that is, to judge the probability that the sample data comes from the real training data. Step 4: Calculate the adversarial loss value, assuming that the data distribution of the real sample is p data , initialized to follow the Gaussian distribution p z The random noise of (z) is z, and the generating network maps G(z; θ g ) transforms the input random noise z to generate data distribution p g , discriminator D(x;θ d ) The cross entropy loss function outputs a scalar [0,1] to represent the probability that the input sample comes from the real data; Step 5: Calculate the distribution distance loss value. The adversarial loss proposed in step 4 enables the generative network to generate test data that conforms to the data distribution of the real test scenario, but it may cause the generated data to collapse in mode, that is, the generated data is relatively single and limited to a small interval of the real distribution. In view of this situation, a distribution distance loss value is proposed. By reducing the distribution distance loss, the generated data distribution is driven to continuously approach the real data distribution, making the generated data more comprehensive and coherent, ensuring that the generated data can meet the requirements of the real test task. Step 6: Calculate the distribution intersection loss. The loss values ​​designed in steps 4 and 5 make the generated data distribution have a good real data distribution. However, when the distribution distance loss value obtained in step 5 is reduced during the training process, only (μ t -μ g ) 2 Or just reduce (σ t -σ g ) 2 ,In order to constrain the comprehensiveness and availability of generated data, a distribution intersection loss value is proposed; Step 7: Calculate the total loss value of the model according to the loss function designed in steps 4, 5, and 6, reduce the loss by gradient descent, and continuously update the model parameters until the model converges and stabilizes; Step 8: Input the noise data that conforms to the Gaussian distribution into the trained generative network, and the generative network will automatically generate rich test data, which will be used for testing in real test tasks.

2. The method for automatically generating test data based on GAN according to claim 1, characterized in that: The specific training data in step 1 should be correct, comprehensive, coherent and determinable, and include correct data, incorrect data and boundary data in the test scenario.

3. The method for automatically generating test data based on GAN according to claim 1, characterized in that: In step 2, for the correct value class, the error value class, and the boundary value class in the training data, the preprocessing adopts a combination of random sampling and resampling to perform single-case sampling from the three classes, thereby obtaining a set of comprehensive triples: i =(t i ,f i ,m i ) (1) Where: Triplet i is the i-th triplet, and the same applies to the subsequent i letters; t is the abbreviation of true, which is the correct value class; f is the abbreviation of false, which is the error value class; m is the abbreviation of margin, which is the boundary value class; the overall formula means: the i-th triplet is composed of the i-th element in the correct value class, the error value class and the boundary value class.

4. The method for automatically generating test data based on GAN according to claim 1, characterized in that: The generating network in step 3 continuously learns and strives to generate data consistent with the distribution of real data to confuse the discriminator, and the discriminator network evolves its own discriminative ability, so that it can correctly discriminate the real data and discriminate the generated data as wrong; The two continue to compete against each other and finally converge to the Nash equilibrium, so that the generative network can generate data with the same distribution as the real data.

5. The method for automatically generating test data based on GAN according to claim 4, characterized in that: The step 3 also includes: Step 3.

1. Design a generative network in the form of a codec. The up and down sampling layers in the network are constructed in the form of residual blocks to reduce information loss in the encoding and decoding process and gradient vanishing and performance degradation in the training process.

6. The method for automatically generating test data based on GAN according to claim 5, characterized in that: In step 3.1, for the input noise z~p that satisfies a certain probability distribution z (z), the encoder extracts the features of the input noise through the downsampling layer and obtains its features in the latent feature space m is the feature embedding in the latent space dimension The feature embedding passes through a common residual module containing a convolutional layer and an average pooling layer, and then is input into the decoder, which performs multi-layer upsampling decoding on it to obtain the generated test sample. When the generated data needs to be rounded, the rounding function can be used to round the generated data up / down with a probability of 1 / 2, where the rounding function is calculated as: Wherein: Ceil function is rounding up, floor function is rounding down, and p is probability; the meaning of this formula is: when rounding the generated data, the rounding function is used to round the generated data up or down with a probability of 1 / 2.

7. The method for automatically generating test data based on GAN according to claim 6, characterized in that: The step 3 also includes: Step 3.2: The task of the discriminant network is to distinguish between generated samples and real samples. The discriminant network consists of multiple down-sampling residual modules and a fully connected layer. In order to make the performance of the discriminator more stable, spectral normalization is used to impose Lipschitz constraints on each layer of the discriminator weights to reduce the oscillation of the loss function and make the model converge faster and more stable.

8. The method for automatically generating test data based on GAN according to claim 1, characterized in that: In step 4, the least squares adversarial loss function is used to calculate the adversarial loss value. The purpose of the generative network in the generative adversarial network is to generate data that is consistent with the distribution of the real data, so that the discriminant network believes that the generated data is the real data, that is, the discriminant D(G(z))→1. The adversarial loss of the generative network is calculated as: Where c represents the discriminant value of the generated data that the generated network expects the discriminant network to discriminate, and its value is 1; The purpose of the discriminant network is to distinguish between generated data and real data, that is, to discriminate D(x)→1, D(G(z))→0. The adversarial loss of the discriminant network is calculated as: Where a and b represent the expected discrimination values ​​of the discriminant network for the generated samples and the real samples, which are 0 and 1 respectively; is the loss of the generated network; G(z) is the generated data of the generated network according to the input noise z; D(G(z)) is the discriminant network's discriminant value of G(z). The discriminant value of D(G(z)) can be understood as a probability. D(G(z)) = 1 means that the discriminant network believes that the probability that the generated value G(z) of the generated network is a true value is 1, that is, the generated value G(z) of the generated network is real enough and consistent with the real sample distribution; if it is 0, the probability is 0 otherwise; c represents the discriminant network's expected discriminant value of the generated data, which is 1; that is, the generated network expects the generated value to be very real, and the discriminant network can judge its true probability to be 1; E is the expected calculation. The meaning of this formula is that the purpose of the generated network in the generative adversarial network is to generate data that is consistent with the real data distribution. Therefore, for the generated network, it expects the discriminant network to judge its generated value to be a true value with a probability of 1; so the loss function is designed, as the loss is reduced, that is The value decreases and approaches 0, D(G(z))-c approaches 0, that is, D(G(z)) approaches c, c is 1; x is the value in the real sample, D(x) is the probability that the discriminant network discriminates the real sample value as the true value; as the loss is reduced, Approaches 0, D(x)-b approaches 0, that is, D(x) approaches b, b is 1; D(G(z))-a approaches 0, D(G(z)) approaches a, a is 0.

9. The method for automatically generating test data based on GAN according to claim 1, characterized in that: The specific method of distributing the distance loss value in step 5 includes: denoting the probability distribution of the real data as X t ~N(μ t ,σ t ), the probability distribution of generated data is X g ~N(μ g ,σ g ), where (μ,σ) are the mean and standard deviation of the distribution respectively; For the test task scenario, the requirement is not only to expect the generated data to meet the real data distribution, but also to expect the overall distribution of the generated data to be close to the real data distribution, that is, μ g →μ t , σ g → t , the distribution distance loss value is specifically calculated as: L2=(μ t -m g ) 2 +(s t -s g ) 2 (5)。 10. The method for automatically generating test data based on GAN according to claim 1, characterized in that: The method for distributing the intersection-and-union loss value in step 6 includes: for the probability distribution X of the real data t ~N(μ t ,σ t ), the probability distribution X of the generated data g ~N(μ g ,σ g ), and its distribution intersection loss is calculated as The distribution intersection loss value further constrains the distribution distance loss value, avoiding it from falling into the constraint of a single variable, so that the generated data can perfectly fit the real data distribution.

Citation Information

Patent Citations

  • Super-resolution imaging method based on oral cavity CBCT reconstruction point cloud

    CN112184556A

  • Missile-borne image deblurring method based on generative adversarial network

    CN113947589A

  • Scale adaptive dense crowd counting method based on adversarial learning network

    CN114973112A

  • Automatic test data generation method based on GAN

    CN117827645A

  • Method for producing low calorie gummy jelly using Ainsliaea acerifolia extract powder

    KR1020210087736A

Cited By

  • Data-mechanism hybrid-driven mixed-flow manufacturing system performance evaluation method

    CN121187237A

  • Power transmission line project quality defect acceptance method based on generative adversarial and reinforcement learning

    CN121353736A