Data set amplification method and system

By alternating training of the generator and discriminator, perturbation samples are generated in the data and feature space, which solves the problem of insufficient generalization of existing dataset augmentation methods, achieves high-quality augmented sample sets, and improves the adaptability and generalization ability of the model.

CN120997062APending Publication Date: 2025-11-21STATE GRID HEBEI ELECTRIC POWER RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510933654.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing dataset augmentation methods rely on manually designed transformation rules, which are difficult to cover all potential distributions in the real world, have insufficient generalization, and may disrupt the natural manifold of the data distribution, introducing invalid or misleading samples.

Method used

By alternating training of the generator and discriminator, perturbation samples are generated in the data space and feature space. A high-quality augmented sample set is generated by adopting dynamic adversarial training and progressive hybrid strategies, thereby optimizing the manifold alignment of the data distribution.

Benefits of technology

It enhances the model's performance in small sample sizes, class imbalance, and complex scenarios, ensuring high-quality and highly adaptable generated samples, and improving the model's generalization ability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997062A_ABST
    Figure CN120997062A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data set amplification method and system. The method is applied to the technical field of data enhancement and comprises the step of generating disturbance samples of a data space and a feature space from a real sample set to enhance data diversity. And then, mapping the random noise into a training sample through a generator, discriminating the sample by means of a discriminator, and continuously improving the quality of the generated sample through alternate training of the generator and the discriminator until convergence. After training is completed, the generator generates an amplified sample set based on random noise. Then, determining a contraction starting point step number and a training step number in an alternate training process, and integrating a real sample, an amplification sample and a mixed sample into a multi-source enhanced data set according to a mixing formula; according to the scheme, data distribution asymptotic approximation is realized through the dynamic mixing coefficient, manifold alignment of the generated sample and the real sample is optimized, the performance of the model in small sample, class imbalance and complex scenes is enhanced, and high quality and high adaptability of the generated sample are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data augmentation, and particularly relates to a data set expansion method and system. BACKGROUND

[0002] Data set expansion is a key means to improve the generalization ability of a model. Due to the complex data distribution in real scenes and the high annotation cost, the original data set often has problems such as insufficient sample size, class imbalance, or limited scene coverage. Through expansion, the diversity of data (such as rotation, scaling, noise injection, etc.) can be simulated, the risk of overfitting can be effectively alleviated, and the robustness of the model to input changes can be enhanced. Especially in the fields of medical images and autonomous driving where annotation is scarce, expansion technology has become a core method to make up for data defects.

[0003] Nowadays, data set expansion is performed manually by preprocessing the original data (such as normalization, alignment), and then applying geometric transformations (rotation, flipping, cropping), photometric transformations (brightness adjustment, contrast enhancement), or noise injection (Gaussian noise, salt and pepper noise); then generating synthetic samples by combining multiple transformations, mixing augmented data with original data to form a training set. Some methods also introduce a manual annotation verification step to ensure the label accuracy of the transformed samples (such as adjusting the bounding box coordinates in the target detection task).

[0004] However, the transformation rules depend on manual design and are difficult to cover all potential distributions in the real world (such as extreme lighting or deformation); and they rely too much on domain priori and lack generalization (such as rotation expansion of medical images may destroy the rationality of anatomical structure); and the combined transformation may cause the data distribution to deviate from the natural manifold, introducing invalid or misleading samples. SUMMARY

[0005] To solve the problems of the prior art, the present disclosure provides a data set expansion method and system. The present disclosure solves the problem that the transformation rules of the current data set expansion method depend on manual design and are difficult to cover all potential distributions in the real world (such as extreme lighting or deformation); and they rely too much on domain priori and lack generalization (such as rotation expansion of medical images may destroy the rationality of anatomical structure); and the combined transformation may cause the data distribution to deviate from the natural manifold, introducing invalid or misleading samples.

[0006] According to a first aspect of the present disclosure, a data set expansion method is provided, comprising: obtaining a real sample set and a first label of the real sample set, generating a first perturbed sample in a data space according to the real sample set and the first label of the real sample set, and generating a second perturbed sample in a feature space according to the real sample set and the first label of the real sample set;

[0007] obtaining first random noise, mapping the first random noise to a training sample by using a generator, inputting the training sample into a discriminator to obtain a discrimination result;

[0008] performing alternating training of the generator and the discriminator according to the training sample and the discrimination result until the generator and the discriminator reach a preset first convergence condition, and generating a first augmented sample set according to the trained generator and the first random noise;

[0009] obtaining a contraction start step number and an alternating training step number of the alternating training process, and generating a mixed sample set according to the alternating training step number, the contraction start step number, the real sample set, the first augmented sample set, and a preset sample mixing formula;

[0010] generating a multi-source enhanced data set according to the real sample set, the first augmented sample set, and the mixed sample set.

[0011] According to a second aspect of the present disclosure, a data set augmentation system is provided for performing the method as described in the first aspect, comprising: a perturbed sample generation module for obtaining a real sample set and a first label of the real sample set, generating a first perturbed sample in a data space according to the real sample set and the first label of the real sample set, and generating a second perturbed sample in a feature space according to the real sample set and the first label of the real sample set;

[0012] a discrimination result generation module for obtaining first random noise, mapping the first random noise to a training sample by using a generator, inputting the training sample into a discriminator to obtain a discrimination result;

[0013] an alternating training module for performing alternating training of the generator and the discriminator according to the training sample and the discrimination result until the generator and the discriminator reach a preset first convergence condition, and generating a first augmented sample set according to the trained generator and the first random noise;

[0014] a mixed sample set generation module for obtaining a contraction start step number and an alternating training step number of the alternating training process, and generating a mixed sample set according to the alternating training step number, the contraction start step number, the real sample set, the first augmented sample set, and a preset sample mixing formula;

[0015] a sample augmentation module for generating a multi-source enhanced data set according to the real sample set, the first augmented sample set, and the mixed sample set.

[0016] According to a third aspect of the present disclosure, an electronic device is provided, comprising a memory and a processor, the memory having a computer program stored thereon, and the processor implementing the method as described above when executing the program.

[0017] In the data set augmentation method and system provided above, the embodiments of the present disclosure generate perturbation samples in data space and feature space, and realize data distribution progressive approximation through dynamic mixing coefficients based on the multi-modal feature space perturbation and progressive mixing strategy and mixed sample set of dynamic adversarial training, optimize the manifold alignment of generated samples and real samples, and enhance the performance of the model in small sample, class imbalance and complex scenarios, thereby ensuring high quality and strong adaptability of the generated samples. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor.

[0019] Figure 1 A data set augmentation method flowchart according to an embodiment of the present disclosure is shown;

[0020] Figure 2 A data set augmentation method flowchart according to an embodiment of the present disclosure is shown;

[0021] Figure 3 A data set augmentation system schematic block diagram according to an embodiment of the present disclosure is shown;

[0022] Figure 4 A block diagram of an exemplary electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0023] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. Note that the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.

[0024] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor indicate their logical order. It should also be understood that in the embodiments of the present disclosure, "multiple" can mean two or more, and "at least one" can mean one, two or more. It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, without explicit limitation or in the context of the preceding and following, it can be understood as one or more. In addition, the term "and / or" in the present disclosure is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the front and rear associated objects. It should also be understood that the description of various embodiments of the present disclosure emphasizes the differences between various embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, they will not be repeated.

[0025] It should be understood that the dimensions of the various parts shown in the drawings are not necessarily shown to scale. The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way limiting. Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but where appropriate, the techniques, methods, and devices should be considered as part of the specification. It should be noted that like reference numerals and letters refer to like items in the following drawings, and therefore, once an item is defined in one drawing, it need not be discussed further in subsequent drawings.

[0026] To make the purposes, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present disclosure.

[0027] Figure 1 A data set augmentation method flowchart is provided for the embodiments of the present disclosure. The method of the embodiments of the present disclosure aims to realize accurate detection of large and small targets of pictures.

[0028] S101, obtaining a real sample set and a first label thereof, generating a first perturbation sample in a data space according to the real sample set and the first label thereof, and generating a second perturbation sample in a feature space according to the real sample set and the first label thereof.

[0029] In a non-multimodal data scenario, the real sample set refers to single-modal data, such as numerical data: temperature, humidity, barometric pressure, and sensor data. Classification data: user behavior records, disease diagnosis labels. Time series data: stock prices, traffic data.

[0030] The first label of the real sample set can be the main label or target label corresponding to the sample, used for classification or regression tasks. Classification task: label is a category, such as "normal / abnormal" and "pass / reject". Regression task: label is a numerical value, such as temperature prediction value and load power.

[0031] The data space can be the representation space of the sample in the original dimension, i.e., the original numerical form of the data. In a non-multimodal scenario, the data space is a feature dimension space.

[0032] The first perturbed sample can be an augmented sample generated in the data space. Based on the real sample set and the first label, a random perturbation term is added or data transformation is performed to generate a new sample in the data space.

[0033] The feature space can be a high-dimensional space in which the sample is represented after feature extraction or mapping.

[0034] The second perturbed sample can be an augmented sample generated in the feature space. Based on the real sample set and the first label, a feature perturbation or interpolation is added to generate a new sample in the feature space.

[0035] Real sample data and its corresponding labels can be obtained from datasets or sensors, and data sources can include historical datasets in fields such as power, transportation, finance, medicine, etc. Real-time collected sensor data such as temperature, humidity, voltage, current, etc. User behavior data such as click rate, dwell time, conversion rate, etc. The sample set contains multi-dimensional features, usually represented in matrix form, with each row representing a sample and each column representing a feature. The label is the category or numerical value corresponding to the sample. Generating perturbed samples in data space means performing numerical perturbation on the original feature dimension. Perturbation methods include additive perturbation: adding random noise to the original sample data. The noise can be subject to Gaussian distribution or uniform distribution. By adding a random offset to each feature dimension, a new sample is generated. Scaling perturbation: scale the original sample by a certain proportion. The sample features are enlarged or reduced by a fixed proportion to generate new samples with changes. Specifically, each sample in the real sample set can be traversed. Apply additive perturbation or scaling perturbation to the sample features. Then generate the same number or more first perturbed samples as the original samples. Then map the real samples to the feature space: map the samples to the feature space through a pre-trained model (such as a deep neural network) to generate feature vectors. Or project the samples into the feature space through a dimension reduction algorithm (such as PCA). Perturbation methods include feature perturbation: adding small perturbations to the feature vector, such as adding noise or scaling features in the feature dimension. Interpolation perturbation: interpolate or interpolate with noise operations on sample vectors within the feature space to generate new samples. Specifically, the real sample set can be mapped to the feature space. Perturb the feature vectors within the feature space. Then generate the same number or more second perturbed samples as the original samples.

[0036] In S102, a first random noise is obtained, the first random noise is mapped to a training sample by using a generator, the training sample is input into a discriminator, and a discrimination result is obtained.

[0037] The first random noise can be a random variable input into a generator in a generative adversarial network (GAN).

[0038] The generator can be a model in the GAN, responsible for mapping the random noise to a fake sample.

[0039] The sample output by the generator is called a training sample, also known as a fake sample or generated sample. The training sample is similar to the real sample in the data space, but is actually a fake sample generated by the generator.

[0040] The discriminator is a model in the GAN used to distinguish between real samples and generated samples.

[0041] The discriminator outputs the probability that the sample is a real sample, with a value range of [0, 1]. The discrimination result is used to evaluate the generation ability of the generator.

[0042] The first random noise can be a vector randomly sampled from a specific distribution, usually a standard normal distribution or a uniform distribution. The generator is a neural network whose task is to take the first random noise as input and output a training sample after multiple layers of nonlinear transformation of the network. This training sample is a sample simulated by the generator according to the noise, aiming to be as close as possible to the distribution of the real sample. The specific process is: input the first random noise into the forward propagation network of the generator, and the generator generates a new training sample. The training sample output by the generator needs to be input into the discriminator for judgment. The discriminator is a neural network responsible for identifying whether the input sample is a real sample or a generated sample. The goal of the discriminator is to judge whether the input training sample comes from the real data set or is a pseudo sample generated by the generator. The discriminator inputs the training sample and outputs a probability value, which represents the probability that the sample training sample is a real sample, i.e., the discrimination result.

[0043] S103, according to the training sample and the discrimination result, the generator and the discriminator are alternately trained until the generator and the discriminator reach a preset first convergence condition, and a first augmented sample set is generated according to the trained generator according to the first random noise.

[0044] The preset first convergence condition can be a criterion for judging whether the training of the generator and the discriminator is completed. These conditions can include the following forms: convergence of the loss function: the loss function of the generator and the discriminator tends to be stable or no longer changes significantly in multiple training iterations. For example, the loss of the generator and the loss of the discriminator no longer fluctuate greatly, or their changes are less than a certain preset threshold. Similarity between generated samples and real samples: the performance of the generator can be evaluated by calculating the difference between generated samples and real samples (such as using a specific metric, such as Fréchet Inception Distance, FID). The convergence condition can be that the difference between the generated sample and the real sample is less than a certain threshold. Prediction accuracy of the discriminator: during training, if the discriminator becomes no longer significant in distinguishing between real samples and generated samples (for example, the output of the discriminator for generated samples is close to 0.5, indicating that it cannot distinguish between real and fake samples), it is considered to be converged. Specific training rounds: sometimes, the convergence condition is determined by a fixed number of training rounds or a specified maximum number of iterations.

[0045] The first augmented sample set can refer to a new sample set generated according to the ability of the generator after the training of the generator is completed.

[0046] First, the generator uses random noise as input to generate a batch of fake samples. These fake samples are used by the discriminator to determine whether they are real samples. The discriminator receives both real samples and generated fake samples and outputs its prediction of whether they are real or not. Then, based on the discriminator's output (i.e., its prediction of real or fake samples), the discriminator's weights are adjusted based on the error so that it can better distinguish between real and generated samples. The generator is updated based on the feedback from the discriminator. The goal of the generator is to make the discriminator unable to correctly distinguish between the generated fake samples and the real samples, so the generator adjusts its own weights based on the discriminator's judgment to generate more realistic samples. After the generator generates a batch of samples, the discriminator judges them, and the generator adjusts its parameters based on the discriminator's feedback. Then, new samples are generated by the generator, and the discriminator continues to classify. This process is repeated multiple times until the generator and discriminator reach the preset convergence condition. This convergence condition is usually that the loss function of the generator and discriminator converges, or the difference between the generated samples and the real samples becomes small enough. Once the training of the generator and discriminator reaches the convergence condition, i.e., the generator generates samples that can well simulate real data, and the discriminator cannot distinguish between real and fake samples, the training is complete. On this basis, the trained generator will use new random noise as input to generate an augmented sample set. This augmented sample set can contain new fake samples, which should be similar in features to real samples, and thus can be used to expand the original data set or for other tasks.

[0047] In S104, the contraction starting step number of the alternating training process and the alternating training step number are obtained, and a mixed sample set is generated according to the alternating training step number, the contraction starting step number, the real sample set, the first augmented sample set, and a preset sample mixing formula.

[0048] The contraction starting step number can be the time when the training strategy starts to be adjusted in the alternating training process. Generally, in adversarial training, the generator may generate low-quality samples in the early stage of alternating training of the generator and the discriminator. The contraction starting step number refers to the stage when the generator starts to be gradually adjusted and optimized, usually after the performance of the discriminator and the generator has reached a certain stage and the generator has been able to generate relatively realistic samples.

[0049] The alternating training step number can be the number of times the generator and the discriminator are alternately updated. Each training involves optimization steps of the generator and the discriminator, and the alternating training step number defines the duration of this training process. In each alternating training step, the generator improves the judgment of the discriminator by adjusting the generated samples, while the discriminator adjusts its parameters by discriminating between real samples and generated samples.

[0050] The mixed sample set can be a new sample set obtained by combining the real sample set and the first augmented sample set. In actual applications, the mixed sample set is usually used to improve the generalization ability of the model or further improve the performance of the model. According to a preset sample mixing formula, the mixed sample set will contain both real data and augmented samples generated by the generator.

[0051] The convergence starting point can be considered to have arrived when the loss function of the generator and the discriminator changes and the loss function reaches a certain preset threshold or shows signs of gradual stabilization. At this time, the performance of the discriminator and the generator is already good enough to start a more stable adversarial training phase. At this moment, the current training step number can be recorded as the contraction starting point step number. This step number marks the initial stable training phase of the generator and the discriminator, and subsequent training will focus more on optimization and fine-tuning. The alternating training step number can be dynamically adjusted according to the performance indicators of the generator and the discriminator, such as the quality of the generated samples and the discrimination accuracy. For example, after several alternating training, if the performance of the discriminator no longer improves significantly, the alternating training step number can be reduced or the training can be stopped. At this time, the current training step number can be obtained as the final alternating training step number. Then, the alternating training step number, the contraction starting point step number, the real sample set, and the first augmented sample set are substituted into the preset sample mixing formula to generate the mixed sample set.

[0052] On the basis of the above technical solutions, optionally, the preset sample mixing formula is:

[0053] X m =λ(k)·X r +(1-λ(k))·X g ;

[0054] Wherein, X m is the mixed sample set; λ(k) is the boundary mixing coefficient; X r is the real sample set; X g is the first augmented sample set.

[0055] Correspondingly, the calculation formula of λ(k) is:

[0056]

[0057] Wherein, γ is a preset contraction speed parameter; k is the alternating training step number; k0 is the contraction starting point step number.

[0058] In this scheme, a preliminary value of g is set before the training begins. The experiment can start from a small value (e.g., 0.01 or 0.05) and observe the training effect of the generator and the discriminator. By evaluating the performance of the generator and the discriminator, g is gradually adjusted so that the weighted proportion of generated samples and real samples gradually changes. For example, whether g needs to be adjusted can be determined by the performance of the validation set. After the experiment ends, the best g value is determined according to the final effect. Cross-validation or hyperparameter optimization techniques can be used to further fine-tune the adjustment.

[0059] If g is 0.1 and k0 is 100, k is 120.

[0060]

[0061] Calculate e -2 ≈0.1353, so:

[0062]

[0063] If the real sample X r = [1, 2, 3] mixed sample set X m = [4, 5, 6]

[0064] X m = 0.8808·[1, 2, 3] + (1-0.8808)·[4, 5, 6] = [1.400, 2.4096, 3.4192]

[0065] S105, according to the real sample set, the first augmented sample set and the mixed sample set, a multi-source augmented data set is generated.

[0066] The multi-source augmented data set can be expanded and enhanced by combining sample data from different sources (e.g., real sample set, first augmented sample set, mixed sample set, etc.), thereby improving the generalization ability and robustness of the model.

[0067] All data sets can be uniformly pre-processed to have the same scale, format and dimension in the feature space. For example, image data may need to be uniform in size, text data may need to be processed with word embedding, sensor data may need to be standardized, etc. Specifically, simple splicing (e.g., directly splicing data sets by rows or columns) can be used, or more complex techniques such as mixing generated samples and real samples in proportion by adversarial training, or using hybrid methods (such as weighted average, data interpolation, etc.) to combine data from different sources. Then combine data from different sources to form a new multi-source augmented data set. For example, assuming that each data source (real, augmented, mixed sample set) has different features and sample distribution, a more rich and diverse training set can be obtained by reasonable fusion. Each data source can be weighted or specific samples can be sampled according to actual needs.

[0068] In the embodiments of the application, the real sample set and the first label of the real sample set are obtained, the first perturbation sample is generated in the data space according to the real sample set and the first label of the real sample set, and the second perturbation sample is generated in the feature space according to the real sample set and the first label of the real sample set; the first random noise is obtained, the first random noise is mapped to a training sample by using a generator, the training sample is input into a discriminator to obtain a discrimination result; the generator and the discriminator are alternately trained according to the training sample and the discrimination result until the generator and the discriminator reach a preset first convergence condition, and the first augmented sample set is generated according to the first random noise by using the trained generator; the contraction start step number and the alternation training step number of the alternation training process are obtained, and the mixed sample set is generated according to the alternation training step number, the contraction start step number, the real sample set, the first augmented sample set and a preset sample mixing formula; and the multi-source augmented data set is generated according to the real sample set, the first augmented sample set and the mixed sample set. Through the above data set augmentation method, the perturbation sample is generated in the data space and the feature space, and the multi-modal feature space perturbation and the progressive mixing strategy based on the dynamic adversarial training and the mixed sample set are used to realize the data distribution progressive approximation by using the dynamic mixing coefficient, optimize the manifold alignment of the generated sample and the real sample, and enhance the performance of the model in the small sample, class imbalance and complex scene, so as to ensure that the generated sample has high quality and strong adaptability.

[0069] On the basis of the above technical solutions, after the multi-source augmented data set is generated according to the real sample set, the first augmented sample set and the mixed sample set, the method further includes:

[0070] According to the second label annotated by the trained discriminator, a pre-training model is obtained, and the pre-training model is trained according to the second label and the multi-source augmented dataset until the pre-training model reaches a preset model training standard.

[0071] In this scheme, the second label refers to the label generated when the multi-source augmented dataset is annotated based on the trained discriminator. It is usually used to improve the quality of the dataset or provide additional supervision information for the model.

[0072] The pre-training model can be a model that has not been trained and is usually used to initialize parameters. In this process, the pre-training model will be trained on the multi-source augmented dataset with the second label.

[0073] The preset model training standard can be a standard for judging whether the pre-training model reaches convergence or ideal performance, and usually includes the following indicators: accuracy (Accuracy): such as ≥95%, F1 score: such as ≥0.92, loss value (Loss): such as ≤0.01, convergence on the validation set: the loss value of the model on the validation set no longer decreases, generalization ability: stable performance on the test set or the validation set.

[0074] First, the multi-source augmented dataset is automatically annotated using the already trained discriminator model. The discriminator will generate a new label for each sample based on the features or categories of the data sample, called the second label. The second label can reflect more fine-grained information of the data, such as subcategories, attribute features, or the discriminator's confidence distribution. After annotation, each sample in the multi-source augmented dataset will have both the original label and the second label. Then a model that has not been trained yet, namely a pre-training model, is introduced. The model only contains randomly initialized or initialized based on existing model framework parameters, but has not been trained on the target dataset. The pre-training model will serve as the starting point for training, and will be gradually optimized through the multi-source augmented dataset and the second label. The pre-training model is trained using the multi-source augmented dataset with the second label. The features of the samples in the multi-source augmented dataset are used as the input of the model, and the second label is used as the training target of the model. During the training process: the model parameters are constantly adjusted to better fit the relationship between the features and the labels of the data. Through multiple rounds of training, the model's classification, recognition, or regression ability for samples is gradually improved. During the model training process, the model performance is continuously evaluated on the validation set or the test set. The model needs to reach the preset model training standard before stopping training.

[0075] In this scheme, after the model is trained on the multi-source augmented dataset, it can maintain stable performance on a wider range of data distribution, and the training efficiency and performance are significantly improved.

[0076] Figure 2A flowchart of a data set augmentation method provided by embodiments of the present disclosure. The method can include the following steps:

[0077] S201, input the first perturbed sample and the second perturbed sample into the trained discriminator to obtain a first output probability of the first perturbed sample and a second output probability of the second perturbed sample, and calculate a discriminator adversarial loss according to the alternating training step number, the contraction start step number, the first output probability, the second output probability, and a preset discriminator adversarial loss calculation formula.

[0078] The first output probability can be the probability value output by the discriminator after inputting the first perturbed sample into the trained discriminator. This probability represents the confidence of the discriminator that the first perturbed sample is a real sample. If the output of the discriminator is close to 1, it means that the discriminator is more inclined to judge the sample as a real sample. If the output is close to 0, it means that the discriminator is more inclined to judge the sample as a generated sample or a perturbed sample.

[0079] The second output probability can be the probability value output by the discriminator after inputting the second perturbed sample into the discriminator. This probability represents the confidence of the discriminator that the second perturbed sample is a real sample. Similarly, an output close to 1 indicates that the discriminator considers the sample to be more real, and an output close to 0 indicates that the discriminator considers the sample to be more like a generated sample.

[0080] The discriminator adversarial loss can be used to measure the performance of the discriminator in distinguishing real samples and generated samples (or perturbed samples) during adversarial training. When the discriminator can accurately distinguish real samples from generated samples, the adversarial loss is lower. If the discriminator cannot distinguish real samples from generated samples, the adversarial loss is higher. The discriminator adversarial loss is calculated based on the alternating training step number, the contraction start step number, the first output probability, the second output probability, and a preset loss calculation formula.

[0081] The first perturbed sample can be obtained from the data space, and the second perturbed sample can be obtained from the feature space. A discriminator that has been trained is used, which has learned to distinguish real samples from generated samples in previous alternating training. The first perturbed sample is sent into the discriminator model one by one or in batches. The discriminator performs feature extraction and judgment on each sample and outputs the corresponding confidence probability, i.e., the first output probability. The first output probability represents the probability that the discriminator considers the first perturbed sample to be a real sample. The second perturbed sample is input into the discriminator one by one or in batches. The discriminator discriminates the sample and outputs the second output probability. The second output probability represents the probability that the discriminator considers the second perturbed sample to be a real sample. After inputting the first perturbed sample and the second perturbed sample into the discriminator, the corresponding first output probability and second output probability are obtained, respectively. Then, the alternating training step number, the contraction start step number, the first output probability, the second output probability are substituted into the preset discriminator adversarial loss calculation formula to calculate the discriminator adversarial loss.

[0082] In the above technical solution, optionally, the preset discriminator adversarial loss calculation formula is:

[0083]

[0084] Wherein, L adv is the discriminator adversarial loss; a is a dynamic weight factor; E represents the mathematical expectation;

[0085] is the expectation of the loss of the first output probability for the first perturbed sample;

[0086] is the expectation of the loss of the second output probability for the second perturbed sample;

[0087] D(X d ) is the first output probability; D(F a ) is the second output probability;

[0088] Wherein, the calculation formula of a is:

[0089]

[0090] Wherein, β is a preset smoothing coefficient; k is the number of alternating training steps; k0 is the contraction starting step number.

[0091] In this scheme, the value of β generally depends on the experience of training or the setting in the past literature. Generally, experiments can be started from a relatively small value (such as 0.01 to 0.1), and gradually adjusted until a suitable value is found. When starting training, a smaller β value is used to allow the model to adjust the adversarial loss smoothly in the initial stage. As the training progresses, β can be gradually increased to force the model to adjust the weight between the generated sample and the adversarial sample more quickly, thereby accelerating the training process.

[0092] If β = 0.05, k is 100, and k0 is 50,

[0093]

[0094] D(X d ) is [0.2, 0.4, 0.3], D(F a ) is [0.7, 0.6], and for each X d , the loss is:

[0095] log(1-0.2) = log(0.8) ≈ -0.2231

[0096] log(1-0.4) = log(0.6) ≈ -0.5108

[0097] log(1-0.3) = log(0.7) ≈ -0.3567

[0098] For each F a , the loss is:

[0099] log(1-0.7) = log(0.3) ≈ -1.204

[0100] log(1-0.6) = log(0.4) ≈ -0.9163

[0101]

[0102] L adv = 0.925 * (-0.3635) + (1-0.925) * (-1.0602) = -0.4164

[0103] In GAN, the discriminator (D) outputs a probability value, usually between 0 and 1. For expressions such as log(1-D(X)), if the output of the discriminator is close to 1 (i.e. the discriminator considers the sample to be "real"), then 1-D(X) will be close to 0, causing the value of the logarithm to become very small (negative infinity). When the loss function is related to the logarithm, the result is usually negative. When the adversarial loss function returns a negative value, it means that the direction in which the generator is optimizing is to reduce the loss. Minimizing this negative value is the goal of the generator optimization, so that the generated samples are closer and closer to the real samples. In adversarial training, the generator is trained to minimize the logarithmic loss, so as to push the output of the discriminator on the generated samples close to 0. Finally, the value of the loss function tends to be negative, and the game between the generator and the discriminator gradually converges.

[0104] S202, mapping the real sample set to the feature space by a preset feature extraction function to obtain a first feature representation of the real sample set, and mapping the augmented sample set to the feature space by the preset feature extraction function to obtain a second feature representation of the augmented sample set.

[0105] The preset feature extraction function can be a function defined in advance for mapping a sample set to a feature space. These functions can convert the original sample set into a feature representation with fixed dimensions, thereby facilitating subsequent comparison, evaluation or model training. Specifically, numerical data can use PCA, LDA, statistical feature extraction. Categorical data can use one-hot encoding, target encoding. Time series data can use sliding window statistical feature extraction.

[0106] The first feature representation can be the feature representation obtained after mapping the real sample set to the feature space. Each real sample corresponds to a feature vector or a feature matrix in the feature space. The first feature representation reflects the distribution of the real sample set in the feature space. It retains the key features of the real sample, such as texture, shape, color distribution, pattern or statistical characteristics.

[0107] The second feature representation can be the feature representation obtained after mapping the augmented sample set to the feature space. Each augmented sample also corresponds to a feature vector or a feature matrix in the feature space. The second feature representation reflects the feature distribution of the augmented sample set in the feature space. It is used for comparison with the first feature representation to verify the similarity or difference between the augmented sample and the real sample in the feature space.

[0108] The preset feature extraction function can be selected or loaded. According to the data type of the sample set, select the appropriate feature extraction function: numerical sample: use dimension compression or statistical feature extraction method, principal component analysis (PCA): map high-dimensional numerical sample to low-dimensional feature space, while keeping the main trend of sample interval. Input: original sample set; output: sample representation in low-dimensional feature space, dimension m x km x k (k < n k < n k < n). Linear discriminant analysis (LDA): feature mapping through sample category information, maximize inter-class distance, minimize intra-class distance. Statistical feature extraction: calculate the statistical features of each sample, such as: mean, variance, skewness, kurtosis, maximum, minimum, etc., to form a feature vector. Category sample: use encoding mapping or discrete feature extraction method. One-hot encoding: map category sample to high-dimensional sparse vector. Target encoding: map category features to average or probability of target variable. Frequency encoding: feature representation based on category frequency distribution. Time series or event sequence: use time series feature extraction. Time series feature mapping: smoothing processing: smooth transformation of time series data, extract trend features. Sliding window statistics: calculate statistical features (such as mean, variance, rate of change) according to time window. Event sequence mapping: encode event sequence into feature vector (such as frequency or time interval based feature calculation). Then input the real sample set one by one or in batches into the feature extraction function. The feature extraction function maps each sample and outputs the feature representation. After mapping all real samples to the feature space, the first feature representation is generated. The first feature representation is a matrix, each row corresponds to a sample vector representation in the feature space. Dimension example: if the original sample dimension is m x n (m is the number of samples, n is the original feature number), after PCA mapping to 10-dimensional feature space, the first feature representation dimension is m x 10. In order to ensure the consistency of the feature space, use the same feature extraction function as the real sample set. This ensures that the augmented samples are compared with the real samples in the same dimensional feature space. Then input the augmented sample one by one or in batches into the feature extraction function. The feature extractor maps each augmented sample to generate a feature vector or feature matrix. After mapping all augmented samples to the feature space, the second feature representation is generated. The second feature representation has the same dimension as the first feature representation, which is used for subsequent feature comparison.

[0109] S203, according to the first feature representation, the second feature representation and the preset feature consistency loss calculation formula, calculate the feature consistency loss of the real sample set and the augmented sample set in the feature space.

[0110] The feature consistency loss can be a loss function that measures the similarity or difference between the real sample set and the augmented sample set in the feature space. It compares the distance or distribution difference of the two groups of samples in the feature space to evaluate the similarity of the augmented samples and the real samples.

[0111] The first feature representation and the second feature representation can be substituted into the preset feature consistency loss calculation formula to obtain the feature consistency loss of the real sample set and the augmented sample set in the feature space.

[0112] On the basis of the above technical solution, optionally, the preset feature consistency loss calculation formula is:

[0113]

[0114] wherein, L feat is the feature consistency loss; f feat (X r ) is the first feature representation; f feat (X g ) is the second feature representation; f feat () is a preset feature extraction function.

[0115] In the present solution, if X r = [X1, X2, X3], f feat (X r ) = [f1, f2, f3], X g =

[0116] [X4, X5, X6], f feat (X g ) = [f4, f5, f6]

[0117] If f1 = 2, f2 = 3, f3 = 4, f4 = 3, f5 = 2, f6 = 5

[0118] For the first sample pair X1 and X4:

[0119]

[0120] For the second sample pair X2 and X5:

[0121]

[0122] For the third sample pair X3 and X6:

[0123]

[0124] S204, according to the adversarial loss, the feature consistency loss, the real sample set and the first augmented sample set, the generator and the discriminator are optimized until the generator and the discriminator reach the preset second convergence condition.

[0125] The preset second convergence condition can be an optimization stopping criterion satisfied by the discriminator and the generator when trained to a certain stage. Compared with the initial training stage, the second convergence condition is usually more stringent, aiming to ensure that the performance of the generator and the discriminator reaches a higher level, the generated samples are more consistent with the real samples, and the generator will not fail due to overfitting or mode collapse. For example, it can be that the adversarial loss is stable and close to balance: when the adversarial loss of the generator and the discriminator converges to a stable range and has small fluctuations within multiple iterations, it means that the two have reached an adversarial balance. The discrimination accuracy of the discriminator on the real samples and the generated samples is close to 50% (i.e. unable to distinguish between true and false), indicating that the generated samples are already realistic enough. The game between the discriminator and the generator enters a relatively balanced state, and there is no longer obvious performance fluctuations.

[0126] The generator optimization uses the adversarial loss: the goal of the generator is to deceive the discriminator so that it cannot accurately distinguish between generated samples and real samples. The feature consistency loss is used: by comparing the distance between the generated samples and the real samples in the feature space, it is ensured that the generated samples are consistent with the real samples in the feature distribution. Using the real sample set and the first augmented sample set as a reference, the generator is guided to generate more realistic samples through feature comparison. The random gradient descent or adaptive moment estimation algorithm is used to update the weights of the generator to minimize the loss function. The discriminator optimization uses the adversarial loss: the discriminator improves the discrimination accuracy by calculating the binary classification error of the real samples and the generated samples. The feature consistency loss is used: by calculating the difference between the generated samples and the real samples in the feature space, the discriminator can more accurately identify the generated samples. The real sample set and the first augmented sample set are input into the discriminator for training at the same time, ensuring that the discriminator can maintain consistent judgment between the real samples and the augmented samples. The discriminator parameters are continuously updated to maintain a stable discrimination ability between the generated sample set and the real sample set. In the optimization process of the generator and the discriminator, the parameters of the two are alternately updated: the generator training includes generating new samples and inputting them into the discriminator. The generator parameters are optimized so that the generated samples can deceive the discriminator. The discriminator training includes inputting the generated samples and the real samples into the discriminator at the same time. The discriminator parameters are optimized to improve the discrimination accuracy. In the optimization process of the generator and the discriminator, the feature consistency loss is introduced to ensure that the generated samples are consistent with the real samples in the feature distribution. Then the generator and the discriminator are alternately updated until the preset second convergence condition is reached.

[0127] S205, generating a second augmented sample set according to the generator whose optimization is completed, and updating the multi-source enhanced data set according to the second augmented sample set.

[0128] The second augmented sample set can refer to a batch of samples generated by the generator based on new random noise after the generator and the discriminator are jointly optimized. This batch of samples is similar to the first augmented sample set, but since the generator is optimized, the sample quality and diversity are higher, and are closer to the real sample distribution. It has stronger feature consistency and is closer to the real sample in the feature space. The multi-source augmented dataset can be further expanded to improve the coverage and expressiveness of the dataset.

[0129] The first random noise can be input to the optimized generator. The generator maps the noise to new samples, which have a distribution closer to the real sample set while retaining a certain diversity. The generated samples are collected as the second augmented sample set. The second augmented sample set is merged with the existing samples in the multi-source augmented dataset. When merging, the label information and data distribution of the sample set are kept consistent to ensure that the augmented dataset does not become unbalanced due to the addition of generated samples. The merged dataset is quality-screened to remove abnormal or low-quality samples. The augmented samples are quality-evaluated based on the discriminator or pre-set evaluation indicators to remove unqualified samples and ensure the quality of the dataset. All samples in the multi-source augmented dataset are remapped to the feature space to ensure that the generated samples are consistent with the real samples in the feature space. The feature consistency loss is calculated to evaluate the distance between the generated samples and the real samples, and to further verify the overall consistency of the dataset. The cleaned and feature-aligned samples are added to the multi-source augmented dataset. The multi-source augmented dataset is further expanded in size and diversity while maintaining consistency with the real sample set and similarity in feature distribution.

[0130] In this embodiment, through the joint optimization of the adversarial loss and the feature consistency loss, the generator can gradually generate samples that are more similar to the real samples. The optimized generator not only generates higher-quality samples, but also avoids generating mode collapse, ensuring that the generated samples have diversity and representativeness. By continuously generating and optimizing samples, the quality and diversity of the augmented sample set are improved. The new augmented samples better capture the features of real data.

[0131] On the basis of the above technical solutions, optionally, after updating the multi-source augmented dataset according to the second augmented sample set, the method further includes:

[0132] determining the local density data of the first augmented sample set, and determining the perturbation coefficient according to the local density data;

[0133] calculating the Gaussian distribution variance of the first augmented sample set, and calculating the perturbation noise of the first augmented sample set according to the perturbation coefficient, the Gaussian distribution variance, and a pre-set perturbation noise generation formula; wherein the pre-set noise generation formula is:

[0134] Z i = ∈i • N(0, σ 2 );

[0135] wherein Z i is noise data; ∈ i is a perturbation coefficient; σ 2 is a Gaussian distribution variance; N(0, σ 2 ) is random noise with a Gaussian distribution variance;

[0136] obtaining a first random sample and a second random sample of a first augmented sample set, obtaining an interpolated enhanced data set according to the first random sample, the second random sample, noise data, and a preset Gaussian noise interpolation formula, and updating a multi-source enhanced data set according to the interpolated enhanced data set; wherein the preset Gaussian noise interpolation formula is:

[0137] X inter = λ1·X1+ (1- λ1)·X2+ Z i ;

[0138] wherein X inter is the interpolated enhanced data set; λ1 is a preset interpolation coefficient; X1 is the first random sample; X2 is the second random sample; and Z i is perturbation noise.

[0139] In the scheme, the local density data can be distribution density information of the sample in the feature space, i.e., the number or distribution of samples in the neighborhood of a certain sample.

[0140] The perturbation coefficient can be a scale factor for adjusting the amplitude of noise during data enhancement, which is dynamically generated according to the local density data.

[0141] The Gaussian distribution variance can be the fluctuation amplitude of the sample in the feature space subject to Gaussian distribution, reflecting the dispersion degree of the sample distribution. The greater the variance, the more dispersed the data distribution; the smaller the variance, the more concentrated the data distribution.

[0142] The perturbation noise can be random noise generated based on the perturbation coefficient, the Gaussian distribution variance, and the noise generation formula in the data enhancement process. It usually exists in the form of Gaussian noise or uniform noise.

[0143] The first random sample and the second random sample can be two groups of samples randomly selected from the first augmented sample set.

[0144] The interpolation enhanced dataset can be a new sample set generated by random samples and noise interpolation. The local density data measures the density of the sample distribution in the feature space. Common methods include: K-neighbor density estimation: the density is calculated based on the distance between the sample and its nearest neighbor sample. Kernel density estimation: the kernel function is used to estimate the probability density of the sample distribution. Local outlier factor: the density is compared by comparing the density difference between the sample and its neighborhood.

[0145] First, the distance between any two points in the sample set is calculated. Common distance measurement methods include Euclidean distance, Manhattan distance, or cosine similarity. Using the K-neighbor algorithm, for each sample point, find the K samples closest to it. The local density value is calculated by accumulating the distance between the sample and the K neighbors. The larger the K value, the smoother the density calculation. The larger the local density value, the higher the sample density in the data-intensive area; the smaller the density value, the lower the sample density in the sparse area. For each sample, based on the difference between its local density and average density, the perturbation coefficient is dynamically calculated. When the sample density is higher than the average density, the perturbation coefficient is close to 1, indicating that the generated sample has less perturbation; when the sample density is lower than the average density, the perturbation coefficient tends to 0, indicating that the generated sample has more perturbation. Then calculate the mean of the augmented sample set. Square the deviation between each sample and the mean. Average all deviation squares to get the Gaussian distribution variance of the augmented sample set. The larger the Gaussian distribution variance, the more dispersed the sample distribution in the feature space; the smaller the Gaussian distribution variance, the more concentrated the sample distribution. Then substitute the perturbation coefficient and the Gaussian distribution variance into the preset perturbation noise generation formula to calculate the perturbation noise of the first augmented sample set. The first random sample and the second random sample are two sample points randomly selected from the first augmented sample set. Random selection can be achieved by the following methods: randomly selecting two groups of samples from the dataset without replacement. Use the k-neighbor method to select two groups of adjacent samples. Sample the dense and sparse areas separately to maintain the diversity of the data distribution. Then substitute the first random sample, the second random sample, and the noise data into the preset Gaussian noise interpolation formula to obtain the interpolation enhanced dataset. Finally, the generated interpolation enhanced dataset is added to the multi-source enhanced dataset.

[0146] In this scheme, by reasonably using perturbation noise, local density data and Gaussian noise interpolation method, a diversified dataset can be generated.

[0147] On the basis of the above technical scheme, optionally, after updating the multi-source enhanced dataset according to the interpolation enhanced dataset, the method further comprises:

[0148] Obtain the image dataset, the text dataset, and the sensor dataset, and determine the image features of the image dataset;

[0149] determine text features of the text dataset, and determine sensor features of the sensor dataset;

[0150] concatenate the image features, the text features, and the sensor features to obtain a multi-modal feature space;

[0151] generate second random noise of the multi-modal feature space by using a standard normal distribution, input the second random noise and the multi-modal feature space into the optimized generator to obtain a third augmented sample set, and update the multi-source enhanced dataset according to the third augmented sample set.

[0152] In the scheme, the image dataset can be a data collection containing a large number of image samples, used for extracting visual features. For example, it can include: face recognition: CelebA, LFW, object detection: COCO, PascalVOC, medical imaging: ChestX-ray, ISIC skin cancer dataset, industrial scene: equipment failure image, workpiece defect detection image, etc.

[0153] The text dataset can be a data collection containing a large number of text samples, used for extracting semantic, sentiment or keyword features. For example, it can include news text: AGNews, THUCNews, sentiment analysis: IMDB movie reviews, SST-2, industrial text: equipment alarm logs, maintenance records, operation manual text.

[0154] The sensor dataset can be a time series or environmental state data collection collected by a sensor, commonly used for device monitoring or state evaluation. For example, it can include industrial Internet of Things: temperature, humidity, current, voltage, vibration sensor data, traffic monitoring: accelerometer, GPS position, vehicle speed, lane deviation data, medical monitoring: heart rate, blood pressure, blood oxygen, EEG data, smart home: smoke, infrared, door magnetic, temperature, humidity sensor data.

[0155] Image features can refer to quantifiable descriptions extracted from image data, used to represent visual information of images, including low-order features: color histogram, RGB mean value, gray value distribution, edge detection features (such as Sobel, Canny), texture features (such as LBP, HOG); high-order features: CNN feature map (such as ResNet, VGG extracted deep features), SIFT, SURF, ORB key point features, semantic features: visual feature vectors extracted by pre-trained models (such as CLIP); Feature representation of image features is usually represented as a high-dimensional vector, for example: ResNet50 output feature dimension is 2048, VGG16 output feature dimension is 4096.

[0156] Text features can refer to the numerical representation extracted from text data to represent the semantic or structural information of the text, including lexical-level features: N-gram features: unigram / bigram / trigram statistics, term frequency-inverse document frequency (TF-IDF); semantic features: word embeddings: word vectors generated by Word2Vec, GloVe, FastText, contextual encoding features: feature vectors extracted by BERT, RoBERTa, etc. Transformer model; sentiment and topic features: sentiment scores: positive, negative, neutral scores, topic distribution: topic distribution features generated by LDA model; feature representation: text features are usually represented as vectors or matrices: TF-IDF feature dimension may be 1000-5000 dimensions, BERT text feature dimension is 768 dimensions, and after multi-document splicing, a [sample number x feature dimension] matrix is formed.

[0157] Sensor features can be features extracted from sensor data that reflect physical states or trends, including time domain features: mean, variance, skewness, kurtosis, peak-to-peak value, root mean square value (RMS), time series trend features (such as autocorrelation coefficients); frequency domain features: fast Fourier transform (FFT) features: frequency distribution, energy spectrum, wavelet transform coefficients: decomposition features of signals at multiple scales of frequency; statistical features: maximum value, minimum value, median, signal energy, instantaneous amplitude, periodicity; feature representation: sensor features are usually represented as time series matrices: [time step x feature dimension] represent single-channel sensor features, and after multi-sensor splicing, a [sample number x feature dimension] matrix is formed.

[0158] Multi-modal feature space can be a high-dimensional space formed by splicing image, text and sensor features together, used to represent the fusion representation of different modal features.

[0159] Second random noise can refer to a random vector or matrix generated from a standard normal distribution (mean 0, variance 1), which is used to mix with the multi-modal feature space input into the generator to enhance sample diversity.

[0160] The third expanded sample set can refer to the new sample set generated by inputting the multi-modal feature space and random noise into the optimized generator. These samples belong to augmented data and are used to enhance the model training data set.

[0161] Existing datasets can be downloaded from public data platforms, comments, titles, or articles can be scraped from social platforms, forums, or news websites to generate text datasets. Real-time sensor data can be collected using IoT devices or embedded devices such as Raspberry Pi to generate sensor datasets. Then use pre-trained deep learning models to extract features, such as: ResNet, VGG, EfficientNet, etc. Input the image into the model and extract the feature vector through the intermediate layer. Image features are usually high-dimensional vectors of fixed length (such as 512 dimensions or 1024 dimensions). Perform word vector representation on the text: use pre-trained models (such as Word2Vec, GloVe, BERT) to convert text into vectors. Extraction method: map each word or sentence to a fixed-dimensional vector. For example, BERT maps a sentence to a 768-dimensional feature vector. Time and frequency domain feature extraction: time domain features: mean, variance, standard deviation, skewness, kurtosis. Frequency domain features: extract frequency features using Fast Fourier Transform (FFT). Feature vector representation: divide the sensor data into intervals within a time window and calculate the feature vector. The feature dimension is determined by the sampling frequency and window size, such as 256 dimensions or 512 dimensions. Concatenate image, text, and sensor features into a multi-dimensional vector. Concatenation rule: concatenate feature vectors of different modalities in order of dimension. For example: image features are 512-dimensional. Text features are 768-dimensional. Sensor features are 256-dimensional. After concatenation, a 1536-dimensional multi-modal feature vector is formed. Then standardize the features of different modalities to make their scales consistent, using batch normalization or regularization. Then map the concatenated multi-modal features to a unified feature space. The dimension of the multi-modal feature space is equal to the dimension of the concatenated features. Use dimension reduction algorithms (t-SNE, PCA) to visualize the multi-modal feature space. Samples of different categories show different distributions in the feature space. Then generate random noise using standard normal distribution: noise follows a normal distribution with mean 0 and standard deviation 1. The dimension of the noise is the same as the dimension of the multi-modal feature space. For example, the dimension of the multi-modal feature is 1536, and the dimension of the noise is also 1536. Input the second random noise and the multi-modal feature space into the optimized generator to generate a third augmented sample set. Finally, add the third augmented sample set to the multi-source augmented dataset.

[0162] In this scheme, multi-modal data augmentation significantly improves the diversity of data by generating new augmented samples, effectively alleviating the small sample problem.

[0163] Figure 3 A data set expansion system schematic diagram is provided for the embodiments of the present disclosure. The system comprises:

[0164] The disturbance sample generation module 301 is configured to obtain a real sample set and a first label of the real sample set, generate a first disturbance sample in a data space according to the real sample set and the first label of the real sample set, and generate a second disturbance sample in a feature space according to the real sample set and the first label of the real sample set.

[0165] The discrimination result generation module 302 is configured to obtain a first random noise, map the first random noise to a training sample by using a generator, input the training sample into a discriminator to obtain a discrimination result.

[0166] The alternating training module 303 is configured to perform alternating training of the generator and the discriminator according to the training sample and the discrimination result until the generator and the discriminator reach a preset first convergence condition, and generate a first augmented sample set according to the generator after the training.

[0167] The mixed sample set generation module 304 is configured to obtain a contraction start step and an alternating training step of an alternating training process, and generate a mixed sample set according to the alternating training step, the contraction start step, the real sample set, the first augmented sample set, and a preset sample mixing formula.

[0168] The sample augmentation module 305 is configured to generate a multi-source enhanced data set according to the real sample set, the first augmented sample set, and the mixed sample set.

[0169] Figure 4 A schematic block diagram of an electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0170] The electronic device 400 includes a computing unit 401 that can perform various appropriate actions and processes according to a computer program stored in a ROM 402 or a computer program loaded into a RAM 403 from a storage unit 408. Various programs and data required for the operation of the electronic device 400 can also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An I / O interface 405 is also connected to the bus 404.

[0171] A plurality of components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0172] The computing unit 401 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs various methods and processes described above, such as the one data set augmentation method. For example, in some embodiments, the one data set augmentation method can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded onto the RAM 403 and executed by the computing unit 401, one or more steps of the one data set augmentation method described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the one data set augmentation method by any other appropriate means, such as by means of firmware.

[0173] Various implementations of the systems and techniques described above herein can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0174] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0175] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0176] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0177] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0178] The computer system can include clients and servers. This relationship can be

[0179] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the disclosure can be executed in parallel, executed in series, or executed in different orders, as long as the desired results of the technology disclosed in the present disclosure are achieved, and the present disclosure is not limited herein.

[0180] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and scope of the disclosure. Any modifications, equivalent substitutions, improvements, and the like, made within the spirit and principles of the disclosure, should be included in the scope of the disclosure.​​​

Claims

1. A method of dataset augmentation, the method comprising: The method comprises: obtaining a real sample set and a first label thereof, generating first perturbed samples in a data space according to the real sample set and the first label thereof, and generating second perturbed samples in a feature space according to the real sample set and the first label thereof; obtaining first random noise, mapping the first random noise to training samples by using a generator, inputting the training samples into a discriminator to obtain a discrimination result; performing alternating training of the generator and the discriminator according to the training samples and the discrimination result until the generator and the discriminator reach a preset first convergence condition, generating a first augmented sample set according to the generator after training according to the first random noise; obtaining a contraction start step number and an alternating training step number of the alternating training process, generating a mixed sample set according to the alternating training step number, the contraction start step number, the real sample set, the first augmented sample set, and a preset sample mixing formula; and generating a multi-source enhanced data set according to the real sample set, the first augmented sample set, and the mixed sample set.

2. The method of claim 1, wherein, Wherein, After the multi-source enhanced data set is generated according to the real sample set, the first augmented sample set, and the mixed sample set, the method further comprises: inputting the first perturbed samples and the second perturbed samples into the trained discriminator to obtain a first output probability of the first perturbed samples and a second output probability of the second perturbed samples, calculating a discriminator adversarial loss according to the alternating training step number, the contraction start step number, the first output probability, the second output probability, and a preset discriminator adversarial loss calculation formula; mapping the real sample set to the feature space by using a preset feature extraction function to obtain a first feature representation of the real sample set, and mapping the augmented sample set to the feature space by using the preset feature extraction function to obtain a second feature representation of the augmented sample set; calculating a feature consistency loss of the real sample set and the augmented sample set in the feature space according to the first feature representation, the second feature representation, and a preset feature consistency loss calculation formula; optimizing the generator and the discriminator according to the adversarial loss, the feature consistency loss, the real sample set, and the first augmented sample set until the generator and the discriminator reach a preset second convergence condition; generating a second augmented sample set according to the first random noise by using the optimized generator, and updating the multi-source enhanced data set according to the second augmented sample set.

3. The method of claim 1, wherein, Wherein, the preset sample mixing formula is: X m = λ(k) · X r + (1 - λ(k)) · X g ; where X m is the mixed sample set; λ(k) is the boundary mixing coefficient; X r is the real sample set; X g is the first augmented sample set; correspondingly, the calculation formula of λ(k) is: wherein γ is a preset contraction speed parameter; k is the alternating training step number; k0 is the contraction start step number.

4. The method of claim 1, wherein, Wherein, the preset discriminator adversarial loss calculation formula is: Wherein, Ladv is the discriminator adversarial loss; a is a dynamic weight factor; denotes the mathematical expectation; taking an expectation over the loss for the first output probability of the first perturbed sample; to take the expectation over the second loss of the second output probability of the second perturbed sample; D(X d ) is the first output probability; D(F a ) is the second output probability; wherein the calculation formula of α is: wherein β is a preset smoothing coefficient; k is the alternating training step number; k0 is the contraction start step number.

5. The method of claim 2, wherein, Wherein, the preset feature consistency loss calculation formula is: wherein L feat is a feature consistency loss; f feat (X r ) is a first feature representation; f feat (X g ) is a second feature representation; f feat () is a pre-set feature extraction function.

6. The method of claim 1, wherein, wherein After the multi-source enhanced data set is generated according to the real sample set, the first augmented sample set, and the mixed sample set, the method further comprises: annotating a second label of the multi-source enhanced data set by using the trained discriminator, obtaining a pre-trained model, training the pre-trained model according to the second label and the multi-source enhanced data set until the pre-trained model reaches a preset model training standard.

7. The method of claim 2, wherein, Wherein, After updating the multi-source augmented dataset according to the second augmented sample set, the method further comprises: determining local density data of the first augmented sample set, determining a perturbation coefficient according to the local density data; calculating a Gaussian distribution variance of the first augmented sample set, and calculating a perturbation noise of the first augmented sample set according to the perturbation coefficient, the Gaussian distribution variance, and a preset perturbation noise generation formula; wherein the preset noise generation formula is: Z i = ε i · N(0, σ 2 ); where Z i is noise data; ∈ i is a perturbation coefficient; σ 2 is a Gaussian distribution variance; N(0, σ 2 ) is random noise with Gaussian distribution variance; obtaining a first random sample and a second random sample of the first augmented sample set, obtaining an interpolation augmented dataset according to the first random sample, the second random sample, the noise data, and a preset Gaussian noise interpolation formula, and updating the multi-source augmented dataset according to the interpolation augmented dataset; wherein the preset Gaussian noise interpolation formula is: X inter = λ1 · X1+ (1 - λ1) · X2+ Z i ; wherein X inter is an interpolated enhanced dataset; λ1 is a preset interpolation coefficient; X1 is a first random sample; X2 is a second random sample; Z i is a disturbance noise.

8. The method of claim 7, wherein, wherein, After updating the multi-source augmented dataset according to the interpolation augmented dataset, the method further comprises: obtaining an image dataset, a text dataset, and a sensor dataset, and determining image features of the image dataset; determining text features of the text dataset, and determining sensor features of the sensor dataset; concatenating the image features, the text features, and the sensor features to obtain a multi-modal feature space; generating a second random noise of the multi-modal feature space using a standard normal distribution, inputting the second random noise and the multi-modal feature space into the optimized generator to obtain a third augmented sample set, and updating the multi-source augmented dataset according to the third augmented sample set.

9. A data set augmentation system for performing the method of any one of claims 1-8, characterized by, The system comprises: a perturbation sample generation module configured to obtain a real sample set and a first label thereof, generate a first perturbation sample in a data space according to the real sample set and the first label thereof, and generate a second perturbation sample in a feature space according to the real sample set and the first label thereof; a discrimination result generation module configured to obtain a first random noise, map the first random noise to a training sample using a generator, input the training sample into a discriminator to obtain a discrimination result; an alternating training module configured to perform alternating training of the generator and the discriminator according to the training sample and the discrimination result until the generator and the discriminator reach a preset first convergence condition, and generate a first augmented sample set according to the first random noise using the trained generator; a mixed sample set generation module configured to obtain a contraction start step number and an alternating training step number of an alternating training process, generate a mixed sample set according to the alternating training step number, the contraction start step number, the real sample set, the first augmented sample set, and a preset sample mixing formula; a sample augmentation module configured to generate a multi-source augmented dataset according to the real sample set, the first augmented sample set, and the mixed sample set.

10. An electronic device comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.