Method and apparatus for augmenting training data or retraining neural networks

Through the combination of neural networks and classifiers, analysis and mutation verification data sets are solved, and the problem of data expansion in the existing technology depends on artificial or randomness, achieving effective expansion of neural network training data and improving model performance.

CN120105085APending Publication Date: 2025-06-06INFINEON TECHNOLOGIES AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411767912.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing data augmentation methods rely on artificial intuition or random processes, and cannot effectively understand the improvement needs of data, resulting in insufficient training data in neural networks and affecting model performance.

Method used

Run the validation dataset through a neural network, analyze the output using a classifier to determine the correct and incorrect predictions, mutations are made against the correctly predicted seeds, and run the mutated seeds through a neural network, analyze their output to determine the increase in coverage, and iteratively generates an augmented dataset.

Benefits of technology

Automatically generate an expanded data set for neural network training and verification, increase verification coverage, improve the robustness of retraining, guide the data expansion process through neural network coverage analysis, select the most suitable data mutations and generate new samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105085A_ABST
    Figure CN120105085A_ABST
Patent Text Reader

Abstract

The invention relates to a method and apparatus for augmenting training data or retraining a neural network. According to an embodiment, a method for augmenting training data for a neural network includes: running a verification data set by the neural network to provide a first output; analyzing the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction; mutating a seed of the verification data set corresponding to the first correct prediction; running the abrupt seed through the neural network to provide a second output; analyzing the second output of the neural network using the classifier to determine a second correct prediction and a second incorrect prediction; determining whether there is an increase in neural network coverage for the seed producing the second correctly predicted mutation; and performing the steps of mutating the seed, running the mutated seed through the neural network, and analyzing a second output of the neural network for a second correctly predicted mutated seed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to electronic systems and, in particular embodiments, to systems and methods for augmenting training data or retraining neural networks. Background Art

[0002] Machine learning, and neural networks in particular, represent an area of ​​research and development with pervasive impact on everything from healthcare to self-driving vehicles. Neural networks consist of interconnected layers of algorithms, called "neurons," that "learn" from data inputs. However, a common obstacle encountered in unifying training methods across landscapes is the limited availability of highly diverse data fed into the system for the purpose of reinforcement learning.

[0003] The task of training a neural network is a repetitive task where each input of data results in a slight adjustment of the neuron's internal parameters, gradually improving the network's performance. However, when training on limited and sometimes scarce data, the challenge is to further improve model performance. Unsurprisingly, machine learning models perform better when more training data is available.

[0004] In response to this, data augmentation techniques have been developed to artificially increase the size of training data by creating modified versions of already available data. However, existing data augmentation methods have their shortcomings: because they are artificial (relying on human intuition) or performed randomly without a real understanding of what improvements are needed in the data to provide richer and more numerous datasets for more effective neural network training. Summary of the invention

[0005] According to an embodiment, a method for augmenting training data for a neural network includes: running a validation data set through a neural network to provide a first output; analyzing the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction; mutating a seed of the validation data set corresponding to the first correct prediction; running the mutated seed through a neural network to provide a second output; analyzing the second output of the neural network using a classifier to determine a second correct prediction and a second incorrect prediction; for the mutated seed that produces the second correct prediction, determining whether there is an increase in neural network coverage; and performing the steps of mutating the seed, running the mutated seed through a neural network, and analyzing the second output of the neural network for the mutated seed that produces the second correct prediction.

[0006] According to another embodiment, an apparatus for augmenting training data for a neural network includes: a processor; and a memory with program instructions stored on the memory, the memory being coupled to the processor, wherein the program instructions, when executed by the processor, enable the apparatus to: run a validation data set through a neural network to provide a first output, analyze the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction, mutate a seed of the validation data set corresponding to the first correct prediction, run the mutated seed through the neural network to provide a second output, analyze the second output of the neural network using the classifier to determine a second correct prediction and a second incorrect prediction, determine whether there is an increase in neural network coverage for the mutated seed that produces the second correct prediction, and perform the steps of: mutating the seed, running the mutated seed through the neural network, and analyzing the second output of the neural network for the mutated seed that produces the second correct prediction.

[0007] According to yet another embodiment, a method for retraining a neural network includes: providing a first set of seeds to the neural network to provide a first output; applying a classifier to the first output to determine, by the classifier, a first seed of the first set of seeds corresponding to a first correct prediction; mutating the first seed to provide a first mutated seed; running the first mutated seed through the neural network to provide a second output; applying the classifier to the second output to determine, by the classifier, a second seed of the first set of seeds corresponding to a second correct prediction; for the determined second seed, determining whether there is an increase in neural network coverage; and in response to determining that a second seed of a plurality of second seeds causes an increase in neural network coverage, retraining the neural network using at least one of the second seeds. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] For a more complete understanding of the present invention and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which:

[0009] Figure 1 A system for generating augmented verification data according to an embodiment is illustrated;

[0010] Figure 2 illustrates a process flow diagram of a data augmentation algorithm according to an embodiment;

[0011] Figure 3A illustrates a neural network according to an embodiment; and Figure 3B , 3C , 3D, 3E, 3F, 3G and 3H are instructions and Figure 3A A table of neural network related embodiments covering the operations of the algorithm;

[0012] Figure 4 is a flow chart of a method according to an embodiment; and

[0013] Figure 5 is a block diagram of a processing system that can be used to implement embodiment systems and algorithms.

[0014] Unless otherwise indicated, corresponding numbers and symbols in different figures generally refer to corresponding parts. The accompanying drawings are drawn to clearly illustrate the relevant aspects of the preferred embodiments and are not necessarily drawn to scale. In order to more clearly illustrate certain embodiments, letters indicating variations of the same structure, material, or process step may follow the figure number. DETAILED DESCRIPTION

[0015] The manufacture and use of the current preferred embodiment are discussed in detail below. However, it should be understood that the present invention provides many applicable inventive concepts that can be implemented in a wide variety of specific contexts. The specific embodiments discussed are merely illustrative of the specific ways of making and using the present invention, and do not limit the scope of the present invention.

[0016] Embodiments of the present invention relate to a system and method for augmenting validation and retraining data for a neural network. In some embodiments, a validation system provides an initial validation data set as input to a neural network and mutates seeds corresponding to the correctly predicted initial validation data set. Next, for the mutated seeds that produced the correct predictions, the system determines whether there is an increase in neural network coverage. These mutated seeds corresponding to the increased neural network coverage can then be further and / or iteratively mutated, evaluated for accuracy, and evaluated for increased coverage to provide an increased number of mutated seeds. The mutated seeds from each cycle of seed mutation and evaluation can then be used to augment the training data. In some embodiments, the resulting seeds of mutations correctly identified by the neural network and / or seeds of mutations incorrectly identified by the neural network can be used as training data to retrain the neural network.

[0017] Advantages of embodiments of the present invention include the ability to automatically generate augmented data sets for neural network training and validation, which increases validation coverage and makes retraining more robust. Embodiment systems and algorithms can advantageously be configured to guide the data augmentation process via neural network coverage analysis, which enables the algorithm to select the data most suitable for mutation and generation of new samples.

[0018] Embodiments of the invention are summarized here. Other embodiments may also be understood from the entirety of the specification and claims submitted here.

[0019] Figure 1A system 100 for generating augmented training data is illustrated. As shown, the system 100 includes a neural network 102, a data augmentation system 110, and a validation data set 104. In an embodiment, the neural network 102 can be a deep learning system designed to process and analyze data through an adaptive algorithm and multiple layers of processing units. The neural network 102 can adopt a deep neural network (DNN) architecture including multiple interconnected layers, including an input layer, one or more hidden layers, and an output layer, wherein each layer includes multiple nodes or neurons. These neurons are designed to process input data and perform various transformations, including activation functions, weighted connections, and biases, to produce accurate outputs.

[0020] In the neural network 102, the layers are interconnected in a hierarchical structure, where the input layer receives the raw data and the output layer provides the final prediction or classification. As the data passes through the network, each hidden layer gradually transforms the data, enabling the detection and extraction of increasingly complex features and patterns. The inclusion of multiple hidden layers allows the DNN to learn complex nonlinear relationships within the data, thereby significantly improving the accuracy of prediction and classification compared to shallow neural networks.

[0021] The neural network 102 is designed to adapt and optimize its internal parameters during a training phase, in which a supervised learning algorithm adjusts the weights and biases of the connections between neurons to minimize the error between the network output and the desired target. This optimization process, often referred to as backpropagation, involves propagating an error signal through the network, updating the parameters to minimize the overall cost function.

[0022] Although embodiments of the present invention are described with respect to DNNs, it should be understood that other types of neural networks may be used to implement the neural network 102, including but not limited to feedforward neural networks, recurrent neural networks (RNNs), convolutional neural networks (CNNs), and radial basis function networks (RBFNs).

[0023] In one embodiment, a hardware implementation of the neural network 102 may be implemented using a dedicated custom integrated circuit (IC), such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). The ASIC or FPGA may be designed to perform parallel computations of neuron activation functions and weight updates, thereby increasing the processing speed and efficiency of the neural network 102. In addition, memory components for storing weights, biases, and intermediate data associated with the NN 102 may be implemented using embedded memory blocks, such as static random access memory (SRAM) cells or flip-flops, to facilitate low-latency data access and efficient operation.

[0024] In another embodiment, the neural network 102 may be implemented using a combination of a general-purpose processor (e.g., a central processing unit, a GPU) and a dedicated hardware accelerator specifically designed to perform neural network operations. The hardware accelerator may include a matrix multiplication unit, a convolution engine unit, and an activation function calculation unit, etc., which can be effectively used to propagate input data through the neural network 102 and update weights and biases during training. The dedicated hardware accelerator may be interconnected to the general-purpose processor via a high-speed data bus or interconnect structure, thereby allowing efficient data exchange and parallel processing capabilities. In addition, the hardware implementation of the neural network 102 can be further optimized by customizing the precision of arithmetic operations (such as using reduced precision arithmetic or quantization techniques) to balance the computational complexity, energy efficiency, and accuracy of the neural network. Alternatively, the neural network 102 may be implemented using one or more general-purpose processors without additional acceleration hardware.

[0025] As shown, validation dataset 104 includes initial validation dataset 106 and augmented dataset 108. In various embodiments, initial validation dataset 106 represents an initial set of validation data, and augmented dataset 108 represents additional validation data generated by data augmentation system 110.

[0026] In the context of implementing a neural network, an initial validation data set 106 may be generated during the process of training the neural network 102. For example, relevant data corresponding to the functionality of the neural network 102 is accumulated, which may extend from structured data (such as database and spreadsheet data) to unstructured data (such as images, text, audio, and video). Subsequently, the assembled data set undergoes a preprocessing phase that includes removing irrelevant or duplicate information, addressing gaps and outliers in the data, and normalizing and scaling the data to standardize the range of values.

[0027] After preprocessing, the collated data set can be divided into subsets, each of which plays a specific role in the neural network implementation. These subsets may include training sets, validation sets, and test sets. The training set helps to adjust the weights of the neural network during the training phase, while the validation set facilitates unbiased evaluation and changes to the model fit during training. The test set used after training measures the overall efficiency of the neural network. The initial validation data set 106 can be obtained by random partitioning of the data to offset potential biases. Alternatively, the initial validation data set 106 may include some or all of the training set and / or validation set. In some embodiments, each data point (also referred to as a "seed") in the validation data set 104 set is accompanied by its expected output so that the data augmentation system 110 can evaluate the accuracy of the prediction provided by the neural network 102.

[0028] The augmented data set 108 is generated by the data augmentation system 110 and can be used for subsequent validation, testing, and retraining of the neural network 102, as described below with respect to embodiments.

[0029] The data augmentation system 110 includes an output evaluation module 112, a coverage evaluation module 114, a seed mutation module 116, and a model retraining module 118. In various embodiments, the data augmentation system 110 may be implemented using software code executed on one or more processors, which is configured to perform data augmentation according to embodiments of the present invention.

[0030] Output evaluation block 112 is configured to evaluate the output of neural network 102 to determine whether neural network 102 correctly classified the seed provided by validation dataset 104. For example, output evaluation block 112 may compare the output of neural network 102 to an expected output associated with a particular seed provided by validation dataset 104. In some embodiments, the expected output is stored with the seed in validation dataset 104 and provided to output evaluation block 112 during operation or execution of data augmentation system 110.

[0031] The coverage assessment block 114 is configured to determine the coverage (e.g., utilization) of one or more neurons in the neural network 102 based on a neural network coverage metric. In general, a neural network coverage metric is a measurement used when evaluating the performance of a neural network to quantify the extent to which neurons in the network have been activated or utilized during the evaluation process. In one embodiment, the k-multi-slice neuron coverage (KMNC) metric is used as further described below; however, other metrics known in the art may be used, including but not limited to neuron coverage (NC), neuron boundary coverage (NBC), strong neuron activation coverage (SNAC), and neural coverage (NLC). During operation, the coverage assessment block 114 may access weight values ​​and other variables associated with the neural network 102 via a digital interface (not shown) and derive a neural network coverage metric therefrom.

[0032] The seed mutation block 116 is configured to mutate the seeds provided by the validation dataset 104 according to high-dimensional and / or low-dimensional mutations. For high-dimensional mutations, the original seeds are subjected to noise, perturbations, and adversarial attacks while still retaining the same signature. On the other hand, for low-dimensional seed mutations, the seeds are manipulated in a latent space, which is a compressed representation of the data containing basic features or patterns and typically has a lower dimensionality than the original data. Techniques such as noise addition, interpolation, extrapolation, linear interpolation, and resampling can be used for low-dimensional mutations, and the mutated version is then transformed back to a high-dimensional space using, for example, an autoencoder such as a conditional variational autoencoder (CVAE).

[0033] High-dimensional transformations for various data types can be used to enhance the functionality and performance of systems utilizing such data. For example, for infrared applications, infrared data can be transformed, such as flipping and rotating, brightness and contrast changes, blurring, scaling and cutting, resolution changes, data mixing, and / or the addition of Gaussian noise. On the other hand, for time-of-flight (ToF) applications, ToF data can benefit from time shifting, scaling, noise addition, time jitter, depth cutting and flipping, data interpolation, resolution changes, and / or outlier injection. For radar applications, the transformation of radar data can include compression (such as range compression), time shifting, Doppler shift, range scaling, noise addition, clutter addition, azimuth and elevation variation, and / or resolution change. Similarly, for audio applications, audio data can be improved by time stretching, pitch shifting, background noise addition, volume variation, time and frequency domain variation, cutting and distortion, audio cascading, velocity disturbance, time change, and / or echo generation. For ultrasound applications, ultrasound data may be transformed by introducing at least one of flipping and rotation, zooming and scaling, noise addition (e.g., speckle noise), contrast and brightness changes, shadow and artifact simulation, texture variation, and / or resolution variation. Finally, for WiFi applications, WiFi data may be modified using at least one of RSSI scaling, signal loss, signal interpolation, noise addition, time jitter, position perturbations, access point (AP) loss and rotation, data splitting, or AP density variation to optimize the quality and utility of the data in its respective application. It should be noted that these examples are by no means exhaustive, and other applications and transformations may be employed depending on the particular system and its specifications.

[0034] The model retraining block 118 is configured to control the neural network 102 to retrain based on the augmented data set 108 generated by the data augmentation system 110. In some embodiments, the model retraining block 118 can be used to retrain the neural network 102 on seeds of mutations that the neural network 102 provides incorrect classifications of. The model retraining block 118 can retrain the neural network 102, for example, by adjusting the weights and biases of the neural network 102 based on the error gradient between the predicted results and the actual results, and iteratively improving the performance of the model during multiple training epochs.

[0035] Figure 2 A process flow diagram of a data augmentation algorithm 200 according to an embodiment of the present invention is illustrated.

[0036] Initially, the algorithm is provided with a validation dataset 202 and a model (neural network) as input. During an empty run 204, the algorithm creates a dictionary of existing states, capturing the distribution of coverage by examining the range of weights assigned to each neuron in the different layers of the neural network. Neuron coverage can be measured at different levels of granularity, such as, but not limited to, neuron coverage (NC) (indicating the proportion of neurons of a neural network that have been activated) and k-multi-section coverage (KMNC) (indicating which k sub-portions of the neuron weight range have been activated). To achieve this, the algorithm runs the validation dataset through the model during a classifier step 206 to measure neuron coverage and identify, for example, which portions of the weight range are activated. This analysis helps determine the degree of neuron coverage achieved using the validation dataset 202. The algorithm then selects a subset of the validation dataset that meets certain criteria (such as correct predictions with high confidence) to form a seed set.

[0037] about Figure 1 As illustrated, validation dataset 202 corresponds to initial validation dataset 106 stored in validation dataset 104. Neuron coverage can be determined using coverage evaluation block 114 to measure the stage of neural network 102.

[0038] After the neural network 102 processes the validation data set 202 in the classifier step 206, the algorithm determines which outputs of the classifier step 206 constitute correct predictions, in which case the seeds corresponding to the correct predictions are entered into the seed queue 232. Samples corresponding to incorrect predictions are discarded in step 208. In some embodiments, when the prediction has a high confidence level or a predetermined confidence threshold, the seeds corresponding to the correct predictions are entered into the seed queue. For example, the seed queue 232 may be used. Figure 1 An output evaluation block 112 is shown to perform an evaluation of the output of the classifier step 206 .

[0039] Next, during step 214, the above Figure 1 The high-dimensional perturbations 216 and / or low-dimensional latent space mutations 218 discussed in the seed mutation block 116 mutate the seeds corresponding to the correct predictions residing in the seed queue 232. In various embodiments, these mutations maintain the semantic integrity of the seeds. For example, if the seed is an image depicting a cat, it is still an image depicting a cat after its transformation.

[0040] As shown, latent space mutation is achieved by applying a variational autoencoder (VAE) and / or a generative adversarial network (GAN) 212 to a training dataset 210 to generate a latent space formula. VAE and GAN are two types of generative models that can learn to represent complex data distributions by discovering latent spaces (low-dimensional representations of high-dimensional data) in the input data.

[0041] The VAE model maps inputs to distributions in latent space rather than points. Given some input or seed, the VAE encodes it into a latent space representation. Introducing slight randomness or bias in this latent representation enables the mutation process, decoding this latent representation back into data space, and producing slightly altered or mutated versions of the original input.

[0042] The GAN model involves a generator network and a discriminator network. The generator creates data that cannot be distinguished from real data, while the discriminator attempts to tell the difference between real data and generated data. A random noise seed is usually the input to the generator network. Mutations can be achieved by changing this seed or introducing randomness, which results in the generation of varying output data. Additionally, mutations can be performed by adding noise or transformations directly to the output of the trained generator.

[0043] High-dimensional mutations can be specific to a data type or domain, while low-dimensional mutations can be applied to any data type and model, involving the training of different autoencoders. In some embodiments, combining both high-dimensional and low-dimensional approaches advantageously reveals different missing data in training, identifies corner cases, and provides a more comprehensive assessment of model performance. In some embodiments, these mutations can be applied to different classes separately, allowing automatic differentiation in augmentation based on class-specific requirements. Therefore, the user only needs to decide whether to apply augmentation to all data or selectively to specific classes.

[0044] The seeds of the mutations generated in mutation step 214 are used as new data samples, which are then fed back into the neural network during classifier step 220. If the neural network does not correctly predict a new data sample, the new data sample is stored in fault pool 222, for example, to avoid wasting time and further resources. In some embodiments, the seeds of the samples stored in fault pool 222 are later analyzed and / or used as data for retraining the neural network. In some embodiments, by providing augmented data set 108 to neural network 102 and using Figure 1 The output evaluation block 112 in the system evaluates the output of the neural network to perform the classifier step 220.

[0045] For correctly predicted tests, during coverage evaluation step 224, the impact of the test on the neural network coverage is evaluated by measuring the neural network coverage of the neural network using the neural network coverage metric of the mutated seed and comparing the current neural network coverage to the previous neural network coverage before the mutation. If the current coverage exceeds the previous coverage of the particular mutated seed, the mutated seed is provided back to the seed queue for further mutation and / or analysis. On the other hand, if the current coverage does not exceed the previous coverage, the mutated seed is discarded in step 230, for example, to avoid wasting time and resources. In some embodiments, the mutated seed may be used Figure 1 The coverage assessment module 114 is shown to perform the coverage assessment step 224 .

[0046] When using this algorithm as a data augmentation technique, newly generated data samples with increased neural network coverage are added to the training dataset, and the model is retrained during the model training and evaluation step 234. For example, Figure 1 Model retraining is performed by navigating to the model retraining block 118 shown. Next, a validation data set is executed to see if accuracy and robustness have improved. Alternatively, an embodiment data augmentation algorithm may be used to generate new data to test the model. In such an embodiment, the newly generated tests, updated accuracy results, and coverage increase reports may be provided to the developer for further analysis. This allows the developer to retrain the neural network by adjusting weights or modifying the architecture of the neural network itself.

[0047] In various embodiments, the seed mutation process is performed in an iterative manner until a user-defined end criterion is met (step 228), and the data augmentation process is terminated at step 236. The user-defined criterion may include, but is not limited to, a predefined accuracy, a specific number of samples generated, or the end of an allotted time period.

[0048] Embodiment augmentation algorithms provide many advantages. For example, in the case of direct augmentation, more data samples lead to improved models with better accuracy and robustness. When augmented data is used for testing purposes, the number of test data samples, prediction results, and coverage improvement reports are advantageously increased to enhance the overall effectiveness of the testing process.

[0049] FIG. 3A to FIG. 3F The diagram of FIG provides an example of a k-multi-section neuron coverage (KMNC) measure according to an embodiment of the present invention. In some embodiments, for example, Figure 1 A coverage assessment block 114 is shown to perform this measurement.

[0050] Figure 3A A simple neural network 300 is illustrated that is representative of many types of embodiment neural networks. Neural network 300 includes an input 302 with input data values ​​x1 and x2, an output 304, and two hidden layers (layer 1 and layer 2), each of which includes three neurons. As shown, layer 1 includes neurons n1, n2, and n3, and layer 2 includes neurons n4, n5, and n6. Each neuron is associated with a set of weights. For example, neuron n1 is associated with weights w11 and w21, neuron n2 is associated with weights w12 and w22, and so on. Therefore, for the example inputs of x1=0.1 and x2=0.5, the sum of the weighted signals entering each neuron and output 304 can be expressed as follows:

[0051] n1=(0.2*w11)+(0.5*w21)

[0052] n2=(0.2*w12)+(0.5*w22)

[0053] n3=(0.2*w13)+(0.5*w23)

[0054] n4=(n1*w31)+(n2*w41)+(n3*w51)

[0055] n5=(n1*w32)+(n2*w42)+(n3*w52)

[0056] n6=(n1*w33)+(n2*w43)+(n3*w53)

[0057] Output = (n4*w61)+(n5*w71)+(n6*w81).

[0058] To simplify the description, a linear activation function is used in this example. However, in the embodiment neural network, any activation function can be applied to the weighted sum of each input. Such activation functions can include, but are not limited to, Sigmoid (Logistic activation), hyperbolic tangent (Tanh), rectified linear unit (ReLU), leaky rectified linear unit (leaky ReLU), parameterized rectified linear unit (PReLU), exponential linear unit (ELU), Swish, Softmax, Softplus, and Maxout.

[0059] Before applying the embodiment data augmentation algorithm, the neural network is trained. This example will assume that each hidden layer exhibits the following weights during training:

[0060] Layer 1 weights :

[0061] w11=0.3, w12=-0.7, w13=0.5

[0062] w21=-0.1, w22=0.8, w23=-0.4

[0063] Layer 2 weights :

[0064] w31=0.6, w32=-0.2, w33=0.4

[0065] w41=-0.5, w42=0.9, w43=-0.7

[0066] w51=0.2, w52=-0.3, w53=0.1.

[0067] It should be understood that these weights are merely illustrative examples, as the weights assigned to an embodiment neural network will depend on the specific training data applied to the neural network and the specific architecture of the neural network.

[0068] In an embodiment of the present invention, the data augmentation algorithm first initializes the coverage criterion, as described above with respect to Figure 2 As described in step 204 of . When using the KMNC neural network coverage metric, the embodiment algorithm obtains the range of values ​​that each neuron n maintains for all training samples it encounters during training in order to generate a histogram interval associated with each neuron. The number of intervals K is selected to suit the accuracy desired for a particular application. Generally, the higher the number K, the higher the accuracy of the coverage. In alternative embodiments utilizing activation functions other than linear activation functions, the histogram of the KMNC neural network coverage metric (or other coverage metric) can be based on the summed weighted inputs of each neuron before or after applying the activation function.

[0069] Figure 3B A table showing the output n1, n2, n3, n4, n5 and n6 of each neuron on three different training sample sets [x1, x2] = [0.2, 0.5], [0.6, 0.1], [0.9, 0.3] is illustrated. The highest and lowest output values ​​of each neuron for a given training data are underlined, which represent their respective output ranges. From here, a histogram interval can be assigned to each neuron. For example, neuron n1, which has the highest and lowest output values ​​of 0.39 and -0.28, respectively, is divided into five histogram intervals: a first interval with an interval boundary between -0.28 and -0.146, a second interval with an interval boundary between -0.146 and -0.012, a third interval with an interval boundary between -0.012 and 0.122, a fourth interval with an interval boundary between 0.122 and 0.256, and a fifth interval with an interval boundary between 0.256 and 0.39, as shown in FIG. Figure 3C As shown. It should be understood that for illustrative purposes, five evenly spaced intervals are selected. In alternative embodiments, more or less than five intervals may be used, and / or the intervals may be spaced non-linearly depending on the particular embodiment and its specifications.

[0070] In the above about Figure 2 During the depicted classifier step 206, the validation data set 202 is applied to the neural network, during which the weighted sum values ​​are monitored for each neuron. Figure 3DA table showing the outputs n1, n2, n3, n4, n5 and n6 of the validation samples [x1, x2] = [0.3, 0.7], [0.8, 0.2], [0.1, 0.4], [0.5, 0.6], [0.7, 0.9] according to this example is illustrated. From this table, the outputs n1, n2, n3, n4, n5 and n6 for neuron n1 can be explained as follows Figure 3E , 3F A histogram is formed as shown in the histograms of 3G and 3H.

[0071] like Figure 3D As shown in the first row of the table, the weighted sum of the inputs to neuron n1 is 0.16, which falls on Figure 3E 4 of the KMNC histogram shown. A check mark is placed in the fourth interval covering values ​​between 0.122 and 0.256 to indicate that the fourth interval is "covered" by the first validation sample [x1, x2] = [0.3, 0.7].

[0072] Figure 3D The next validation sample in the second row of the table provides a value of 0.29 for neuron n1, which falls on Figure 3F Therefore, a check mark is added to the fifth interval. Next, Figure 3D The third row of the table provides a value of 0.29 for neuron n1, which falls in Figure 3G In the third bin of the KMNC histogram shown, a check mark is therefore added to the third bin. At this point, the third bin, the fourth bin, and the fifth bin contain check marks and are considered to be "covered".

[0073] Figure 3D The fourth and fifth validation samples listed in the table provide respective values ​​of 0.19 and 0.31 corresponding to already covered bins 4 and 5 of the KMNC histogram. Therefore, in addition to the third, fourth, and fifth histogram bins, the validation data does not cover additional histogram bins. From here, the KMNC neural network coverage metric can be calculated as follows:

[0074]

[0075] Since neuron n1 is covered by three intervals and K = 5, the coverage metric of neuron n1 is 3 / 5*100% = 60%. This coverage metric can also be applied to all neurons n1, n2, n3, n4, n5 and n6. For example, if a total of 20 intervals are covered in six neurons, the total coverage will be:

[0076]

[0077] In an embodiment, once the initial coverage metric is determined, further coverage can be evaluated for the mutated sample, such as during a coverage evaluation step 224, to evaluate the coverage of the mutated sample and determine whether the mutated seed (providing the correct output) increases the neural network coverage metric. For example, if the mutated seed has a value of [x1, y2] = [0, 1] and produces a value of n1 = -0.1, then the second interval of the KMNC histogram will also be covered, such as Figure 3H As shown, this increases the number of intervals covered by neuron n1 from 3 to 4. Therefore, the coverage metric for neuron n1 becomes:

[0078]

[0079] This is an increase of 60% based on the initial validation data. If the total number of covered intervals increases from 20 to 22 for all six neurons n1, n2, n3, n4, n5, and n6, the total coverage metric for all six neurons becomes:

[0080]

[0081] Coverage increased from 66.67% based on the initial validation data.

[0082] Although the above examples are specifically directed to the KMNC neural network coverage metric, it should be understood that in alternative embodiments of the present invention, other neural network coverage metrics may be used. For example, neuron coverage (NC) is a metric in which a neuron is considered "activated" if its value exceeds a user-specified threshold. The threshold is typically set based on the accuracy of coverage required for a particular application. Coverage is then determined as the ratio of "activated" neurons to the total number of neurons in the network. Similarly, neuron boundary coverage (NBC) analyzes the range of values ​​of neurons covered by the training data. If the value of a neuron is not within this range of values, the neuron is considered "covered." Coverage in this article is defined as the ratio of covered neurons to all neurons.

[0083] There is also the Strong Neuron Activation Coverage (SNAC) metric, which, similar to NBC, considers a neuron to be "covered" if its value is above the maximum value in a range of values. Coverage is measured as the ratio of covered neurons to all neurons. Neural Coverage (NLC), on the other hand, is slightly different in that it considers a single hidden layer as a basic computational unit rather than a single neuron. NLC captures four key properties of the distribution of neuron outputs - divergence, correlation, density, and shape, providing an accurate description of how a neural network understands its inputs via an approximate distribution rather than a neuron. It should be understood that these examples of neural network coverage metrics are non-limiting examples, as other neural network coverage metrics may also be used.

[0084] Figure 4 is a flow chart of a method 400 according to an embodiment. According to an example, the system 100 may perform Figure 4 One or more processing blocks.

[0085] like Figure 4 As shown, method 400 may include: running a validation data set through a neural network to provide a first output (block 402). For example, as described above, system 100 may run initial validation data set 106 of validation data set 104 through neural network 102 to provide a first output. Figure 4 As further shown in FIG. 4 , method 400 may include analyzing the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction (block 404). For example, as described above, output evaluation block 112 of system 100 may analyze the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction. Figure 4 As shown, method 400 may include mutating a seed of the validation data set corresponding to the first correct prediction (block 406). For example, as described above, seed mutation block 116 of system 100 may mutate a seed of validation data set 104 corresponding to the first correct prediction. Figure 4 As further shown in , method 400 may include: running the mutated seed through the neural network to provide a second output (block 408). For example, as described above, system 100 may run the mutated seed through neural network 102 to provide a second output.

[0086] Method 400 also includes analyzing the second output of the neural network using a classifier to determine a second correct prediction and a second incorrect prediction (block 410). For example, as described above, output evaluation block 112 of system 100 can analyze the second output of the neural network using a classifier to determine the second correct prediction and the second incorrect prediction. For the mutant seed that produced the second correct prediction, determining whether there is an increase in neural network coverage (block 412). For example, as described above, coverage evaluation block 114 of system 100 can determine whether there is an increase in neural network coverage for the mutant seed that produced the second correct prediction. Figure 4 As further shown in FIG. 4 , method 400 may include performing the steps of mutating a seed, running the mutated seed through a neural network, and analyzing a second output of the neural network for the mutated seed that produced a second correct prediction (block 414 ).

[0087] It should be noted that although Figure 4 Example blocks of method 400 are shown, but in some implementations, method 400 may include Figure 4Additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in. Additionally or alternatively, two or more blocks of method 400 may be performed in parallel.

[0088] Reference now Figure 5 , provides a block diagram of a processing system 500 according to an embodiment of the present invention. The processing system 500 depicts several parts (such as Figure 1 The system 100 shown in Figure 2 The algorithm 200 shown in FIG. 200 , and / or FIG. 3A to FIG. 3H The common platform and common components and functionality of the algorithms detailed in .

[0089] The processing system 500 may include, for example, a central processing unit (CPU) 502 and a memory 504 connected to a bus 508, and may be configured to perform the processes discussed above according to program instructions stored in the memory 504 or on other non-transitory computer-readable media. If desired or necessary, the processing system 500 may also include: a display adapter 510 for providing a connection to a local display 512; and an input-output (I / O) adapter 514 for providing an input / output interface for one or more input / output devices 516 (such as a mouse, keyboard, flash drive, etc.).

[0090] The processing system 500 may also include a network interface 518, which may be implemented using a network adapter configured to couple to a wired link (such as a network cable, USB interface, etc.) and / or a wireless / cellular link for communicating with the network 520. The network interface 518 may also include a suitable receiver and transmitter for wireless communication. It should be noted that the processing system 500 may include other components. For example, if implemented externally, the processing system 500 may include hardware components power supply, cables, motherboard, removable storage media, housing, etc. These other components (although not shown) are considered to be part of the processing system 500. In some embodiments, the processing system 500 may be implemented on a single monolithic semiconductor integrated circuit and / or on the same monolithic semiconductor integrated circuit as other disclosed system components.

[0091] Embodiments of the invention are summarized here. Other embodiments may also be understood from the entirety of the specification and claims submitted here.

[0092] Example 1. A method for augmenting training data for a neural network, the method comprising: running a validation data set through a neural network to provide a first output; analyzing the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction; mutating a seed of the validation data set corresponding to the first correct prediction; running the mutated seed through a neural network to provide a second output; analyzing the second output of the neural network using a classifier to determine a second correct prediction and a second incorrect prediction; for the mutated seed that produced the second correct prediction, determining whether there is an increase in neural network coverage; and performing the steps of mutating the seed, running the mutated seed through a neural network, and analyzing the second output of the neural network for the mutated seed that produced the second correct prediction.

[0093] Example 2. The method of Example 1, further comprising: using the seeds of the mutation that produced the second correct prediction as training data to further train the neural network.

[0094] Example 3. The method of one of Examples 1 or 2, further comprising: determining a neural network coverage metric for a validation dataset combined with the mutated seed.

[0095] Example 4. The method of one of Examples 1 to 3, wherein mutating the seed comprises: performing a high-dimensional perturbation or a latent space mutation.

[0096] Example 5. The method of Example 4, wherein the latent space mutation is performed using a conditional variational autoencoder (CVAE).

[0097] Example 6. The method of one of Examples 4 or 5, wherein the high-dimensional perturbations include: for infrared applications: for infrared applications: at least one of flipping and rotation, brightness and contrast changes, Gaussian noise, blurring, scaling and cropping, resolution changes, or data mixing; for time-of-flight (TOF) applications: at least one of time shifting, scaling, noise addition, time jittering, depth cropping and flipping, data interpolation, resolution changes, or outlier injection; for radar applications: range compression, time shifting, Doppler shift, range scaling, noise addition, clutter addition, azimuth and elevation variations, or resolution changes for audio applications: at least one of time change, pitch shifting, background noise addition, volume variation, time and frequency domain variation, clipping and distortion, audio concatenation, velocity disturbance, or echo generation; for ultrasound applications: at least one of flipping and rotation, zooming and scaling, noise addition, contrast and brightness variation, shadow and artifact simulation, texture variation, or resolution variation; or for WiFi applications: at least one of RSSI scaling, signal loss, signal interpolation, noise addition, time jitter, position disturbance, access point (AP) loss and rotation, data splitting, or AP density variation.

[0098] Example 7. The method of one of Examples 4 to 6, wherein the latent space mutation includes interpolation, extrapolation, linear interpolation, and resampling.

[0099] Example 8. The method of one of Examples 1 to 7, wherein determining whether there is an increase in neural network coverage includes determining a k-multi-faceted neuron cover (KMNC).

[0100] Example 9. The method of Example 8, wherein determining the KMNC includes: determining an output range of each neuron based on a validation data set, dividing the output range into K intervals, determining coverage of each interval with respect to the validation data set, and determining whether an increased number of intervals are covered when using a corresponding mutated seed.

[0101] Example 10. The method of one of Examples 1 to 9, wherein determining whether there is an increase in neural network coverage includes determining neuron coverage (NC), neuron boundary coverage (NBC), strong neuron activation coverage (SNAC), or neural coverage (NLC).

[0102] Example 11. The method of one of Examples 1 to 11, wherein the neural network is a deep neural network.

[0103] Example 12. An apparatus for augmenting training data for a neural network, the apparatus comprising: a processor; and a memory with program instructions stored on the memory, which is coupled to the processor, wherein the program instructions, when executed by the processor, enable the apparatus to: run a validation data set through a neural network to provide a first output, analyze the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction, mutate a seed of the validation data set corresponding to the first correct prediction, run the mutated seed through the neural network to provide a second output, analyze the second output of the neural network using the classifier to determine a second correct prediction and a second incorrect prediction, for the mutated seed that produced the second correct prediction, determine whether there is an increase in neural network coverage, and perform the steps of mutating the seed, running the mutated seed through the neural network, and analyzing the second output of the neural network for the mutated seed that produced the second correct prediction.

[0104] Example 13. The apparatus of Example 12, wherein the program instructions further enable the apparatus to further train the neural network using the seeds of the mutation that produced the second correct prediction as training data.

[0105] Example 14. The apparatus of one of Examples 12 or 13, wherein the program instructions further enable the apparatus to mutate the seeds by performing a high-dimensional perturbation or a latent space mutation.

[0106] Example 15. The apparatus of Example 14, wherein the latent space mutation is performed using a conditional variational autoencoder (CVAE).

[0107] Example 16. The device of one of Examples 14 or 15, wherein the high-dimensional perturbations include: for infrared applications: for infrared applications: at least one of flipping and rotation, brightness and contrast changes, Gaussian noise, blurring, scaling and cropping, resolution changes, or data mixing; for time-of-flight (TOF) applications: at least one of time shifting, scaling, noise addition, time jittering, depth cropping and flipping, data interpolation, resolution changes, or outlier injection; for radar applications: range compression, time shifting, Doppler shift, range scaling, noise addition, clutter addition, azimuth and elevation variations, or resolution for audio applications: at least one of time change, pitch shifting, background noise addition, volume variation, time and frequency domain variation, clipping and distortion, audio concatenation, velocity disturbance, or echo generation; for ultrasound applications: at least one of flipping and rotation, zooming and scaling, noise addition, contrast and brightness variation, shadow and artifact simulation, texture variation, or resolution variation; or for WiFi applications: at least one of RSSI scaling, signal loss, signal interpolation, noise addition, time jitter, position disturbance, access point (AP) loss and rotation, data splitting, or AP density variation.

[0108] Example 17. The apparatus of one of Examples 14 to 16, wherein the latent space mutation includes interpolation, extrapolation, linear interpolation, and resampling.

[0109] Example 18. The apparatus of one of Examples 12 to 17, wherein the program instructions further enable the apparatus to determine whether there is an increase in neural network coverage by determining a k-multi-faceted neuron coverage (KMNC).

[0110] Example 19. The apparatus of Example 18, wherein determining the KMNC comprises: determining an output range of each neuron based on a validation data set, dividing the output range into K intervals, determining coverage of each interval with respect to the validation data set, and determining whether an increased number of intervals are covered when using a corresponding mutated seed.

[0111] Example 20. The device of one of Examples 12 to 19, wherein the program instructions further enable the device to determine whether there is an increase in neural network coverage by determining neuron coverage (NC), neuron boundary coverage (NBC), strong neuron activation coverage (SNAC), or neural coverage (NLC).

[0112] Example 21. The apparatus of one of Examples 12 to 20, wherein the neural network is a deep neural network.

[0113] Example 22. The apparatus of one of Examples 12 to 21, further comprising a neural network.

[0114] Example 23. A method for retraining a neural network, the method comprising: providing a first set of seeds to the neural network to provide a first output; applying a classifier to the first output to determine, by the classifier, a first seed of the first set of seeds that corresponds to a first correct prediction; mutating the first seed to provide a first mutated seed; running the first mutated seed through the neural network to provide a second output; applying the classifier to the second output to determine, by the classifier, a second seed of the first set of seeds that corresponds to a second correct prediction; for the determined second seed, determining whether there is an increase in neural network coverage; and in response to determining that a second seed of a plurality of second seeds causes an increase in neural network coverage, retraining the neural network using at least one of the second seeds.

[0115] Example 24. The method of Example 23, wherein: applying the classifier to the second output further includes: applying the classifier to the second output to determine, by the classifier, a third seed of the first set of seeds that corresponds to an incorrect prediction; and using at least one of the third seeds to constrain the neural network.

[0116] Although the present invention has been described with reference to illustrative embodiments, this description is not intended to be interpreted in a limiting sense. Various modifications and combinations of the illustrative embodiments and other embodiments of the present invention will be apparent to those skilled in the art, based on the reference specification. Therefore, the appended claims are intended to cover any such modifications or embodiments.

Claims

1. A method for augmenting training data for a neural network, the method comprising: running a validation data set through the neural network to provide a first output; analyzing the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction; mutating a seed of the validation dataset corresponding to the first correct prediction; running the mutated seed through the neural network to provide a second output; analyzing the second output of the neural network using the classifier to determine a second correct prediction and a second incorrect prediction; for the seed that produced the second correctly predicted mutation, determining whether there is an increase in neural network coverage; as well as The steps of mutating the seed, running the mutated seed through the neural network, and analyzing the second output of the neural network for the mutated seed that produced the second correct prediction are performed.

2. The method according to claim 1, further comprising: The neural network is further trained using the mutant seeds that produced the second correct prediction as training data.

3. The method according to claim 1, further comprising: A neural network coverage metric is determined for the validation dataset combined with the mutated seed.

4. The method of claim 1, wherein mutating the seed comprises: Perform high-dimensional perturbations or mutations of the latent space.

5. The method of claim 4, wherein the latent space mutation is performed using a conditional variational autoencoder (CVAE).

6. The method according to claim 4, wherein the high-dimensional perturbation comprises: For infrared applications: at least one of flipping and rotation, brightness and contrast changes, Gaussian noise, blurring, scaling and cropping, resolution changes, or data blending; For time-of-flight TOF applications: at least one of time shifting, scaling, noise addition, time jittering, depth cropping and flipping, data interpolation, resolution change, or outlier injection; For radar applications: at least one of range compression, time shift, Doppler shift, range scaling, noise addition, clutter addition, azimuth and elevation variation, or resolution change; For audio applications: at least one of time alteration, pitch shifting, background noise addition, volume variation, time and frequency domain variation, clipping and distortion, audio concatenation, speed perturbation, or echo generation; For ultrasound applications: at least one of flipping and rotating, zooming and scaling, noise addition, contrast and brightness variation, shadow and artifact simulation, texture variation, or resolution variation; or For WiFi applications: at least one of RSSI scaling, signal loss, signal interpolation, noise addition, time jitter, position disturbance, access point AP loss and rotation, data fragmentation, or AP density variation.

7. The method of claim 4, wherein the latent space mutation comprises interpolation, extrapolation, linear interpolation and resampling.

8. The method of claim 1, wherein determining whether there is an increase in neural network coverage comprises determining a k-multi-section neuron coverage (KMNC).

9. The method of claim 8, wherein determining the KMNC comprises: An output range of each neuron is determined based on the validation data set, the output range is divided into K intervals, coverage of each interval with respect to the validation data set is determined, and whether an increased number of intervals are covered when using a corresponding mutated seed.

10. The method of claim 1 , wherein determining whether there is an increase in neural network coverage comprises: Determine whether a neuron covers NC, a neuron border covers NBC, a strong neuron activation covers SNAC, or a neuron covers NLC.

11. The method of claim 1, wherein the neural network is a deep neural network.

12. An apparatus for augmenting training data for a neural network, the apparatus comprising: processor; as well as a memory with program instructions stored on the memory, the memory being coupled to the processor, wherein the program instructions, when executed by the processor, enable the apparatus to: running a validation data set through the neural network to provide a first output, analyzing the first output of the neural network using a classifier to determine a first correct prediction and a first incorrect prediction, mutating a seed of the validation dataset corresponding to the first correct prediction, running the mutated seed through the neural network to provide a second output, analyzing the second output of the neural network using the classifier to determine a second correct prediction and a second incorrect prediction, for the seed that produced the second correctly predicted mutation, determining whether there is an increase in neural network coverage, and The steps of mutating the seed, running the mutated seed through the neural network, and analyzing the second output of the neural network for the mutated seed that produced the second correct prediction are performed.

13. The apparatus of claim 12, wherein the program instructions further enable the apparatus to further train the neural network using the seeds of the mutation that produced the second correct prediction as training data.

14. The apparatus of claim 12, wherein the program instructions further enable the apparatus to mutate the seed by performing a high-dimensional perturbation or a latent space mutation.

15. The apparatus of claim 14, wherein the latent space mutation is performed using a conditional variational autoencoder (CVAE).

16. The apparatus of claim 14, wherein the high-dimensional perturbation comprises: For infrared applications: at least one of flipping and rotation, brightness and contrast changes, Gaussian noise, blurring, scaling and cropping, resolution changes, or data blending; For time-of-flight TOF applications: at least one of time shifting, scaling, noise addition, time jittering, depth cropping and flipping, data interpolation, resolution change, or outlier injection; For radar applications: at least one of range compression, time shifting, Doppler shifting, range scaling, noise addition, clutter addition, azimuth and elevation variation, or resolution change; For audio applications: at least one of time alteration, pitch shifting, background noise addition, volume variation, time and frequency domain variation, clipping and distortion, audio concatenation, speed perturbation, or echo generation; For ultrasound applications: at least one of flipping and rotating, zooming and scaling, noise addition, contrast and brightness variation, shadow and artifact simulation, texture variation, or resolution variation; or For WiFi applications: at least one of RSSI scaling, signal loss, signal interpolation, noise addition, time jitter, position disturbance, access point AP loss and rotation, data fragmentation, or AP density variation.

17. The apparatus of claim 14, wherein the latent space mutation comprises interpolation, extrapolation, linear interpolation, and resampling.

18. The apparatus of claim 12, wherein the program instructions further enable the apparatus to determine whether there is an increase in neural network coverage by determining a k-multi-section neuron coverage (KMNC).

19. The apparatus of claim 18, wherein determining the KMNC comprises: An output range of each neuron is determined based on the validation data set, the output range is divided into K intervals, coverage of each interval with respect to the validation data set is determined, and whether an increased number of intervals are covered when using a corresponding mutated seed.

20. The apparatus of claim 12, wherein the program instructions further enable the apparatus to determine whether there is an increase in neural network coverage by determining a neuron coverage NC, a neuron boundary coverage NBC, a strong neuron activation coverage SNAC, or a neural coverage NLC.

21. The apparatus of claim 12, wherein the neural network is a deep neural network.

22. The apparatus of claim 12, further comprising the neural network.

23. A method for retraining a neural network, the method comprising: providing a first set of seeds to the neural network to provide a first output; applying a classifier to the first output to determine, by the classifier, a first seed of the first set of seeds that corresponds to a first correct prediction; mutating the first seed to provide a first mutated seed; running the first mutated seed through the neural network to provide a second output; applying the classifier to the second output to determine, by the classifier, a second seed of the first set of seeds corresponding to a second correct prediction; For the determined second seed, determining whether there is an increase in neural network coverage; and In response to determining that a second seed of the plurality of second seeds results in an increase in coverage of the neural network, the neural network is retrained using at least one of the second seeds.

24. The method of claim 23, wherein: Applying the classifier to the second output further comprises: applying the classifier to the second output to determine, by the classifier, a third seed of the first set of seeds that corresponds to an incorrect prediction; and The neural network is retrained using at least one of the third seeds.