Iterative small sample refinement micro-training for neural networks

Through the iterative small sample refinement micro-training method, using multiple sets of hyperparameters and data sets to fine-tune the neural network, solving the shortcomings of conventional training techniques in accuracy and quality, achieving higher neural network accuracy and fewer visual artifacts.

CN120218191APending Publication Date: 2025-06-27NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510161032.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-03-13
Filing Date
2020-10-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Conventional neural network training techniques have shortcomings in accuracy and quality, especially when training based on insufficient, one-sided or combined data sets, resulting in inaccurate training losses and insufficient data, which in turn limits the effectiveness of retraining.

Method used

Using the iterative small sample refinement micro-training method, the first micro-training neural network is generated by using the first set of hyperparameters and the first training data set to receive the neural network trained to satisfy the loss function, and receive the second set of hyperparameters and the second training data set, and limiting the neural network to adjust the weights, the first micro-training neural network is generated to reduce visual artifacts.

Benefits of technology

Improve the accuracy and quality of the neural network, reduce visual artifacts, and gradually improve accuracy without significantly changing the computational structure of the trained neural network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218191A_ABST
    Figure CN120218191A_ABST
Patent Text Reader

Abstract

The invention discloses micro-training for iterative small sample refinement of a neural network. The disclosed micro-training techniques improve the accuracy of trained neural networks by performing iterative refinement at a low learning rate using a relatively short series of micro-training steps. The neural network training framework receives the trained neural network and a second training data set and a hyper-parameter set. The neural network training framework promotes an increase in incremental accuracy by using a lower learning rate to adjust one or more weights of the trained neural network without essentially changing the computational structure of the trained neural network, resulting in a micro-trained neural network. Changes in accuracy and / or quality of the micro-trained neural network may be evaluated. Other micro-training sessions may be performed on the micro-trained neural network to further improve accuracy or quality.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This divisional patent application is a divisional application of the patent application with the filing date of October 12, 2020, the application number of 202011083316.3, and the invention title of "Iterative Small-Sample Refinement Micro-Training for Neural Networks". Technical Field

[0002] This disclosure relates to neural network training, and more particularly, to iterative small-sample refinement micro-training of neural networks. Background Art

[0003] Conventional neural network training techniques sometimes produce suboptimal results in terms of accuracy or quality. This is especially true when training is based on a dataset that may be insufficient, one-sided, or a combination thereof. In addition, traditional training techniques typically do not provide additional improvement opportunities in situations where the training loss is inaccurate or the data is insufficient such that retraining is impractical or ineffective. In generative neural network image synthesis applications, the results may be suboptimal in the form of image artifacts in the generated images. These problems and / or other problems associated with the prior art need to be addressed. Summary of the Invention

[0004] Disclosed is a method, computer-readable medium, and system for micro-training a neural network to improve accuracy and / or quality. The method includes receiving a neural network trained to satisfy a loss function using a first set of hyperparameters and a first training dataset, receiving a second training dataset, and receiving a second set of hyperparameters. In one embodiment, the second learning parameters specified in the second set of hyperparameters limit the adjustment of one or more weights used by the neural network as compared to the corresponding first learning parameters in the first set of hyperparameters. The method further includes applying the second training dataset to the neural network according to the second set of hyperparameters to produce a first micro-trained neural network by adjusting one or more weights used by the neural network to process the second training dataset. In certain applications, the trained neural network generates output data including visual artifacts, and the first micro-trained neural network produced according to the method reduces the visual artifacts. Brief Description of the Drawings

[0005] Figure 1A A flowchart illustrating a method for micro-training a neural network according to one embodiment is shown.

[0006] Figure 1B Micro-training within the entire hypothesis space according to one embodiment is shown.

[0007] Figure 1C A neural network framework according to one embodiment is shown.

[0008] Figure 2AA flowchart of a method for improving neural network training using micro-training according to one embodiment is shown.

[0009] Figure 2B A graph showing the average difference between layers of various micro-training networks according to one embodiment is shown.

[0010] Figure 3 A parallel processing unit according to one embodiment is shown.

[0011] Figure 4A A general processing cluster within the parallel processing unit of Figure 3 is shown according to one embodiment.

[0012] Figure 4B A memory partition unit of the parallel processing unit of Figure 3 is shown according to one embodiment.

[0013] Figure 5A A streaming multiprocessor of Figure 4A is shown according to one embodiment.

[0014] Figure 5B is a conceptual diagram of a processing system implemented using a Figure 3 PPU according to an embodiment.

[0015] Figure 5C An exemplary system in which various architectures and / or functions of the various previous embodiments can be implemented is shown. Detailed Description

[0016] The disclosed technique, referred to herein as micro-training, improves the accuracy of a trained neural network by performing iterative refinement at a low learning rate using a series of few-shot micro-training steps. The micro-training steps include far fewer training iterations than the initial training of the trained neural network. In such cases, the lower learning rate helps to gradually improve accuracy without significantly changing the computational structure of the trained neural network. In this context, the computational structure refers to both the neural network topology and the various distributions of internal representations therein (e.g., via activation weights, activation functions, etc.). A given network topology can specify how internal artificial neuron nodes are organized into layers and interconnected. After each micro-training step, there can be an evaluation step (e.g., input from an operator via a user interface) to evaluate the incremental quality change. For example, a small number of pixels associated with a thin line (e.g., in an outdoor scene, a dark telephone line against a bright sky) may exhibit aliasing artifacts visible to a human operator (viewer), which were ignored in conventional automated training; however, these pixels can be optimized during micro-training to have proper antialiasing. In such cases, micro-training refines the previously trained network to reduce or eliminate such visually significant artifacts (e.g., aliasing).

[0017] Figure 1A FIG. 4 shows a flowchart of a method 110 for micro-training a neural network according to one embodiment. Although method 110 is described in the context of a processing unit, method 110 can also be performed by a program, a custom circuit, or a combination of a custom circuit and a program. For example, method 110 can be performed by a GPU (Graphics Processing Unit), a CPU (Central Processing Unit), or any processor capable of performing operations for evaluating and training a neural network. Additionally, those of ordinary skill in the art will understand that any system that performs method 110 is within the scope and spirit of the embodiments of the present disclosure. In one embodiment, the processing unit performs method 110 in conjunction with various operations of a neural network training framework and / or a neural network runtime system. In certain embodiments, the processing unit includes one or more instances of a parallel processing unit, such as Figure 3 parallel processing unit 300.

[0018] Method 110 begins at step 111, where the processing unit receives a neural network (G S ) trained to satisfy a loss function (L S ) using a first set of hyperparameters (H S ) and a first training dataset (D S)。In one embodiment, the neural network is a deep generative neural network configured to generate images. In one embodiment, the first set of hyperparameters includes at least one model scale parameter, such as the epoch count, batch size, training iteration count, learning rate, and loss function. In one embodiment, the epoch count specifies the number of times of training on all specified training samples. Each training pass on a given training sample includes a forward pass and a backward pass. The specified training samples can be organized into batches, where the batch size specifies the number of training samples per batch. The number of training iterations specifies the number of training passes performed on different batches to train a given neural network once on all available training samples. For example, for 1000 training samples and a batch size of 200, 5 iterations are required to complete one epoch. In one embodiment, a given set of hyperparameters can refer to one or more sets of training samples. Additionally, the learning rate is a value that scales the speed at which a given neural network adjusts its weights during a given pass. Further, the loss function can specify the difference between the predicted output and the actual output computed by the neural network. In the case of hyperparameters, the loss function can specify the function used to compute the difference.

[0019] In certain use cases, by using the first set of hyperparameters (H S ) and the first training dataset (D S ) to optimize the loss function (L S ), the neural network (G S ) is trained to generate new images. However, when evaluating the neural network using a different test dataset (D T ), the results may be unsatisfactory (e.g., artifacts visible in the generated images). The results may be unsatisfactory due to one or more reasons. A first exemplary reason occurs when the loss function L S is different from the test loss function (L T ). Thus, when evaluating against the test loss function L T , the training optimized for the loss function L S may be insufficient. In this case, the loss function (L S ) may provide insufficient loss feedback to train the neural network G S in a way that avoids visual artifacts, while the visual artifacts may only be important for L T . This situation is particularly challenging when the test loss function involves a subjective human viewer.

[0020] When the distribution of the first training dataset (D S ) is different from the test dataset (D TA second exemplary reason for unsatisfactory results may occur when the distribution of ) is sufficiently different. In this case, the first training dataset may lack sufficient representative data to train the neural network G in a way that avoids visual artifacts. S When the first set of hyperparameters (H S A third exemplary reason for unsatisfactory results may occur when suboptimal tuning is performed. However, optimizing only the hyperparameters (H S ) to overcome training deficiencies may often be impractical.

[0021] When any of the above three reasons for unsatisfactory results is operative in a neural network training use case, it is conventional to simply retrain the neural network G S It may not necessarily improve the quality of the evaluation results. S To match L T may be impractical; capturing a sufficiently large training dataset may be impractical; and optimizing H S However, the micro-training techniques disclosed herein provide a mechanism to improve results without overcoming the impracticality barrier.

[0022] In one embodiment, S is equal to zero, and the neural network G S is a trained neural network (G0) trained using a first training data set (D0) and a first set of hyperparameters (H0). In various use cases, the trained neural network can generate output data including visual artifacts. The artifacts can include, but are not limited to, geometric aliasing artifacts (e.g., jagged edges, blocky appearance), noise artifacts (e.g., rendering noise artifacts), lighting effect artifacts (e.g., water reflection artifacts), and temporal artifacts (e.g., flickering, swimming appearance).

[0023] In step 113, the processing unit receives a second training dataset (D1). The second training dataset D1 may include additional training samples that are selected to specifically train the neural network to suppress visual artifacts. For example, to improve anti-aliasing quality, additional images depicting thin high-contrast lines can be obtained and mixed with the second training dataset (D1) for use during micro-training, thereby guiding the neural network G1 to produce more continuous and aesthetically pleasing anti-aliased lines without disturbing other valuable training. In step 115, the processing unit receives a second set of hyperparameters (H1). In one embodiment, a second learning parameter is specified in the second set of hyperparameters to limit the adjustment of one or more weights used by the neural network compared to the corresponding first learning parameter in the first set of hyperparameters. In one embodiment, the first learning parameter includes a first learning rate, and the second learning parameter includes a second learning rate that is less than the first learning rate. In certain embodiments, the second learning rate is ten times to a thousand times lower than the first learning rate. For example, the first learning rate can be in the range of 1e-3 to 1e-5, while the second learning rate can be in the range of 1e-4 to 1e-8.

[0024] In one embodiment, the first set of hyperparameters includes a first training iteration count, and the second set of hyperparameters includes a second training iteration count that is less than the first training iteration count. In certain embodiments, the second training iteration count is a thousand times (or more) smaller than the first training iteration count. More generally, the second set of hyperparameters can specify the total amount of computation used for training, which can be hundreds to thousands of times (or more times) smaller than the total amount of computation specified by the first set of hyperparameters.

[0025] In step 117, the processing unit applies the second training dataset to the neural network according to the second set of hyperparameters while adjusting one or more weights that the neural network uses to process the second training dataset to produce a first micro-trained neural network. In this way, the first micro-trained neural network (G1) represents an additional training instance of the trained neural network (G0).

[0026] In one embodiment, the processing unit combines and applies the second training dataset with at least a portion of the first training dataset to produce a first micro-trained neural network. For example, the entire second training dataset and the entire first training dataset can be used for training and producing the first micro-trained neural network. In another example, the entire second training dataset and approximately half of the first training dataset can be used. Alternatively, various other combinations of the second training dataset and the first training dataset can be applied for training and producing the first micro-trained neural network. In one embodiment, the second training iteration count is used for training and producing the first micro-trained neural network.

[0027] In one embodiment, each weight of the first micro-training neural network can be adjusted during micro-training. In an alternative embodiment, certain weights, such as those associated with a particular layer, can be locked and not adjusted during micro-training.

[0028] In one embodiment, the trained neural network implements a U-Net architecture with a first set of activation function weights, and the first micro-training neural network implements a corresponding U-Net architecture with a second, different set of activation function weights. In various embodiments, the trained neural network and the first micro-training neural network are networks within a generative adversarial neural network (GAN) system. A GAN generally includes a generator network and a discriminator network, each of which can be a deep neural network, such as a U-Net with an arbitrary depth architecture. The GAN structure pits the generator network against the discriminator network, where the generator network learns to generate synthetic data that is indistinguishable from natural data, and the discriminator network learns to distinguish the synthetic data from the natural data. In certain applications, the generator network can be trained to produce high-quality synthetic data, such as synthetic fictional images. In other applications, the discriminator network learns to generalize its recognition scope beyond natural or initially trained data. In the context of the present disclosure, any technically feasible training mechanism (e.g., backpropagation) can be performed during training without departing from the scope and spirit of the various embodiments.

[0029] More illustrative information regarding various alternative architectures and features that can be used to implement the foregoing framework will now be given according to the user's requirements. It should be noted in particular that the following information is presented for illustrative purposes and should not be construed in any way as limiting. Any of the following features can optionally be combined or not exclude the other features described.

[0030] Figure 1B Micro-training within the entire hypothesis space 140 according to one embodiment is shown. As shown, the untrained neural network G UTraverse the initial training path 142, thereby generating a trained neural network G0. The initial training path 142 can be traversed according to any technically feasible training technique. The trained neural network G0 is within the local optimization region 144, but the trained neural network G0 may not actually provide the desired result 146 based on the first training dataset D0 and the first set of hyperparameters H0. The disclosed methods 110 and 200 refine the trained neural network G0 to make it closer to the desired result 146. In this example, the trained neural network G0 is refined through a path from the trained neural network G0 to the micro-trained neural networks G1, G2, and finally G3. Additionally, the technique provides subjective human input to better align the automated training results with human perception, thereby improving quality in a visually significant and different way from human perception, but it is difficult to algorithmically model in the form of an automated loss function.

[0031] As shown, the initial training result uses the training dataset D0, the loss function, and the hyperparameters H0 to generate a trained neural network G0. The improved training result using the disclosed micro-training technique generates a refined neural network G3, which is closer to the desired result 146. Making small changes to the trained neural network G0 during micro-training preserves the benefits of the original training using the training dataset D0 while allowing for smaller modifications that can improve quality. For example, the refined neural network G1 can generally replicate the trained neural network G0, but making small changes to the activation function weights can improve quality.

[0032] The disclosed micro-training technique includes: receiving the trained neural network G0 (G S , S=0 ); receiving a second training dataset (e.g., D1); receiving a second set of hyperparameters H1; and training a new micro-trained neural network G S based on the neural network G S+1 . In the first micro-training session, the neural network G1 is generated from the neural network G0. In one embodiment, additional training samples can be added to subsequent second training datasets (e.g., D2, D3, etc.), and each subsequent micro-training session (e.g., iteration) can produce subsequent neural networks G2, G3, etc. Multiple micro-training sessions can be performed to further refine the subsequent neural networks G S+n . Micro-training generally preserves the internal computational structure of the trained neural network, thereby allowing for a comparison between the originally trained neural network (G S ) and the subsequently micro-trained neural networks G S+1Perform comparison and interpolation operations on the output. As shown, the disclosed technique allows the micro-trained neural network G3 to provide results closer to the ideal result 146 than the conventionally trained neural network G0. Additionally, the disclosed technique provides an improvement in neural network quality while advantageously requiring only a modest amount of additional computational effort beyond the initial training, as the number of training iterations required for micro-training is orders of magnitude less than traditional training.

[0033] In an exemplary use case, after a micro-trained neural network is generated, certain training data can be processed by the micro-trained neural network and the results presented to a viewer for evaluation. If the results are evaluated as acceptable, the viewer can provide input to a user interface to indicate that the completion requirements have been met. In this example, the viewer may be evaluating visual artifacts related to anti-aliasing, noise reduction, lighting effects, etc. Such visual artifacts may be difficult to algorithmically quantify as better or worse relative to a previous training session, but the viewer can easily provide a subjective evaluation based on human perception of the artifacts. As a further example, a second training dataset can be constructed to include training data that specifically addresses the visual artifacts targeted by the micro-training. In a particular application of anti-aliasing, a small fraction of the total screen pixels may have artifacts, such as those associated with thin, high-contrast lines (e.g., in an outdoor scene, a dark telephone line against a bright sky). Since only a few pixels are affected by certain aliasing artifacts, traditional training techniques may not be able to reliably produce high-quality results for these few pixels; however, these aliasing artifacts may be very noticeable to the viewer and significantly degrade the image quality.

[0034] Figure 1C A neural network framework 170 according to one embodiment is shown. As shown, the neural network framework 170 includes a discriminator 178 that is configured to receive a reference sample 176 that includes reference image data or a synthetic sample 186 that includes synthetic image data. The discriminator 178 generates a loss output that is used by a parameter adjustment unit 180 to calculate adjustments to the respective neural network parameters. In the context described below, the loss represents the confidence that the selected sample 176 or 186 is a reference sample rather than a synthetic sample. The parameter adjustment unit 180 also receives hyperparameters as input. The reference sample 176 can be selected from a training dataset 174 that includes captured images from real-world scenes to be used as reference sample images 175. The generator 184 synthesizes the sample 186 based on previous training and a latent random variable 182 and / or other inputs. In one embodiment, the generator 184 includes a first neural network and the discriminator 178 includes a second neural network.

[0035] In one embodiment, the neural network framework 170 is configured to operate in a Generative Adversarial Network (GAN) mode, where the discriminator 178 is trained to identify "real" reference sample images 175, and the generator 184 is trained to synthesize "fake" samples 186. In one embodiment, the discriminator 178 is trained on samples 176, and each training pass includes a forward pass that evaluates the samples 176 and a backward pass that adjusts the weights and / or biases within the discriminator 178 using, for example, backpropagation techniques. Additionally, the generator 184 is then trained to synthesize samples 186 that can deceive the discriminator 178. Each training pass includes a forward pass in which the samples 186 are synthesized, and a backward pass in which the weights and / or biases within the generator 184 are adjusted (e.g., using backpropagation). In one embodiment, the parameter adjustment unit 180 performs backpropagation to compute new neural network parameters (e.g., weights and / or biases) resulting from a given training pass.

[0036] During the adversarial training process, the discriminator 178 can learn to generalize better, while the generator 184 can learn to synthesize better. Both improvements can be useful separately. In some use cases, such as image enhancement (e.g., super-resolution / upsampling, anti-aliasing, denoising, etc.), training optimization may be required to overcome artifacts in the images synthesized by the initially trained neural network G0 within the generator 184. Such training refinement can be provided when the neural network framework 170 is configured to execute Figure 1A the micro-training method 110 described in Figure 2A and / or the method 200 described in

[0037] In one embodiment, the neural network framework 170 is configured to operate in a micro-training mode, where the sample image 175 is selected to specifically target the deficiencies in the initially trained neural network G0. In the micro-training mode, the generator 184 generates samples 186, which are displayed by the user interface 188 on a display device. The samples 186 can be displayed next to previously generated samples, and a viewer can determine whether the samples 186 are an improvement over the previously generated samples. Additionally, the user interface 188 can display a set of samples 186 on the display device and receive input from the viewer indicating whether the generator 184 has been sufficiently trained during the micro-training. In one embodiment, the neural network framework 170 is configured to execute Figure 1A the method 110 described in Figure 2A and the method 200 described in

[0038] Figure 2A FIG. 140 shows a flowchart of a method 200 for using micro-training to improve neural network training according to one embodiment. Although method 200 is described in the context of a processing unit, method 200 can also be performed by a program, a custom circuit, or a combination of a custom circuit and a program. For example, method 200 can be performed by a GPU (graphics processing unit), a CPU (central processing unit), or any processor capable of performing operations for evaluating and training neural networks. In addition, those of ordinary skill in the art will understand that any system that performs method 200 is within the scope and spirit of the embodiments of the present disclosure. In one embodiment, the processing unit performs method 200 in conjunction with various operations of a neural network training framework and / or a neural network runtime system. In certain embodiments, the processing unit includes one or more instances of a parallel processing unit, such as Figure 3 the parallel processing unit 300. In one embodiment, Figure 1C the neural network framework 170 described in FIG. 142 is at least partially implemented on the processing unit and configured to perform method 200.

[0039] Method 200 begins at step 201, where the processing unit synthesizes a first set of data using a generator neural network. In one embodiment, the generator neural network includes the trained neural network of method 110. In one embodiment, the synthesized data includes one or more images (e.g., video frames). Images can be generated according to any technically feasible technique, including techniques known in the art for deep learning super sampling (DLSS), super resolution / upsampling, and / or anti-aliasing, denoising, and neural networks configured to act as generator networks.

[0040] In step 203, it is determined whether a completion requirement is met. Any technically feasible technique can be performed to determine that the completion requirement is met. In one embodiment, the synthesized one or more images are presented to a human viewer on a display device, and if the viewer evaluates the quality of the one or more images as being good enough, the completion requirement is met. For example, a user interface, such as user interface 188, can receive input from the viewer indicating that the result is acceptable and thus the completion requirement is met. In one embodiment, the user interface is executed on the processing unit, and the image and user interface tools are presented on the display device.

[0041] If the completion requirement is met in step 204, method 200 terminates. Otherwise, if the completion requirement is not met, method 200 proceeds to 205. To complete step 204, the processing unit receives an indication that the completion requirement is met. In one embodiment, the completion requirement is met when the user interface receives an input indication that the micro-training has produced good enough results.

[0042] In step 205, the processing unit prepares a second training dataset. In one embodiment, preparing the second training dataset may include receiving an input via a user interface to select images to be included in the second training dataset. Images may be selected to better align the distribution of the target output data included in the training dataset Ds with the test requirements of the generator neural network, where the training dataset Ds is used during the micro-training of the generator neural network for the test requirements of the generator neural network represented by the test dataset D T The preparation of the second training dataset may include, but is not limited to, capturing additional training samples that are specifically targeted at visual artifacts and / or image features identified by viewers to be removed through micro-training. The preparation of the second training dataset may further include, but is not limited to, removing samples that may be incorrect or missing from the first training dataset, recapturing the incorrect samples, and adding / modifying / enhancing the first training dataset to more closely align the training distribution of the second training dataset with the test dataset. Then, method 200 continues to execute Figure 1A method 110 to produce a micro-trained generator network. After method 110 is completed, method 200 proceeds to step 207.

[0043] In step 207, the processing unit uses the micro-trained generator network to synthesize a second set of data. In one embodiment, the synthesized data includes one or more images (e.g., video frames). Images may be generated according to any technically feasible technique, including techniques known in the art for deep learning super sampling (DLSS), super resolution / upsampling, and / or anti-aliasing, denoising, techniques provided by neural networks configured to act as generator networks, etc.

[0044] In step 209, it is determined whether the result is improved between the first set of data and the second set of data. In one embodiment, the images including the first set of data are compared with the corresponding images including the second set of data on a display device for use by a human viewer. The viewer may evaluate the quality of the displayed images. For example, it may be determined that the result is improved by receiving an input indicating an improvement in the result from the viewer via a user interface. In one embodiment, the user interface is executed on the processing unit, and the images and user interface tools are presented on the display device.

[0045] If the result is improved in step 210, the method returns to step 203. Otherwise, the method proceeds to step 211. In step 211, the processing unit adjusts one or more micro-training parameters. Additionally, the processing unit may discard the micro-training neural network previously generated by method 110. Adjusting one or more micro-training parameters may include, but is not limited to, adding training samples (e.g., images) to the second training dataset, removing training samples from the second training dataset, and adjusting one or more hyperparameters such as the learning rate, number of iterations, etc. In one embodiment, the viewer performs the adjustment of one or more micro-training parameters via a user interface. After completing step 211, the method returns to step 205.

[0046] Multiple traversals of method steps 203 to 211 may be performed until the completion requirement is met in step 204 and the user interface receives an input indication that the micro-training has produced a good enough result. During each micro-training of method 110, a subsequent new neural network (e.g., G1, G2, G3, etc.) is generated. Depending on whether the new neural network improves the result, each new neural network may be retained or discarded.

[0047] In one embodiment, method 110 and / or method 200 may perform transfer learning to produce a new neural network G optimized for an application different from the initially trained neural network G0. S+n In another embodiment, for example, in a discriminator network, method 110 and / or method 200 may be performed to improve generalization.

[0048] More generally, the disclosed techniques provide for fast refinement training of existing (e.g., pre-trained) neural networks, fast refinement for new applications using only a small training set for the new application, and a loop mechanism among operators in the training loop.

[0049] Figure 2B FIG. 250 shows a graph of the average differences between layers of various micro-training networks according to one embodiment. As shown, the vertical axis 252 indicates the overall differences between the layer coefficients (weights and biases) of various micro-trained neural networks (G1, G2, etc.) that are from the same parent (i.e., the initially trained neural network G0) but have different micro-training or micro-training depths represented by lines 255, 256, 257, and 258. The horizontal axis 254 contains discrete markers, each marker representing the weights and biases of a different neural network layer of a specific neural network topology. As shown, the difference in layer coefficients indicated by line 255 is generally greater than the difference in layer coefficients indicated by line 258. Additionally, the neural network associated with line 255 has been micro-trained to be further away from the parent neural network than the neural network associated with line 258.

[0050] As shown by the overall shape of the weight and bias differences of various micro-trained neural networks, the small iterative steps and low learning rates associated with micro-training do not change the overall computational structure of the micro-trained neural network. Preserving the computational structure between neural networks enables operations such as comparison and interpolation between a parent network and different networks generated using micro-training. For example, an image sharpening neural network can be trained to increase the sharpness of a synthesized output image, but the resulting output image may be evaluated as oversharpened; thus, the sharpness can be reduced using the average or interpolation of the weights between the parent neural network and the image sharpening neural network. Such an interpolation step only requires interpolation of the weights and biases and does not require any additional training. More generally, computational synthesis can be performed between and within a parent neural network and a micro-trained network generated from the parent neural network.

[0051] Parallel processing architecture

[0052] Figure 3 FIG. shows a parallel processing unit (“PPU”) 300 according to at least one embodiment. In one embodiment, the PPU 300 is a multi-threaded processor implemented on one or more integrated circuit devices. The PPU 300 is a latency hiding architecture designed to process many threads in parallel. A thread (i.e., an execution thread) is an instance of a set of instructions configured to be executed by the PPU 300. In one embodiment, the PPU 300 is a graphics processing unit (GPU), which is configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data in order to generate two-dimensional (2D) image data for display on a display device (such as a liquid crystal display (LCD) device). In other embodiments, the PPU 300 is used for general-purpose computing. Although an exemplary parallel processor is provided here for illustrative purposes, it should be strongly noted that such a processor is presented only for illustrative purposes, and any processor can be employed to supplement and / or replace it.

[0053] One or more PPU 300s can be configured to accelerate high-performance computing (HPC), data center, and machine learning applications. The PPU 300 can be configured to accelerate many deep learning systems and applications, including autonomous vehicle platforms, deep learning, high-precision speech, image, text recognition systems, intelligent video analysis, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, etc.

[0054] As Figure 3As shown, the PPU 300 includes an input / output (I / O) unit 305, a front-end unit 315, a scheduler unit 320, a work distribution unit 325, a hub 330, a crossbar (Xbar) 370, one or more general processing clusters (GPCs) 350, and one or more memory partition units 380. The PPU 300 is interconnected via one or more high-speed NV Links 310 to a host processor or other PPU 300. The PPU 300 is connected to a local memory 304 that includes multiple memory devices. In one embodiment, the local memory includes multiple dynamic random access memory (DRAM) devices. The DRAM devices can be configured as a high-bandwidth memory (HBM) subsystem, and multiple DRAM dies are stacked within each device.

[0055] The NVLink 310 interconnect enables the system to scale and include one or more PPUs in combination with one or more CPUs 300, supports cache coherence between the PPU 300 and the CPU, and CPU master control. Data and / or commands can be transferred by the NVLink 310 through the hub 330 to or from other units of the PPU 300, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown) of the PPU 300. NVLink 310 is described in more detail in conjunction with Figure 5B NVLink 310 is described in more detail.

[0056] The I / O unit 305 is configured to send and receive communications (e.g., commands, data) from a host processor (not shown) via an interconnect 302. The I / O unit 305 communicates directly with the host processor via the interconnect 302 or via one or more intermediate devices (such as a memory bridge). In one embodiment, the I / O unit 305 can communicate with one or more other processors (such as one or more PPU 300s) via the interconnect 302. In one embodiment, the I / O unit 305 implements a peripheral component interconnect express (PCIe) interface for communication via a PCIe bus and the interconnect 302 is a PCIe bus. In an alternative embodiment, the I / O unit 305 implements any other well-known type of interface for communicating with external devices.

[0057] The I / O unit 305 decodes packets received via the interconnect 302. In one embodiment, the packets represent commands configured to cause the PPU 300 to perform various operations. The I / O unit 305 sends the decoded commands to various other units of the PPU 300 as specified by the commands. For example, some commands are sent to the front-end unit 315. Other commands are sent to the hub 330 or other units of the PPU 300, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, the I / O unit 305 is configured to route communications between the various logical units of the PPU 300.

[0058] In one embodiment, a program executed by a host processor encodes a command stream in a buffer that provides a workload to the PPU 300 for processing. The workload includes instructions and data to be processed by those instructions. The buffer is an area in a memory that is accessible (e.g., read / write) by both the host processor and the PPU 300. For example, the I / O unit 305 may be configured to access the buffer in the system memory connected to the interconnect 302 via a memory request transmitted through the interconnect 302. In one embodiment, the host processor writes the command stream to the buffer and then transmits a pointer to the start of the command stream to the PPU 300. The front-end unit 315 receives pointers to one or more command streams. The front-end unit 315 manages one or more streams, reads commands from those streams, and forwards the commands to the various units of the PPU 300.

[0059] The front-end unit 315 is coupled to a scheduler unit 320 that configures the various GPCs 350 to process tasks defined by one or more streams. The scheduler unit 320 is configured to track status information related to the various tasks managed by the scheduler unit 320. This status may indicate which GPC 350 a task is assigned to, whether the task is active or inactive, the priority associated with the task, etc. The scheduler unit 320 manages the execution of multiple tasks on one or more GPCs 350.

[0060] The scheduler unit 320 is coupled to a work distribution unit 325 that is configured to dispatch tasks for execution on the GPCs 350. The work distribution unit 325 keeps track of a plurality of scheduled tasks received from the scheduler unit 320. In one embodiment, the work distribution unit 325 manages a pending task pool and an active task pool for each GPC 350. The pending task pool includes a plurality of time slots (e.g., 32 time slots) that contain tasks assigned to be processed by a particular GPC 350. The active task pool may include a plurality of time slots (e.g., 4 time slots) for tasks actively being processed by the GPC 350. As the GPC 350 completes the execution of a task, the task is evicted from the active task pool of the GPC 350 and one of the other tasks from the pending task pool is selected and scheduled for execution on the GPC 350. If an active task is idle on the GPC 350, e.g., while waiting for a data dependency to be resolved, the active task is evicted from the GPC 350 and returned to the pending task pool while another task from the pending task pool is selected and scheduled for execution on the GPC 350.

[0061] The work distribution unit 325 communicates with one or more GPCs 350 via an XBar 370. The XBar 370 is an interconnect network that couples many of the units of the PPU 300 to other units of the PPU 300. For example, the XBar 370 can be configured to couple the work distribution unit 325 to a particular GPC 350. Although not explicitly shown, other units of one or more PPUs 300 can also be connected to the XBar 370 via a hub 330.

[0062] Tasks are managed by the scheduler unit 320 and dispatched by the work distribution unit 325 to one of the GPCs 350. The GPC 350 is configured to process the tasks and produce results. The results can be consumed by other tasks in the GPC 350, routed to a different GPC 350 via the XBar 370, or stored in the memory 304. The results can be written to the memory 304 by a memory partitioning unit 380 that implements a memory interface for writing data to or reading data from the memory 304. The results can be sent via an NVLink 310 to another PPU 300 or CPU. In one embodiment, the PPU 300 includes U partitioning units 380 that are equal to the number of separate and distinct memory devices 304 coupled to the PPU 300. The memory partitioning unit 380 is described in more detail below in connection with Figure 4B Describe the memory partitioning unit 380 in more detail.

[0063] In one embodiment, a host processor executes a driver kernel that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 300. In one embodiment, multiple compute applications are executed simultaneously by the PPU 300, and the PPU 300 provides isolation, quality of service (“QoS”), and independent address spaces for the multiple compute applications. The application generates instructions (e.g., API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 300. The driver kernel outputs the tasks to one or more streams processed by the PPU 300. Each task includes one or more related thread groups, which may be referred to as warps. In one embodiment, a warp includes 32 related threads that may be executed in parallel. Cooperative threads may refer to multiple threads, including instructions for executing a task and exchanging data through shared memory. Threads and cooperative threads are described in more detail in conjunction with Figure 5A Threads and cooperative threads are described in more detail.

[0064] Figure 4A Shows a Figure 3 GPC 350 of the PPU 300 according to one embodiment. As Figure 4A shown, each GPC 350 includes multiple hardware units for processing tasks. In one embodiment, each GPC 350 includes a pipeline manager 410, a pre-raster operation unit (PROP) 415, a raster engine 425, a work distribution crossbar (WDX) 480, a memory management unit (MMU) 490, and one or more data processing clusters (DPCs) 420. It will be understood that Figure 4A the GPC 350 of Figure 4A may include units other than those shown in Figure 4A or additional hardware units in addition to those shown.

[0065] In one embodiment, the operation of GPC 350 is controlled by pipeline manager 410. Pipeline manager 410 manages the configuration of one or more DPCs 420 to process tasks assigned to GPC 350. In one embodiment, manager 410 may configure at least one of one or more DPCs 420 to implement at least a portion of a graphics rendering pipeline. For example, DPC 420 may be configured to execute a vertex shader program on programmable streaming multiprocessors (SMs) 440. Pipeline manager 410 may also be configured to route packets received from work assignment unit 325 to appropriate logic units within GPC 350. For example, some packets may be routed to fixed function hardware units in PROP 415 and / or raster engine 425, while other packets may be routed to DPC 420 for processing by primitive engine 435 or SM 440. In one embodiment, pipeline manager 410 may configure at least one of one or more DPCs 420 to implement a neural network model and / or a compute pipeline.

[0066] PROP unit 415 is configured to route data generated by raster engine 425 and DPC 420 to raster operation (ROP) units, as described in more detail Figure 4B below. PROP unit 415 may also be configured to perform optimizations for color blending, organize pixel data, perform address translation, etc.

[0067] Raster engine 425 includes a plurality of fixed function hardware units configured to perform various raster operations. In one embodiment, raster engine 425 includes a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, and a tile merge engine. The setup engine receives transformed vertices and generates plane equations associated with the geometric primitives defined by those vertices. The plane equations are transmitted to the coarse raster engine to generate coverage information for the primitives (e.g., x, y coverage masks for tiles). The output of the coarse raster engine is transmitted to the culling engine where fragments associated with primitives that fail a z-test are culled and transmitted to the clipping engine where fragments that lie outside the view frustum are clipped. Those fragments that have been clipped and culled may be passed to the fine raster engine to generate the attributes of the pixel fragments based on the plane equations generated by the setup engine. The output of raster engine 425 includes, for example, fragments that will be processed by a fragment shader implemented within DPC 420.

[0068] Each DPC 420 included in GPC 350 includes an M-Pipeline Controller (MPC) 430, a primitive engine 435, and one or more SMs 440. The MPC 430 controls the operation of the DPC 420 and routes packets received from the pipeline manager 410 to the appropriate units in the DPC 420. For example, packets associated with a vertex can be routed to the primitive engine 435, which is configured to fetch vertex attributes associated with the vertex from the memory 304. In contrast, packets associated with a shader program can be sent to the SM 440.

[0069] The SM 440 includes programmable stream processors configured to process tasks represented by multiple threads. Each SM 440 is multi-threaded and configured to execute multiple threads (e.g., 32 threads) from a particular thread group simultaneously. In one embodiment, the SM 440 implements a SIMD (Single Instruction Multiple Data) architecture, where each thread in a group of threads (e.g., a warp) is configured to process different data sets based on the same set of instructions. All threads in the thread group execute the same instructions. In another embodiment, the SM 440 implements a SIMT (Single Instruction Multiple Threads) architecture, where each thread in a group of threads is configured to process different data sets based on the same set of instructions, but individual threads within the thread group are allowed to diverge during execution. In one embodiment, a program counter, a call stack, and an execution state are maintained for each warp, enabling concurrency between the warp and serial execution within the warp when threads within the warp diverge. In another embodiment, a program counter, a call stack, and an execution state are maintained for each individual thread, enabling equal concurrency between all threads, within and between warps. When the execution state is maintained for each individual thread, threads executing the same instructions can converge and execute in parallel to achieve maximum efficiency. The SM 440 will be described in more detail below in conjunction with Figure 5A the SM 440 will be described in more detail.

[0070] The MMU 490 provides an interface between the GPC 350 and the memory partition unit 380. The MMU 490 can provide virtual address to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, the MMU 490 provides one or more Translation Lookaside Buffers (TLBs) for translating virtual addresses to physical addresses in the memory 304.

[0071] Figure 4B illustrates the memory partition unit 380 of the PPU 300 according to one embodiment. As Figure 3 shown in Figure 4BAs shown, the memory partitioning unit 380 includes a raster operation (ROP) unit 450, a level 2 (L2) cache 460, and a memory interface 470. The memory interface 470 is coupled to the memory 304. The memory interface 470 can implement a 32, 64, 128, 1024-bit data bus, etc. for high-speed data transfer. In one embodiment, the PPU 300 includes U memory interfaces 470, one memory interface 470 for each pair of memory partitioning units 380, where each pair of memory partitioning units 380 is connected to a corresponding memory device of the memory 304. For example, the PPU 300 can be connected to up to Y storage devices, such as a high-bandwidth memory stack or a graphics double data rate version 5, synchronous dynamic random access memory, or other types of persistent storage.

[0072] In one embodiment, the memory interface 470 implements an HBM2 memory interface, and Y is equal to half of U. In one embodiment, the HBM2 memory stack is on the same physical package as the PPU 300, providing significant power and area savings compared to a traditional GDDR5 SDRAM system. In one embodiment, each HBM2 stack includes four memory dies, and Y is equal to 4, while each die of the HBM2 stack includes two 128-bit channels, for a total of 8 channels and a 1024-bit data bus width.

[0073] In one embodiment, the memory 304 supports single error correction double error detection (SECDED) error correction code (ECC) to protect data. ECC provides higher reliability for computing applications that are sensitive to data corruption. Reliability is particularly important in large-scale cluster computing environments where the PPU300 processes very large data sets and / or long-running applications.

[0074] In one embodiment, the PPU 300 implements a multi-level memory hierarchy. In one embodiment, the memory partitioning unit 380 supports unified memory to provide a single unified virtual address space for the CPU and PPU 300 memories, enabling data sharing between virtual memory systems. In one embodiment, the access frequency of the PPU 300 to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU 300 that accesses the pages more frequently. In one embodiment, the NVLink 310 supports address translation services, allowing the PPU 300 to directly access the CPU's page table and providing the PPU 300 with full access to the CPU memory.

[0075] In one embodiment, the copy engine transfers data between multiple PPUs 300 or between a PPU 300 and a CPU. The copy engine can generate a page fault for an address not mapped to a page table. Then, the memory partitioning unit 380 can service the page fault, map the address to the page table, and then the copy engine can perform the transfer. In a conventional system, memory is pinned (e.g., non-pageable) for multiple copy engine operations between multiple processors, thereby substantially reducing the available memory. When a hardware page fault occurs, the address can be passed to the copy engine without worrying about whether the memory page is resident, and the copy process is transparent.

[0076] Data from the memory 304 or other system memory can be fetched by the memory partitioning unit 380 and stored in the L2 cache 460, which is on-chip and shared among the respective GPCs 350. As shown, each memory partitioning unit 380 includes a portion of the L2 cache 460 associated with the corresponding memory 304. Then, lower-level caches can be implemented in the respective units within the GPC 350. For example, each SM 440 can implement a level-1 (L1) cache. The L1 cache is dedicated memory for a particular SM 440. Data from the L2 cache 460 can be fetched and stored in each L1 cache for processing in the functional units of the SM 440. The L2 cache 460 is coupled to the memory interface 470 and the XBar 370.

[0077] The ROP unit 450 performs graphics raster operations related to pixel color, such as color compression, pixel blending, etc. The ROP unit 450 also implements a depth test together with the raster engine 425, receiving the depth of the sample position associated with the pixel fragment from the culling engine of the raster engine 425. The depth is tested against the corresponding depth in the depth buffer for the sample position associated with the fragment. If the fragment passes the depth test for the sample position, the ROP unit 450 updates the depth buffer and sends the result of the depth test to the raster engine 425. It will be appreciated that the number of memory partitioning units 380 can be different from the number of GPCs 350, and thus, each ROP unit 450 can be coupled to each GPC 350. The ROP unit 450 tracks the data packets received from different GPCs 350 and determines which GPC 350 to route the results generated by the ROP unit 450 through the Xbar 370. Although in Figure 4B the ROP unit 450 is included within the memory partitioning unit 380, in other embodiments, the ROP unit 450 can be external to the memory partitioning unit 380. For example, the ROP unit 450 can reside in the GPC 350 or other units.

[0078] Figure 5A illustrates a streaming multiprocessor 440 according to one embodiment. As Figure 4A shown, the SM 440 includes an instruction cache 505, one or more scheduler units 510, a register file 520, one or more processor cores 550, one or more special function units (SFUs) 552, one or more load / store units (LSUs) 554, an interconnect network 580, and a shared memory / L1 cache 570. Figure 5A As described above, the work distribution unit 325 dispatches tasks to be executed on the GPC 350 of the PPU 300. The task is assigned to a specific DPC 420 within the GPC 350, and if the task is associated with a shader program, the task can be assigned to the SM440. The scheduler unit 510 receives tasks from the work distribution unit 325 and manages the instruction scheduling of one or more thread blocks assigned to the SM 440. The scheduler unit 510 schedules the thread blocks to execute as warps of parallel threads, where each thread block is assigned at least one warp. In one embodiment, each warp executes 32 threads. The scheduler unit 510 can manage multiple different thread blocks, assign warps to different thread blocks, and then schedule instructions from multiple different cooperating groups to respective functional units (e.g., cores 550, SFUs 552, and LSUs 554) within each clock cycle.

[0079] A cooperating group is a programming model for organizing groups of communicating threads that allows developers to express the granularity at which threads are communicating, enabling the expression of richer and more efficient parallel decompositions. The cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. Traditional programming models provide a single simple construct for synchronizing cooperating threads: a barrier that spans all threads of a thread block (e.g., the syncthreads() function). However, programmers often want to define thread groups at a size smaller than the thread block granularity and synchronize within the defined groups to achieve higher performance, design flexibility, and reuse software in the form of collective-scoped functional interfaces.

[0080] Cooperating groups enable programmers to explicitly define thread groups at sub-block (e.g., as small as a single thread) and multi-block granularities and perform collective operations such as synchronizing the threads within the cooperating group. The programming model supports clean composition across software boundaries, so library and utility functions can synchronize safely within their local contexts without having to make assumptions about convergence. The cooperating group primitives enable new cooperative parallel patterns, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire thread block grid.

[0081]

[0082] ​The dispatch unit 515 is configured to send instructions to one or more functional units. In this embodiment, the scheduler unit 510 includes two dispatch units 515, which enables scheduling of two different instructions from the same warp per clock cycle. In an alternative embodiment, each scheduler unit 510 may include a single dispatch unit 515 or additional dispatch units 515.

[0083] Each SM 440 includes a register file 520, which provides a set of registers for the functional units of the SM 440. In one embodiment, the register file 520 is partitioned among each functional unit such that each functional unit is assigned a dedicated portion of the register file 520. In another embodiment, the register file 520 is partitioned among different threads executed by the SM 440. The register file 520 provides temporary storage for operands of data paths connected to the functional units.

[0084] Each SM 440 includes L processing cores 550. In one embodiment, the SM 440 includes a large number (e.g., 128, etc.) of different processing cores 550. Each core 550 may include fully pipelined, single-precision, double-precision, and / or mixed-precision processing units, and the mixed-precision processing unit includes a floating-point arithmetic logic unit and an integer arithmetic logic unit. In one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In one embodiment, the core 550 includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0085] The tensor cores are configured to perform matrix operations, and in one embodiment, one or more tensor cores are included in the core 550. In particular, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In one embodiment, each tensor core operates on a 4×4 matrix and performs matrix multiplication and accumulation operations D = A×B + C, where A, B, C, and D are 4×4 matrices.

[0086] In one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, while the accumulation matrices C and D can be 16-bit floating-point or 32-bit floating-point matrices. The tensor core operates on 16-bit floating-point input data with 32-bit floating-point accumulation. The 16-bit floating-point multiplication requires 64 operations and produces a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point addition for 4x4x4 matrix multiplication. In practice, tensor cores are used to perform larger two-dimensional or higher-dimensional matrix operations that are composed of these smaller elements. APIs such as the CUDA 9 C++ API expose specialized matrix load, matrix multiplication and accumulation, and matrix store operations to efficiently use the tensor cores in CUDA-C++ programs. At the CUDA level, the warp-level interface assumes that matrices of size 16×16 span all 32 threads of a warp.

[0087] Each SM 440 also includes M SFUs 552 that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, the SFU 552 can include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, the SFU 552 can include a texture unit configured to perform texture map filtering operations. In one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texture pixels) from the memory 304 and sample the texture map to produce sampled texture values for use in a shader program executed by the SM 440. The texture map is then stored in the shared memory / L1 cache 470. The texture unit uses mip-maps (e.g., texture maps with different levels of detail) to implement texture operations such as filtering operations. In one embodiment, each SM 340 includes two texture units.

[0088] Each SM 440 also includes N LSUs 554 that implement load and store operations between the shared memory / L1 cache 570 and the register file 520. Each SM 440 includes an interconnect network 580 that connects each functional unit to the register file 520 and connects the LSU 554 to the register file 520 and the shared memory / L1 cache 570. In one embodiment, the interconnect network 580 is a crossbar switch that can be configured to connect any functional unit to any register in the register file 520 and connect the LSU 554 to the register file and memory locations in the shared memory / L1 cache 570.

[0089] The shared memory / L1 cache 570 is an array of on-chip memory that allows data storage and communication between the SM 440 and the primitive engine 435, and between threads within the SM 440. In one embodiment, the shared memory / L1 cache 570 includes a storage capacity of 128 KB and is in the path from the SM 440 to the memory partition unit 380. The shared memory / L1 cache 570 can be used for caching reads and writes. One or more of the shared memory / L1 cache 570, the L2 cache 460, and the memory 304 are backing stores.

[0090] Combining the data cache and shared memory functions into a single memory block provides optimal overall performance for both types of memory access. This capacity can be used as a cache for programs that do not use shared memory. For example, if the shared memory is configured to use half of its capacity, texture and load / store operations can use the remaining capacity. The integration within the shared memory / L1 cache 570 enables the shared memory / L1 cache 570 to be used as a high-throughput pipeline for streaming data, while providing high bandwidth and low-latency access to frequently reused data.

[0091] When configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. Specifically, Figure 3 the fixed-function graphics processing units shown in are bypassed, creating a simpler programming model. In the general-purpose parallel computing configuration, the work distribution unit 325 directly assigns and distributes thread blocks to the DPC 420. Threads within a block execute the same program using unique thread IDs in the computation to ensure that each thread uses the SM 440 to execute the program and perform the computation to generate a unique result, communicate between threads using the shared memory / L1 cache 570, and read and write to global memory using the LSU 554 through the shared memory / L1 cache 570 and the memory partition unit 380. When configured for general-purpose parallel computing, the SM 440 can also write commands that the scheduler unit 320 can use to start new work on the DPC 420.

[0092] The PPU 300 can be included in a desktop computer, laptop computer, tablet computer, server, supercomputer, smart phone (e.g., wireless, handheld device), personal digital assistant (PDA), digital camera, vehicle, head-mounted display, handheld electronic device, etc. In one embodiment, the PPU 300 is embodied on a single semiconductor substrate. In another embodiment, the PPU 300 is included in a system-on-chip (SoC) together with one or more other devices (e.g., additional PPU 300s, memory 304, reduced instruction set computer (RISC) CPU, memory management unit (MMU), digital-to-analog converter (DAC), etc.).

[0093] In one embodiment, the PPU 300 can be included on a graphics card that includes one or more storage devices. The graphics card can be configured to interact with a PCIe slot on a desktop computer motherboard. In yet another embodiment, the PPU 300 can be an integrated graphics processing unit (iGPU) or parallel processor in a chipset of a motherboard.

[0094] Exemplary computing system

[0095] As developers expose and utilize more parallelism in applications such as artificial intelligence computing, systems with multiple GPUs and CPUs are being used in various industries. High-performance GPU-accelerated systems with tens to thousands of computing nodes have been deployed in data centers, research institutions, and supercomputers to solve increasingly large problems. As the number of processing devices in high-performance systems increases, communication and data transfer mechanisms need to scale to support the increased bandwidth.

[0096] Figure 5B is a conceptual diagram of a processing system 500 implemented using Figure 3 the PPU 300 according to one embodiment. The exemplary system 565 can be configured to implement Figure 1A the method 110 shown and / or Figure 2A the method 200 shown. The processing system 500 includes a CPU 530, a switch 510, and multiple PPU 300s, as well as corresponding memory 304. The NVLink 310 provides a high-speed communication link between each PPU 300. Although a specific number of NVLink 310 and interconnect 302 connections are shown as Figure 5B shown, the number of connections to each PPU 300 and the CPU 530 can vary. The switch 510 interacts between the interconnect 302 and the CPU 530. The PPU 300, memory 304, and NVLink 310 can be located on a single semiconductor platform to form a parallel processing module 525. In one embodiment, the switch 510 supports two or more protocols that interact between various different connections and / or links.

[0097] In another embodiment (not shown), NVLink 310 provides one or more high-speed communication links between each PPU 300 and CPU 530, and switch 510 interfaces between interconnect 302 and each PPU 300. PPU 300, memory 304, and interconnect 302 may be located on a single semiconductor platform to form a parallel processing module 525. In yet another embodiment (not shown), interconnect 302 provides one or more communication links between each PPU 300 and CPU 530, and switch 510 interfaces between each PPU 300 using NVLink 310 to provide one or more high-speed communication links between the PPUs 300. In another embodiment (not shown), NVLink 310 provides one or more high-speed communication links between PPU 300 and CPU 530 through switch 510. In yet another embodiment (not shown), interconnect 302 directly provides one or more communication links between each PPU 300. One or more of the NVLink 310 high-speed communication links may be implemented as a physical NVLink interconnect or an on-chip or die interconnect using the same protocol as NVLink 310.

[0098] In the context of this specification, a single semiconductor platform may refer to a single, solitary semiconductor-based integrated circuit fabricated on a die or chip. It should be noted that the term "single semiconductor platform" may also refer to a multi-chip module with increased connectivity that emulates on-chip operation and represents a substantial improvement over traditional bus implementations. Of course, various circuits or devices may be placed separately or in various combinations of semiconductor platforms according to the user's requirements. Alternatively, parallel processing module 525 may be implemented as a circuit board substrate, and each PPU 300 and / or memory 304 may be a packaged device. In one embodiment, CPU 530, switch 510, and parallel processing module 525 are located on a single semiconductor platform.

[0099] In one embodiment, the signaling rate of each NVLink 310 is 20 to 25 Gigabits / s, and each PPU 300 includes six NVLink 310 interfaces (as Figure 5B shown, each PPU 300 includes five NVLink 310 interfaces). Each NVLink 310 provides a data transfer rate of 25 Gigabytes / s in each direction, where six links provide 300 Gigabytes / s. When CPU 530 also includes one or more NVLink 310 interfaces, the NVLinks 310 may be dedicated to PPU-to-PPU communication (as Figure 5Bshown), or for communication for some combination of PPU-to-PPU and PPU-to-CPU.

[0100] In one embodiment, NVLink 310 allows direct load / store / atomic access from CPU 530 to each PPU 300 memory 304. In one embodiment, NVLink 310 supports coherence operations, allowing data read from memory 304 to be stored in the cache hierarchy of CPU 530, which reduces the cache access latency of CPU 530. In one embodiment, NVLink 310 includes support for address translation services (ATS), allowing PPU 300 to directly access the page tables within CPU 530. One or more of NVLink 310 can also be configured to operate in a low power mode.

[0101] Figure 5C An exemplary system 565 is shown in which various architectures and / or functions of the various previous embodiments can be implemented. Exemplary system 565 can be configured to implement Figure 1A method 110 as shown and Figure 2A method 200 as shown.

[0102] As shown, a system 565 is provided that includes at least one central processing unit 530 connected to a communication bus 575. The communication bus 575 can be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect), PCI-Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol. System 565 also includes a main memory 540. Control logic (software) and data are stored in the main memory 540, which can take the form of random access memory (RAM).

[0103] System 565 also includes an input device 560, a parallel processing system 525, and a display device 545, such as a conventional CRT (Cathode Ray Tube), LCD (Liquid Crystal Display), LED (Light Emitting Diode), plasma display, etc. User input can be received from the input device 560, which can be, for example, a keyboard, mouse, touchpad, microphone, etc. Each of the foregoing modules and / or devices can even be located on a single semiconductor platform to form system 565. Alternatively, depending on the user's needs, the various modules can also be located separately in semiconductor platforms or in various combinations of semiconductor platforms.

[0104] In addition, system 565 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) such as the Internet, a peer-to-peer network, a cable network interface, etc.) through a network interface 535 for communication purposes.

[0105] The system 565 may also include auxiliary memory (not shown). The auxiliary memory 610 includes, for example, a hard disk drive and / or a removable storage drive, which represent a floppy disk drive, a magnetic tape drive, an optical disk drive, a digital versatile disk (DVD) drive, a recording device, and a universal serial bus (USB) flash drive. The removable storage drive reads from and / or writes to a removable storage unit in a well-known manner.

[0106] A computer program or a computer control logic algorithm may be stored in the main memory 540 and / or the auxiliary memory. When such a computer program is executed, it enables the system 565 to perform various functions. The memory 540, the memory, and / or any other memory are possible examples of computer-readable media.

[0107] The architecture and / or functionality of each of the previous figures may be implemented in the context of a general-purpose computer system, a circuit board system, a game console system dedicated to entertainment purposes, a dedicated system, and / or other required systems. For example, the system 565 may take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smart phone (e.g., a wireless, handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, a mobile phone device, a television, a workstation, a game console, an embedded system, and / or any other type of logic.

[0108] Although various embodiments have been described above, it should be understood that they are merely exemplary and not restrictive. Therefore, the breadth and scope of the preferred embodiments should not be limited by any of the above exemplary embodiments, but should be defined only in accordance with the appended claims and their equivalents.

[0109] Machine learning

[0110] Deep neural networks (DNNs) developed on processors such as the PPU 300 have been used in a variety of use cases, from self-driving cars to faster drug development, from automatic image captioning in online image databases to translation in intelligent real-time language video chat applications. Deep learning is a technique that mimics the neural learning process of the human brain, continuously learning, becoming smarter over time, and providing more accurate results more quickly. Initially, an adult teaches a child to correctly identify and classify various shapes, and eventually the child can identify shapes without any guidance. Similarly, a deep learning or neural learning system needs to be trained in object recognition and classification as it becomes smarter and more efficient at recognizing basic objects, occluded objects, etc., while also assigning context to the objects.

[0111] At the simplest level, neurons in the human brain view the various inputs they receive, assign an importance level to each of these inputs, and pass the output to other neurons for them to operate on. An artificial neuron or perceptron is the most basic model of a neural network. In one example, a perceptron can receive one or more inputs representing the various features of an object that the perceptron is trained to recognize and classify, and assign a certain weight to each of these features based on the importance of that feature in defining the shape of the object.

[0112] A deep neural network (DNN) model consists of many layers of connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.) that can be trained with large amounts of input data to quickly solve complex problems with high precision. In one example, the first layer of a DNN model breaks down an input image of a car into its individual parts and looks for basic patterns such as lines and angles. The second layer assembles the lines to look for higher-level patterns, such as wheels, windshields, and rearview mirrors. The next layer identifies the type of vehicle, and the last few layers generate a label for the input image, identifying the model of a specific car brand.

[0113] Once the DNN is trained, it can be deployed and used to identify and classify objects or models in a process called inference. Examples of inference (the process by which the DNN extracts useful information from a given input) include identifying the handwritten numbers on a check deposited in an ATM, identifying the image of a friend in a photo, providing movie recommendations to over fifty million users, identifying and classifying cars, pedestrians, and road hazards in different types of self-driving cars, or translating human speech in real time.

[0114] During training, data flows through the DNN in the forward propagation phase until a prediction indicating the label corresponding to the input is produced. If the neural network does not correctly label the input, the error between the correctly labeled and the predicted label is analyzed and the weights of each feature are adjusted in the backpropagation phase until the DNN correctly labels the input and other inputs in the training dataset. Training complex neural networks requires a large amount of parallel computing performance, including floating-point multiplication and addition supported by the PPU 300. Inference is less computationally intensive than training, which is a latency-sensitive process where the trained neural network is applied to new inputs that have never been seen before for tasks such as classifying images, translating speech, and inferring new information.

[0115] Neural networks rely heavily on matrix mathematical operations, and complex multi-layer networks require a large amount of floating-point performance and bandwidth to improve efficiency and speed. The PPU 300 has thousands of processing cores, is optimized for matrix mathematical operations, and provides performance in the tens to hundreds of TFLOPS, making it a computing platform capable of providing the performance required for deep neural network-based artificial intelligence and machine learning applications.

[0116] Note that the techniques described herein (e.g., methods 110 and 200) may be embodied in executable instructions stored on a computer-readable medium for use by or in conjunction with a processor-based instruction execution machine, system, apparatus, or device. Those skilled in the art will recognize that for some embodiments, various types of computer-readable media may be included to store data. As used herein, "computer-readable medium" includes one or more of any suitable medium for storing executable instructions of a computer program such that the instruction execution machine, system, apparatus, or device can read (or retrieve) the instructions from the computer-readable medium and execute the instructions for performing the described embodiments. Suitable storage formats include one or more of electronic, magnetic, optical, and electromagnetic formats. A non-exhaustive list of conventional exemplary computer-readable media includes: portable computer floppy disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash devices, and optical storage devices including portable compact discs (CDs), portable digital video discs (DVDs), etc.

[0117] It should be understood that the arrangement of the components shown in the figures is for illustrative purposes and that other arrangements are possible. For example, one or more of the elements described herein may be implemented in whole or in part as electronic hardware components. Other elements may be implemented as software, hardware, or a combination of software and hardware. Moreover, some or all of these other elements may be combined, some other elements may be entirely omitted, and additional components may be added while still implementing the functions described herein. Thus, the subject matter described herein may be embodied in many different variations, and all such variations are considered to be within the scope of the claims.

[0118] To assist in understanding the subject matter described herein, many aspects are described in terms of sequences of actions. Those skilled in the art will recognize that the various actions may be performed by dedicated circuitry or circuits, program instructions executed by one or more processors, or a combination of both. The description of any sequence of actions herein is not intended to imply that a particular order must be followed to perform that sequence. Unless otherwise indicated herein or clearly contradicted by the context, all of the methods described herein may be performed in any suitable order.

[0119] In the context of the described subject matter, particularly in the context of the appended claims, the use of the terms "a", "an", and "the" and similar references should be construed to cover both the singular and plural forms unless otherwise specified herein or clearly contradicted by the context. The phrase "at least one" followed by a list of one or more items (e.g., "at least one of A and B") should be construed to mean either one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise specified herein or clearly contradicted by the context. Further, the foregoing description is for illustrative purposes only and not for purposes of limitation, as the scope of the sought protection is defined by the claims set forth below and their equivalents. The use of any and all examples or exemplary language (e.g., "such as") provided herein is for illustrative purposes only and is not intended to limit the scope of the subject matter unless otherwise required. In the claims and the written description, the use of the term "based on" and other similar phrases indicates the conditions that result in the outcome and is not intended to exclude any other conditions that result in that outcome. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the claimed invention.

Claims

1. A method for improving a neural network, comprising: Receiving a neural network trained using a first set of hyperparameters and a first training dataset to satisfy a loss function, wherein the trained neural network generates output data including visual artifacts; Receiving a second training dataset, wherein the second training dataset includes additional training samples selected to train the trained neural network to suppress the visual artifacts; Receiving a second set of hyperparameters, wherein a second learning parameter specified in the second set of hyperparameters restricts adjustment of one or more weights used by the neural network as compared to a corresponding first learning parameter in the first set of hyperparameters; and Applying the second training dataset to the neural network according to the second set of hyperparameters while adjusting the one or more weights, the neural network using the one or more weights to process the second training dataset to produce a first micro-trained neural network.

2. The method according to claim 1, wherein, The first learning parameter includes a first learning rate, and the second learning parameter includes a second learning rate less than the first learning rate.

3. The method according to claim 2, wherein, The second learning rate is at least ten times lower than the first learning rate.

4. The method according to claim 1 further comprises: Determining that a completion requirement has been satisfied.

5. The method according to claim 4, wherein determining includes receiving an input indication from a user interface.

6. The method according to claim 1 further includes: Generating and displaying test images from corresponding training images in the second training dataset using the first micro-trained neural network, wherein visual artifacts in the test images are reduced as compared to second test images generated by the neural network for the corresponding training images.

7. The method according to claim 1, wherein The visual artifacts include geometric aliasing artifacts.

8. The method according to claim 1, wherein, The visual artifacts include rendering noise artifacts.

9. A system for improving a neural network, comprising: A storage circuit storing programming instructions; A parallel processing unit coupled to the storage circuit, wherein the parallel processing unit retrieves and executes the programming instructions to: Receive a neural network trained using a first set of hyperparameters and a first training dataset to satisfy a loss function, wherein the trained neural network generates output data including visual artifacts; Receive a second training dataset, wherein the second training dataset includes additional training samples selected to train the trained neural network to suppress the visual artifacts; Receive a second set of hyperparameters, wherein a second learning parameter specified in the second set of hyperparameters restricts adjustment of one or more weights used by the neural network as compared to a corresponding first learning parameter in the first set of hyperparameters; and Apply the second training dataset to the neural network according to the second set of hyperparameters while adjusting the one or more weights used by the neural network to process the second training dataset to produce a first micro-trained neural network.

10. A non-transitory computer-readable medium storing computer instructions for facial analysis, which when executed by one or more processors cause the one or more processors to perform the following operations: Receive a neural network trained using a first set of hyperparameters and a first training dataset to satisfy a loss function, where, The trained neural network generates output data including visual artifacts; Receive a second training dataset, where the second training dataset includes additional training samples selected to train the trained neural network to suppress the visual artifacts; Receive a second set of hyperparameters, where a second learning parameter specified in the second set of hyperparameters restricts adjustment of one or more weights used by the neural network compared to a corresponding first learning parameter in the first set of hyperparameters; and Apply the second training dataset to the neural network according to the second set of hyperparameters while adjusting the one or more weights used by the neural network to process the second training dataset to produce a first micro-trained neural network.