Cell image synthesis using one or more neural networks

Through multi-condition generation adversarial networks, the integration of gene expression and medical image data is generated to generate high-quality synthetic images, solving the problem of artificial interdependence in the prior art and achieving high efficiency and accuracy of medical image analysis.

CN112102329BActive Publication Date: 2025-09-02NVIDIA CORP

Patent Information

Application Number
CN202010529158.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-17
Filing Date
2020-06-11
Publication Date
2025-09-02
Estimated Expiration
2041-03-05

AI Technical Summary

Technical Problem

The prior art requires human interaction in the process of generating medical image analysis in the medical field, resulting in limited accuracy.

Method used

Multi-conditional generative adversarial network (GAN) is used to integrate gene expression data and medical image data. Through a combined model of generator and discriminator, synthetic images with discriminant radiogenomic maps are generated to achieve end-to-end image and gene feature association.

Benefits of technology

It realizes efficient generation and analysis of medical images, improves the authenticity and quality of image synthesis, enhances the correlation between image features and gene data, and supports more accurate medical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112102329B_ABST
    Figure CN112102329B_ABST
Patent Text Reader

Abstract

The present invention discloses apparatus, systems, and techniques for synthesizing cell images using one or more neural networks, generating synthesized images comprising digital representations of cell groups realistically blended with an appropriate background image. In at least one embodiment, one or more neural networks are used to fuse background image data and gene expression data to generate such synthesized images.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Advances in computer technology have led to improved capabilities for object recognition and analysis. For example, in the medical field, computer technology can provide increasing accuracy in analyzing patients and diagnosing various diseases or conditions. However, the processes used to generate such analyses, which require at least some degree of human interaction or judgment, may have limited accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Various embodiments according to the present disclosure will be described with reference to the accompanying drawings, in which:

[0003] Figure 1 shows an example image that may be processed or generated using one or more neural networks in at least one embodiment;

[0004] Figure 2 Components of an example generator in at least one embodiment are shown;

[0005] Figure 3A and Figure 3B Components of an example fusion block and an example discriminator, respectively, are shown in at least one embodiment;

[0006] Figure 4 shows example inputs and example outputs of a composition architecture in at least one embodiment;

[0007] Figure 5 shows example genetic data clustering in at least one embodiment;

[0008] Figure 6A and Figure 6B illustrates example training and inference portions of a process for generating a synthetic image using at least one trained neural network that may be utilized in at least one embodiment;

[0009] Figure 7 An example environment is shown that may be used for implementation in at least one embodiment;

[0010] Figure 8 An example system for training an image synthesis network that may be utilized in at least one embodiment is shown;

[0011] Figure 9 illustrates layers of an example statistical model that may be utilized in at least one embodiment;

[0012] Figure 10 Inference and / or training logic in at least one embodiment is shown;

[0013] Figure 11Inference and / or training logic in at least one embodiment is shown;

[0014] Figure 12 The training and deployment of a deep neural network in at least one embodiment is shown;

[0015] Figure 13 An example data center system in at least one embodiment is shown;

[0016] Figure 14 is a block diagram illustrating a computer system in at least one embodiment;

[0017] Figure 15 is a block diagram illustrating a computer system in at least one embodiment;

[0018] Figure 16 A computer system in at least one embodiment is shown;

[0019] Figure 17 A computer system according to at least one embodiment is shown;

[0020] Figure 18 An exemplary integrated circuit and associated graphics processor that may be fabricated using one or more IP cores in at least one embodiment are shown;

[0021] Figures 19A-19B An exemplary integrated circuit and associated graphics processor that may be fabricated using one or more IP cores in at least one embodiment are shown;

[0022] Figure 20A-Figure 20B shows additional exemplary graphics processor logic in at least one embodiment;

[0023] Figure 21 A computer system in at least one embodiment is shown;

[0024] Figure 22 A parallel processor in at least one embodiment is shown;

[0025] Figure 23 A partition unit in at least one embodiment is shown;

[0026] Figure 24 illustrates a processing cluster in at least one embodiment;

[0027] Figure 25 A graphics multiprocessor in at least one embodiment is shown;

[0028] Figure 26is a block diagram illustrating a processor microarchitecture for a processor in at least one embodiment;

[0029] Figure 27 A deep learning application processor in at least one embodiment is shown;

[0030] Figure 28 is a block diagram illustrating an example neuromorphic processor in at least one embodiment;

[0031] Figure 29 At least a portion of a graphics processor in at least one embodiment is shown;

[0032] Figure 30 is a block diagram of at least a portion of a graphics processor core in at least one embodiment;

[0033] Figure 31A and Figure 31B Thread execution logic in at least one embodiment is shown;

[0034] Figure 32 A parallel processing unit ("PPU") in at least one embodiment is shown;

[0035] Figure 33 illustrates a general processing cluster ("GPC") in at least one embodiment;

[0036] Figure 34 A memory partitioning unit of a parallel processing unit ("PPU") in at least one embodiment is shown; and

[0037] Figure 35 A streaming multiprocessor in at least one embodiment is shown. DETAILED DESCRIPTION

[0038] Figure 1 A set of example images 100 that can be utilized in at least one embodiment is shown. In at least one embodiment, the neural network used to generate the synthetic image can take the form of a multi-conditional GAN ​​with a style specification. In at least one embodiment, foreground and background fusion can be modeled within the network, and the image and genetic code can be used for synthesis. In at least one embodiment, as Figure 1As shown, the network can accept a background image 102 as input, which represents the type of tissue or location where such a nodule may be located. Nodules are used as an example for illustration, but any grouping of cells or other such materials, as well as objects that may be completely unrelated to humans or living cells, can also be used. In at least one embodiment, in addition to other such options, the background image to be used for training can be selected from a set of background images, or used as a random portion of a larger background image. In at least one embodiment, the network can process the input image and gene data to generate a segmentation mask 104 indicating the background and foreground areas (nodules), as well as a composite image 106 showing the nodules mixed with the background image through a fusion process. In at least one embodiment, the gene expression data comes from actual training data. In at least one embodiment, a set of nodules can be analyzed, each of which will correspond to a specific vector of gene expression data and will be associated with image features. In at least one embodiment of the training and inference time, the gene expression data can be used together with the background image. Furthermore, in at least one embodiment, interpolated gene expression data may be used, which is generated using different types of gene codes for different nodules.

[0039] In at least one embodiment, imaging genomics can be used to determine correlations between cancer imaging features and gene expression. In at least one embodiment, imaging genomics can be used to diagnose disease using different imaging techniques, including magnetic resonance imaging (MR) for capturing image data related to potential brain tumors and computed tomography (CT) for capturing image data related to potential non-small cell lung cancer (NSCLC). In at least one embodiment, related tasks can be handled in a holistic, end-to-end manner. In at least one embodiment, image features can be learned from relevant training data and optimized for a specific task. In at least one embodiment, a network such as a generative adversarial network (GAN) can be used, which can fuse information from different sources to generate a desired output. In at least one embodiment, a multi-conditional GAN ​​can be used to holistically analyze gene expression data and medical image data. In at least one embodiment, by combining expression data and images for new sample generation, image features and gene embeddings can be learned directly from the data in an end-to-end manner. In at least one embodiment, the gene expression data can be any suitable dataset, for example, a public NSCLC dataset with gene expression profiles from RNA sequencing.

[0040] In at least one embodiment, image-gene correlations can be formulated by solving a multi-conditional GAN. In at least one embodiment, a GAN architecture and fusion block can be used to combine image (as background) and gene (as object and "style") data. In at least one embodiment, a smooth object / background fusion that can be modeled within the network can be provided. In at least one embodiment, this synthesis strategy can also be used to generate discriminative radiogenomic maps. In at least one embodiment, radiogenomic map generation can be formulated as an image synthesis task.

[0041] Figure 2 Components of a GAN-based architecture that, in at least one embodiment, can be used to perform such synthesis are shown. Figure 2 The structure of the generator portion 200 of the architecture in at least one embodiment is shown. In at least one embodiment and as shown, the generator can accept a background image 102 and gene expression data 202 as input training data. In at least one embodiment, the generator can generate a composite image 106 from the background image and the gene expression data, the composite image 106 including the nodule characterized by the genomic data and located in the background image. In at least one embodiment, the generator also produces a binary segmentation mask 104 representing the area or boundary of the generated nodule. In at least one embodiment, the generator performs at least three main tasks, including encoding the background image on the left path, encoding the gene expression data on the right path, and information fusion for the composite image and mask generation along the center path, with the information fusion taking the results from the left path and the right path as input.

[0042] In at least one embodiment, a GAN used as such a generator can perform separation and blending of objects and background, as well as fusion of image and gene representations. With respect to the blending task, in at least one embodiment, the network does not remove any portion of the background image. Instead, in at least one embodiment, the network models the objects and background within the network using two strategies. In at least one embodiment, a fusion block is used at each resolution level to control the overlap between the generated objects (e.g., nodules or cell groups) and the reference background image data. In at least one embodiment, a segmentation mask 104 can be generated as an auxiliary output of the segmentation mask to help guide this separation. In at least one embodiment, at each stage, the network can perform a "soft" blending of the image data for the object and background. In at least one embodiment, this iterative blending approach can help ensure spatial continuity of the inferred synthetic image. An advantage in at least one embodiment is that it produces a segmentation mask 104 together with the inferred image 106, which, if used for data augmentation techniques, can help make the segmentation mask useful for other tasks such as detection and segmentation.

[0043] In at least one embodiment, the example GAN can also use word embeddings, such as those used in computer vision, to generate a base image that can be combined with background image data at the bottleneck layer of the encoder-decoder network. In at least one embodiment, a significant difference between word embeddings and gene representations relates to the fact that words are much more closely related to images than to gene representations. In at least one embodiment, it can be determined to model gene information as an abstract "style" of an image and use style transfer techniques to guide the synthesis process. Specifically, in at least one embodiment, high-dimensional gene expression data can be encoded using a mapping network. In at least one embodiment, the mapping network can include several fully connected (FC) layers or more complex conditional enhancement blocks, among other options. In at least one embodiment, to provide improved interpretability of the gene encoding, two FC layers can be used to encode the raw gene data g into a vector (g). In at least one embodiment, the vector (g) can be further concatenated with a noise vector n to generate a lower-dimensional gene code 204, which can be used as a base style map. In at least one embodiment, an image encoder consisting of multiple convolutional layers can be used to encode the background image. In at least one embodiment, a series of fusion blocks can be used to combine image features and gene map data. In at least one embodiment, the fusion block can obtain image features from both the background and the previous steps, combined with the genetic "style" map, to achieve appropriate blending of objects and background, as well as appropriate fusion of image information and genetic information.

[0044] Figure 3AComponents of an example fusion block 300 in at least one embodiment are shown. In at least one embodiment, there can be a fusion block 300 at each resolution level or at least a subset of the resolution levels. At each of these resolution levels, there can be three inputs to the fusion block, including background image features, gene maps, and synthetic image features from the previous layer. In at least one embodiment, since the image data contains information about both the object and the background, the synthetic features can be further encoded via two layers of convolution 302 and batch normalization 304. In at least one embodiment, the number of channels is doubled during this process. In at least one embodiment, the resulting code is divided into two parts. In at least one embodiment, the first half is used as a weight map to control how much object and / or background information will be passed for further processing at this layer, and the other half will be used as an object feature map.

[0045] In at least one embodiment, and as Figure 3A As shown in [ 3 ], both object and background feature maps are controlled by element-wise multiplication with a weight map (+) and its inverse (-). In at least one embodiment, the weight map (+) suppresses background information, primarily through nodule features to be normalized by the genetic code. In at least one embodiment, this is because the genetic code is less relevant to the background and more relevant to controlling the appearance of the generated nodules. In at least one embodiment, the inverse map (-) can be used to suppress information about where nodules will be generated, thereby enhancing background information that will be aligned with the input image. In at least one embodiment, the genetic code can control the "style" of the synthesized nodules via the Adaptive Instance Normalization (AdaIN) layer 306. In at least one embodiment, these two components are summed together and fed into the upsampling / decoding layer 308. In at least one embodiment, and compared to completely erasing or otherwise discarding pixel values ​​from a portion of the image as "inpainting," the weight map is a learned probability that retains the information necessary for smooth object-background fusion. In at least one embodiment, and compared to word embedding synthesis, this approach provides stronger object-background separation because the genetic map is primarily applied to the object region and has little effect on the background.

[0046] In at least one embodiment, as described above, such a GAN can encode genomic features as a vector, outputting both a synthetic image and a segmentation mask. Figure 3B An example discriminator in at least one embodiment is shown. In at least one embodiment, the input to the discriminator is a tuple of image segmentation gene codes. In at least one embodiment, two encoders are used for the discriminator 356, the first encoder 352 is the discriminator D IEncode the image, and the second encoder 354 is the discriminator D IS Encode the image segmentation pair. In at least one embodiment, the output of the second encoder output is further combined with the gene code φ(g) 204 and further encoded through convolution, batch normalization, and leaky ReLU activation layers for the discriminator D ISG . In at least one embodiment, for the discriminator, three different loss functions are used to enhance learning of genetic information and images. In at least one embodiment, the first discriminator loss is related to whether the image is real or fake. In at least one embodiment, the second loss is related to how well the segmentation matches the image. In at least one embodiment, the third loss is related to how well all three inputs match. In at least one embodiment, the discriminator is trained with a least squares loss function. In at least one embodiment, given an image x, a matching genetic code g, and a matching segmentation mask m, the tuples to be distinguished include tuples containing non-matching genetic codes Mismatched segmentation masks Synthetic image G x and the composite mask G m Let p d and p G Representing the distribution of real and synthetic data, we have x, g, m, and G x , G m ~p G In at least one embodiment, and in various combinations, this results in:

[0047]

[0048]

[0049]

[0050] In at least one embodiment, to train the generator, a background reconstruction loss can be added to guide feature extraction of the background image during synthesis. In at least one embodiment, the losses are all optimized together. In at least one embodiment, let is the segmentation mask G m (e.g., background area) is the reverse morphological erosion version, ⊙ represents element-wise multiplication, in the composite image G x Calculate L on the background between the base image x I loss:

[0051]

[0052] Figure 4A pair of example image sets 400 and 410 that can be used or generated with a GAN in at least one embodiment are shown. In at least one embodiment, each set includes a background image provided as input and a synthesized image generated by the GAN. In at least one embodiment, each set also displays a background weight map and the resulting segmentation map. In at least one embodiment, the background weight image controls how the background and foreground are blended together. In at least one embodiment, when blending a background image with a computer-generated nodule, it may be desirable to include as much of the input background image outside the nodule region as possible, even if the majority of the generated image is focused on the nodule or genetic code portion. In at least one embodiment, the background weight image is used to control the blending, providing a softer blend than using a segmentation mask alone. In at least one embodiment, the second set shows some variations for a case with ground-glass opacity. In at least one embodiment, it can be observed that the original background image is not significantly altered because the main structural or background features are preserved. In at least one embodiment, the synthesized nodule is naturally blended with the background image, even for the case of ground-glass. In at least one embodiment, it can also be seen that for two background images at different reconstructions, one smoother than the other, the sharpness of the resulting nodule image is still well aligned with the background image. Figure 4 Also shown is a visual representation of a gene expression data set, which in at least one embodiment can be a large multidimensional data set. In at least one embodiment, the dimensionality of the gene expression data 420 can be reduced to generate a gene code 430, which also includes data for accounting for a certain amount of noise. In at least one embodiment, in order to provide improved interpretability of the gene encoding, the original gene data can be encoded into a vector and then further concatenated with the noise vector to generate a lower dimensional gene code 430, which can be used as a base pattern map. In at least one embodiment, the gene expression data can be a vector of approximately 30,000 feature points in length, while the lower dimensional gene code can be a vector of dimension 128, etc.

[0053] In at least one embodiment, a background image can be created by first segmenting the lung region for each image, where the nodule region is excluded from the lung mask. In at least one embodiment, a distance transform is calculated for the resulting mask, and a center is selected at a random location, for example, 5mm to 25mm from the mask boundary. In at least one embodiment, a 60×60×60mm image is cropped around each center. 3 A volume of interest (VOI) is defined and many random slices (e.g., 20) are extracted from each VOI.

[0054] Although other types of neural networks and machine learning can be utilized, in at least one embodiment, a multi-conditional GAN ​​can be used, combined with a new structure for style control and fusion, to effectively generate realistic nodules whose appearance is controlled by their genomic features. In at least one embodiment, such a GAN can be adjusted on both a background image and a gene expression profile to synthesize the corresponding image. In at least one embodiment, the image and gene features can be fused at different proportions to ensure the authenticity and quality of the synthesized image. In at least one embodiment, an end-to-end mechanism is implemented to model and associate features holistically. In at least one embodiment, the method proposed herein can not only provide an effective and controllable means of generating various nodules, but also provide a discriminative radiogenomic map that links genomic and image features.

[0055] Figure 5 Shown are raw genetic data 500 and associated images in at least one embodiment. In at least one embodiment, the raw data illustrates how different types of genes are fused together. In at least one embodiment, genetic data points are distributed on the genetic map based at least in part on their corresponding image features. In at least one embodiment and utilizing an example solution, the converted genetic code can be mapped to a 2D plane. Once mapping has been performed, in at least one embodiment, any appropriate clustering algorithm or method can be used to perform clustering. In at least one embodiment, the nodules corresponding to a given cluster will then have similar image features, thereby allowing the genetic code falling within a given cluster to have those image features for generating nodules of that type in a synthetic image. In at least one embodiment, this method can be used to determine the correlation between genetic data and image feature data because these data types are weakly combined. In at least one embodiment, when generating a nodule corresponding to a given genetic code, the image feature data can be used as a loose constraint type.

[0056] Figure 6AAn example process 600 for training a generative adversarial network to infer synthetic images in at least one embodiment is shown. It should be understood that for this and other processes discussed herein, unless otherwise noted, there may be additional, alternative, or fewer steps that may be performed in a similar or alternative order or in parallel. Furthermore, this example discusses training a generative adversarial network (GAN) using text data, but as discussed elsewhere herein, in at least one embodiment, there may be model types that are trained using a variety of different types of data. In at least one embodiment, gene expression data is obtained 602, which may indicate a specific grouping of cells, such as a type of lung nodule. In at least one embodiment, this gene expression data may be mapped to specific groupings of features exhibited by the cell groups associated with the gene expression data. In at least one embodiment, background image data is also obtained 604, which, as mentioned, may include obtaining the background image data and determining one or more portions that lack other features of the cell grouping or size. In at least one embodiment, the gene expression data and background image data may be fed into the GAN as training data.

[0057] In at least one embodiment, the GAN may have at least three logical branches that perform three functions. In at least one embodiment, in the first pair of branches, the GAN will encode 606 gene expression data and background image data. In at least one embodiment, the data will be encoded into separate vectors. In at least one embodiment, dimensionality reduction can be performed in layers, as discussed herein. In at least one embodiment, the third logical branch of the GAN can take the encoded image and gene data and generate 608 a synthetic image and a segmentation mask. In at least one embodiment, the synthetic image can include a representation of a cell group having features determined by the gene code data, and the cell group can be fused into the background image to provide a realistic fusion of the cells and the background. In at least one embodiment, the synthetic image and segmentation mask output by the generator can be fed 610 to the discriminator along with the gene data to determine a loss function. In at least one embodiment, the loss can treat the three inputs as three separate losses that are optimized together. In at least one embodiment, appropriate network parameters of the GAN are then updated 612 based in part on these determined loss values.

[0058] Figure 6BAn example process 650 for inferring a synthetic image using such a trained model in at least one embodiment is shown. In at least one embodiment, gene expression data and background image data to be used to infer the synthetic image are obtained 652. In at least one embodiment, the data can be provided 654 as input to the trained model. In at least one embodiment, the trained model can process the data and infer 656 a synthetic image that includes a realistic digital representation of a group of cells, such as a nodule, mixed into the background image. In at least one embodiment, the synthetic image can be used for various purposes, such as can include use as training data for other neural networks. For example, in at least one embodiment, the synthetic image can be used to train a neural network to infer whether the group of cells represented in the input image is malignant or benign, among other such diagnoses.

[0059] As mentioned above, a growing number of industries and applications are leveraging machine learning. For example, deep neural networks (DNNs) developed on processors are being used in a variety of use cases, from self-driving cars to faster drug development, from automated image analysis for security systems to intelligent real-time language translation in video chat applications. Deep learning is a technology that models the neural learning processes of the human brain, continuously learning, getting smarter, and delivering faster and more accurate results over time. Initially, adults teach children how to correctly identify and classify shapes, eventually enabling them to recognize shapes without any instruction. Similarly, deep learning or neural learning systems designed for similar tasks will need to be trained to become smarter and more efficient at recognizing basic objects, occluded objects, and other aspects, while also assigning context to these objects.

[0060] At the simplest level, neurons in the human brain look at the inputs they receive, assign a level of importance to each of these inputs, and pass outputs to other neurons to act on them. Artificial neurons, or perceptrons, are the basic model of neural networks. In one example, a perceptron can receive one or more inputs representing features of the object it is being trained to recognize and classify, and assign a certain weight to each of these features based on their importance in defining the object's shape.

[0061] Deep neural network (DNN) models consist of multiple layers of many connected perceptrons (e.g., nodes) and can be trained with large amounts of input data to quickly solve complex problems with high accuracy. In one example, the first layer of a DNN model breaks down an input image of a car into its components and looks for basic patterns such as lines and angles. The second layer assembles these lines to look for higher-level patterns, such as wheels, windshields, and rearview mirrors. The next layer identifies the type of vehicle, and the final layers generate labels for the input image to identify the model of a specific car brand. Once trained, a DNN can be deployed and used to recognize and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include recognizing handwritten digits on checks deposited at an ATM, identifying images of friends in photos, providing movie recommendations, identifying and classifying different types of cars, pedestrians, and road hazards in self-driving cars, or translating human speech in near real time.

[0062] During training, data flows through the DNN in a forward propagation phase until a prediction is produced indicating the label corresponding to the input. If the neural network does not correctly label the input, the error between the correct and predicted labels is analyzed, and the weights of each feature are adjusted in a backpropagation phase until the DNN correctly labels the input and other inputs in the training dataset. Training complex neural networks requires a large amount of parallel computing performance, including supporting floating-point multiplication and addition. Inference, which is less computationally intensive than training, is a latency-sensitive process in which the trained neural network is applied to new, previously unseen inputs to classify images, translate speech, and infer new information.

[0063] Neural networks rely heavily on matrix math operations, and complex, multi-layer networks require significant floating-point performance and bandwidth for efficiency and speed. Computing platforms with thousands of processing cores optimized for matrix math operations and delivering tens to hundreds of TFLOPS can deliver the performance required for deep neural network-based artificial intelligence and machine learning applications.

[0064] Figure 7Components of an example system 700 that can be used to train and utilize machine learning, according to at least one embodiment, are shown. In at least one embodiment, the components can be provided by a combination of computing devices and resources, or a single computing system, that can be under the control of a single entity or multiple entities. Furthermore, in at least one embodiment, aspects can be triggered, initiated, or requested by different entities. For example, in at least one embodiment, training of a neural network can be directed by a vendor associated with a vendor environment 706, while in at least one embodiment, training of a neural network can be requested by a customer or other user who can access the vendor environment via a client device 702 or other such resource. In at least one embodiment, training data (or data to be analyzed by a trained neural network) can be provided by a vendor, a user, or a third-party content provider 724. In at least one embodiment, client device 702 can be a vehicle or object that can navigate on behalf of a user, for example, the user can submit requests and / or receive instructions to assist in navigating the device.

[0065] In this example, a request can be submitted over at least one network 704 to be received by a provider environment 706. The client device can be any suitable electronic and / or computing device that enables a user to generate and send such a request, such as may include a desktop computer, a laptop computer, a computer server, a smartphone, a tablet computer, a gaming console (portable or otherwise), a computer processor, computing logic, and a set-top box, etc. The network 704 may include any suitable network for transmitting requests or other such data, such as may include the Internet, an intranet, an Ethernet network, a cellular network, a local area network (LAN), a network for making direct wireless connections between nodes, etc.

[0066] In this example, a request may be received by an interface layer 708, which may forward the data to a training and inference manager 710. In at least one embodiment, the manager may be a system or service comprising hardware and software for managing services and requests for data or content. The manager may receive a request to train a neural network and may provide the requested data to a training manager 712. If the request is unspecified, the training manager 712 may select an appropriate model or network to use and train the model using the relevant training data. In at least one embodiment, the training data may be a batch of data received from the client device 702 or obtained from a third-party provider 724 and stored in a training data repository 714. The training manager 712 may be responsible for training the data, for example, using the LARC-based methods discussed herein. The network may be any suitable network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN). Once the network is trained and successfully evaluated, the trained network may be stored in a model repository 716, which may, for example, store different models or networks for a user, application, or service. As described above, in at least one embodiment, multiple models may exist for a single application or entity, such that multiple models may be utilized based on a number of different factors.

[0067] At a later point in time, a request for content (e.g., a path determination) or data determined or influenced at least in part by a trained neural network may be received from client device 702 (or another such device). The request may include, for example, input data to be processed using the neural network to obtain one or more inferences or other output values, classifications, or predictions. Although a different system or service may be used in at least one embodiment, the input data may be received by interface layer 708 and directed to inference module 718. If not already stored locally in inference module 718, inference module 718 may retrieve an appropriate trained network, such as a trained deep neural network (DNN) as described herein, from model repository 716. Inference module 718 may provide data as input to the trained network and may then generate one or more inferences as output. For example, this may include a classification of an instance of the input data. The inferences may then be sent to client device 702 for display to the user or other communication with the user. Contextual data for the user may also be stored in user context data repository 722, which may include data about the user that may be used as network input for generating inferences or determining instances of the data returned to the user. Related data, including at least a portion of the input or inference data, may also be stored in a local database 720 for use in processing future requests. In at least one embodiment, a user may use account or other information to access resources or functionality of the provider environment. If permitted and available, user data may also be collected and used to further train models to provide more accurate inferences for future requests. In at least one embodiment, a request for a machine learning application 726 executed on a client device 702 may be received through a user interface, and the results may be displayed through the same interface. The client device may include resources such as a processor 728 and memory 730 for generating the request and processing the results or responses, as well as at least one data storage element 732 for storing data for the machine learning application 726.

[0068] In at least one embodiment, processor 728 (or the processor of training manager 712 or inference module 718) will be a central processing unit (CPU). However, as described above, resources in such environments may utilize GPUs to process data for at least certain types of requests. GPUs, with their thousands of cores and designed to handle massively parallel workloads, have become popular in deep learning for training neural networks and generating predictions. While using GPUs for offline building allows for faster training of larger and more complex models, generating predictions offline means either the request-time input features cannot be used or predictions must be generated for all feature permutations and stored in lookup tables to service real-time requests. If the deep learning framework supports CPU mode and the model is small and simple enough to be run feedforward on a CPU with reasonable latency, a service on a CPU instance can host the model. In this case, training can be performed offline on the GPU, and inference can be performed in real time on the CPU. If a CPU approach is not a viable option, the service can run on a GPU instance. However, because GPUs have different performance and cost characteristics than CPUs, running a service that offloads runtime algorithms to a GPU may require a different design than a CPU-based service.

[0069] Figure 8An example system 800 is shown, according to at least one embodiment, that can be used to classify data or generate inferences. Based on the teachings and suggestions contained herein, it should be apparent that various types of predictions, labels, or other outputs can also be generated for input data. Furthermore, supervised and unsupervised training can be used in at least one embodiment discussed herein. In this example, a set of training data 802 (e.g., classified or labeled data) is provided as input to serve as training data. The training data can include instances of at least one type of object on which a neural network is to be trained, as well as information identifying objects of that type. For example, the training data might include a set of images, each image containing a representation of an object type, wherein each image also contains or is associated with a label, metadata, classification, or other information identifying the type of object represented in the respective image. Other types of data can also be used as training data, including text data, audio data, video data, and the like. In this example, the training data 802 is provided as training input to a training manager 804. The training manager 804 can be a system or service comprising hardware and software, such as one or more computing devices executing a training application for training a neural network (or other model or algorithm, etc.). In this example, the training manager 804 receives an instruction or request indicating the type of model to be used for training. The model can be any appropriate statistical model, network, or algorithm that can be used for such purposes, and may include, for example, an artificial neural network, a deep learning algorithm, a learning classifier, a Bayesian network, etc. The training manager 804 can select an initial model or other untrained model from an appropriate repository 806 and train the model using the training data 802 to generate a trained model 808 (e.g., a trained deep neural network) that can be used to classify or analyze similar types of data, or generate other such inferences. In at least one embodiment where training data is not used, the input data can still be trained based on the selection of an appropriate initial model by the training manager 804.

[0070] The model can be trained in a variety of different ways, which may depend in part on the type of model selected. For example, in at least one embodiment, a set of training data can be provided to a machine learning algorithm, where the model is a model artifact created by the training process. Each instance of the training data contains a correct answer (e.g., a classification), which can be referred to as a target or target attribute. The learning algorithm finds patterns in the training data that map the input data attributes to the target, the answer to be predicted, and outputs a machine learning model that captures these patterns. The machine learning model can then be used to obtain predictions for new data for which the target is not specified.

[0071] In one example, the training manager 804 can select from a set of machine learning models, including binary classification, multi-class classification, and regression models. The type of model to be used can depend at least in part on the type of target to be predicted. Machine learning models for binary classification problems predict a binary outcome, such as one of two possible classifications. Learning algorithms such as logistic regression can be used to train binary classification models. Machine learning models for multi-class classification problems allow predictions to be generated for multiple classes, such as predicting one of more than two outcomes. Multinomial logistic regression can be useful for training multi-class models. Machine learning models for regression problems predict numerical values. Linear regression is useful for training regression models.

[0072] To train a machine learning model according to at least one embodiment, a training manager must determine the input training data source, as well as other information, such as the name of the data attribute containing the target to be predicted, required data transformation instructions, and training parameters to control the learning algorithm. During the training process, the training manager 804 in at least one embodiment can automatically select an appropriate learning algorithm based on the type of target specified in the training data source. The machine learning algorithm can accept parameters that control certain properties of the training process and the resulting machine learning model. These are referred to herein as training parameters. If no training parameters are specified, the training manager can utilize known default values ​​that work well for a wide range of machine learning tasks. Examples of training parameters for which values ​​can be specified include the maximum model size, the maximum number of passes through the training data, the type of shuffle, the type of regularization, the learning rate, and the amount of regularization. Default settings can be specified, and adjustment values ​​can be selected to fine-tune performance.

[0073] The maximum model size is the total size, in bytes, of the patterns created during model training. By default, a model of the specified size is created, for example, a 100MB model. If the training manager cannot identify enough patterns to fill the model size, a smaller model is created. If the training manager discovers more patterns than can be accommodated within the specified size, it implements a maximum cutoff by trimming those patterns that have the least impact on the quality of the learned model. Choosing a model size controls the tradeoff between the model's predictive quality and cost. A smaller model may cause the training manager to remove many patterns to fit within the maximum size limit, affecting prediction quality. Larger models may be more expensive to query for real-time predictions. Larger input datasets do not necessarily result in larger models because the model stores the patterns, not the input data. If the patterns are few and simple, the resulting model will be smaller. Input data with a large number of raw attributes (input columns) or derived features (output of data transformations) may discover and store more patterns during training.

[0074] In at least one embodiment, the training manager 804 may perform multiple passes or iterations on the training data to attempt to discover patterns. There may be a default number of passes, such as ten, and in at least one embodiment, a maximum number of passes may be set, such as up to one hundred passes. In at least one embodiment, there may not be a maximum set, or there may be a set of convergence criteria or other factors that trigger the end of the training process. In at least one embodiment, the training manager 804 may monitor the quality of the patterns during training and may automatically stop training when there are no more data points or patterns to discover. Data sets with only a small number of observations may require more passes over the data to achieve a sufficiently high model quality. Larger data sets may contain many similar data points, which may reduce the need for a large number of passes. A potential impact of selecting more passes over the data is that model training may take longer and cost more in terms of resources and system utilization.

[0075] In at least one embodiment, the training data is shuffled before training or between training passes. In at least one embodiment, the shuffling is a random or pseudo-random shuffling to produce a truly random ordering, although there may be constraints to ensure that certain types of data are not grouped, or if such grouping occurs, the shuffled data can be reshuffled, etc. Shuffling changes the sequence or arrangement of the data used for training so that the training algorithm does not encounter groupings of similar types of data or too many consecutive observations of a single type of data. For example, a model may be trained to predict objects. Before uploading, the data may be sorted by object type. The algorithm can then process the data in alphabetical order by object type, initially encountering only data of a specific object type. The model will begin to learn clusters of objects of that type. The model will then encounter only data of a second object type and will attempt to adjust the model to fit that object type, potentially degrading images that were well-suited for the first object type. This abrupt switch between object types may result in a model that is unable to learn how to accurately predict object types. In at least one embodiment, shuffling can be performed before partitioning the training dataset into training and evaluation subsets, thereby utilizing a relatively even distribution of data types for both phases. In at least one embodiment, training manager 804 may automatically shuffle the data using, for example, a pseudo-random shuffling technique.

[0076] When creating a machine learning model, the training manager 804 in at least one embodiment can enable the user to specify settings. For example, the user can specify one or more evaluation settings to indicate a portion of the input data to be retained for evaluating the predictive quality of the machine learning model. The user can also specify a policy that indicates which attributes and attribute transformations are used for model training. The user can also specify various training parameters that control the training process and certain properties of the resulting model.

[0077] Once the training manager determines that model training is complete, for example by using at least one of the final criteria discussed herein, the trained model 808 can be provided to the classifier 814 for use in classifying (or otherwise generating inferences about) validation data 812. As shown, this involves a logical transition between the model's training mode and the model's inference mode. However, in at least one embodiment, the trained model 808 will first be passed to an evaluator 810, which can include an application, process, or service executed on at least one computing resource (e.g., a CPU or GPU of at least one server) for evaluating the quality (or other aspects) of the trained model. The model is evaluated to determine whether it provides at least a minimum acceptable or threshold level of performance when predicting targets for new and future data. If not, the training manager 804 can continue training the model. Since future data instances will typically have unknown target values, it may be desirable to examine machine learning accuracy metrics on data for which the target answers are known and use this evaluation as a proxy for predictive accuracy for future data.

[0078] In at least one embodiment, a model is evaluated using a subset of the training data 802 provided for training. The subset can be determined using the shuffling and splitting methods described above. This evaluation data subset will be labeled with a target and can therefore serve as a resource for evaluating ground truth. Using the same data used for training to evaluate the predictive accuracy of a machine learning model is not useful because a model that memorizes the training data rather than generalizing from it may generate a positive evaluation. Once training is complete, the evaluation data subset is processed using the trained model 808, and an evaluator 810 can determine the accuracy of the model by comparing the ground truth data with the model's corresponding output (or predicted / observed values). The evaluator 810 in at least one embodiment can provide a summary or performance metric that indicates the degree of match between the predicted and true values. If the trained model does not meet at least a minimum performance standard or other such accuracy threshold, the training manager 804 can be instructed to perform further training or, in some cases, attempt to train a new or different model. If the trained model 808 meets the relevant standards, the trained model can be provided for use by the classifier 814.

[0079] When creating and training a machine learning model, in at least one embodiment, it may be desirable to specify model settings or training parameters that will result in a model that makes the most accurate predictions. Example parameters include the number of forward and / or backward passes to be performed, regularization, model size, and shuffling type. However, as described above, selecting the model parameter settings that produce the best predictive performance on the evaluation data may lead to model overfitting. Overfitting occurs when the model memorizes patterns present in both the training and evaluation data sources but fails to generalize to patterns in the data. Overfitting often occurs when the training data includes all the data used in the evaluation. An overfitted model may perform well during evaluation but may not make accurate predictions on new or other validation data. To avoid selecting an overfitted model as the best model, the training manager may retain additional data to validate the model's performance. For example, the training dataset may be split 60% for training and 40% for evaluation or validation, possibly in two or more phases. After selecting the model parameters that best fit the evaluation data, resulting in convergence on a subset of the validation data (e.g., half of the validation data), a second validation run can be performed using the remaining validation data to ensure the model's performance. If the model meets expectations on the validation data, then the model is not overfitting the data. Alternatively, a test or holdout set can be used to test parameters. Using a second validation or testing step can help choose appropriate model parameters to prevent overfitting. However, taking more data out of the training process for validation results in less data available for training. This can be problematic for smaller datasets, as there may not be enough data available for training. One approach in this situation is to perform cross-validation, as described elsewhere in this article.

[0080] There are many metrics or insights that can be used to review and evaluate the predictive accuracy of a given model. A sample evaluation result includes a predictive accuracy metric that reports the model's overall success and visualizations that help explore the model beyond the predictive accuracy metric. The results can also provide the ability to view the impact of setting score thresholds (e.g., for binary classification) and can generate alerts based on criteria used to check the effectiveness of the evaluation. The choice of metric and visualization may depend, at least in part, on the type of model being evaluated.

[0081] Once a trained machine learning model has been satisfactorily trained and evaluated, it can be used to build or support a machine learning application. In at least one embodiment, building a machine learning application is an iterative process involving a series of steps. The core machine learning problem can be constructed based on what is observed and the answer that the model is to predict. Data can then be collected, cleaned, and prepared to make it suitable for use by the machine learning model training algorithm. This data can be visualized and analyzed for integrity checks to verify the quality of the data and understand the data. The original data (e.g., input variables) and the answer data (e.g., target) may not be represented in a way that can be used to train a highly predictive model. Therefore, it may be desirable to build a more predictive input representation or feature from the original variables. The resulting features can be input into a learning algorithm to build a model and evaluate the quality of the model based on the data retained from model construction. The model can then be used to generate predictions of the target answer for new data instances.

[0082] exist Figure 8 In the example system 800, after providing an evaluation, the trained model 810 is provided to or made available to a classifier 814, which is capable of using the trained model to process validation data. For example, this may include data received from a user or an unclassified third party, such as query images, seeking information about representations in those images. The trained model can be used by the classifier to process the validation data, and the resulting results 816 can be sent back to the corresponding source or otherwise processed or stored. In at least one embodiment, and where such use is permitted, the now-classified data instances can be stored in a training data repository, which can be used by a training manager for further training of the trained model 808. In at least one embodiment, the model is trained continuously as new data becomes available, but in at least one embodiment, the model is trained periodically, such as daily or weekly, depending on factors such as the size of the dataset or the complexity of the model.

[0083] The classifier 814 may include appropriate hardware and software for processing the validation data 812 using the trained model. In at least one embodiment, the classifier will include one or more computer servers, each having one or more graphics processing units (GPUs) capable of processing data. The configuration and design of the GPUs may make them more suitable for processing machine learning data than CPUs or other such components. In at least one embodiment, the trained model can be loaded into the GPU memory, and the received data instances are provided to the GPU for processing. The GPU can have many more cores than the CPU, and the GPU cores can be less complex. Therefore, a given GPU may be able to process thousands of data instances simultaneously through different hardware threads. The GPU can also be configured to maximize floating point throughput, which can provide significant additional processing advantages for large data sets.

[0084] Even when GPUs, accelerators, and other such hardware are used to accelerate tasks such as training models or classifying data using such models, such tasks can still require significant time, resource allocation, and cost. For example, if a machine learning model is to be trained using 800 passes, and the dataset includes 1,000,000 data instances to be used for training, each pass will need to process all million instances. Different parts of the architecture can also be supported by different types of devices. For example, training can be performed using a set of servers at a logically centralized location, such as can be provided as a service, while classification of the raw data can be performed by such a service or on a client device. In at least one embodiment, these devices can also be owned, operated, or controlled by the same entity or multiple entities.

[0085] Figure 9 An example neural network 900 that can be trained or otherwise utilized in accordance with at least one embodiment is shown. In this example, the statistical model is an artificial neural network (ANN) that includes multiple layers of nodes, including an input layer 902, an output layer 906, and multiple layers 904 of intermediate nodes, typically referred to as "hidden" layers because the internal layers and nodes are typically not visible or accessible in various neural networks. Although several intermediate layers are shown for illustrative purposes only, it should be understood that there is no limit to the number of intermediate layers that can be utilized, and any limit on layers will generally be a factor of the resources or time required to process using this model. As discussed elsewhere herein, other types of models, networks, algorithms, or processes may also be used because they may include other numbers or selections of nodes and layers. Validation data may be processed by the layers of the network to generate a set of inference or reasoning scores, which may then be fed into a loss function 908.

[0086] In this example network 900, all nodes in a given layer are interconnected to all nodes in adjacent layers. As shown in the figure, the nodes in the middle layer are then connected to the nodes in the two adjacent layers respectively. In some models, nodes are also called neurons or connected units, and the connections between nodes are called edges. Each node can perform a function for the input it receives, for example by using a specified function. Nodes and edges can be assigned different weights during training, and each layer of nodes can perform specific types of transformations on the input it receives, where these transformations can also be learned or adjusted during training. Learning can be supervised or unsupervised, which may depend at least in part on the type of information contained in the training dataset. Various types of neural networks can be used, for example, including convolutional neural networks (CNNs), which include multiple convolutional layers and a set of pooling layers, and various types of neural networks have been shown to be beneficial for applications such as image recognition. CNNs are also easier to train than other networks due to the relatively small number of parameters to be determined.

[0087] In at least one embodiment, various tuning parameters can be used to train such complex machine learning models. Selecting parameters, fitting the model, and evaluating the model are part of the model tuning process, often referred to as hyperparameter optimization. In at least one embodiment, such tuning can include introspection of the underlying model or data. In training or production settings, a robust workflow is important to avoid overfitting of hyperparameters, as described elsewhere herein. Cross-validation and adding Gaussian noise to the training dataset are useful techniques to avoid overfitting to any one dataset. For hyperparameter optimization, in at least one embodiment, it may be desirable to keep the training set and validation set fixed. In at least one embodiment, hyperparameters can be tuned in certain categories, which can include, for example, data preprocessing, CNN architecture definition (e.g., filter size, number of filters), stochastic gradient descent (SGD) parameters (e.g., learning rate), and regularization (e.g., dropout probability).

[0088] In an example preprocessing step, instances in a dataset can be embedded into a lower-dimensional space of a specific size. The size of this space is a parameter to be adjusted. The architecture of this CNN contains many adjustable parameters. The filter size parameter can represent the interpretation of information corresponding to the size of the instance to be analyzed. In computational linguistics, this is called the n-gram size. The example CNN uses three different filter sizes, which represent different possible n-gram sizes. The number of filters in each filter size can correspond to the depth of the filter. Each filter attempts to learn something different from the instance structure, such as the sentence structure of text data. In the convolutional layer, the activation function can be a rectified linear unit, and the pooling type is set to max pooling. The results can then be concatenated into a one-dimensional vector, and the final layer is fully concatenated to the two-dimensional output. This corresponds to binary classification, to which an optimization function can be applied. One such function is the root mean square (RMS) propagation method that implements gradient descent, where example hyperparameters can include the learning rate, batch size, maximum gradient normal, and epochs. Regularization can be a very important consideration when using neural networks. As mentioned above, in at least one embodiment, the input data can be relatively sparse. In this case, the main hyperparameters can be dropped at the penultimate layer, which means that a certain proportion of nodes will not "fire" at each training cycle. The example training process can recommend different hyperparameter configurations based on feedback on the performance of previous configurations. The model can be trained using the recommended configurations, evaluated on a specified validation set, and a performance report is provided. This process can be repeated, such as the trade-off between exploration (learning more about different configurations) and development (leveraging previous knowledge to achieve better results)

[0089] Because CNN training can be parallelized and GPU-powered computing resources can be utilized, multiple optimization strategies can be tried for different scenarios. Complex scenarios allow for tuning of the model architecture, preprocessing, and stochastic gradient descent parameters. This expands the model configuration space. In the basic scenario, only preprocessing and stochastic gradient descent parameters are tuned. Compared to the basic scheme, this complex scheme allows for many more configuration parameters. Tuning of the joint space can be performed using linear or exponential steps and iterated through the model's optimization loop. This tuning process can be significantly less expensive than tuning procedures such as random search and grid search, without any noticeable performance loss.

[0090] In at least one embodiment, back propagation can be used to calculate the gradient for determining the weights of a neural network. Back propagation is a form of differentiation, and as described above, a gradient descent optimization algorithm can use it to adjust the weights applied to a node or neuron. In at least one embodiment, the gradient of a related loss function can be used to determine the weights. Back propagation can utilize the derivative of the loss function with respect to the output generated by the statistical model. As described above, each node can have an associated activation function that defines the output of each node. Various activation functions can be used appropriately, for example, radial basis functions (RBFs) and sigmoids can be included, which can be used for data conversion by various support vector machines (SVMs). The activation function of the intermediate layer of a node is referred to as an inner product core in this article. These functions can include, for example, a recognition function, a step function, a sigmoidal function, a ramp function, etc. The activation function can also be linear or nonlinear.

[0091] Reasoning and training logic

[0092] Figure 10 Inference and / or training logic 1015 is shown for performing inference and / or training operations in at least one embodiment. Figure 10 and / or Figure 11 Provides details regarding the inference and / or training logic 1015 .

[0093] In at least one embodiment, inference and / or training logic 1015 may include, but is not limited to, data storage 1001 to store forward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in accordance with aspects of at least one embodiment. In at least one embodiment, data storage 1001 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with at least one embodiment during forward propagation of input / output data and / or weight parameters during inference and / or training using aspects of one or more embodiments. In at least one embodiment, any portion of data storage 1001 may be included with other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.

[0094] In at least one embodiment, any portion of data storage 1001 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, data storage 1001 may be a cache memory, dynamic random access memory ("DRAM"), static random access memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage device. In at least one embodiment, the choice of whether data storage 1001 is internal or external to a processor, for example, whether it is composed of DRAM, SRAM, flash memory, or other types of memory, depends on the available storage space on or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inferencing and / or training the neural network, or some combination of these factors.

[0095] In at least one embodiment, inference and / or training logic 1015 may include, but is not limited to, data storage 1005 to store backpropagation and / or output weights and / or input / output data corresponding to a neural network or layer of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, data storage 1005 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during backpropagation of input / output data and / or weight parameters. In at least one embodiment, any portion of data storage 1005 may be included with other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of data storage 1005 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, data storage 1005 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the data storage 1005 is selected to be internal or external to the processor, e.g., comprised of DRAM, SRAM, flash memory, or other memory types, depending on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inference and / or training the neural network, or some combination of these factors.

[0096] In at least one embodiment, data store 1001 and data store 1005 may be separate storage structures. In at least one embodiment, data store 1001 and data store 1005 may be the same storage structure. In at least one embodiment, data store 1001 and data store 1005 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of data store 1001 and data store 1005 may be included with other on-chip or off-chip data stores, including a processor's L1, L2, or L3 cache or system memory.

[0097] In at least one embodiment, the inference and / or training logic 1015 may include, but is not limited to, one or more arithmetic logic units ("ALUs") 1010 to perform logic and / or arithmetic operations based at least in part on or directed by training and / or inference code, the results of which may result in activations (e.g., output values ​​from layers or neurons within a neural network) stored in activation memory 1020 as a function of input / output and / or weight parameter data stored in data store 1001 and / or data store 1005. In at least one embodiment, the activations stored in activation memory 1020 are generated by linear algebra and / or matrix-based mathematics performed by ALU 1010 in response to executing instructions or other code, wherein weight values ​​stored in data store 1005 and / or data store 1001 are operands with other values ​​(e.g., bias values, gradient information, momentum values, or other parameters or hyperparameters), any or all of which may be stored in data store 1005 or data store 1001 or other on-chip or off-chip memory. In at least one embodiment, one or more ALUs 1010 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 1010 may be external to the processor or other hardware logic device or circuit (e.g., a coprocessor) in which they are used. In at least one embodiment, ALUs 1010 may be included within the execution units of the processor or otherwise included in a group of ALUs accessible to the execution units of the processor, which group of ALUs may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, data storage 1001, data storage 1005, and activation storage 1020 may be on the same processor or other hardware logic device or circuit, while in another embodiment, they may be on different processors or other hardware logic devices or circuits, or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 1020 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry and may be retrieved and / or processed using the processor's fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0098] In at least one embodiment, activation storage 1020 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, activation storage 1020 may be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether activation storage 1020 is internal or external to the processor, for example, or composed of DRAM, SRAM, flash memory, or some other memory type, may be based on the memory available on or off chip, the latency requirements for performing training and / or inference functions, the batch size of data used in inferring and / or training neural networks, or some combination of these factors. In at least one embodiment, Figure 10 The inference and / or training logic 1015 shown in FIG can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 10 The illustrated inference and / or training logic 1015 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).

[0099] Figure 11 Inference and / or training logic 1015 in at least one embodiment is shown. In at least one embodiment, inference and / or training logic 1015 may include, but is not limited to, hardware logic where computing resources are dedicated or otherwise used exclusively in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 11 The inference and / or training logic 1015 shown in FIG can be used in conjunction with an application specific integrated circuit (ASIC), such as Google's Processing unit, Inference Processing Unit (IPU) from GraphcoreTM or from Intel Corp (e.g., "Lake Crest") processor. In at least one embodiment, Figure 11The inference and / or training logic 1015 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 1015 includes, but is not limited to, data storage 1001 and data storage 1005, which can be used to store weight values ​​and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 11 In at least one embodiment shown in FIG, data storage 1001 and data storage 1005 are each associated with dedicated computing resources, such as computing hardware 1002 and computing hardware 1006, respectively. In at least one embodiment, computing hardware 1002 and computing hardware 1006 each include one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) solely on information stored in data storage 1001 and data storage 1005, respectively, with the results stored in activation memory 1020.

[0100] In at least one embodiment, each of the data stores 1001 and 1005 and the corresponding computing hardware 1002 and 1006 corresponds to a different layer of a neural network, thereby providing activations from one "storage / compute pair 1001 / 1002" of the data store 1001 and computing hardware 1002 as inputs to the next "storage / compute pair 1005 / 1006" of the data store 1005 and computing hardware 1006, reflecting the conceptual organization of the neural network. In at least one embodiment, each storage / compute pair 1001 / 1002 and 1005 / 1006 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in the inference and / or training logic 1015 after or in parallel with the storage / compute pairs 1001 / 1002 and 1005 / 1006.

[0101] Neural network training and deployment

[0102] Figure 12The training and deployment of a deep neural network according to at least one embodiment is shown. In at least one embodiment, an untrained neural network 1206 is trained using a training dataset 1202. In at least one embodiment, the training framework 1104 is the PyTorch framework, while in other embodiments, the training framework 1104 is Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 1104 trains the untrained neural network 1106 and enables it to be trained using the processing resources described herein to generate a trained neural network 1108. In at least one embodiment, the weights can be randomly selected or through the use of a deep belief network. In at least one embodiment, the training can be performed in a supervised, partially supervised, or unsupervised manner.

[0103] In at least one embodiment, untrained neural network 1106 is trained using supervised learning, where training dataset 1102 includes inputs paired with expected outputs for those inputs, or where training dataset 1102 includes inputs with known outputs, and the outputs of the neural network are manually graded. In at least one embodiment, untrained neural network 1106 is trained in a supervised manner to process inputs from training dataset 1102 and compare the resulting outputs to a set of expected or anticipated outputs. In at least one embodiment, errors are then propagated back through untrained neural network 1106. In at least one embodiment, training framework 1104 adjusts the weights that control untrained neural network 1106. In at least one embodiment, training framework 1104 includes tools for monitoring the extent to which untrained neural network 1106 is converging toward a model, such as trained neural network 1108, suitable for generating correct answers, such as results 1114, based on known input data, such as new data 1112. In at least one embodiment, the training framework 1104 iteratively trains the untrained neural network 1106 while adjusting the weights to refine the output of the untrained neural network 1106 using a loss function and an adjustment algorithm (e.g., stochastic gradient descent). In at least one embodiment, the training framework 1104 trains the untrained neural network 1106 until the untrained neural network 1106 reaches a desired accuracy. In at least one embodiment, the recurrent neural network 1108 can then be deployed to perform any number of machine learning operations.

[0104] In at least one embodiment, untrained neural network 1106 is trained using unsupervised learning, wherein untrained neural network 1106 attempts to train itself using unlabeled data. In at least one embodiment, unsupervised learning training dataset 1102 will include input data without any associated output data or "ground truth" data. In at least one embodiment, untrained neural network 1106 can learn groupings within training dataset 1102 and can determine how individual inputs relate to untrained dataset 1102. In at least one embodiment, unsupervised training can be used to generate a self-organizing map, a trained neural network 1108 capable of performing operations useful for reducing the dimensionality of new data 1112. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for identifying data points in new dataset 1112 that deviate from the normal pattern of new dataset 1112.

[0105] In at least one embodiment, semi-supervised learning can be used, a technique in which the training dataset 1102 includes a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 1104 can be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 1108 to adapt to new data 1112 without forgetting the knowledge infused into the network during initial training.

[0106] Data Center

[0107] Figure 13 An example data center 1300 is shown in which at least one embodiment may be used. In at least one embodiment, the data center 1300 includes a data center infrastructure layer 1310 , a framework layer 1320 , a software layer 1330 , and an application layer 1340 .

[0108] In at least one embodiment, Figure 13As shown, the data center infrastructure layer 1310 may include a resource coordinator 1312, grouped computing resources 1314, and node computing resources ("node CRs") 1316(1)-1316(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 1316(1)-1316(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), storage devices (e.g., dynamic read-only memories), memory devices (e.g., solid-state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules. In at least one embodiment, one or more of the node CRs 1316(1)-1316(N) may be servers having one or more of the above-mentioned computing resources.

[0109] In at least one embodiment, grouped computing resources 1314 may include individual groups of node CRs housed in one or more racks (not shown), or individual groups of multiple racks housed in data centers at various geographic locations (also not shown). Individual groups of node CRs within grouped computing resources 1314 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs including CPUs or processors may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and any combination of network switches.

[0110] In at least one embodiment, resource orchestrator 1322 may be configured or otherwise configured to control one or more node CRs 1316(1)-1316(N) and / or grouped computing resources 1314. In at least one embodiment, resource orchestrator 1322 may comprise a software design infrastructure ("SDI") management entity for data center 1300. In at least one embodiment, a resource orchestrator may comprise hardware, software, or some combination thereof.

[0111] In at least one embodiment, Figure 13As shown, the framework layer 1320 includes a job scheduler 1332, a configuration manager 1334, a resource manager 1336, and a distributed file system 1338. In at least one embodiment, the framework layer 1320 may include a framework for supporting the software 1332 of the software layer 1330 and / or one or more applications 1342 of the application layer 1340. In at least one embodiment, the software 1332 or the application 1342 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1320 may be, but is not limited to, a free and open source software web application framework, such as Apache Spark, which can utilize the distributed file system 1338 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, job scheduler 1332 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of data center 1300. In at least one embodiment, configuration manager 1334 may be capable of configuring different layers (e.g., software layer 1330 and framework layer 1320 including Spark) and a distributed file system 1338 for supporting large-scale data processing. In at least one embodiment, resource manager 1336 may be capable of managing clustered or grouped computing resources mapped to or allocated to support distributed file system 1338 and job scheduler 1332. In at least one embodiment, clustered or grouped computing resources may include grouped computing resources 1314 at data center infrastructure layer 1310. In at least one embodiment, resource manager 1336 may coordinate with resource coordinator 1312 to manage these mapped or allocated computing resources.

[0112] In at least one embodiment, the software 1332 included in the software layer 1330 may include software used by the node CRs 1316(1)-1316(N), the grouped computing resources 1314, and / or at least a portion of the distributed file system 1338 of the framework layer 1320. The one or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.

[0113] In at least one embodiment, the applications 1342 included in the application layer 1340 may include one or more types of applications used by at least a portion of the node CRs 1316(1)-1316(N), the grouped computing resources 1314, and / or the distributed file system 1338 of the framework layer 1320. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0114] In at least one embodiment, any of configuration manager 1334, resource manager 1336, and resource coordinator 1312 can implement any number and type of self-modification actions based on any number and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of data center 1300 from making potentially poor configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.

[0115] In at least one embodiment, data center 1300 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 1300. In at least one embodiment, using the weight parameters calculated using one or more training techniques described herein, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1300.

[0116] In at least one embodiment, a data center can use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as image recognition, speech recognition, or other artificial intelligence services.

[0117] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11Provides details about the reasoning and / or training logic 1015. In at least one embodiment, the reasoning and / or training logic 1015 may be Figure 13 In some embodiments, the present invention relates to a method for performing inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0118] In accordance with at least one embodiment, data center infrastructure 1310 can receive input text and target the input to corresponding components of application layer 1340 and software layer 1330 for training and / or inference as discussed herein.

[0119] Computer system

[0120] Figure 14 1400, which may include a processor 1400 that may include execution units to execute instructions. In at least one embodiment, computer system 1400 may include, but is not limited to, components such as processor 1402, whose execution units include logic to execute algorithms for processing data in accordance with the present disclosure, such as the embodiments described herein. In at least one embodiment, computer system 1400 may include, but is not limited to, a processor 1402, whose execution units include logic to execute algorithms for processing data. In at least one embodiment, computer system 1400 may include a processor such as the Intel® processor 1402 available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM 、 XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessor, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, computer system 1400 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0121] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications may include microcontrollers, digital signal processors ("DSPs"), system-on-chips, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system that can execute one or more instructions according to at least one embodiment.

[0122] In at least one embodiment, the computer system 1400 may include, but is not limited to, a processor 1402, which may include, but is not limited to, one or more execution units 1408 to perform machine learning model training and / or reasoning according to the techniques described herein. In at least one embodiment, the computer system 1400 is a single-processor desktop or server system, but in another embodiment, the computer system 1400 may be a multi-processor system. In at least one embodiment, the processor 1402 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor that implements a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 1402 may be coupled to a processor bus 1410 that can transmit data signals between the processor 1402 and other components in the computer system 1400.

[0123] In at least one embodiment, processor 1402 may include, but is not limited to, a level 1 ("L1") internal cache memory ("cache") 1404. In at least one embodiment, processor 1402 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 1402. Other embodiments may include a combination of internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, register file 1406 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and an instruction pointer register.

[0124] In at least one embodiment, an execution unit 1408, including but not limited to logic for performing integer and floating-point operations, is also located within processor 1402. In at least one embodiment, processor 1402 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, execution unit 1408 may include logic for processing a packed instruction set 1409. In at least one embodiment, by including packed instruction set 1409 in the instruction set of general-purpose processor 1402 and associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data within general-purpose processor 1402. In one or more embodiments, many multimedia applications may be accelerated and executed more efficiently by using the full width of the processor's data bus to perform operations on the packed data, which may eliminate the need to transfer smaller units of data across the processor's data bus to perform one or more operations one data element at a time.

[0125] In at least one embodiment, execution unit 1408 may also be used in a microcontroller, an embedded processor, a graphics device, a DSP, and other types of logic circuits. In at least one embodiment, computer system 1400 may include, but is not limited to, memory 1420. In at least one embodiment, memory 1420 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 1420 may store instructions 1419 and / or data 1421 represented by data signals that may be executed by processor 1402.

[0126] In at least one embodiment, a system logic chip can be coupled to the processor bus 1410 and the memory 1420. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub ("MCH") 1416, and the processor 1402 can communicate with the MCH 1416 via the processor bus 1410. In at least one embodiment, the MCH 1416 can provide a high-bandwidth memory path 1418 to the memory 1420 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 1416 can initiate data signals between the processor 1402, the memory 1420, and other components in the computer system 1400, and bridge data signals between the processor bus 1410, the memory 1420, and the system I / O 1422. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 1416 may be coupled to the memory 1420 via a high-bandwidth memory path 1218 , and the graphics / video card 1412 may be coupled to the MCH 1416 via an Accelerated Graphics Port (“AGP”) interconnect 1414 .

[0127] In at least one embodiment, computer system 1400 may use system I / O 1422 as a proprietary hub interface bus to couple MCH 1416 to I / O controller hub ("ICH") 1430. In at least one embodiment, ICH 1430 may provide direct connection to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to memory 1420, chipset, and processor 1402. Examples may include, but are not limited to, an audio controller 1429, a firmware hub ("flash BIOS") 1428, a wireless transceiver 1426, data storage 1424, a traditional I / O controller 1423 including user input and keyboard interface, a serial expansion port 1427 (e.g., a universal serial bus (USB)), and a network controller 1434. Data storage 1424 may include a hard drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0128] In at least one embodiment, Figure 14 The system is shown as comprising interconnected hardware devices or "chips", while in other embodiments, Figure 14 An exemplary system on a chip ("SoC") may be shown. In at least one embodiment, Figure 14The devices shown in FIG1400 may be interconnected with a proprietary interconnect, a standardized interconnect (eg, PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 1400 are interconnected using a Compute Express Link (CXL) interconnect.

[0129] Reasoning and / or training logic 1015 is used to perform reasoning and / or training operations related to one or more embodiments. Figure 10 and / or Figure 11 Provides details about the reasoning and / or training logic 1015. In at least one embodiment, the reasoning and / or training logic 1015 may be in the system Figure 14 for use in performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0130] In some embodiments, for example, a video data stream may be received via expansion port 1427 or wireless transceiver 1426 and then directed to processor 1402 and / or video graphics card 1412 for processing. Depending on whether the component is part of a device (e.g., an autonomous vehicle) or a separate device, the output may then be transmitted via I / O to a control system or to the vehicle via a wireless transceiver.

[0131] Figure 15 1 is a block diagram illustrating an electronic device 1500 for utilizing a processor 1510 according to at least one embodiment. In at least one embodiment, the electronic device 1500 may be, for example, but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0132] In at least one embodiment, system 1500 may include, but is not limited to, a processor 1510 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1510 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, Figure 15 shows a system comprising interconnected hardware devices or "chips", while in other embodiments, Figure 15 An exemplary system on a chip ("SoC") may be shown. In at least one embodiment, Figure 15 The devices shown in can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 13 One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.

[0133] In at least one embodiment, Figure 15 It may include a display 1524, a touch screen 1525, a touchpad 1530, a near field communication unit ("NFC") 1545, a sensor hub 1540, a thermal sensor 1546, a fast chipset ("EC") 1535, a trusted platform module ("TPM") 1538, a BIOS / firmware / flash memory ("BIOS, FW Flash") 1522, a DSP 1560, a drive 1520 (e.g., a solid state disk ("SSD") or a hard disk drive ("HDD")), a wireless local area network unit ("WLAN") 1550, a Bluetooth unit 1552, a wireless wide area network unit ("WWAN") 1556, a global positioning system (GPS) 1555, a camera ("USB 3.0 camera") 1554 (e.g., a USB 3.0 camera), and / or a low power double data rate ("LPDDR") memory unit ("LPDDR3") 1515 implemented in, for example, the LPDDR3 standard. These components may each be implemented in any suitable manner.

[0134] In at least one embodiment, other components may be communicatively coupled to processor 1510 via the components discussed above. In at least one embodiment, accelerometer 1541, ambient light sensor (“ALS”) 1542, compass 1543, and gyroscope 1544 may be communicatively coupled to sensor hub 1540. In at least one embodiment, thermal sensor 1539, fan 1537, keyboard 1546, and touchpad 1530 may be communicatively coupled to EC 1535. In at least one embodiment, speaker 1563, earphone 1564, and microphone (“mic”) 1565 may be communicatively coupled to audio unit (“audio codec and class-D amplifier”) 1564, which in turn may be communicatively coupled to DSP 1560. In at least one embodiment, audio unit 1564 may include, for example, but not limited to, an audio codec / decoder (“codec”) and a class-D amplifier. In at least one embodiment, SIM card (“SIM”) 1557 may be communicatively coupled to WWAN unit 1556. In at least one embodiment, components such as the WLAN unit 1550 and the Bluetooth unit 1552 and the WWAN unit 1556 may be implemented as a next generation form factor (NGFF).

[0135] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Provides details about the reasoning and / or training logic 1015. In at least one embodiment, the reasoning and / or training logic 1015 may be Figure 15 for use in performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0136] Figure 16 Illustrated is a computer system 1600 in accordance with at least one embodiment. In at least one embodiment, the computer system 1600 is configured to implement the various processes and methods described throughout this disclosure.

[0137] In at least one embodiment, computer system 1600 includes, but is not limited to, at least one central processing unit ("CPU") 1602 connected to a communication bus 1610 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1600 includes, but is not limited to, main memory 1604 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data may be stored in main memory 1604 in the form of random access memory ("RAM"). In at least one embodiment, a network interface subsystem ("network interface") 1622 provides an interface to other computing devices and networks for receiving data from computer system 1600 and transmitting data to other systems.

[0138] In at least one embodiment, computer system 1600 includes, but is not limited to, input device 1608, parallel processing system 1612, and display device 1606, which may be implemented using a cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light emitting diode ("LED"), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from input device 1608 (such as a keyboard, mouse, touchpad, microphone, etc.). In at least one embodiment, each of the aforementioned modules may be located on a single semiconductor platform to form a processing system.

[0139] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Provides details about the reasoning and / or training logic 1015. In at least one embodiment, the reasoning and / or training logic 1015 may be Figure 16 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0140] Figure 17 A computer system 1700 is shown according to at least one embodiment. In at least one embodiment, computer system 1700 includes, but is not limited to, a computer 1710 and a USB drive 1720. In at least one embodiment, computer 1710 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, computer 1710 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.

[0141] In at least one embodiment, the USB disk 1720 includes, but is not limited to, a processing unit 1730, a USB interface 1740, and USB interface logic 1750. In at least one embodiment, the processing unit 1730 can be any instruction execution system, device, or apparatus capable of executing instructions. In at least one embodiment, the processing unit 1730 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1730 includes an application-specific integrated circuit ("ASIC") that is optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 1730 is a tensor processing unit ("TPC") that is optimized to perform machine learning inference operations. In at least one embodiment, the processing core 1730 is a vision processing unit ("VPU") that is optimized to perform machine vision and machine learning inference operations.

[0142] In at least one embodiment, USB interface 1740 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 1740 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1740 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 1750 can include any number and type of logic that enables processing unit 1730 to connect to a device (e.g., computer 1710) via USB connector 1740.

[0143] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Provides details about the reasoning and / or training logic 1015. In at least one embodiment, the reasoning and / or training logic 1015 may be Figure 17 In some embodiments, the present invention relates to a method for performing inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0144] Figure 18 is a block diagram illustrating an exemplary system on a chip integrated circuit 1800 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1800 includes one or more application processors 1805 (e.g., CPUs), at least one graphics processor 1810, and may additionally include an image processor 1815 and / or a video processor 1820, any of which may be modular IP cores. In at least one embodiment, integrated circuit 1800 includes peripheral or bus logic, including a USB controller 1825, a UART controller 1830, an SPI / SDIO controller 1835, and an I.sup.2S / I.sup.2C controller 1840. In at least one embodiment, integrated circuit 1800 may include a display device 1845 coupled to one or more of a High-Definition Multimedia Interface (HDMI) controller 1850 and a Mobile Industry Processor Interface (MIPI) display interface 1855. In at least one embodiment, memory may be provided by a flash memory subsystem 1860, including flash memory and a flash memory controller. In at least one embodiment, a memory interface for accessing SDRAM or SRAM memory devices may be provided via a memory controller 1865. In at least one embodiment, some integrated circuits also include an embedded security engine 1870.

[0145] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Details are provided regarding inference and / or training logic 1015. In at least one embodiment, inference and / or training logic 1015 may be used in integrated circuit 1800 to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0146] For example, inference and / or training logic 1015 can accept an input video stream and generate inferences for objects represented in the video stream, as described herein. In at least some embodiments, image processor 1815 can be used to process video frames as they are received.

[0147] Figures 19A-19B An exemplary integrated circuit and associated graphics processor according to various embodiments described herein are shown, which can be manufactured using one or more IP cores. In addition to the illustrations, other logic and circuits may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0148] Figures 19A-19B is a block diagram illustrating an exemplary graphics processor for use within a SoC according to embodiments described herein. Figure 19A An exemplary graphics processor 1910 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Figure 19B Another exemplary graphics processor 1940 of a system on a chip integrated circuit is shown, which can be manufactured using one or more IP cores according to at least one embodiment. In at least one embodiment, Figure 19A The graphics processor 1910 is a low power graphics processor core. In at least one embodiment, Figure 19B The graphics processor 1940 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1910, 1940 can be Figure 18 A variant of the graphics processor 1810.

[0149] In at least one embodiment, the graphics processor 1910 includes a vertex processor 1905 and one or more fragment processors 1915A-1915N (e.g., 1915A, 1915B, 1915C, 1915D through 1915N-1 and 1915N). In at least one embodiment, the graphics processor 1910 can execute different shader programs via separate logic, such that the vertex processor 1905 is optimized to perform operations for the vertex shader program, while the one or more fragment processors 1915A-1915N perform fragment (e.g., pixel) shading operations for the fragment or pixel or shader program. In at least one embodiment, the vertex processor 1905 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the one or more fragment processors 1915A-1915N use the primitives and vertex data generated by the vertex processor 1905 to generate a frame buffer for display on a display device. In at least one embodiment, one or more fragment processors 1915A-1915N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.

[0150] In at least one embodiment, graphics processor 1910 additionally includes one or more memory management units (MMUs) 1920A-1920B, one or more caches 1925A-1925B, and one or more circuit interconnects 1930A-1930B. In at least one embodiment, one or more MMUs 1920A-1920B provide a mapping of virtual to physical addresses for graphics processor 1910, including for vertex processor 1905 and / or fragment processors 1915A-1915N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 1925A-1925B. In at least one embodiment, one or more MMUs 1920A-1920B may synchronize with other MMUs within the system, including with other MMUs. Figure 18 One or more MMUs associated with one or more application processors 1805, graphics processor 1815, and / or video processor 1820 enable each processor 1805-1820 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1930A-1930B enable graphics processor 1910 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.

[0151] In at least one embodiment, graphics processor 1940 includes Figure 19A One or more MMUs 1920A-1920B, caches 1925A-1925B, and circuit interconnects 1930A-1930B of graphics processor 1910. In at least one embodiment, graphics processor 1940 includes one or more shader cores 1955A-1955N (e.g., 1955A, 1955B, 1955C, 1955D, 1955E, 1955F, through 1955N-1 and 1955N), which provide a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1940 includes an inter-core task manager 1945 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1955A-1955N and a tiling unit 1958 to accelerate tile-based rendering operations in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within the scene or to optimize use of internal caches.

[0152] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Provides details about the inference and / or training logic 1015. In at least one embodiment, the inference and / or training logic 1015 may be implemented in an integrated circuit. Figure 19A and / or Figure 19B Inference and / or training logic 1015 may be used to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. For example, inference and / or training logic 1015 may accept an input video stream and generate inferences for objects represented in the video stream, as described herein.

[0153] Figure 20A-Figure 20B Additional exemplary graphics processor logic according to embodiments described herein is shown. In at least one embodiment, Figure 20A Shows that can be included in Figure 18 The graphics core 2000 within the graphics processor 1810 may, in at least one embodiment, be Figure 19B Unified shader cores 1955A-1955N. Figure 20B A highly parallel, general-purpose graphics processing unit 2030 suitable for deployment on a multi-chip module in at least one embodiment is shown.

[0154] In at least one embodiment, graphics core 2000 includes a shared instruction cache 2002, texture unit 2018, and cache / shared memory 2020, which are shared by execution resources within graphics core 2000. In at least one embodiment, graphics core 2000 may include multiple slices 2001A-2001N, or partitions of each core, and a graphics processor may include multiple instances of graphics core 2000. Slices 2001A-2001N may include support logic including local instruction caches 2004A-2004N, thread schedulers 2006A-2006N, thread dispatchers 2008A-2008N, and a set of registers 2010A-2010N. In at least one embodiment, slices 2001A-2001N may include a set of additional function units (AFU2012A-2012N), floating point units (FPU2014A-2014N), integer arithmetic logic units (ALU2016-2016N), address calculation units (ACU2013A-2013N), double precision floating point units (DPFPU2015A-2015N) and matrix processing units (MPU2017A-2017N).

[0155] In at least one embodiment, the FPU 2014A-2014N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the FPU 2015A-2015N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU 2016A-2016N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPU 2017A-2017N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPU 2017A-2017N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFU 2012A-2012N can perform additional logical operations not supported by the floating-point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).

[0156] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Details are provided regarding inference and / or training logic 1015. In at least one embodiment, inference and / or training logic 1015 may be used in graphics core 2000 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0157] Figure 20BA general purpose processing unit (GPGPU) 2030 is shown in at least one embodiment, which can be configured to enable highly parallel computational operations to be performed by an array of graphics processing units. In at least one embodiment, GPGPU 2030 can be directly linked to other instances of GPGPU 2030 to create a multi-GPU cluster to increase the training speed for deep neural networks. In at least one embodiment, GPGPU 2030 includes a host interface 2032 to enable connection to a host processor. In at least one embodiment, host interface 2032 is a PCI Express interface. In at least one embodiment, host interface 2032 can be a vendor-specific communication interface or communication structure. In at least one embodiment, GPGPU 2030 receives commands from the host processor and dispatches the execution threads associated with those commands to a set of compute clusters 2036A-2036H using a global scheduler 2034. In at least one embodiment, compute clusters 2036A-2036H share a cache memory 2038. In at least one embodiment, cache memory 2038 may serve as a higher level cache for cache memories within compute clusters 2036A-2036H.

[0158] In at least one embodiment, GPGPU 2030 includes memory 2044A-2044B coupled to compute clusters 2036A-2036H via a set of memory controllers 2042A-2042B. In at least one embodiment, memory 2044A-2044B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.

[0159] In at least one embodiment, computing clusters 2036A-2036H each include a set of graphics cores, such as Figure 20A The graphics core 2000 may include multiple types of integer and floating-point logic units, including those for performing computational operations within a range of precision suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each of the compute clusters 2036A-2036H may be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of the floating-point units may be configured to perform 64-bit floating-point operations.

[0160] In at least one embodiment, multiple instances of GPGPU 2030 can be configured to operate as a compute cluster. In at least one embodiment, the communications used by compute clusters 2036A-2036H for synchronization and data exchange vary between embodiments. In at least one embodiment, multiple instances of GPGPU 2030 communicate via host interface 2032. In at least one embodiment, GPGPU 2030 includes an I / O hub 2039 that couples GPGPU 2030 to GPU link 2040, enabling direct connections to other instances of GPGPU 2030. In at least one embodiment, GPU link 2040 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 2030. In at least one embodiment, GPU link 2040 is coupled to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 2030 are located in separate data processing systems and communicate via a network device accessible via host interface 2032. In at least one embodiment, GPU link 2040 may be configured to enable connection to a host processor, in addition to or in place of host interface 2032 .

[0161] In at least one embodiment, GPGPU 2030 can be configured to train neural networks. In at least one embodiment, GPGPU 2030 can be used within an inference platform. In at least one embodiment where GPGPU 2030 is used for inference, the GPGPU can include fewer compute clusters 2036A-2036H than when the GPGPU is used to train a neural network. In at least one embodiment, the memory technology associated with memories 2044A-2044B can differ between the inference and training configurations, with higher-bandwidth memory technology being dedicated to the training configuration. In at least one embodiment, the inference configuration of GPGPU 2030 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can provide support for one or more 8-bit integer dot product instructions, which can be used during inference operations of a deployed neural network.

[0162] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11Details are provided regarding inference and / or training logic 1015. In at least one embodiment, inference and / or training logic 1015 may be used in GPGPU 2030 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0163] Figure 21 2 is a block diagram illustrating a computing system 2100 according to at least one embodiment. In at least one embodiment, computing system 2100 includes a processing subsystem 2101 having one or more processors 2102 and a system memory 2104 communicating via an interconnect path that may include a memory hub 2105. In at least one embodiment, memory hub 2105 may be a separate component within a chipset assembly or integrated within one or more processors 2102. In at least one embodiment, memory hub 2105 is coupled to an I / O subsystem 2111 via a communication link 2106. In one embodiment, I / O subsystem 2111 includes an I / O hub 2107, which enables computing system 2100 to receive input from one or more input devices 2108. In at least one embodiment, I / O hub 2107 may enable a display controller, included in one or more processors 2102, to provide output to one or more display devices 2110A. In at least one embodiment, the one or more display devices 2110A coupled to the I / O hub 2107 may include local, internal, or embedded display devices.

[0164] In at least one embodiment, the processing subsystem 2101 includes one or more parallel processors 2112 coupled to the memory hub 2105 via a bus or other communication link 2113. In at least one embodiment, the communication link 2113 can be one of many standard-based communication link technologies or protocols, such as, but not limited to, PCI Express, or can be a vendor-specific communication interface or communication structure. In at least one embodiment, the one or more parallel processors 2112 form a computationally focused parallel or vector processing system that can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 2112 form a graphics processing subsystem that can output pixels to one of one or more display devices 2110A coupled via the I / O hub 2107. In at least one embodiment, the one or more parallel processors 2112 can also include a display controller and display interface (not shown) to enable direct connection to the one or more display devices 2110B.

[0165] In at least one embodiment, a system storage unit 2114 can be connected to the I / O hub 2107 to provide a storage mechanism for the computing system 2100. In at least one embodiment, an I / O switch 2116 can be used to provide an interface mechanism to enable connections between the I / O hub 2107 and other components, such as a network adapter 2118 and / or a wireless network adapter 2119 that can be integrated into the platform, as well as various other devices that can be added via one or more add-on devices 2120. In at least one embodiment, the network adapter 2118 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 2119 can include one or more of Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more radios.

[0166] In at least one embodiment, computing system 2100 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to I / O hub 2107. Figure 21 The communication paths interconnecting the various components in the system may be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI-Express), or other bus or point-to-point communication interfaces and / or protocols (e.g., NV-Link high-speed interconnect or interconnect protocol).

[0167] In at least one embodiment, one or more parallel processors 2112 include circuits optimized for graphics and video processing (including, for example, video output circuits) and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 2112 include circuits optimized for general-purpose processing. In at least one embodiment, the components of the computing system 2100 can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 2112, memory hub 2105, one or more processors 2102, and I / O hub 2107 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of the computing system 2100 can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computing system 2100 can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules into a modular computing system.

[0168] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Details are provided regarding inference and / or training logic 1015. In at least one embodiment, inference and / or training logic 1015 can be used in system diagram 2100 to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0169] processor

[0170] Figure 22 2200 in accordance with at least one embodiment. In at least one embodiment, the various components of the parallel processor 2200 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the parallel processor 2200 is shown as a processor according to an exemplary embodiment. Figure 21 A variation of the one or more parallel processors 2112 is shown.

[0171] In at least one embodiment, parallel processor 2200 includes parallel processing unit (PPU) 2202. In at least one embodiment, PPU 2202 includes an I / O unit 2204 that enables communication with other devices, including other instances of PPU 2202. In at least one embodiment, I / O unit 2204 can be directly connected to other devices. In at least one embodiment, I / O unit 2204 connects to other devices using a hub or switch interface (e.g., memory hub 2105). In at least one embodiment, the connection between memory hub 2105 and I / O unit 2204 forms communication link 2113. In at least one embodiment, I / O unit 2204 connects to a host interface 2206 and a memory crossbar switch 2216, where host interface 2206 receives commands for performing processing operations, and memory crossbar switch 2216 receives commands for performing memory operations.

[0172] In at least one embodiment, when host interface 2206 receives command buffers via I / O unit 2204, host interface 2206 can direct work operations to execute those commands to front-end 2208. In at least one embodiment, front-end 2208 is coupled to scheduler 2210, which is configured to distribute commands or other work items to processing cluster array 2212. In at least one embodiment, scheduler 2210 ensures that processing cluster array 2212 is properly configured and in a valid state before distributing tasks to processing cluster array 2212. In at least one embodiment, scheduler 2210 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 2210 can be configured to perform complex scheduling and work distribution operations at both coarse and fine granularity, thereby enabling rapid preemption and context switching of threads executing on processing array 2212. In at least one embodiment, host software can authenticate workloads for scheduling on processing array 2212 through one of multiple graphics processing doorbells. In at least one embodiment, the workload may then be automatically distributed across the processing array 2212 by scheduler 2210 logic within a microcontroller that includes scheduler 2210 .

[0173] In at least one embodiment, processing cluster array 2212 may include up to "N" processing clusters (e.g., cluster 2214A, cluster 2214B, through cluster 2214N). In at least one embodiment, each cluster 2214A-2214N of processing cluster array 2212 may execute a large number of concurrent threads. In at least one embodiment, scheduler 2210 may allocate work to clusters 2214A-2214N of processing cluster array 2212 using various scheduling and / or work distribution algorithms, which may vary depending on the workload generated by each program or computation type. In at least one embodiment, scheduling may be handled dynamically by scheduler 2210 or may be assisted in part by compiler logic during the compilation of program logic configured to be executed by processing cluster array 2212. In at least one embodiment, different clusters 2214A-2214N of processing cluster array 2212 may be assigned to process different types of programs or to perform different types of computations.

[0174] In at least one embodiment, processing cluster array 2212 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster array 2212 can be configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, processing cluster array 2212 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.

[0175] In at least one embodiment, processing cluster array 2212 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 2212 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing cluster array 2212 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing units 2202 may transfer data from system memory via I / O units 2204 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 2222) during processing and then written back to system memory.

[0176] In at least one embodiment, when parallel processing unit 2202 is used to perform graphics processing, scheduler 2210 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 2214A-2214N of processing cluster array 2212. In at least one embodiment, portions of processing cluster array 2212 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen-space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of clusters 2214A-2214N can be stored in a buffer to allow the intermediate data to be transferred between clusters 2214A-2214N for further processing.

[0177] In at least one embodiment, the processing cluster array 2212 can receive processing tasks to be executed via the scheduler 2210, which receives commands defining the processing tasks from the front end 2208. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 2210 can be configured to obtain the index corresponding to the task, or can receive the index from the front end 2208. In at least one embodiment, the front end 2208 can be configured to ensure that the processing cluster array 2212 is configured in a valid state before starting the workload specified by the incoming command buffer (e.g., batch buffer, push buffer, etc.).

[0178] In at least one embodiment, each of one or more instances of parallel processing unit 2202 can be coupled to parallel processor memory 2222. In at least one embodiment, parallel processor memory 2222 can be accessed via memory crossbar 2216, which can receive memory requests from processing cluster array 2212 and I / O unit 2204. In at least one embodiment, memory crossbar 2216 can access parallel processor memory 2222 via memory interface 2218. In at least one embodiment, memory interface 2218 can include multiple partition units (e.g., partition unit 2220A, partition unit 2220B, through partition unit 2220N), each of which can be coupled to a portion of parallel processor memory 2222 (e.g., a memory unit). In at least one embodiment, the plurality of partition units 2220A-2220N are configured to be equal to the number of memory cells, such that the first partition unit 2220A has a corresponding first memory cell 2224A, the second partition unit 2220B has a corresponding memory cell 2224B, and the Nth partition unit 2220N has a corresponding Nth memory cell 2224N. In at least one embodiment, the number of partition units 2220A-2220N may not be equal to the number of memory devices.

[0179] In at least one embodiment, memory units 2224A-2224N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 2224A-2224N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, render targets, such as frame buffers or texture maps, may be stored across memory units 2224A-2224N, allowing partition units 2220A-2220N to write portions of each render target in parallel to efficiently use the available bandwidth of parallel processor memory 2222. In at least one embodiment, local instances of parallel processor memory 2222 may be eliminated to facilitate a unified memory design that utilizes system memory in combination with local cache memory.

[0180] In at least one embodiment, any of the clusters 2214A-2214N in the processing cluster array 2212 can process data to be written to any memory unit 2224A-2224N within the parallel processor memory 2222. In at least one embodiment, the memory crossbar 2216 can be configured to transmit the output of each cluster 2214A-2214N to any partition unit 2220A-2220N or another cluster 2214A-2214N, which can perform other processing operations on the output. In at least one embodiment, each cluster 2214A-2214N can communicate with a memory interface 2218 via the memory crossbar 2216 to read from or write to various external storage devices. In at least one embodiment, memory crossbar switch 2216 has connections to memory interface 2218 for communicating with I / O unit 2204, as well as connections to local instances of parallel processor memory 2222, thereby enabling processing units within different processing clusters 2214A-2214N to communicate with system memory or other memory that is not local to parallel processing unit 2202. In at least one embodiment, memory crossbar switch 2216 can use virtual channels to separate traffic flows between clusters 2214A-2214N and partition units 2220A-2220N.

[0181] In at least one embodiment, multiple instances of parallel processing unit 2202 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 2202 can be configured to interoperate with each other, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of parallel processing unit 2202 can include a higher precision floating point unit than other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 2202 or parallel processor 2200 can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0182] Figure 23 is a block diagram of a partition unit 2320 according to at least one embodiment. In at least one embodiment, the partition unit 2320 is Figure 22 220N。In at least one embodiment, the partition unit 2320 includes an L2 cache 2321, a frame buffer interface 2325, and a raster operations unit ("ROP") 2326. The L2 cache 2321 is a read / write cache that is configured to perform load and store operations received from the memory crossbar 2316 and the ROP 2326. In at least one embodiment, the L2 cache 2321 outputs read misses and urgent writeback requests to the frame buffer interface 2325 for processing. In at least one embodiment, updates can also be sent to the frame buffer via the frame buffer interface 2325 for processing. In at least one embodiment, the frame buffer interface 2325 interacts with one of the memory units in the parallel processor memory, such as the memory units 2224A-2224N of FIG. 22 (e.g., within the parallel processor memory 2222).

[0183] In at least one embodiment, ROP 2326 is a processing unit that performs raster operations such as stenciling, z-testing, blending, etc. In at least one embodiment, ROP 2326 then outputs the processed graphics data, which is stored in graphics memory. In at least one embodiment, ROP 2326 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic may be lossless compression logic utilizing one or more of a variety of compression algorithms. The type of compression performed by ROP 2326 may vary depending on the statistical characteristics of the data to be compressed. For example, in at least one embodiment, incremental color compression is performed on the depth and color data on a per-tile basis.

[0184] In at least one embodiment, ROP 2326 is included within each processing cluster (e.g., clusters 2214A-2214N of FIG. 22 ), rather than within partition unit 2320. In at least one embodiment, read and write requests for pixel data are transferred through memory crossbar 2316 rather than pixel fragment data transfers. In at least one embodiment, processed graphics data may be displayed on a display device such as a Figure 21 2100), routed by processor 2102 for further processing, or by Figure 22 One of the processing entities within parallel processor 2200 is routed for further processing.

[0185] Figure 24 is a block diagram of a processing cluster 2414 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is Figure 22 In at least one embodiment, one or more of the one or more processing clusters 2414 can be configured to execute many threads in parallel, where a "thread" refers to an instance of a particular program executed on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of synchronized threads, which uses an instruction unit that is configured to issue instructions to a set of processing engines within each processing cluster.

[0186] In at least one embodiment, the operation of the processing cluster 2214 can be controlled by a pipeline manager 2432 that assigns processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 2432 Figure 22The scheduler 2210 receives instructions and manages the execution of these instructions through the graphics multiprocessor 2434 and / or the texture unit 2436. In at least one embodiment, the graphics multiprocessor 2434 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 2414. In at least one embodiment, one or more instances of the graphics multiprocessor 2434 may be included within the processing cluster 2414. In at least one embodiment, the graphics multiprocessor 2434 may process data, and the data crossbar 2440 may be used to distribute the processed data to one of multiple possible destinations (including other shader units). In at least one embodiment, the pipeline manager 2432 may facilitate the distribution of processed data by specifying the destination of the processed data to be distributed to the data crossbar 2440.

[0187] In at least one embodiment, each graphics multiprocessor 2434 within a processing cluster 2414 may include the same set of function execution logic (e.g., arithmetic logic units, load-store units, etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, where new instructions may be issued before previous instructions have completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.

[0188] In at least one embodiment, instructions transmitted to processing cluster 2414 constitute threads. In at least one embodiment, a group of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program on different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 2434. In at least one embodiment, a thread group can include fewer threads than the number of processing engines within graphics multiprocessor 2434. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines, one or more processing engines may be idle during the processing of a loop within the thread group. In at least one embodiment, a thread group can also include more threads than the number of processing engines within graphics multiprocessor 2044. In at least one embodiment, when a thread group includes more threads than the number of processing engines within graphics multiprocessor 2434, processing can be performed within consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 2434.

[0189] In at least one embodiment, the graphics multiprocessor 2434 includes an internal cache memory to perform load and store operations. In at least one embodiment, the graphics multiprocessor 2434 can abandon the internal cache and use cache memory within the processing cluster 2414 (e.g., L1 cache 2448). In at least one embodiment, each graphics multiprocessor 2434 can also access a partition unit (e.g., Figure 22 L2 cache within partition units 2220A-2220N) is shared across all processing clusters 2414 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 2434 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processing unit 2402 can be used as global memory. In at least one embodiment, processing cluster 2414 includes multiple instances of graphics multiprocessor 2434, which can share instructions and data that can be stored in L1 cache 2448.

[0190] In at least one embodiment, each processing cluster 2414 may include a memory management unit ("MMU") 2445 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2445 may reside in Figure 22 2218. In at least one embodiment, the MMU 2445 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles and to cache memory lines. In at least one embodiment, the MMU 2445 may include an address translation lookaside buffer (TLB) or a cache that may reside within the graphics multiprocessor 2434 or L1 cache or processing cluster 2414. In at least one embodiment, the physical address is processed to assign surface data access locality to allow for efficient request interleaving between partition units. In at least one embodiment, a cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0191] In at least one embodiment, the processing clusters 2414 can be configured such that each graphics multiprocessor 2434 is coupled to a texture unit 2436 to perform texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 2434, and retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 2434 outputs processed tasks to a data crossbar 2440 to provide the processed tasks to another processing cluster 2414 for further processing or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via the memory crossbar 2416. In at least one embodiment, a preROP 2442 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 2434 and direct the data to a ROP unit, which can communicate with a partition unit (e.g., a partition unit) as described herein. Figure 22 In at least one embodiment, the PreROP 2442 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.

[0192] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Details are provided regarding inference and / or training logic 1015. In at least one embodiment, inference and / or training logic 1015 may be used in graphics processing cluster 2214 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0193] Figure 25A graphics multiprocessor 2534 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 2534 is coupled to a pipeline manager 2532 of the processing cluster 2514. In at least one embodiment, the graphics multiprocessor 2534 has an execution pipeline that includes, but is not limited to, an instruction cache 2552, an instruction unit 2554, an address mapping unit 2556, a register file 2558, one or more general purpose graphics processing unit (GPGPU) cores 2562, and one or more load / store units 2566. The GPGPU cores 2562 and the load / store units 2566 are coupled to a cache memory 2572 and a shared memory 2570 via a memory and cache interconnect 2568.

[0194] In at least one embodiment, the instruction cache 2552 receives a stream of instructions to be executed from the pipeline manager 2532. In at least one embodiment, the instructions are cached in the instruction cache 2552 and dispatched for execution by the instruction unit 2554. In one embodiment, the instruction unit 2054 can dispatch instructions as thread groups (e.g., warps), assigning each thread group to a different execution unit within the GPGPU core 2562. In at least one embodiment, instructions can access any local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 2556 can be used to convert addresses in the unified address space into different memory addresses that can be accessed by the load / store unit 2566.

[0195] In at least one embodiment, register file 2558 provides a set of registers for the functional units of graphics multiprocessor 2534. In at least one embodiment, register file 2558 provides temporary storage for operands for the data paths of the functional units (e.g., GPGPU core 2562, load / store unit 2566) connected to graphics multiprocessor 2534. In at least one embodiment, register file 2558 is divided between each functional unit such that a dedicated portion of register file 2558 is allocated to each functional unit. In at least one embodiment, register file 2558 is divided between the different warps being executed by graphics multiprocessor 2534.

[0196] In at least one embodiment, the GPGPU cores 2562 may each include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessor 2534. The GPGPU cores 2562 may be architecturally similar or may differ in architecture. In at least one embodiment, a first portion of the GPGPU core 2562 includes a single-precision FPU and integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 2534 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed-function or special-function logic.

[0197] In at least one embodiment, the GPGPU core 2562 includes SIMD logic capable of executing a single instruction on multiple sets of data. In at least one embodiment, the GPGPU core 2562 can physically execute SIMD4, SIMD8, and SIMD16 instructions, and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.

[0198] In at least one embodiment, the memory and cache interconnect 2568 is an interconnect network that connects each functional unit of the graphics multiprocessor 2534 to the register file 2558 and shared memory 2570. In at least one embodiment, the memory and cache interconnect 2568 is a crossbar interconnect that allows the load / store unit 2566 to perform load and store operations between the shared memory 2570 and the register file 2558. In at least one embodiment, the register file 2558 can operate at the same frequency as the GPGPU core 2562, resulting in very low latency for data transfers between the GPGPU core 2562 and the register file 2558. In at least one embodiment, the shared memory 2570 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 2534. In at least one embodiment, the cache memory 2572 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 2536. In at least one embodiment, the shared memory 2570 can also be used as a program-managed cache. In at least one embodiment, in addition to automatically cached data stored in cache memory 2572, threads executing on GPGPU core 2562 may programmatically store data in shared memory.

[0199] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated on the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (such as can be internal to the package or chip). In at least one embodiment, regardless of the manner in which the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0200] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11Details are provided regarding inference and / or training logic 1015. In at least one embodiment, inference and / or training logic 1015 may be used in graphics multiprocessor 2234 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0201] Figure 26 is a block diagram illustrating a microarchitecture for processor 2600, which may include logic circuitry for executing instructions, according to at least one embodiment. In at least one embodiment, processor 2600 may execute instructions including x86 instructions, ARM instructions, specialized instructions for application-specific integrated circuits (ASICs), and the like. In at least one embodiment, processor 2610 may include registers for storing packed data, such as the 64-bit-wide MMX™ registers in microprocessors enabled with MMX technology from Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers, available in integer and floating-point form, may operate with packed data elements with Single Instruction Multiple Data ("SIMD") and Streaming SIMD Extensions ("SSE") instructions. In at least one embodiment, 128-bit-wide XMM registers associated with SSE2, SSE3, SSE4, AVX, or later (generally referred to as "SSEx") technology may store such packed data operands. In at least one embodiment, processor 2110 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.

[0202] In at least one embodiment, processor 2600 includes an in-order front end ("Front End") 2601 to fetch instructions for execution and prepare them for later use in the processor pipeline. In at least one embodiment, Front End 2601 may include several units. In at least one embodiment, instruction prefetcher 2626 retrieves instructions from memory and provides them to instruction decoder 2628, which in turn decodes or interprets the instructions. For example, in at least one embodiment, instruction decoder 2628 decodes received instructions into one or more machine-executable operations called "microinstructions" or "micro-operations" (also referred to as "micro-ops" or "micro-instructions"). In at least one embodiment, instruction decoder 2628 parses the instructions into opcodes and corresponding data and control fields, which may be used by the microarchitecture to perform the operations according to at least one embodiment. In at least one embodiment, trace cache 2630 may assemble the decoded microinstructions into a program-ordered sequence or trace in microinstruction queue 2634 for execution. In at least one embodiment, when trace cache 2630 encounters a complex instruction, microcode ROM 2632 provides the microinstructions necessary to complete the operation.

[0203] In at least one embodiment, some instructions may be converted into a single micro-op, while other instructions may require several micro-ops to complete the entire operation. In at least one embodiment, if more than four micro-ops are required to complete an instruction, the instruction decoder 2628 may access the microcode ROM 2632 to execute the instruction. In at least one embodiment, an instruction may be decoded into a smaller number of micro-ops for processing at the instruction decoder 2628. In at least one embodiment, if multiple micro-ops are required to complete the operation, the instruction may be stored in the microcode ROM 2632. In at least one embodiment, the trace cache 2630 references the entry point programmable logic array ("PLA") to determine the correct micro-op pointer for reading the microcode sequence from the microcode ROM 2632 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after the microcode ROM 2632 completes the micro-op sequencing for the instruction, the front end 2601 of the machine may resume fetching micro-ops from the trace cache 2630.

[0204] In at least one embodiment, an out-of-order execution engine ("OOO engine") 2603 can prepare instructions for execution. In at least one embodiment, the OOO logic has multiple buffers to smooth and reorder the instruction flow to optimize performance as instructions flow down the pipeline and are scheduled for execution. In at least one embodiment, the OOO engine 2603 includes, but is not limited to, an allocator / register renamer 2640, a memory microinstruction queue 2642, an integer / floating-point microinstruction queue 2644, a memory scheduler 2646, a fast scheduler 2602, a slow / general purpose floating-point scheduler ("slow / general purpose FP scheduler") 2604, and a simple floating-point scheduler ("simple FP scheduler") 2606. In at least one embodiment, the fast scheduler 2602, the slow / general purpose floating-point scheduler 2104, and the simple floating-point scheduler 2606 are also collectively referred to as "microinstruction schedulers 2602, 2604, 2606." In at least one embodiment, the allocator / register renamer 2640 allocates the machine buffers and resources required for each microinstruction to execute in order. In at least one embodiment, the allocator / register renamer 2640 renames logical registers into entries in the register file. In at least one embodiment, the allocator / register renamer 2640 also allocates an entry for each microinstruction in one of two microinstruction queues: a memory microinstruction queue 2642 for memory operations and an integer / floating point microinstruction queue 2644 for non-memory operations, preceding the memory scheduler 2646 and the microinstruction schedulers 2602, 2604, 2606. In at least one embodiment, the microinstruction schedulers 2602, 2604, 2606 determine when a microinstruction is ready to execute based on the readiness of its dependent input register operand sources and the availability of execution resource microinstructions that need to be completed. In at least one embodiment, the fast scheduler 2602 of at least one embodiment can schedule on every half of the main clock cycle, while the slow / general floating point scheduler 2604 and the simple floating point scheduler 2606 can schedule once per main processor clock cycle. In at least one embodiment, the microinstruction schedulers 2602, 2604, 2606 arbitrate on the dispatch ports to schedule microinstructions for execution.

[0205] In at least one embodiment, execution block b11 includes, but is not limited to, integer register file / branch network 2608, floating-point register file / branch network ("FP register file / branch network") 2610, address generation units ("AGUs") 2612 and 2614, fast arithmetic logic units ("fast ALUs") 2616 and 2618, slow arithmetic logic unit ("slow ALU") 2620, floating-point ALU ("FP") 2622, and floating-point move unit ("FP move") 2624. In at least one embodiment, integer register file / branch network 2608 and floating-point register file / bypass network 2610 are also referred to herein as "register files 2608, 2610." In at least one embodiment, AGUs 2612 and 2614, fast ALUs 2616 and 2618, slow ALU 2620, floating-point ALU 2622, and floating-point move unit 2624 are also referred to herein as "execution units 2612, 2614, 2616, 2618, 2620, 2622, and 2624." In at least one embodiment, execution block b11 may include, but is not limited to, any number (including zero) and type of register files, branch networks, address generation units, and execution units (in any combination).

[0206] In at least one embodiment, register files 2608 and 2610 may be arranged between microinstruction schedulers 2602, 2604, and 2606 and execution units 2612, 2614, 2616, 2618, 2620, 2622, and 2624. In at least one embodiment, integer register file / branch network 2608 performs integer operations. In at least one embodiment, floating-point register file / branch network 2610 performs floating-point operations. In at least one embodiment, each of register files 2608 and 2610 may include, but is not limited to, a branch network that can bypass or forward recently completed results that have not yet been written to the register file to new dependent objects. In at least one embodiment, register files 2608 and 2610 can communicate data with each other. In at least one embodiment, integer register file / branch network 2608 may include, but is not limited to, two separate register files: one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, floating point register file / branch network 2610 may include, but is not limited to, 128-bit wide entries, as floating point instructions typically have operands that are 64 to 128 bits wide.

[0207] In at least one embodiment, execution units 2612, 2614, 2616, 2618, 2620, 2622, and 2624 can execute instructions. In at least one embodiment, register files 2608 and 2610 store integer and floating-point data operand values ​​required for microinstructions to execute. In at least one embodiment, processor 2600 can include, but is not limited to, any number of execution units 2612, 2614, 2616, 2618, 2620, 2622, and 2624, and combinations thereof. In at least one embodiment, floating-point ALU 2622 and floating-point move unit 2624 can perform floating-point, MMX, SIMD, AVX, SSE, or other operations, including specialized machine learning instructions. In at least one embodiment, floating-point ALU 2622 can include, but is not limited to, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro-operations. In at least one embodiment, floating-point hardware can be used to process instructions involving floating-point values. In at least one embodiment, ALU operations can be passed to fast ALUs 2616 and 2618. In at least one embodiment, fast ALUs 2616 and 2618 can perform fast operations with an effective latency of half a clock cycle. In at least one embodiment, most complex integer operations go to slow ALU 2620, as slow ALU 2620 may include, but is not limited to, integer execution hardware for long-latency operations, such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations can be performed by ALUs 2612 and 2614. In at least one embodiment, fast ALU 2616, fast ALU 2618, and slow ALU 2620 can perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 2616, fast ALU 2618, and slow ALU 2620 can be implemented to support various data bit sizes, including 16, 32, 128, 256, and the like. In at least one embodiment, the floating point ALU 2622 and floating point shift unit 2624 can be implemented to support a range of operands having bits of various widths. In at least one embodiment, the floating point ALU 2622 and floating point shift unit 2624 can operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.

[0208] In at least one embodiment, microinstruction schedulers 2602, 2604, and 2606 schedule dependent operations before the parent load completes execution. In at least one embodiment, because microinstructions can be speculatively scheduled and executed in processor 2600, processor 2600 may also include logic for handling memory misses. In at least one embodiment, if a data load misses in the data cache, there may be dependent operations running in the pipeline, temporarily preventing the scheduler from having the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, it may be necessary to replay dependent operations to allow independent operations to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.

[0209] In at least one embodiment, the term "register" may refer to an on-board processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, registers may be those that can be used from outside the processor (from a programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuit. Instead, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented using a variety of different techniques by circuits within the processor, such as dedicated physical registers, physical registers dynamically allocated using register renaming, a combination of dedicated and dynamically allocated physical registers, and the like. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packing data.

[0210] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Details are provided regarding the inference and / or training logic 1015. In at least one embodiment, some or all of the inference and / or training logic 1015 may be incorporated into the EXE block 2611 and other memories or registers (not shown). For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more ALUs shown in the EXE block 2611. Additionally, weight parameters may be stored in on-chip or off-chip memory and / or registers (not shown), which configure the ALUs of the EXE block 2611 to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0211] Figure 27 A deep learning application processor 2700 is shown in accordance with at least one embodiment. In at least one embodiment, the deep learning application processor 2700 uses instructions that, if executed by the deep learning application processor 2700, cause the deep learning application processor 2700 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 2700 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 2700 performs matrix multiplication operations or is "hardwired" into hardware as a result of executing one or more instructions, or both. In at least one embodiment, the deep learning application processor 2700 includes, but is not limited to, processing clusters 2710(1)-2710(12), inter-chip links (“ICLs”) 2720(1)-2720(12), inter-chip controllers (“ICCs”) 2730(1)-2730(2), memory controllers (“Mem Ctrlrs”) 2742(1)-2742(4), high bandwidth memory physical layer (“HBM PHY”) 2744(1)-2744(4), a management controller central processing unit (“management controller CPU”) 2750, serial peripheral interface, inter-integrated circuit, and general purpose input / output blocks (“SPI, I2C, GPIO”) 2760, peripheral component interconnect express controller and direct memory access block (“PCIe controller and DMA”) 2770, and sixteen-lane peripheral component interconnect express port (“PCI Express x 16”) 2780.

[0212] In at least one embodiment, the processing cluster 2710 can perform deep learning operations, including inference or prediction operations based on weight parameters calculated based on one or more training techniques, including those described herein. In at least one embodiment, each processing cluster 2710 can include, but is not limited to, any number and type of processors. In at least one embodiment, the deep learning application processor 2700 can include any number and type of processing clusters 2700. In at least one embodiment, the inter-chip link 2720 is bidirectional. In at least one embodiment, the inter-chip link 2720 and the inter-chip controller 2730 enable multiple deep learning application processors 2700 to exchange information, including activation information generated from executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2700 can include any number (including zero) and type of ICL 2720 and ICC 2730.

[0213] In at least one embodiment, HBM2 2740 provides a total of 32GB of memory. HBM2 2740(i) is associated with both a memory controller 2742(i) and an HBM PHY 2744(i). In at least one embodiment, any number of HBM2 2740 can provide any type and total amount of high-bandwidth memory and can be associated with any number (including zero) and type of memory controllers 2742 and HBM PHYs 2744. In at least one embodiment, SPI, I2C, GPIO 2760, PCIe controller and DMA 2770, and / or PCIe 2780 can be replaced with any number and type of blocks to implement any number and type of communication standards in any technically feasible manner.

[0214] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Details are provided regarding the inference and / or training logic 1015. In at least one embodiment, the deep learning application processor is used to train a machine learning model (e.g., a neural network) to predict or infer information provided to the deep learning application processor 2700. In at least one embodiment, the deep learning application processor 2700 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 2700. In at least one embodiment, the processor 2700 can be used to perform one or more of the neural network use cases described herein.

[0215] Figure 28is a block diagram of a neuromorphic processor 2800 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2800 can receive one or more inputs from a source external to the neuromorphic processor 2800. In at least one embodiment, these inputs can be transmitted to one or more neurons 2802 within the neuromorphic processor 2800. In at least one embodiment, the neurons 2802 and their components can be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2800 can include, but is not limited to, thousands of instances of neurons 2802, although any suitable number of neurons 2802 can be used. In at least one embodiment, each instance of a neuron 2802 can include a neuron input 2804 and a neuron output 2806. In at least one embodiment, a neuron 2802 can generate an output that can be transmitted to the inputs of other instances of the neuron 2802. In at least one embodiment, the neuron input 2804 and the neuron output 2806 can be interconnected via a synapse 2808.

[0216] In at least one embodiment, neurons 2802 and synapses 2808 can be interconnected so that the neuromorphic processor 2800 operates to process or analyze information received by the neuromorphic processor 2800. In at least one embodiment, when an input received via a neuron input 2804 exceeds a threshold, the neuron 2802 can send an output pulse (or "trigger" or "spike"). In at least one embodiment, the neuron 2802 can sum or integrate the signal received at the neuron input 2804. For example, in at least one embodiment, the neuron 2802 can be implemented as a leaky integrate-and-trigger neuron, where if the sum (referred to as the "membrane potential") exceeds a threshold, the neuron 2802 can generate an output (or "trigger") using a transfer function such as a sigmoid or threshold function. In at least one embodiment, the leaky integrate-and-trigger neuron can sum the signal received at the neuron input 2804 into a membrane potential and can apply an attenuation factor (or leakage) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-trigger neuron may trigger if multiple input signals are received at neuron input 2804 quickly enough to exceed a threshold, for example, before the membrane potential decays too low to trigger. In at least one embodiment, neuron 2802 may be implemented using circuitry or logic that receives input, integrates the input into a membrane potential, and decays the membrane potential. In at least one embodiment, the inputs may be averaged, or any other suitable transfer function may be used. Furthermore, in at least one embodiment, neuron 2802 may include, but is not limited to, comparator circuitry or logic that generates an output spike at neuron output 2806 when the result of applying the transfer function to neuron input 2804 exceeds a threshold. In at least one embodiment, once neuron 2802 triggers, it may ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, neuron 2802 may resume normal operation after a suitable period of time (or recovery period).

[0217] In at least one embodiment, neurons 2802 can be interconnected via synapses 2808. In at least one embodiment, synapses 2808 can be operable to transmit a signal from the output of a first neuron 2802 to the input of a second neuron 2802. In at least one embodiment, a neuron 2802 can transmit information across more than one instance of synapse 2808. In at least one embodiment, one or more instances of a neuron output 2806 can be connected to an instance of a neuron input 2804 in the same neuron 2802 via an instance of synapse 2808. In at least one embodiment, an instance of a neuron 2802 that generates an output to be transmitted across an instance of synapse 2808 can be referred to as a "presynaptic neuron" relative to that instance of synapse 2808. In at least one embodiment, an instance of a neuron 2802 that receives an input transmitted across an instance of synapse 2808 can be referred to as a "postsynaptic neuron" relative to an instance of synapse 2808. In at least one embodiment, with respect to various instances of synapses 2808, because an instance of neuron 2802 can receive input from one or more instances of synapses 2808 and can also transmit output through one or more instances of synapses 2808, a single instance of neuron 2802 can be both a "pre-synaptic neuron" and a "post-synaptic neuron."

[0218] In at least one embodiment, neurons 2802 may be organized into one or more layers. Each instance of a neuron 2802 may have a neuron output 2806 that may fan out to one or more neuron inputs 2804 via one or more synapses 2808. In at least one embodiment, the neuron output 2806 of a neuron 2802 in a first layer 2810 may be connected to the neuron input 2804 of a neuron 2802 in a second layer 2812. In at least one embodiment, layers 2810 may be referred to as "feed-forward layers." In at least one embodiment, each instance of a neuron 2802 in an instance of the first layer 2810 may fan out to each instance of a neuron 2802 in the second layer 2812. In at least one embodiment, the first layer 2810 may be referred to as a "fully connected feed-forward layer." In at least one embodiment, each instance of a neuron 2802 in each instance of the second layer 2812 may fan out to fewer than all instances of a neuron 2802 in the third layer 2814. In at least one embodiment, the second layer 2812 may be referred to as a "sparsely connected feed-forward layer." In at least one embodiment, neurons 2802 in the second layer 2812 may fan out to neurons 2802 in multiple other layers, including neurons 2802 in the (same) second layer 2812. In at least one embodiment, the second layer 2812 may be referred to as a "recurrent layer." In at least one embodiment, the neuromorphic processor 2800 may include, but is not limited to, any suitable combination of recurrent layers and feed-forward layers, including, but not limited to, sparsely connected feed-forward layers and fully connected feed-forward layers.

[0219] In at least one embodiment, the neuromorphic processor 2800 may include, but is not limited to, a reconfigurable interconnect fabric or a dedicated hardwired interconnect to connect synapses 2808 to neurons 2802. In at least one embodiment, the neuromorphic processor 2800 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 2802 as needed based on the neural network topology and neuron fan-in / fan-out. For example, in at least one embodiment, synapses 2808 may be connected to neurons 2802 using an interconnect fabric such as a network on a chip or through dedicated connections. In at least one embodiment, the synaptic interconnect and its components may be implemented using circuitry or logic.

[0220] Figure 29is a block diagram of a graphics processor 2900, which may be a discrete graphics processing unit or a graphics processor integrated with multiple processing cores. In at least one embodiment, the graphics processor 2900 communicates with registers on the graphics processor 2900 via a memory-mapped I / O interface using commands stored in memory. In at least one embodiment, the graphics processor 2900 includes a memory interface 2914 for accessing memory. In at least one embodiment, the memory interface 2914 is an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.

[0221] In at least one embodiment, the graphics processor 2900 also includes a display controller 2902 for driving display output data to a display device 2920. In at least one embodiment, the display controller 2902 includes hardware for one or more overlay planes and the composition of multiple layers of video or user interface elements for the display device 2920. In at least one embodiment, the display device 2920 can be an internal or external display device. In at least one embodiment, the display device 2920 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In at least one embodiment, the graphics processor 2400 includes a video codec engine 2406 to encode, decode, or transcode media to, from, or between one or more media coding formats, including but not limited to Moving Picture Experts Group (MPEG) formats such as MPEG-2, Advanced Video Coding (AVC) formats such as H.264 / MPEG-4 AVC, and Society of Motion Picture and Television Engineers (SMPTE) 421M / VC-1, and Joint Photographic Experts Group (JPEG) formats (such as JPEG and Motion JPEG (MJPEG) formats).

[0222] In at least one embodiment, the graphics processor 2900 includes a block image transfer (BLIT) engine 2904 to perform two-dimensional (2D) rasterizer operations, including, for example, bit-boundary block transfers. However, in at least one embodiment, 2D graphics operations are performed using one or more components of a graphics processing engine (GPE) 2910. In at least one embodiment, the GPE 2910 is a compute engine for performing graphics operations, including three-dimensional (3D), and media operations.

[0223] In at least one embodiment, GPE 2910 includes a 3D pipeline 2912 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions that operate on 3D primitive shapes (e.g., rectangles, triangles, etc.). 3D pipeline 2912 includes programmable and fixed-function elements that perform various tasks and / or spawn execution threads to 3D / media subsystem 2915. While 3D pipeline 2912 can be used to perform media operations, in at least one embodiment, GPE 2910 also includes a media pipeline 2916 for performing media operations, such as video post-processing and image enhancement.

[0224] In at least one embodiment, the media pipeline 2916 includes fixed-function or programmable logic units to perform one or more specialized media operations, such as video decode acceleration, video deinterlacing, and video encoding acceleration, instead of or on behalf of the video codec engine 2906. In at least one embodiment, the media pipeline 2916 also includes a thread generation unit to generate threads for execution on the 3D / media subsystem 2915. In at least one embodiment, the generated threads perform computations for media operations on one or more graphics execution units included in the 3D / media subsystem 2915.

[0225] In at least one embodiment, the 3D / media subsystem 2915 includes logic for executing threads generated by the 3D pipeline 2912 and the media pipeline 2916. In at least one embodiment, the 3D pipeline 2912 and the media pipeline 2916 send thread execution requests to the 3D / media subsystem 2915, which includes thread dispatch logic for arbitrating and dispatching the various requests to available thread execution resources. In at least one embodiment, the execution resources include an array of graphics execution units for processing 3D and media threads. In at least one embodiment, the 3D / media subsystem 2915 includes one or more internal caches for thread instructions and data. In at least one embodiment, the subsystem 2915 also includes shared memory, including registers and addressable memory, to share data between threads and store output data.

[0226] Reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11Details regarding the inference and / or training logic 1015 are provided. In at least one embodiment, some or all of the inference and / or training logic 1015 may be incorporated into the graphics processor 2900. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more of the ALUs embodied below. Additionally, in at least one embodiment, the inference and / or training operations described herein may utilize processors other than the ALUs described herein. Figure 10 or Figure 11 In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 2900 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0227] Figure 30 is a block diagram of the hardware logic of a graphics processor core 3000 according to at least one embodiment described herein. In at least one embodiment, the graphics processor core 3000 is included within a graphics core array. In at least one embodiment, the graphics processor core 3000 (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 3000 is an example of a graphics core slice, and the graphics processors described herein can include multiple graphics core slices based on target power and performance envelopes. In at least one embodiment, each graphics core 3000 can include fixed function blocks 3030 coupled to multiple sub-cores 3001A-3001F, also referred to as sub-slices, which include modular blocks of general-purpose and fixed-function logic.

[0228] In at least one embodiment, fixed function block 3030 includes a geometry / fixed function pipeline 3036, which may be shared by all sub-cores in graphics processor 3000, for example, in lower performance and / or lower power graphics processor implementations. In at least one embodiment, geometry / fixed function pipeline 3036 includes a 3D fixed function pipeline, a video front end unit, a thread spawner and thread dispatcher, and a unified return buffer manager that manages a unified return buffer.

[0229] In at least one fixed embodiment, functional block 3030 also includes a graphics SoC interface 3037, a graphics microcontroller 3038, and a media pipeline 3039. In at least one fixed embodiment, graphics SoC interface 3037 provides an interface between graphics core 3000 and other processor cores in the on-chip integrated circuit system. In at least one embodiment, graphics microcontroller 3038 is a programmable subprocessor that can be configured to manage various functions of graphics processor 3000, including thread dispatching, scheduling, and preemption. In at least one embodiment, media pipeline 3039 includes logic that facilitates decoding, encoding, pre-processing, and / or post-processing of multimedia data, including image and video data. In at least one embodiment, media pipeline 3039 implements media operations via requests to computational or sampling logic within sub-cores 3001-3001F.

[0230] In at least one embodiment, the SoC interface 3037 enables the graphics core 3000 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared last-level cache, system RAM, and / or embedded on-chip or packaged DRAM. In at least one embodiment, the SoC interface 3037 may also enable communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline) and enable the use and / or implementation of global memory atomics that can be shared between the graphics core 3000 and the CPU within the SoC. In at least one embodiment, the SoC interface 3037 may also implement power management controls for the graphics core 3000 and enable interfaces between the clock domain of the graphics core 3000 and other clock domains within the SoC. In at least one embodiment, the SoC interface 3037 enables the reception of command buffers from a command stream converter and a global thread dispatcher, which are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, commands and instructions may be dispatched to the media pipeline 3039 when media operations are to be performed, or to the geometry and fixed function pipelines (e.g., geometry and fixed function pipeline 3036, geometry and fixed function pipeline 3014) when graphics processing operations are to be performed.

[0231] In at least one embodiment, the graphics microcontroller 3038 can be configured to perform various scheduling and management tasks for the graphics core 3000. In at least one embodiment, the graphics microcontroller 3038 can perform graphics and / or compute workload scheduling on the various graphics parallel engines within the execution unit (EU) arrays 3002A-3002F, 3004A-3004F in the sub-cores 3001A-3001F. In at least one embodiment, host software executing on a CPU core of a SoC including the graphics core 3000 can submit a workload to one of multiple graphics processor doorbells, which invokes scheduling operations on the appropriate graphics engine. In at least one embodiment, scheduling operations include determining which workload to run next, submitting the workload to the command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, graphics microcontroller 3038 may also facilitate low power or idle states for graphics core 3000, thereby providing graphics core 3000 with the ability to save and restore registers across low power state transitions within graphics core 3000 independent of the operating system and / or graphics driver software on the system.

[0232] In at least one embodiment, graphics core 3000 may have up to N modular sub-cores, more or less than the sub-cores 3001A-3001F shown. For each set of N sub-cores, in at least one embodiment, graphics core 3000 may also include shared function logic 3010, shared and / or cache memory 3012, geometry / fixed function pipelines 3014, and additional fixed function logic 3016 to accelerate various graphics and compute processing operations. In at least one embodiment, shared function logic 3010 may include logic units (e.g., samplers, math, and / or inter-thread communication logic) that may be shared by each of the N sub-cores within graphics core 3000. In at least one embodiment, fixed, shared, and / or cache memory 3012 may be the last level cache for the N sub-cores 3001A-3001F within graphics core 3000 and may also serve as shared memory accessible by multiple sub-cores. In at least one embodiment, geometry / fixed function pipeline 3014 may be included in place of geometry / fixed function pipeline 3036 within fixed function block 3030 and may include the same or similar logic units.

[0233] In at least one embodiment, graphics core 3000 includes additional fixed-function logic 3016, which may include various fixed-function acceleration logic for use by graphics core 3000. In at least one embodiment, additional fixed-function logic 3016 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are at least two geometry pipelines, including a full geometry pipeline and a culling pipeline within geometry / fixed-function pipelines 3016, 3036, which may be included in additional fixed-function logic 3016. In at least one embodiment, the culling pipeline is a modified version of the full geometry pipeline. In at least one embodiment, the full pipeline and the culling pipeline can execute different instances of an application, each with a separate context. In at least one embodiment, position-only shading can hide long culling runs for discarded triangles, allowing shading to complete earlier in some cases. For example, in at least one embodiment, the culling pipeline logic in the additional fixed function logic 3016 can execute position shaders in parallel with the main application and generate critical results faster than the full pipeline because the culling pipeline obtains and masks the position attributes of the vertices without having to perform rasterization and render the pixels to the frame buffer. In at least one embodiment, the culling pipeline can use the generated critical results to calculate visibility information for all triangles, regardless of whether they are culled. In at least one embodiment, the full pipeline (which in this case may be called a replay pipeline) can consume visibility information to skip culled triangles to mask only visible triangles that are ultimately passed to the rasterization stage.

[0234] In at least one embodiment, the additional fixed-function logic 3016 may also include machine learning acceleration logic, such as fixed-function matrix multiplication logic, to implement optimizations including for machine learning training or inference.

[0235] In at least one embodiment, each graphics sub-core 3001A-3001F includes a set of execution resources that can be used to perform graphics, media, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader programs. In at least one embodiment, the graphics sub-core 3001A-3001F includes multiple EU arrays 3002A-3002F, 3004A-3004F, thread dispatch and inter-thread communication (TD / IC) logic 3003A-3003F, 3D (e.g., texture) samplers 3005A-3005F, media samplers 3006A-3006F, shader processors 3007A-3007F, and shared local memory (SLM) 3008A-3008F. Each EU array 3002A-3002F, 3004A-3004F includes multiple execution units, which are general-purpose graphics processing units capable of servicing graphics, media, or compute operations, executing floating-point and integer / fixed-point logic operations, including graphics, media, or compute shader programs. In at least one embodiment, TD / IC logic 3003A-3003F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between threads executing on the execution units of the sub-core. In at least one embodiment, 3D samplers 3005A-3005F can read texture or other 3D graphics-related data into memory. In at least one embodiment, the 3D samplers can read texture data differently based on the configured sampling state and texture format associated with a given texture. In at least one embodiment, media samplers 3006A-3006F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics sub-core 3001A-3001F may alternatively include a unified 3D and media sampler. In at least one embodiment, threads executing on execution units within each sub-core 3001A-3001F may utilize shared local memory 3008A-3008F within each sub-core, enabling threads executing within a thread group to execute using a pool of on-chip memory.

[0236] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Detail is provided regarding the inference and / or training logic 1015. In at least one embodiment, some or all of the inference and / or training logic 1015 may be incorporated into the graphics processor 3010. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize the 3D pipeline 3010, the graphics microcontroller 3038, the geometry and fixed function pipelines 3014 and 3036, or the like. Figure 29 Furthermore, in at least one embodiment, the inference and / or training operations described herein may use the addition Figure 10 or Figure 11 In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 3000 to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0237] Figure 31A-Figure 31B Thread execution logic 3100 is shown for an array of processing elements comprising a graphics processor core, in accordance with at least one embodiment. Figure 31A At least one embodiment is shown in which thread execution logic 3100 is used. Figure 31B Illustrative internal details of an execution unit are shown in accordance with at least one embodiment.

[0238] like Figure 31A As shown in FIG, in at least one embodiment, thread execution logic 3100 includes a shader processor 3102, a thread dispatcher 3104, an instruction cache 3106, a scalable execution unit array including a plurality of execution units 3108A-3108N, a sampler 3110, a data cache 3112, and a data port 3114. In at least one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., execution units 3108A, 3108B, 3108C, 3108D, or any one of 3108N-1 to 3108N), for example, based on the computational requirements of the workload. In at least one embodiment, the scalable execution units are interconnected via an interconnect structure that links to each execution unit. In at least one embodiment, thread execution logic 3100 includes one or more connections to a memory (such as system memory or cache memory) through instruction cache 3106, data port 3114, sampler 3110, and one or more of execution units 3108A-3108N. In at least one embodiment, each execution unit (e.g., 3108A) is an independent programmable general-purpose computing unit capable of executing multiple simultaneous hardware threads while processing multiple data elements in parallel for each thread. In at least one embodiment, the array of execution units 3108A-3108N is scalable to include any number of individual execution units.

[0239] In at least one embodiment, execution units 3108A-3108N are primarily used to execute shader programs. In at least one embodiment, shader processor 3102 can process various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 3104. In at least one embodiment, thread dispatcher 3104 includes logic for arbitrating thread initialization requests from graphics and media pipelines and instantiating requested threads on one or more execution units in execution units 3108A-3108N. For example, in at least one embodiment, a geometry pipeline can dispatch vertex, tessellation, or geometry shaders to thread execution logic for processing. In at least one embodiment, thread dispatcher 3104 can also handle runtime thread generation requests from executing shader programs.

[0240] In at least one embodiment, execution units 3108A-3108N support an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs from graphics libraries (e.g., Direct3D and OpenGL) to execute with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., compute and media shaders). In at least one embodiment, each execution unit 3108A-3108N includes one or more arithmetic logic units (ALUs) capable of multi-issue single instruction, multiple data (SIMD), and multi-threaded operation enables an efficient execution environment despite higher latency memory accesses. In at least one embodiment, each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. In at least one embodiment, execution is multiple issues per clock to the pipeline, which is capable of performing integer, single-precision and double-precision floating-point operations, SIMD branching functions, logical operations, transcendental operations, and other operations. In at least one embodiment, while waiting for data from memory or one of the shared functions, dependency logic within execution units 3108A-3108N causes the waiting thread to sleep until the requested data is returned. In at least one embodiment, while the waiting thread is sleeping, hardware resources can be dedicated to processing other threads. For example, in at least one embodiment, during the delay associated with vertex shader operations, the execution unit can perform operations on a pixel shader, a fragment shader, or another type of shader program (including a different vertex shader).

[0241] In at least one embodiment, each of execution units 3108A-3108N operates on an array of data elements. In at least one embodiment, the number of data elements is the "execution size" or number of lanes of an instruction. In at least one embodiment, an execution lane is the logic used to perform data element access, masking, and flow control within an instruction. In at least one embodiment, the number of lanes can be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) used for a particular graphics processor. In at least one embodiment, execution units 3108A-3108N support integer and floating point data types.

[0242] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements can be stored in registers as packed data types, and the execution unit will process various elements based on the data size of the element. For example, in at least one embodiment, when a 256-bit wide vector is operated, 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four separate 64-bit packed data elements (quad word (QW) size data elements), eight separate 32-bit packed data elements (double word (DW) size data elements), sixteen separate 16-bit packed data elements (word (W) size data elements) or thirty-two separate 8-bit data elements (byte (B) size data elements). However, in at least one embodiment, different vector widths and register sizes are possible.

[0243] In at least one embodiment, one or more execution units can be combined into a fused execution unit 3109A-3109N having thread control logic (3107A-3107N) that executes for the fused EU. In at least one embodiment, multiple EUs can be merged into a EU group. In at least one embodiment, each EU in the fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group can vary according to various embodiments. In at least one embodiment, each EU can execute various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. In at least one embodiment, each fused graphics execution unit 3109A-3109N includes at least two execution units. For example, in at least one embodiment, the fused execution unit 3109A includes a first EU 3108A, a second EU 3108B, and thread control logic 3107A shared by the first EU 3108A and the second EU 3108B. In at least one embodiment, thread control logic 3107A controls threads executing on fused graphics execution unit 3109A, allowing each EU within fused execution units 3109A-3109N to execute using an instruction pointer register.

[0244] In at least one embodiment, one or more internal instruction caches (e.g., 3106) are included in thread execution logic 3100 to cache thread instructions for the execution units. In at least one embodiment, one or more data caches (e.g., 3112) are included to cache thread data during thread execution. In at least one embodiment, a sampler 3110 is included to provide texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, sampler 3110 includes specialized texture or media sampling functionality to process texture or media data during the sampling process before providing the sampled data to the execution units.

[0245] During execution, in at least one embodiment, the graphics and media pipeline sends thread initiation requests to thread execution logic 3100 via thread spawning and dispatching logic. In at least one embodiment, once a set of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 3102 is invoked to further calculate output information and cause the results to be written to output surfaces (e.g., color buffer, depth buffer, stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader calculates the values ​​of various vertex attributes to be interpolated across the rasterized objects. In at least one embodiment, the pixel processor logic within shader processor 3102 then executes the pixel or fragment shader program provided by an application programming interface (API). In at least one embodiment, to execute the shader program, shader processor 3102 dispatches threads to execution units (e.g., 3108A) via thread dispatcher 3104. In at least one embodiment, shader processor 3102 uses texture sampling logic in sampler 3110 to access texture data stored in texture maps in memory. In at least one embodiment, arithmetic operations on texture data and input geometry data calculate pixel color data for each geometry fragment, or discard one or more pixels for further processing.

[0246] In at least one embodiment, the data port 3114 provides a memory access mechanism for the thread execution logic 3100 to output processed data to memory for further processing on the graphics processor output pipeline. In at least one embodiment, the data port 3114 includes or is coupled to one or more cache memories (e.g., data cache 3112) to cache data for memory access via the data port.

[0247] like Figure 31BAs shown, in at least one embodiment, graphics execution unit 3108 may include an instruction fetch unit 3137, a general register file array (GRF) 3124, an architectural register file array (ARF) 3126, a thread arbiter 3122, an issue unit 3130, a branch unit 3132, a set of SIMD floating point units (FPUs) 3134, and, in at least one embodiment, a set of dedicated integer SIMD ALUs 3135. In at least one embodiment, GRF 3124 and ARF 3126 include a set of general register files and architectural register files associated with each simultaneous hardware thread that can be active in graphics execution unit 3108. In at least one embodiment, per-thread architectural state is maintained in ARF 3126, while data used during thread execution is stored in GRF 3124. In at least one embodiment, the execution state of each thread, including the instruction pointer for each thread, may be maintained in thread-specific registers in ARF 3126.

[0248] In at least one embodiment, graphics execution unit 3108 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, where execution unit resources are logically allocated for executing multiple simultaneous threads.

[0249] In at least one embodiment, the graphics execution unit 3108 can collectively issue multiple instructions, each of which can be a different instruction. In at least one embodiment, the thread arbiter 3122 of a graphics execution unit thread 3108 can dispatch instructions to one of the issue unit 3130, branch unit 3142, or SIMD FPU3 134 for execution. In at least one embodiment, each execution thread can access 128 general purpose registers in the GRF 3124, each of which can store 32 bytes and can be accessed as a SIMD 8-element vector of 32-bit data elements. In at least one embodiment, each execution unit thread can access 4KB of GRF 3124, although embodiments are not limited thereto and more or fewer register resources may be provided in other embodiments. In at least one embodiment, a maximum of seven threads can execute simultaneously, although the number of threads per execution unit may vary depending on the embodiment. In at least one embodiment where seven threads have access to 4KB, the GRF 3124 can store a total of 28KB. In at least one embodiment, flexible addressing modes may allow registers to be addressed together to efficiently build wider registers or rectangular block data structures representing strides.

[0250] In at least one embodiment, memory operations, sampler operations, and other longer latency system communications are scheduled via "send" instructions executed by the message passing send unit 3130. In at least one embodiment, dispatching branch instructions to the dedicated branch unit 3132 facilitates SIMD divergence and eventual convergence.

[0251] In at least one embodiment, the graphics execution unit 3108 includes one or more SIMD floating-point units (FPUs) 3134 to perform floating-point operations. In at least one embodiment, the FPUs 3134 also support integer computations. In at least one embodiment, the FPUs 3134 can SIMD up to M 32-bit floating-point (or integer) operations, or SIMD up to 2M 16-bit integer or 16-bit floating-point operations. In at least one embodiment, at least one of the FPUs provides extended math capabilities to support high-throughput transcendental math functions and double-precision 64-bit floating point. In at least one embodiment, a set of 8-bit integer SIMD ALUs 3135 are also present and can be specifically optimized to perform operations related to machine learning computations.

[0252] In at least one embodiment, an array of multiple instances of graphics execution unit 3108 may be instantiated in graphics sub-core groupings (e.g., sub-slices). In at least one embodiment, execution unit 3108 may execute instructions across multiple execution lanes. In at least one embodiment, each thread executing on graphics execution unit 3108 executes on a different lane.

[0253] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Provides details about the reasoning and / or training logic 815. In at least one embodiment, some or all of the reasoning and / or training logic 1015 may be incorporated into the execution logic 3100. Additionally, in at least one embodiment, other than Figure 10 or Figure 11 In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the execution logic 3100 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0254] Figure 32A parallel processing unit ("PPU") 3200 is shown in accordance with at least one embodiment. In at least one embodiment, PPU 3200 is configured with machine-readable code that, if executed by PPU 3200, causes PPU 3200 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, PPU 3200 is a multi-threaded processor implemented on one or more integrated circuit devices and utilizes multithreading as a latency hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) executed in parallel on multiple threads. In at least one embodiment, a thread refers to an execution thread and is an instance of a group of instructions configured to be executed by PPU 3200. In at least one embodiment, PPU 3200 is a graphics processing unit ("GPU") configured to implement a graphics rendering pipeline for processing three-dimensional ("3D") graphics data to generate two-dimensional ("2D") image data for display on a display device, such as a liquid crystal display ("LCD") device. In at least one embodiment, PPU 3200 is used to perform computations such as linear algebra operations and machine learning operations. Figure 32 The example parallel processor is shown for illustrative purposes only and should be construed as a non-limiting example of a processor architecture contemplated within the scope of the present disclosure, and any suitable processor may be employed in addition to and / or in place of it.

[0255] In at least one embodiment, one or more PPUs 3200 are configured to accelerate high-performance computing ("HPC"), data center, and machine learning applications. In at least one embodiment, the PPU 3200 is configured to accelerate deep learning systems and applications, including the following non-limiting examples: autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analysis, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations.

[0256] In at least one embodiment, the PPU 3200 includes, but is not limited to, an input / output ("I / O") unit 3206, a front-end unit 3210, a scheduler unit 3212, a work distribution unit 3214, a hub 3216, a crossbar switch ("Xbar") 3220, one or more general processing clusters ("GPCs") 3218, and one or more partitioning units ("memory partitioning units") 3222. In at least one embodiment, the PPU 3200 is connected to a host processor or other PPUs 3200 via one or more high-speed GPU interconnects ("GPU interconnects") 3208. In at least one embodiment, the PPU 2800 is connected to a host processor or other peripheral devices via an interconnect 3202. In one embodiment, the PPU 3200 is connected to local memory including one or more memory devices ("memory") 3204. In at least one embodiment, the memory devices 3204 include, but are not limited to, one or more dynamic random access memory ("DRAM") devices. In at least one embodiment, one or more DRAM devices are configured and / or configurable as a high bandwidth memory ("HBM") subsystem with multiple DRAM dies stacked within each device.

[0257] In at least one embodiment, the high-speed GPU interconnect 3208 may refer to a wire-based, multi-lane communication link that a system uses to scale and includes one or more PPUs 3200 in conjunction with one or more central processing units ("CPUs"), supporting cache coherence between the PPUs 3200 and the CPUs and CPU mastering. In at least one embodiment, the high-speed GPU interconnect 3208 transmits data and / or commands to other units of the PPU 3200, such as one or more copy engines, video encoders, video decoders, power management units, and / or other processors, via the hub 3216. Figure 32 Other components that may not be explicitly shown.

[0258] In at least one embodiment, the I / O unit 3206 is configured to receive data from the host processor ( Figure 323206 sends and receives communications (e.g., commands, data). In at least one embodiment, the I / O unit 3206 communicates with the host processor directly through the system bus 3202 or through one or more intermediate devices (e.g., a memory bridge). In at least one embodiment, the I / O unit 3206 can communicate with one or more other processors (e.g., one or more PPUs 3200) via the system bus 3202. In at least one embodiment, the I / O unit 3206 implements a Peripheral Component Interconnect Express ("PCIe") interface for communicating over the PCIe bus. In at least one embodiment, the I / O unit 3206 implements an interface for communicating with external devices.

[0259] In at least one embodiment, the I / O unit 3206 decodes packets received via the system bus 3202. In at least one embodiment, at least some of the packets represent commands configured to cause the PPU 3200 to perform various operations. In at least one embodiment, the I / O unit 3206 sends the decoded commands to various other units of the PPU 3200 as specified by the commands. In at least one embodiment, the commands are sent to the front end unit 3210 and / or to the hub 3216 or other units of the PPU 3200, such as one or more copy engines, video encoders, video decoders, power management units, etc. Figure 32 In at least one embodiment, I / O unit 3206 is configured to route communications between the various logical units of PPU 3200.

[0260] In at least one embodiment, a program executed by a host processor encodes a command stream in a buffer that provides a workload to the PPU 3200 for processing. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in memory that is accessible (e.g., read / write) by both the host processor and the PPU 3200—the host interface unit can be configured to access the buffer in system memory connected to the system bus 3202 via memory requests transmitted via the I / O unit 3206 over the system bus 3202. In at least one embodiment, the host processor writes a command stream into the buffer and then sends a pointer indicating the beginning of the command stream to the PPU 3200, causing the front end unit 3210 to receive pointers to one or more command streams and manage the one or more command streams, reading commands from the command streams and forwarding the commands to the various units of the PPU 3200.

[0261] In at least one embodiment, the front end unit 3210 is coupled to a scheduler unit 3212, which configures the various GPCs 3218 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 3212 is configured to track status information related to the various tasks managed by the scheduler unit 3212, where the status information may indicate which GPC 3218 the task is assigned to, whether the task is active or inactive, the priority associated with the task, and the like. In at least one embodiment, the scheduler unit 3212 manages multiple tasks that execute on one or more GPCs 3218.

[0262] In at least one embodiment, the scheduler unit 3212 is coupled to a work distribution unit 3214, which is configured to dispatch tasks for execution on GPCs 3218. In at least one embodiment, the work distribution unit 3214 tracks a plurality of scheduled tasks received from the scheduler unit 3212 and manages a pending task pool and an active task pool for each GPC 3218. In at least one embodiment, the pending task pool includes a plurality of time slots (e.g., 32 time slots) containing tasks assigned to be processed by a particular GPC 3218; the active task pool may include a plurality of time slots (e.g., 4 time slots) for tasks actively being processed by a GPC 3218, such that as a task in a GPC 3218 completes execution, the task is evicted from the active task pool of GPC 3218, and one of the other tasks is selected from the pending task pool and scheduled for execution on GPC 3218. In at least one embodiment, if an active task is idle on a GPC 3218, such as while waiting for data dependencies to be resolved, the active task is evicted from the GPC 3218 and returned to the pending task pool, while another task in the pending task pool is selected and scheduled for execution on the GPC 3218.

[0263] In at least one embodiment, work distribution unit 3214 communicates with one or more GPCs 3218 via XBar 3220. In at least one embodiment, XBar 3220 is an interconnect network that couples many units of PPU 3200 to other units of PPU 3200 and can be configured to couple work distribution unit 3214 to a specific GPC 3218. In at least one embodiment, one or more other units of PPU 3200 can also be connected to XBar 3220 through hub 3216.

[0264] In at least one embodiment, tasks are managed by a scheduler unit 3212 and assigned to one of the GPCs 3218 by a work distribution unit 3214. The GPC 3218 is configured to process tasks and produce results. In at least one embodiment, the results can be consumed by other tasks in the GPC 3218, routed to a different GPC 3218 via an XBar 3220, or stored in memory 3204. In at least one embodiment, the results can be written to memory 3204 via a partition unit 3222, which implements a memory interface for writing data to or reading data from memory 3204. In at least one embodiment, the results can be transferred to another PPU 3204 or a CPU via a high-speed GPU interconnect 3208. In at least one embodiment, the PPU 3200 includes, but is not limited to, U partition units 3222, which is equal to the number of separate and distinct memory devices 3204 coupled to the PPU 3200. In at least one embodiment, the following in conjunction with Figure 34 The partition unit 3222 is described in more detail.

[0265] In at least one embodiment, the host processor executes a driver core that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 3200. In one embodiment, multiple computing applications are executed simultaneously by the PPU 3200, and the PPU 3200 provides isolation, quality of service ("QoS"), and independent address spaces for the multiple computing applications. In at least one embodiment, the application generates instructions (e.g., in the form of API calls) that cause the driver core to generate one or more tasks for execution by the PPU 3200, and the driver core outputs the tasks to one or more streams processed by the PPU 3200. In at least one embodiment, each task includes one or more related groups of threads, which may be referred to as warps. In at least one embodiment, a warp includes multiple related threads (e.g., 32 threads) that can be executed in parallel. In at least one embodiment, a cooperative thread may refer to multiple threads that include instructions for performing tasks and exchanging data through shared memory. In at least one embodiment, in combination Figure 34 Threads and cooperating threads are described in greater detail according to at least one embodiment.

[0266] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11Details are provided regarding the inference and / or training logic 1015. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the PPU 3200. In at least one embodiment, the PPU 3200 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or the PPU 3200. In at least one embodiment, the PPU 3200 can be used to perform one or more of the neural network use cases described herein.

[0267] Figure 33 A general processing cluster ("GPC") 3300 is shown in accordance with at least one embodiment. In at least one embodiment, the GPC 3300 is Figure 32 3300 , and each GPC 3300 includes, but is not limited to, a plurality of hardware units for processing tasks, and each GPC 3300 includes, but is not limited to, a pipeline manager 3302 , a pre-raster operations unit (“PROP”) 3304 , a raster engine 3308 , a work distribution crossbar (“WDX”) 3316 , a memory management unit (“MMU”) 3318 , one or more data processing clusters (“DPCs”) 3306 , and any suitable combination of components.

[0268] In at least one embodiment, the operation of GPC 3300 is controlled by pipeline manager 3302. In at least one embodiment, pipeline manager 3302 manages the configuration of one or more DPCs 3306 to process tasks assigned to GPC 3300. In at least one embodiment, pipeline manager 3302 configures at least one of the one or more DPCs 3306 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, DPC 3306 is configured to execute vertex shader programs on a programmable streaming multiprocessor ("SM") 3314. In at least one embodiment, pipeline manager 3302 is configured to route packets received from a work distribution unit to appropriate logic within GPC 3300, and in at least one embodiment, some packets may be routed to fixed-function hardware units in PROP 3304 and / or raster engine 3308, while other packets may be routed to DPC 3306 for processing by primitive engine 3312 or SM 3314. In at least one embodiment, pipeline manager 3302 configures at least one of DPCs 3306 to implement a neural network model and / or a computational pipeline.

[0269] In at least one embodiment, PROP unit 3304 is configured to route data generated by raster engine 3308 and DPC 3306 to the above combined Figure 32 The ROP unit in the partition unit 3222 is described in greater detail. In at least one embodiment, the PROP unit 3304 is configured to perform optimizations for color blending, organize pixel data, perform address translation, and the like. In at least one embodiment, the raster engine 3308 includes, but is not limited to, a plurality of fixed-function hardware units configured to perform various raster operations, and in at least one embodiment, the raster engine 3308 includes, but is not limited to, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile aggregation engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices; the plane equations are passed to the coarse raster engine to generate coverage information for the primitives (e.g., an x, y coverage mask for the tile); the output of the coarse raster engine is passed to the culling engine, where fragments associated with primitives that fail the z test are culled, and to the clipping engine, where fragments outside the viewing frustum are clipped. In at least one embodiment, the clipped and culled fragments are passed to a fine raster engine to generate properties for the pixel fragments based on a plane equation generated by the setup engine. In at least one embodiment, the output of the raster engine 3308 includes fragments to be processed by any appropriate entity (e.g., by a fragment shader implemented within DPC 3306).

[0270] In at least one embodiment, each DPC 3306 included in a GPC 3300 includes, but is not limited to, an M-pipeline controller ("MPC") 3310; a primitive engine 3312; one or more SMs 3314; and any suitable combination thereof. In at least one embodiment, the MPC 3310 controls the operation of the DPC 3306, routing packets received from the pipeline manager 3302 to appropriate units within the DPC 3306. In at least one embodiment, packets associated with vertices are routed to the primitive engine 3312, which is configured to fetch vertex attributes associated with the vertices from memory; conversely, packets associated with shader programs may be sent to the SM 3314.

[0271] In at least one embodiment, SM 3314 includes, but is not limited to, a programmable streaming processor configured to process tasks represented by multiple threads. In at least one embodiment, SM 3314 is multithreaded and configured to simultaneously execute multiple threads (e.g., 32 threads) from a particular thread group, and implements a single instruction, multiple data ("SIMD") architecture, in which each thread in a group of threads (e.g., a warp) is configured to process a different data set based on the same instruction set. In at least one embodiment, all threads in a thread group execute the same instructions. In at least one embodiment, SM 3314 implements a single instruction, multiple thread ("SIMT") architecture, in which each thread in a group of threads is configured to process a different data set based on the same instruction set, but in which individual threads in a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, thereby enabling concurrency between warps and serial execution within a warp when threads in the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby enabling equal concurrency between all threads within a warp and between warps. In at least one embodiment, execution state is maintained for each individual thread, and threads executing the same instruction can be converged and executed in parallel to improve efficiency. At least one embodiment of SM 3314 is described in more detail below.

[0272] In at least one embodiment, the MMU 3318 provides a communication channel between the GPC 3300 and the memory partition unit (e.g., Figure 32 The MMU 3318 provides an interface between the memory and the partition unit 3222, and provides virtual to physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, the MMU 3318 provides one or more translation lookaside buffers ("TLBs") for performing translation of virtual addresses to physical addresses in memory.

[0273] The reasoning and / or training logic 1015 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 10 and / or Figure 11 Provides details regarding inference and / or training logic 1015. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the GPC 3300. In at least one embodiment, the GPC 3300 is used to infer or predict information based on a machine learning model (e.g., a neural network) that has been trained by another processor or system or the GPC 3300. In at least one embodiment, the GPC 3300 can be used to perform one or more of the neural network use cases described herein.

[0274] Figure 34 A memory partition unit 3400 of a parallel processing unit ("PPU") is shown according to one embodiment. In at least one embodiment, the memory partition unit 3400 includes, but is not limited to, a raster operations ("ROP") unit 3402; a level 2 ("L2") cache 3404; a memory interface 3406; and any suitable combination thereof. In at least one embodiment, the memory interface 3406 is coupled to a memory. The memory interface 3406 can implement a 32-, 64-, 128-, 1024-bit data bus, etc., for high-speed data transfer. In one embodiment, the PPU includes U memory interfaces 3406, one memory interface 3406 for each pair of partition units 3400, where each pair of partition units 3400 is connected to a corresponding memory device. For example, in at least one embodiment, the PPU can be connected to up to Y memory devices, such as a high-bandwidth memory stack or graphics double data rate, version 5, synchronous dynamic random access memory ("GDDR5 SDRAM").

[0275] In one embodiment, the memory interface 3406 implements a High Bandwidth Memory 2 ("HBM2") memory interface, and Y is equal to half of U. In at least one embodiment, the HBM2 memory stack is located on the same physical package as the PPU, saving significant power and area compared to a GDDR5 SDRAM system. In one embodiment, each HBM2 stack includes four memory dies, and Y is equal to 4, and the HBM2 stack includes two 128-bit channels per die, for a total of 8 channels and a data bus width of 1024 bits. In at least one embodiment, the memory supports single-error correction double-error detection ("SECDED") error correction code ("ECC") to protect data. ECC provides higher reliability for computing applications that are sensitive to data corruption.

[0276] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partitioning unit 3400 supports unified memory to provide a single unified virtual address space for CPU and PPU memory, thereby enabling data sharing between virtual memory systems. In at least one embodiment, the frequency of PPU accesses to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the pages more frequently. In one embodiment, the high-speed GPU interconnect 3208 supports address translation services that allow the PPU to directly access the CPU's page tables and provide full access to CPU memory by the PPU.

[0277] In one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In one embodiment, the copy engine can generate a page fault for an address that is not mapped in the page table, and the memory partition unit 3000 then services the page fault, maps the address into the page table, and then the copy engine performs the transfer. In at least one embodiment, fixed (or, non-pageable) memory is operated for multiple copy engines between multiple processors, thereby substantially reducing the available memory. In one embodiment, due to hardware page faults, addresses can be passed to the copy engine regardless of whether the memory page is resident, and the copy process is transparent.

[0278] According to at least one embodiment, Figure 32 Data from memory 3204 or other system memory is retrieved by the memory partition unit 3400 and stored in the L2 cache 3404, which is located on-chip and shared between the various GPCs. In one embodiment, each memory partition unit 3400 includes at least a portion of the L2 cache associated with the corresponding memory device. In at least one embodiment, lower-level caches are implemented in various units within the GPC. In one embodiment, each SM 3314 can implement a level 1 ("L1") cache, where the L1 cache is private memory dedicated to a particular SM 3314, and data from the L2 cache 3404 is retrieved and stored in each L1 cache for processing in the functional units of the SM 3314. In one embodiment, the L2 cache 3404 is coupled to the memory interface 3406 and the XBar 3220.

[0279] In one embodiment, ROP unit 3402 performs pixel color-related graphics raster operations, such as color compression and pixel blending. In one embodiment, ROP unit 3402 performs depth testing in conjunction with raster engine 3308, receiving the depth of sample locations associated with pixel fragments from the culling engine of raster engine 3308. In at least one embodiment, a depth test is performed against the corresponding depth in the depth buffer for the sample locations associated with the fragments. In at least one embodiment, if the fragment passes the depth test for the sample location, ROP unit 3402 updates the depth buffer and sends the result of the depth test to raster engine 3308. It will be appreciated that the number of partition units 3400 can differ from the number of GPCs, and therefore, in at least one embodiment, each ROP unit 3402 can be coupled to each GPC. In at least one embodiment, ROP unit 3402 tracks packets received from different GPCs and determines to which GPC to route the results generated by ROP unit 3402 via Xbar 3320.

[0280] Figure 35 Streaming Multiprocessor ("SM") 3500 is shown according to one embodiment. In at least one embodiment, SM 3500 is Figure 33SM. In at least one embodiment, SM 3500 includes, but is not limited to, an instruction cache 3502; one or more scheduler units 3504; a register file 3508; one or more processing cores ("cores") 3510; one or more special function units ("SFUs") 3512; one or more load / store units ("LSUs") 3514; an interconnect network 3516; a shared memory / level 1 ("L1") cache 3518; and any suitable combination thereof. In at least one embodiment, a work distribution unit dispatches tasks for execution on general processing clusters ("GPCs") of parallel processing units ("PPUs"), with each task being assigned to a specific data processing cluster ("DPC") within the GPC, and, if the task is associated with a shader program, the task is assigned to SM 3500. In one embodiment, scheduler unit 3504 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks assigned to SM 3500. In at least one embodiment, the scheduler unit 3504 schedules thread blocks for execution as warps of parallel threads, where each thread block is assigned at least one warp. In at least one embodiment, each warp executes a thread. In at least one embodiment, the scheduler unit 3504 manages a plurality of different thread blocks, assigns warps to different thread blocks, and then distributes instructions from a plurality of different cooperating groups to various functional units (e.g., core 3510, SFU 3512, and LSU 3514) on each clock cycle.

[0281] In at least one embodiment, cooperative groups may refer to a programming model for organizing groups of communicating threads. This programming model allows developers to express the granularity of the communicating threads, enabling richer, more efficient decompositions of parallelism. In at least one embodiment, a cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. In at least one embodiment, applications of the programming model provide a single, simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, in at least one embodiment, programmers often want to define thread groups at a granularity smaller than that of a thread block and synchronize within the defined group, thereby achieving higher performance, design flexibility, and software reuse in the form of collectively scoped functional interfaces. In at least one embodiment, cooperative groups enable programmers to define thread groups that are explicitly located at sub-block and multi-block granularity and perform collective operations, such as synchronization, on threads within the cooperative group. The programming model supports clear composition across software boundaries, so libraries and utility functions can safely synchronize within their local context without making assumptions about convergence. In at least one embodiment, the cooperative group primitive enables new cooperative parallelism patterns, including but not limited to producer-consumer parallelism, opportunistic parallelism, and global synchronization across an entire grid of thread blocks.

[0282] In at least one embodiment, the dispatch unit 3506 is configured to send instructions to one or more functional units, and the scheduler unit 3504 includes, but is not limited to, two dispatch units 3506 that enable two different instructions from the same warp to be dispatched in each clock cycle. In at least one embodiment, each scheduler unit 3504 includes a single dispatch unit 3506 or additional dispatch units 3506.

[0283] In at least one embodiment, each SM 3500 includes a register file 3508 that provides a set of registers for the functional units of SM 3500. In at least one embodiment, register file 3508 is partitioned between each functional unit, such that each functional unit is allocated a dedicated portion of register file 3508. In at least one embodiment, register file 3508 is partitioned by the different warps executed by SM 3500, and register file 3508 provides temporary storage for operands connected to the data paths of the functional units. In at least one embodiment, each SM 3500 includes a number L of processing cores 3510. In at least one embodiment, SM 3500 includes a large number (e.g., 128 or more) of different processing cores 3510. In at least one embodiment, each core 3510 includes, but is not limited to, a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit, including, but not limited to, a floating-point arithmetic logic unit ("ALU") and an integer arithmetic logic unit ("ALU"). In at least one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In at least one embodiment, processing core 3510 includes, but is not limited to, 64 single-precision (32-bit) floating point cores, 64 integer cores, 32 double-precision (64-bit) floating point cores, and 8 tensor cores.

[0284] According to at least one embodiment, the Tensor Cores are configured to perform matrix operations. In at least one embodiment, core 3510 includes one or more Tensor Cores. In at least one embodiment, the Tensor Cores are configured to perform deep learning matrix arithmetic, such as convolution operations used for neural network training and inference. In at least one embodiment, each Tensor Core operates on a 4×4 matrix and performs a matrix multiplication and accumulation operation D=A×B+C, where A, B, C, and D are 4×4 matrices.

[0285] In at least one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, and the accumulation matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the Tensor Cores operate on 16-bit floating-point input data with 32-bit floating-point accumulation. In at least one embodiment, the 16-bit floating-point multiplication requires 64 operations and produces a full-precision product, which is then accumulated with other intermediate products for a 4×4×4 matrix using 32-bit floating-point addition. In at least one embodiment, the Tensor Cores are used to perform operations on larger two-dimensional or higher-dimensional matrices constructed from these smaller elements. In at least one embodiment, an API such as the CUDA 9 C++ API exposes specialized matrix load, matrix multiplication and accumulation, and matrix store operations to efficiently use the Tensor Cores from CUDA-C++ programs. In at least one embodiment, at the CUDA level, the warp-level interface assumes that a 16×16 matrix spans all 32 threads of the warp.

[0286] In at least one embodiment, each SM 3500 includes, but is not limited to, M SFUs 3512 that perform specialized functions (e.g., attribute evaluation, reciprocal square root, etc.). In at least one embodiment, SFUs 3512 include, but are not limited to, tree traversal units configured to traverse a hierarchical tree data structure. In at least one embodiment, SFUs 3512 include, but are not limited to, texture units configured to perform texture map filtering operations. In at least one embodiment, the texture units are configured to load a texture map (e.g., a 2D array of texels) from memory and sample the texture map to generate sampled texture values ​​for use in a shader program executed by the SM 3500. In at least one embodiment, the texture map is stored in shared memory / LI cache 3518. In at least one embodiment, the texture units perform texture operations, such as filtering operations using mip-maps (e.g., texture maps with varying levels of detail). In at least one embodiment, each SM 3500 includes, but is not limited to, two texture units.

[0287] In at least one embodiment, each SM 3500 includes, but is not limited to, N LSUs 3514 that implement load and store operations between shared memory / L1 cache 3518 and register file 3508. In at least one embodiment, each SM 3500 includes, but is not limited to, an interconnection network 3516 that connects each functional unit to register file 3508 and the LSUs 3514 to register file 3508, and a shared memory / L1 cache 3518. In at least one embodiment, the interconnection network 3516 is a crossbar switch that is configurable to connect any functional unit to any register in register file 3508 and to connect the LSUs 3514 to memory locations in register file and shared memory / cache 3518.

[0288] In at least one embodiment, shared memory / L1 cache 3518 is an array of on-chip memory that, in one embodiment, allows for data storage and communication between the SM 3500 and the primitive engines, as well as between threads within the SM 3500. In at least one embodiment, shared memory / L1 cache 3518 includes, but is not limited to, 128KB of storage capacity and is located in the path from the SM 3500 to the partition unit. In at least one embodiment, shared memory / L1 cache 3518 is used to cache reads and writes. One or more of the shared memory / L1 cache 3518, the L2 cache, and the memory is a backing store.

[0289] In at least one embodiment, combining data cache and shared memory functionality into a single memory block provides improved performance for both types of memory accesses. In at least one embodiment, this capacity is used or utilized as a cache by programs that do not utilize shared memory. For example, if the shared memory is configured to utilize half of its capacity, texture and load / store operations can utilize the remaining capacity. According to at least one embodiment, integration within shared memory / L1 cache 3518 enables shared memory / L1 cache 3518 to serve as a high-throughput pipeline for streaming data, while providing high-bandwidth and low-latency access to frequently reused data. In at least one embodiment, when configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function graphics processing unit is bypassed, creating a simpler programming model. In at least one embodiment, in a general-purpose parallel computing configuration, the work distribution unit directly allocates and distributes threads' blocks to DPCs. In at least one embodiment, threads in a block execute the same program, use unique thread IDs in computations to ensure each thread generates unique results, use SM 3500 to execute the program and perform computations, use shared memory / L1 cache 3518 to communicate between threads, and LSU 3514 reads and writes global memory through shared memory / L1 cache 3518 and a memory partitioning unit. In at least one embodiment, when configured for general-purpose parallel computing, SM 3500 writes commands that scheduler unit 3504 can use to start new work on a DPC.

[0290] In at least one embodiment, the PPU is included in or coupled to a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., wireless, handheld device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, etc. In at least one embodiment, the PPU is implemented on a single semiconductor substrate. In at least one embodiment, the PPU is included in a system-on-chip ("SoC") along with one or more other devices (e.g., an additional PPU, memory, a reduced instruction set computer ("RISC") CPU, one or more memory management units ("MMUs"), a digital-to-analog converter ("DAC"), etc.

[0291] In at least one embodiment, the PPU can be included on a graphics card that includes one or more storage devices. The graphics card can be configured to connect to a PCIe slot on a desktop computer motherboard. In at least one embodiment, the PPU can be an integrated graphics processing unit ("iGPU") included in a chipset on the motherboard.

[0292] Reasoning and / or training logic 1015 is used to perform reasoning and / or training operations related to one or more embodiments. Figure 10 and / or Figure 11 Provides details regarding inference and / or training logic 1015. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the SM 3500. In at least one embodiment, the SM 3500 is used to infer or predict information based on a machine learning model (e.g., a neural network) that has been trained by another processor or system or by the SM 3500. In at least one embodiment, the SM 3500 can be used to perform one or more of the neural network use cases described herein.

[0293] In at least one embodiment, a single semiconductor platform may refer to a single semiconductor-based integrated circuit or chip. In at least one embodiment, a multi-chip module with increased connectivity may be used, which emulates on-chip operations and provides substantial improvements over implementations utilizing a central processing unit ("CPU") and bus. In at least one embodiment, various modules may also be placed separately or in various combinations of semiconductor platforms, depending on user needs.

[0294] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various variations and alternative constructions, certain illustrated embodiments are shown in the drawings and have been described in detail above. However, it should be understood that there is no intention to limit the invention to the particular forms or versions disclosed, but on the contrary, it is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the invention, as defined by the appended claims.

[0295] The use of the terms "a," "an," and "said," and similar references in the context of describing the disclosed embodiments (especially in the context of the appended claims) should be interpreted as covering the singular and the plural, unless otherwise indicated herein or clearly contradicted by the context. Unless otherwise indicated, the terms "comprise," "have," "include," and "include" should be interpreted as open-ended terms (i.e., meaning "including but not limited to"). The term "connected" (unmodified and referring to a physical connection) should be understood to mean fully or partially contained in, attached to, or connected together, even if something intervenes. References to numerical ranges herein are intended to serve only as a shorthand method, and unless otherwise indicated herein, each individual value falling within the range is referred to separately, and each individual value is incorporated into the specification as if separately referenced herein. Use of the term "set" (e.g., "a set of items)" or "subset" should be interpreted as a non-empty set comprising one or more members, unless otherwise indicated by the context or contradicted by it. Furthermore, unless otherwise indicated or contradicted by the context, the term "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but the subset and the corresponding set may be equivalent.

[0296] Connective language, such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C," unless expressly stated otherwise or clearly contradicted by context, can be understood with the context to present an item, clause, or the like, which can be A or B or C, or any non-empty subset of the set of A and B and C. For example, in the illustrative example of a set having three members, the connective phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such connective language is not intended to imply that certain embodiments require at least one A, at least one B, and at least one C, each of which is used for presentation. Additionally, unless expressly stated otherwise or contradicted by context, the term "plurality" denotes plurality (e.g., "a plurality of items" means a plurality of items). The number of items in "plurality" is at least two, but can be more when indicated explicitly or by context. Further, unless specified otherwise or clear from context, the phrase "based on" means "based at least in part on" rather than "based solely on."

[0297] The operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or clearly contradicted by the context. In one embodiment, processes such as those described herein (or variations and / or combinations thereof) are executed by hardware or a combination thereof under the control of one or more computer systems, one or more computer systems configured with executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors. In one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that includes a plurality of instructions that can be executed by one or more processors. In one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that does not include transient signals (e.g., propagated transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuits (e.g., buffers, caches, and queues) within a transceiver for transient signals. In one embodiment, the code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media having executable instructions stored thereon (or other memory to store executable instructions), the executable instructions being executed by one or more processors of the computer system, causing the computer system to perform the operations described herein. In one embodiment, the set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lack all code, while the plurality of non-transitory computer-readable storage media collectively store all code. In one embodiment, the executable instructions are executed such that different instructions are executed by different processors—for example, a non-transitory computer-readable storage medium stores instructions, and a main CPU executes some instructions while a graphics processor unit ("GPU") executes other instructions. In one embodiment, different components of the computer system have independent processors, and the different processors execute different subsets of instructions.

[0298] Thus, in one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software that enables the performance of the operations. Furthermore, a computer system implementing an embodiment of the present disclosure is a single device, and in another embodiment, is a distributed computer system comprising multiple devices that operate in different ways such that the distributed computer system performs the operations described herein and such that no single device performs all of the operations.

[0299] Unless otherwise claimed, the use of any and all examples or exemplary language (e.g., "such as") provided herein is intended merely to better illuminate embodiments of the invention and does not limit the scope of the invention. No language in this specification should be construed as indicating any non-claimed element as essential to the practice of the invention.

[0300] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0301] In the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. Rather, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0302] Unless otherwise indicated, it should be understood that throughout this specification, terms such as "process," "computing," "calculating," "determining," and the like refer to the actions and / or processes of a computer or computing system or similar electronic computing device that operates on and / or transforms data (e.g., electronic) represented as physical quantities (e.g., electronic) in the registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers, or other such information storage, transmission, or display devices of the computing system.

[0303] In a similar manner, the term "processor" may refer to any device or portion of a device that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As non-limiting examples, a "processor" may be a central processing unit (CPU) or a graphics processing unit (GPU). A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Likewise, each process may refer to multiple processes to execute instructions sequentially or in parallel, continuously or intermittently. The terms "system" and "method" are used interchangeably herein to the extent that a system may embody one or more methods and the method may be considered a system.

[0304] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. The process of obtaining, acquiring, receiving, or inputting analog and digital data may be accomplished in a variety of ways, such as by receiving the data as a parameter to a function call or a call to an application programming interface. In some embodiments, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting the data via a serial or parallel interface. In another embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transferring the data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data may be accomplished by transmitting the data as an input or output parameter to a function call, a parameter to an application programming interface, or an inter-process communication mechanism.

[0305] Although the above discussion sets forth example implementations of the described technology, other architectures may be used to implement the described functionality and are intended to be within the scope of this disclosure. In addition, although specific responsibilities are defined above for discussion purposes, the various functions and responsibilities may be allocated and divided in different ways depending on the circumstances.

[0306] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. A processor comprising: One or more circuits for generating one or more images of one or more cells using one or more neural networks based at least in part on inputs thereto including background image data and gene expression data associated with visual features of the one or more cells.

2. The processor of claim 1 , wherein the one or more neural networks: The one or more images are inferred, the one or more neural networks being a multi-conditional generative adversarial network (GAN) trained using medical image data and gene expression data.

3. The processor of claim 1 , wherein the one or more neural networks are trained in part by encoding medical image data and gene expression data and fusing the encoded data to generate a composite image and a segmentation mask, the composite image comprising a representation of a group of cells mixed with a background portion of the medical image data.

4. The processor of claim 1 , wherein the one or more neural networks are further trained by passing the synthetic image, the segmentation mask, and the gene code for the gene expression data to a discriminator to determine a set of loss values, wherein one or more network parameters of the one or more neural networks are updated using the set of loss values.

5. The processor of claim 1 , wherein the one or more neural networks utilize a learned genomic graph between visual features and gene expression data of the one or more cells.

6. A system comprising: one or more memories for storing inputs to the one or more neural networks including background image data and gene expression data associated with visual features of the one or more cells; as well as One or more processors cause one or more circuits to generate one or more images of the one or more cells using the one or more neural networks based at least in part on the gene expression data.

7. The system of claim 6, wherein the one or more neural networks: The one or more images are inferred, the one or more neural networks being a multi-conditional generative adversarial network (GAN) trained using medical image data and gene expression data.

8. The system of claim 6 , wherein the one or more neural networks are trained in part by encoding medical image data and gene expression data and fusing the encoded data to generate a composite image and a segmentation mask, the composite image comprising a representation of a group of cells mixed with a background portion of the medical image data.

9. The system of claim 6 , wherein the one or more neural networks are further trained by passing the synthetic image, the segmentation mask, and the gene code for the gene expression data to a discriminator to determine a set of loss values, wherein the set of loss values ​​is used to update one or more network parameters of the one or more neural networks.

10. The system of claim 6, wherein the one or more neural networks utilize a learned genomic graph between visual features and gene expression data of the one or more cells.

11. A method comprising: generating one or more images of the one or more cells using the one or more neural networks based at least in part on inputs to the one or more neural networks including background image data and gene expression data associated with visual features of the one or more cells; and The one or more images are stored.

12. The method according to claim 11, further comprising: The one or more images are inferred, wherein the one or more neural networks are multi-conditional generative adversarial networks (GANs) trained using medical image data and gene expression data.

13. The method of claim 11 , wherein the one or more neural networks are trained in part by encoding medical image data and gene expression data and fusing the encoded data to generate a composite image and a segmentation mask, the composite image comprising a representation of a cell group mixed with a background portion of the medical image data.

14. The method of claim 11 , wherein the one or more neural networks are further trained by passing the synthetic image, the segmentation mask, and the gene code for the gene expression data to a discriminator to determine a set of loss values, wherein the set of loss values ​​is used to update one or more network parameters of the one or more neural networks.

15. The method of claim 11, wherein the one or more neural networks utilize a learned genomic graph between visual features and gene expression data of the one or more cells.

16. A processor comprising: One or more circuits for training one or more neural networks based at least in part on inputs to the one or more neural networks comprising background image data and genetic information associated with visual features of one or more cells, the one or more neural networks to be used to infer one or more images of the one or more cells.

17. The processor of claim 16, wherein the one or more neural networks are multi-conditional generative adversarial networks (GANs).

18. The processor of claim 16, wherein the one or more neural networks are trained in part by encoding medical image data and genetic information and fusing the encoded data to generate a composite image and a segmentation mask, the composite image comprising a representation of a cell group mixed with a background portion of the medical image data.

19. The processor of claim 16 , wherein the one or more neural networks are further trained by passing the synthetic image, the segmentation mask, and the genetic code for the genetic information to a discriminator to determine a set of loss values, wherein one or more network parameters of the one or more neural networks are updated using the set of loss values.

20. The processor of claim 16, wherein the one or more neural networks are further trained to learn a genomic map between visual features of the one or more cells and the genetic information.

21. A system comprising: one or more memories for storing inputs to the one or more neural networks including background image data and genetic information associated with visual features of the one or more cells; as well as One or more processors cause one or more circuits to train the one or more neural networks to infer one or more images of the one or more cells based at least in part on the genetic information.

22. The system of claim 21, wherein the one or more neural networks are multi-conditional generative adversarial networks (GANs).

23. The system of claim 21 , wherein the one or more neural networks are trained in part by encoding medical image data and genetic information and fusing the encoded data to generate a composite image and a segmentation mask, the composite image comprising a representation of a cell group mixed with a background portion of the medical image data.

24. The system of claim 21 , wherein the one or more neural networks are further trained by passing the synthetic image, the segmentation mask, and the genetic code for the genetic information to a discriminator to determine a set of loss values, wherein one or more network parameters of the one or more neural networks are updated using the set of loss values.

25. The system of claim 21, wherein the one or more neural networks are trained to learn a genomic map between visual features of the one or more cells and the genetic information.

26. A method comprising: one or more circuits for training one or more neural networks to infer one or more images of one or more cells based at least in part on inputs to the one or more neural networks comprising background image data and genetic information associated with visual features of the one or more cells; and The neural network is stored.

27. The method of claim 26, wherein the one or more neural networks are multi-conditional generative adversarial networks (GANs).

28. The method of claim 26, wherein the one or more neural networks are trained in part by encoding medical image data and genetic information and fusing the encoded data to generate a composite image and a segmentation mask, the composite image comprising a representation of a cell group mixed with a background portion of the medical image data.

29. The method of claim 26, wherein the one or more neural networks are further trained by passing the synthetic image, the segmentation mask, and the genetic code for the genetic information to a discriminator to determine a set of loss values, wherein one or more network parameters of the one or more neural networks are updated using the set of loss values.

30. The method of claim 26, wherein the one or more neural networks are further trained to learn genomic maps between visual features of the one or more cells and the genetic information.

Citation Information

Patent Citations

  • Drug indication and response prediction systems and method using ai deep learning based on convergence of different category data

    US20190164632A1

Cited By

  • Feature extraction with three-dimensional information

    US20250131680A1