Training of one or more neural networks using synthetic data
By employing a neural network-based upscaler to upscale low-resolution images generated by a renderer system, the challenges of generating high-resolution video content efficiently are addressed, resulting in high-quality, high-frame-rate output.
Patent Information
- Application Number
- JP2022526732
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-26
- Filing Date
- 2021-10-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-10-25
AI Technical Summary
Existing technologies face challenges in generating high-resolution video content efficiently, particularly in devices with limited resource capacity, as they struggle to maintain high frame rates and often result in low-quality content with artifacts.
The use of a renderer system that generates low-resolution images and then upscales them using a neural network-based upscaler, which incorporates temporal reconstruction and blending to achieve high-quality, high-resolution images in real-time.
This approach enables real-time rendering of high-resolution images with quality equivalent to native resolution, while significantly increasing frame rates and reducing processing resource requirements.
Smart Images

Figure 0007696342000005 
Figure 0007696342000006 
Figure 0007696342000007
Abstract
Description
Technical Field
[0001] This is a PCT application of U.S. Patent Application No. 17 / 080,503, filed on October 26, 2020. The disclosure of that application is hereby incorporated by reference in its entirety for all purposes.
[0002] At least one embodiment relates to processing resources used to execute and facilitate artificial intelligence. For example, at least one embodiment relates to a processor or computing system used to train a neural network according to various novel techniques described herein.
Background Art
[0003] Images and video content are being generated more and more, at higher resolutions, and are being displayed on higher-quality displays. The techniques for generating this content at these higher resolutions are often extremely resource-intensive, which can be a problem for devices with limited resource capacity. Additionally, video content is often required to be displayed at a target frame rate or minimum frame rate, and it can be difficult to generate this high-resolution content at such frame rates. Often, the quality of the resulting content is constrained by these and other limitations. Upscaling may be used to generate higher-resolution content, but this content often has problems with artifacts or is otherwise not of the desired quality. Additionally, it can be difficult to obtain sufficient training data for training a neural network to perform such tasks.
Summary of the Invention
Means for Solving the Problems
[0004] Various embodiments according to the present disclosure will be described with reference to the drawings.
Brief Description of the Drawings
[0005]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12A
Figure 12B
Figure 12C
Figure 12D
Figure 12E
Figure 12F
Figure 13
Figure 14A
Figure 14B
Figure 15A
Figure 15B
Figure 16
Figure 17A
Figure 17B
Figure 17C
Figure 17D
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26A
Figure 26B
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33A
Figure 33B
DETAILED DESCRIPTION OF THE INVENTION
[0006] In at least one embodiment, content such as video game content or animation can be generated using a renderer 102, a rendering engine, or other such content generators. In at least one embodiment, the renderer 102 can receive input for one or more frames of a sequence and generate frames of an image or video using stored content 104 (e.g., maps and graphic assets) that has been modified at least in part based on that input. In at least one embodiment, this renderer 102 can utilize rendering software such as Epic Games' Unreal Engine4, which can provide functions such as deferred shading, global illumination, lit transparency, post-processing, and GPU particle simulation using vector fields, and can be part of a rendering pipeline. In at least one embodiment, the amount of processing required for this complex rendering of full high-resolution images can make it difficult to render these video frames at a current frame rate, such as at least 60 frames per second (fps). In at least one embodiment, the renderer 102 may instead be used to generate a rendered image 106 at a resolution lower than one or more final output resolutions in order to meet timing requirements and reduce processing resource requirements. In at least one embodiment, this low-resolution rendered image 106 may be processed using an upscaler 108 to generate an upscaled image 110 that represents the content of the low-resolution rendered image 106 at a resolution equal to (or at least closer to) the target output resolution.
[0007] In at least one embodiment, an upscaler system 108 (which can take the form of a service, system, module, or device) can be used to upscale individual frames of a video or animation sequence. In at least one embodiment, the amount of upscaling to be performed can depend on the initial resolution of the rendered image and the target resolution of the display, such as transitioning from 1080p to 4K resolution. In at least one embodiment, additional processes can be performed as part of an upsampling process that can include anti-aliasing and temporal smoothing. In at least one embodiment, an appropriate reconstruction filter can be utilized, such as one that can involve a Gaussian filter. In at least one embodiment, the upsampling process can take into account subpixel dithering that can be applied on a per-frame basis.
[0008] In at least one embodiment, deep learning can be used to infer these upsampled video frames of the sequence. In at least one embodiment, temporal reconstruction can be used to provide in a combined form of anti-aliasing and super-resolution. In at least one embodiment, information from the corresponding sequence of video frames can be used to infer a higher quality upsampled image. In at least one embodiment, one or more heuristics based on prior knowledge of the rendering pipeline that do not require learning from data can be used. In at least one embodiment, this can include jitter-aware upsampling and accumulating samples at the upsampled resolution. In at least one embodiment, this previous process data, to infer a higher quality upsampled image 110 than would be generated by the upsampling algorithm alone, can be provided as input to an upscaler 108 including at least one neural network, along with the current input video frame and the previously inferred frame. In at least one embodiment, this upsampling substantially shifts those jitters and per-frame samples such that the jitters and per-frame samples are aligned with a past buffer that can be at a higher resolution.
[0009] In at least one embodiment, this upscaled image 110 can be provided as an input to a neural network 112 to determine one or more blending factors or blending weights. In at least one embodiment, this neural network can also determine at least some filtering to be applied when reconstructing the current image or when blending with a previous image. In at least one embodiment, this information can then be provided along with this upscaled image 110 to a blending component 114 that is to be blended with at least one previous image of this sequence. In at least one embodiment, this blending of the current image with a previous (or past) image of the sequence can help with temporal convergence to a good and sharp high-resolution output image 116, which can then be provided for presentation via a display 120 or other such presentation mechanism. In at least one embodiment, a copy of this high-resolution output image 116 can also be stored in a history buffer 118 or other such storage location for blending with images generated subsequently in this sequence. In at least one embodiment, such a process can utilize deep learning to reconstruct images for real-time rendering at a resolution several times (e.g., 2×, 4×, or 8×) higher than the actual rendering resolution, with a reconstructed image quality that is at least equivalent to native resolution rendering in terms of details, temporal stability, and the absence of common artifacts such as ghosting or latency. In at least one embodiment, the reconstruction speed can be accelerated using tensor cores, and using the techniques presented herein can make this rendering process much more sample efficient, leading to a significant increase in frames per second for various applications.
[0010] In at least one embodiment, such a base approach can be used to reconstruct an image for real-time rendering at a resolution higher than the actual rendering resolution generated by a rendering engine. In at least one embodiment, the resulting reconstructed image quality is at least equivalent to, or even exceeds, that of native resolution rendering in terms of details, temporal stability, and the absence of common artifacts such as ghosting or latency. In at least one embodiment, this reconstruction speed can be accelerated using tensor cores. In at least one embodiment, such an approach can make the rendering process much more sample efficient, leading to a significant increase in the possible frame rate for various applications.
[0011] In at least one embodiment, a set of training data that is completely synthetic can be generated. In at least one embodiment, a network for performing upsampling generally requires at least two images for training, namely, an image at a lower resolution and a version of the same image at a higher resolution that can be used as a reference for the upsampling image generated by this network based on this lower resolution version. In at least one embodiment, a network that also performs tasks such as temporal smoothing may also require at least one sequence of images at one or more of these resolutions. In at least one embodiment, these two different resolution versions of the same image can be rendered separately or independently by a deterministic rendering engine. In at least one embodiment, this deterministic rendering engine can generate the same animation sequence for a scene or clip such that there is little to no difference between these versions other than the resolution difference. In at least one embodiment, this deterministic engine can be obtained, generated, or modified to be deterministic such that there is no (or very little) randomization or other variation between the rendering of these two versions.
[0012] In at least one embodiment, the system 200 for generating this training data can be utilized as shown in FIG. 2. In at least one embodiment, a user can utilize an interface 206 such as a console, command line, script tool, or application programming interface (API) to give commands or selections to the renderer 202. In at least one embodiment, this may include information about the scene to be rendered, as well as two or more resolutions at which the scene is to be rendered therein. In at least one embodiment, via this interface, the user can use a deterministic approach to specify whether these scenes should be equivalent or whether it is possible to have some variance or randomization for performance or other such purposes. In at least one embodiment, via this interface, the user can also specify content 204 such as maps or assets to be used in rendering these images, which can be part of the scene or sequence of images or frames to be rendered.
[0013] In at least one embodiment, the renderer 202 can be instructed to render the image 210 at a first, lower resolution, which can be the native resolution at which the rendering engine renders the image content for the application thereon. In at least one embodiment, this can include rendering a first version of this image, for example, at 1080p, which is the resolution expected to be generated by the rendering engine for a game or animation. In at least one embodiment, the user can also utilize the interface 206 to specify one or more lower and / or higher resolutions to be used to generate versions of this same image, where the versions can relate to different instances having at least one different aspect, such as resolution, color depth, temporal resolution, position in the sequence, etc. In at least one embodiment, this can include each resolution at which it is expected that this content will be upsampled for output. In at least one embodiment, the renderer 206 can then generate one or more output images 214 at those target resolutions. In at least one embodiment, these higher-resolution images can be equivalent, at least to the extent possible or practical for a properly functioning system, for synthetically generating different versions of a single image at different resolutions. In at least one embodiment, a sequence of these images can then be provided as training data for a neural network that, at least, accepts the lower-resolution rendered image as input and upsamples that image in real time to a higher-resolution image at the target output resolution. In at least one embodiment, a sequence will be generated that includes, respectively, the lower-resolution image 210 and one of these higher-resolution images 212, 214 as a reference or ground truth for the upsampled image generated by this network trained based on this lower-resolution image 210.
[0014] In at least one embodiment, using synthetic data instead of "real" data such as actual images generated during a live game session can help achieve generalization. In at least one embodiment, synthetic data can be generated without artifacts, such as by downsampling one of these images, where an image at one resolution may be generated when that image is used to generate at another resolution. In at least one embodiment, the generation of synthetic training data also allows for the intentional inclusion of common rendering artifacts such as aliasing, dithering, moiré patterns, high-frequency specular aliasing, and noise. In at least one embodiment, a dataset can be used to train the network without actual or live data collected from the application for which a network that may be fully synthetic is to be utilized. In at least one embodiment, such an approach may be counterintuitive, but can generate generalized results such that the resulting model can be applied in real time with high quality to actual or live application data.
[0015] In at least one embodiment, such a dataset can be used to train one or more neural networks for real-time rendering super-resolution. In at least one embodiment, the synthetic dataset can include images with intentional rendering artifacts that may be related to noise, specular aliasing, shader aliasing, geometry aliasing, hidden objects, ghosting, or dithering. In at least one embodiment, this data can then be used to train a super-resolution model for image reconstruction. In at least one embodiment, this synthetic dataset can include a sufficient amount of data with sufficient variance such that the model can be trained to avoid generating images with each of the known number of rendering artifacts, particularly those with these artifacts. In at least one embodiment, such techniques can be used to generate high-quality real-time super-resolution images that can be at least partially based on a set of three-dimensional (3D) assets.
[0016] In at least one embodiment, the same set of maps and 3D assets can be used to render images at two or more resolutions. In at least one embodiment, the determined camera path may be used to render a video clip of a determined length at a first resolution and then reused to render a video clip of the determined length at a second resolution. In at least one embodiment, this synthetic dataset can be generated to emulate a diverse set of games in order to enable the network to generalize to different types of content to be rendered. In at least one embodiment, this separate generation can also enable images to be generated at resolutions that are not in integer ratios to each other. In at least one embodiment, the images are generated at the rendered resolution and the output resolution for training the network. In at least one embodiment, a sequence of low-resolution images and a sequence of high-resolution images are used for training to provide temporal smoothing and other such aspects, and the second image of these sequences may correspond to past images of the sequence, as described in more detail elsewhere in this specification. In at least one embodiment, a certain percentage of these training images may have intentionally introduced artifacts, such as including 10% of a given dataset. In at least one embodiment, instead of generating the large dataset offline that will later be used to train the network, all of these artifacts can be synthesized during training. In at least one embodiment, such an approach can greatly reduce the amount of time required to train the network. In at least one embodiment, at least some of these artifacts are generated by the renderer such that they are incorporated into this synthetic dataset without the need for them to be added separately. In at least one embodiment, other augmentations can be added to increase the various artifacts or features in this synthetic dataset. In at least one embodiment, this can include adding noise, particles, or patterns as an overlay augmentation.
[0017] In at least one embodiment, a number of training sequences can be generated. In at least one embodiment, in order to generate a training dataset, random subsampling can then be performed. In at least one embodiment, some or a percentage of the extended frames can be included. In at least one embodiment, these extended frames can be randomly generated for each training epoch to provide improved training and variance in this training data.
[0018] In at least one embodiment, the rendering engine can have various features that help improve the performance of this rendering engine. In at least one embodiment, this may include a random seed or other non-determinism. In at least one embodiment, this non-determinism can be treated as a bug, and a process such as a bug fix can be used to make the rendering engine deterministic. In at least one embodiment, a deterministic rendering engine can repeatedly generate the same scene without variance and output this scene at different resolutions or with other specified output features such as different aspect ratios or color depths. In at least one embodiment, such a rendering engine can be run multiple times to generate images of different resolutions in terms of pixel accuracy response or similarity.
[0019] In at least one embodiment, various frames corresponding to at least two different resolutions are captured at least twice. In at least one embodiment, corresponding frames can be compared to determine whether there are any differences other than resolution, and if there are differences, at least one of these frames can be recaptured. In at least one embodiment, most of the frames match based on being output from a deterministic rendering engine, but there can still sometimes be issues or cases where there are minor differences, such as when the GPU is not functioning properly and there may be random electrons injected into this process. In at least one embodiment, redundancy is provided so that there can be recapture for any inconsistencies, and a safety process can be utilized to identify any of these differences. In at least one embodiment, this process can continue until the pairs match or until the maximum number of retry attempts is reached. In at least one embodiment, this can be important because errors that affect even 5% of the pixel information can sometimes prevent the success of real-time supersampling.
[0020] In at least one embodiment, the noise present in these rendered images may be artificially amplified to better train a real-time super-resolution network for processing the noise. In at least one embodiment, some types of visual features (e.g., specular reflections) may require a large number of samples to appear realistic. In at least one embodiment, the rendering engine utilizes a small number of high-density samples to generate a stable image, which may not appear realistic when displayed. In at least one embodiment, thinning these samples may result in a noisy image. In at least one embodiment, this thinning may be amplified to increase the presence of noise in the training images. In at least one embodiment, higher resolution textures than those supported by the rendering engine may also be used to generate various rendering artifacts to be resolved. In at least one embodiment, a reference is generated using several samples (e.g., 64) per pixel and then reconstructed using a reconstruction filter, such as a Lanczos filter, with the correct dither offset to achieve optimal image quality in these training target images.
[0021] In at least one embodiment, a process 300 for training a network to perform real-time rendering super-resolution can be performed as shown in FIG. 3. In at least one embodiment, the renderer can be made to render 302 an image at a first resolution, such as corresponding to the resolution of an initial rendering image for a target application. In at least one embodiment, this renderer can be made to render 304 the same image at one or more target output resolutions different from this initial rendering resolution. In at least one embodiment, the target output resolution is a resolution higher than this initial rendering resolution. In at least one embodiment, a determination can be made as to whether there are more resolutions to be generated for this image, such as when this image may be upsampled to different resolutions for different instances. In at least one embodiment, this process can continue until a version of this image is rendered at all target output resolutions. In at least one embodiment, a determination 306 can be made as to whether to augment with one or more artifacts. In at least one embodiment, this can include adding an artifact (e.g., noise or particles) as an overlay to at least this lower-resolution image, or can include adding one or more additional images including image artifacts. In at least one embodiment, a pair of resolutions of this image, including the version at this initial rendering resolution and the upsampled output resolution, can be provided 308 as training data for training a real-time rendering super-resolution network. In at least one embodiment, this trained network can then be utilized 310 to upsample application images in real time. In at least one embodiment, these images can be 2D views of a virtual 3D environment from the perspective of a virtual camera. In at least one embodiment, only the generated synthetic data is included in this training data set.
[0022] In at least one embodiment, a process 400 for training a network can be performed as shown in FIG. 4. In at least one embodiment, a first version of training data, such as an image or a sequence of images, is synthetically generated 402, such as by being rendered at a first resolution. In at least one embodiment, a second version of this training data (e.g., an image or sequence) is synthetically generated 404, such as by being generated at a second resolution. In at least one embodiment, one or more neural networks can then be trained 406 using at least these two versions, which may be part of a fully synthetic dataset. In at least one embodiment, these images may be part of an image sequence. In at least one embodiment, these neural networks can be used for real-time rendering super-resolution and upsampling.
[0023] In at least one embodiment, a pre-trained neural network may be used to reconstruct a high-resolution image using a low-resolution image sequence as input. In at least one embodiment, multiple low-resolution rendering images are accumulated over time to reconstruct a high-resolution image with full details that can withstand native resolution rendering. In at least one embodiment, these low-resolution samples should be accumulated at the correct locations to obtain the full details of the high-resolution rendering. In at least one embodiment, a set of blending weights is calculated to achieve correct accumulation at a high resolution. In at least one embodiment, these weights may be at least partially based on the original sample locations in the lower-resolution rendering images.
[0024] In at least one embodiment, the client device 502 can generate content for a session, such as a gaming session or a video viewing session, using components of the content application 504 on the client device 502 and data stored locally on that client device. In at least one embodiment, the content application 524 (e.g., a gaming or streaming media application) running on the content server 520 can initiate at least a session associated with the client device 502, which can utilize a session manager and user data stored in the user database 534, cause the content 532 to be determined by the content manager 526, render it using the rendering engine 528 if required for this type of content or platform, and transmit it to the client device 502 using an appropriate transmission manager 522 for download, streaming, or sending via another such transmission channel. In at least one embodiment, the client device 502 that receives this content can also include or alternatively provide this content to the corresponding content application 504, which includes a rendering engine 510 for rendering at least a portion of this content for presentation via the client device 502, such as video content by the display 506 and audio such as sound and music by at least one audio playback device 508, such as speakers or headphones. In at least one embodiment, at least a portion of this content is already stored, rendered, or accessible to the client device 502 so that transmission via the network 540 is not required for at least that portion of the content, such as when the content was previously downloaded or stored locally on a hard drive or optical disk.In at least one embodiment, a transmission mechanism such as data streaming may be used to transmit this content from server 520 or content database 534 to client device 502. In at least one embodiment, at least a portion of this content may be obtained or streamed from another source, such as third-party content service 550 which may also include content application 552 for generating or providing the content. In at least one embodiment, a portion of this functionality may be executed using multiple processors within one or more computing devices, such as may include multiple computing devices, or a combination of a central processing unit (CPU) and a graphics processing unit (GPU).
[0025] In at least one embodiment, the content application 524 includes a content manager 526 that can determine or analyze the content before the content is sent to the client device 502. In at least one embodiment, the content manager 526 can also include or work with other components that can generate, modify, or enhance the content to be provided. In at least one embodiment, this can include a rendering engine 528 for rendering content such as aliased content at a first resolution. In at least one embodiment, an upsampling or scaling component 530 can generate at least one additional version of this image at a different, higher or lower resolution and can perform at least some processing such as anti-aliasing. In at least one embodiment, a blending component 532, which can include at least one neural network, can perform blending for one or more of those images against one or more previous images as described herein. In at least one embodiment, the content manager 526 can then select an image or video frame at an appropriate resolution for sending to the client device 502. In at least one embodiment, the content application 504 on the client device 502 can also include components such as a rendering engine 510, an upsampling module 512, and a blending module 514 such that any or all of this functionality can be further or alternatively performed on the client device 502. In at least one embodiment, the content application 552 on the third-party content service system 550 can also include such functionality. In at least one embodiment, the location where at least a portion of this functionality is performed can be configurable or can depend on factors such as the type of the client device 502 or the availability of a network connection with an appropriate bandwidth among such factors.In at least one embodiment, the upsampling module 530 or the blending module 532 may include one or more neural networks to perform or assist with this function, and those neural networks (or at least the network parameters for those networks) may be provided by the content server 520 or a third - party system 550. In at least one embodiment, the system for content generation can include any suitable combination of hardware and software at one or more locations. In at least one embodiment, the generated image or video content at one or more resolutions is also provided to or made available for another client device 560, such as for download or streaming from a media source that stores a copy of the image or video content. In at least one embodiment, this can include transmitting an image of game content for a multi - player game if different client devices can display the content at different resolutions, including one or more super - resolutions.
[0026] Inference and training logics FIG. 6A shows inference and / or training logic 615 used to perform inference and / or training operations with respect to one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B.
[0027] In at least one embodiment, the inference and / or training logic 615 may include, without limitation, code and / or data storage 601 for storing forward propagation and / or output weights, and / or input / output data, and / or other parameters for constructing neurons or layers of a neural network that are trained and / or used to infer in one or more embodiments. In at least one embodiment, the training logic 615 may include, or be coupled to, code and / or data storage 601 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 601 has weight and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which this code corresponds. In at least one embodiment, the code and / or data storage 601 stores the weight parameters and / or input / output data of each layer of the neural network that is trained or used in conjunction with one or more embodiments while the input / output data and / or weight parameters are propagated forward during training and / or inference using the aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 601 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.
[0028] In at least one embodiment, any portion of the code and / or data storage 601 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 601 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 601 is internal or external to, for example, a processor, or the choice of including DRAM, SRAM, flash, or some other type of storage, may be determined according to the on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions being executed, the batch size of the data used in neural network inference and / or training, or any combination of these factors.
[0029] In at least one embodiment, the inference and / or training logic 615 may include, without limitation, backward propagation and / or output weights corresponding to neurons or layers of a neural network that are trained and / or used for inference in one or more embodiments, and / or code and / or data storage 605 for storing input / output data. In at least one embodiment, the code and / or data storage 605 stores the weight parameters and / or input / output data of each layer of a neural network that is trained or used in conjunction with one or more embodiments while backpropagating the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 605 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 605 has weights and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage 605 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. In at least one embodiment, any portion of the code and / or data storage 605 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 605 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether the code and / or data storage 605 is, for example, internal or external to the processor, or whether it includes DRAM, SRAM, flash, or some other type of storage, may be determined according to the on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions being executed, the batch size of the data used in the neural network inference and / or training, or any combination of these factors.
[0030] In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be separate storage structures. In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be the same storage structure. In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of the code and / or data storage 601 and the code and / or data storage 605 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.
[0031] In at least one embodiment, the inference and / or training logic 615 includes, without limitation, one or more arithmetic logic units (ALUs) 610 including integer and / or floating point units to perform logical and / or arithmetic operations based at least in part on and / or indicated by training and / or inference code (e.g., graph code), the result of which may generate activations (e.g., output values from layers or neurons within a neural network) stored in the activation storage 620, which are functions of the data of the input / output and / or weight parameters stored in the code and / or data storage 601 and / or the code and / or data storage 605. In at least one embodiment, the activations stored in the activation storage 620 are generated according to linear algebra calculations and / or matrix-based calculations performed by the ALU 610 in response to executing instructions or other code, where the weight values stored in the code and / or data storage 605 and / or the code and / or data storage 601 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data storage 605, or the code and / or data storage 601, or another storage on-chip or off-chip.
[0032] In at least one embodiment, the ALU 610 is included within one or more processors, or other hardware logic devices or circuits, but in another embodiment, the ALU 610 may be external to the processor or other hardware logic device or circuit using them (e.g., a coprocessor). In at least one embodiment, the ALU 610 may be included within the execution unit of a processor, or may be distributed among execution units of processors that are either within the same processor or different processors of a different type (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, the ALU 610 may be otherwise included within an ALU bank accessible by an execution unit of a processor, which may be either within the same processor or different processors of a different type (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, the code and / or data storage 601, the code and / or data storage 605, and the activation storage 620 may be in the same processor or other hardware logic device or circuit, and in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in any combination of the same processor or other hardware logic device or circuit and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activation storage 620 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. Further, the inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuit, and may be fetched and / or processed using the fetch, decode, scheduling, execution, retirement, and / or other logic circuits of the processor.
[0033] In at least one embodiment, the activation storage 620 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation storage 620 may be fully or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the selection of whether the activation storage 620 is, for example, inside or outside a processor, or the selection of including DRAM, SRAM, flash, or some other type of storage, may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being executed, batch size of data used in neural network inference and / or training, or any combination of these factors. In at least one embodiment, the inference and / or training logic 615 shown in FIG. 6A may be used in conjunction with an application-specific integrated circuit (ASIC) such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., “Lake Crest”) processor from Intel Corp. In at least one embodiment, the inference and / or training logic 615 shown in FIG. 6A may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field programmable gate array (FPGA).
[0034] FIG. 6B shows inference and / or training logic 615 according to at least one or more embodiments. In at least one embodiment, the inference and / or training logic 615 may include, without limiting to hardware logic, in which computing resources are dedicated to one or more layers of neurons in a neural network for weight values or other information, or otherwise only used in conjunction with them. In at least one embodiment, the inference and / or training logic 615 shown in FIG. 6B may be used in conjunction with an application specific integrated circuit (ASIC) such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., “Lake Crest”) processor from Intel Corporation. In at least one embodiment, the inference and / or training logic 615 shown in FIG. 6B may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (“GPU”) hardware, or a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 615 includes, without limitation, code and / or data storage 601, and code and / or data storage 605, and uses these to store code (e.g., graph code), weight values, and / or bias values, gradient information, momentum values, and / or other information including other parameters or hyperparameter information. In at least one embodiment shown in FIG. 6B, each of code and / or data storage 601 and code and / or data storage 605 is associated with dedicated computing resources such as computing hardware 602 and computing hardware 606, respectively. In at least one embodiment, each of computing hardware 602 and computing hardware 606 includes one or more ALUs that execute mathematical functions such as linear algebra functions only on the information stored in code and / or data storage 601 and code and / or data storage 605, respectively, and the results are stored in activation storage 620.
[0035] In at least one embodiment, each of code and / or data storage 601 and 605, and corresponding computing hardware 602 and 606, respectively corresponds to different layers of a neural network, such that the activation resulting from one "storage / compute pair 601 / 602" of code and / or data storage 601 and computing hardware 602 is provided as an input to a "storage / compute pair 605 / 606" of code and / or data storage 605 and computing hardware 606 in order to reflect the conceptual organization of the neural network. In at least one embodiment, storage / compute pairs 601 / 602 and 605 / 606 may correspond to two or more layers of a neural network. In at least one embodiment, additional storage / compute pairs (not shown) may be included in inference and / or training logic 615 after or in parallel with storage / compute pairs 601 / 602 and 605 / 606.
[0036] Data center FIG. 7 shows an exemplary data center 700 in which at least one embodiment may be used. In at least one embodiment, data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0037] As shown in FIG. 7, in at least one embodiment, the data center infrastructure layer 710 may include a resource orchestrator 712, grouped computing resources 714, and node computing resources (“node C.R.”) 716(1) to 716(N), where “N” represents any positive integer. In at least one embodiment, the node C.R. 716(1) to 716(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (solid state drive or disk drive), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, but are not limited thereto. In at least one embodiment, one or more of the node C.R. 716(1) to 716(N) may be a server having one or more of the computing resources described above.
[0038] In at least one embodiment, the grouped computing resources 714 may include separate groups of node C.R.s housed within one or more racks (not shown), or multiple racks housed in a data center at various graphical locations (also not shown). Separate groups of node C.R.s within the grouped computing resources 714 may include grouped compute resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s that include a CPU or processor may be grouped within one or more racks to provide compute resources for supporting one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0039] In at least one embodiment, the resource orchestrator 712 may configure or otherwise control one or more node C.R.s 716(1) - 716(N) and / or the grouped computing resources 714. In at least one embodiment, the resource orchestrator 712 may include a software design infrastructure ("SDI") management entity for the data center 700. In at least one embodiment, the resource orchestrator may include hardware, software, or some combination thereof.
[0040] In at least one embodiment shown in FIG. 7, the framework layer 720 includes a job scheduler 722, a configuration manager 724, a resource manager 726, and a distributed file system 728. In at least one embodiment, the framework layer 720 may include a framework for supporting software 732 of the software layer 730 and / or one or more applications 742 of the application layer 740. In at least one embodiment, the software 732 or the application 742 may each include web-based service software or an application, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 may be a kind of free and open-source software web application framework, such as Apache Spark (registered trademark) (hereinafter referred to as "Spark"), which can use the distributed file system 728 for large-scale data processing (for example, "big data"), but is not limited thereto. In at least one embodiment, the job scheduler 722 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 700. In at least one embodiment, the configuration manager 724 may be able to configure different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 728 for supporting large-scale data processing. In at least one embodiment, the resource manager 726 may be able to manage clustered or grouped computing resources mapped or allocated to support the distributed file system 728 and the job scheduler 722. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 714 in the data center infrastructure layer 710.In at least one embodiment, the resource manager 726 may manage these mapped or allocated computing resources in cooperation with the resource orchestrator 712.
[0041] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the node C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distribution file system 728 of the framework layer 720. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.
[0042] In at least one embodiment, the application 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the node C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distribution file system 728 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, recognition computing, and software for training or inference, machine learning applications including machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0043] In at least one embodiment, any one of the configuration manager 724, the resource manager 726, and the resource orchestrator 712 may implement any number and type of self-corrective measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-corrective measures may prevent a data center operator of the data center 700 from determining a configuration that may be defective and may eliminate parts of the data center that are not being fully utilized and / or have low performance.
[0044] In at least one embodiment, the data center 700 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 700. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 700 by using weight parameters calculated by one or more techniques described herein.
[0045] In at least one embodiment, the data center may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the resources described above. Further, the one or more software and / or hardware resources described above may be configured as a service to enable a user to perform training or inference of information, such as image recognition, speech recognition, or other artificial intelligence services.
[0046] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in system FIG. 7 for inference or prediction operations, based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0047] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently.
[0048] Computer system FIG. 8 is a block diagram illustrating an exemplary computer system, which may be a system 800 having interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof, formed with a processor that may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, computer system 800 may include, without limitation, components such as processor 802 for using an execution unit that includes logic for executing an algorithm for processing data according to the present disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 800 may include a processor such as a PENTIUM® processor family, XeonTM, Itanium® , XScaleTM and / or StrongARMTM, Intel® CoreTM, or Intel® NervanaTM microprocessor available from Intel Corporation of Santa Clara, California, although other systems (including PCs having other microprocessors, engineering workstations, set-top boxes, etc.) may be used. In at least one embodiment, computer system 800 may execute a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may be used.
[0049] Embodiments may be used in other devices such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (DSP), a system-on-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of executing one or more instructions according to at least one embodiment.
[0050] In at least one embodiment, computer system 800 may include, without limitation, a processor 802, which may include, without limitation, one or more execution units 808 for performing training and / or inference of a machine learning model according to the techniques described herein. In at least one embodiment, computer system 800 is a single-processor desktop or server system, but in another embodiment, computer system 800 may be a multi-processor system. In at least one embodiment, processor 802 may include, without limitation, a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 802 may be coupled to a processor bus 810, which may transmit digital signals between processor 802 and other components within computer system 800.
[0051] In at least one embodiment, processor 802 may include, without limitation, a level 1 (L1) internal cache memory (cache) 804. In at least one embodiment, processor 802 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to processor 802. Other embodiments may also include a combination of both internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, register file 806 may store different types of data in various registers including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers.
[0052] In at least one embodiment, the execution unit 808, which includes, without limitation, logic for performing integer and floating point operations, is also in the processor 802. In at least one embodiment, the processor 802 may also include a read only memory ("ROM") for storing microcode ("u-code") for certain macro instructions. In at least one embodiment, the execution unit 808 may include logic for handling a packed instruction set 809. In at least one embodiment, by including the packed instruction set 809 in the instruction set of the general purpose processor 802 along with the associated circuitry for executing the instructions, operations used by many multimedia applications can be executed using the packed data of the general purpose processor 802. In one or more embodiments, by performing operations on packed data using the full width of the processor's data bus, many multimedia applications can be accelerated and executed more efficiently, thereby eliminating the need to transfer smaller units of data between the processor's data buses to perform one or more operations on one data element at a time.
[0053] In at least one embodiment, the execution unit 808 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 800 may include, without limitation, a memory 820. In at least one embodiment, the memory 820 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other memory device. In at least one embodiment, the memory 820 may store instructions 819 and / or data 821 represented by data signals that may be executed by the processor 802.
[0054] In at least one embodiment, the system logic chip may be coupled to the processor bus 810 and the memory 820. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (“MCH”) 816, and the processor 802 may communicate with the MCH 816 via the processor bus 810. In at least one embodiment, the MCH 816 may provide a high bandwidth memory path 818 to the memory 820 for storing instructions and data, and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 816 may direct data signals between the processor 802, the memory 820, and other components of the computer system 800, and may bridge data signals between the processor bus 810, the memory 820, and the system I / O 822. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 816 may be coupled to the memory 820 via the high bandwidth memory path 818, and the graphics / video card 812 may be coupled to the MCH 816 via an Accelerated Graphics Port (“AGP”) interconnect 814.
[0055] In at least one embodiment, computer system 800 may use a system I / O 822, which is a proprietary hub interface bus for coupling MCH 816 to an I / O controller hub ("ICH") 830. In at least one embodiment, ICH 830 may provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 820, a chipset, and processor 802. By way of example, a legacy I / O controller 823 including an audio controller 829, a firmware hub ("Flash BIOS") 828, a wireless transceiver 826, a data storage 824, a user input and keyboard interface 825, a serial expansion port 827 such as a Universal Serial Bus ("USB"), and a network controller 834 may be included, without limitation. Data storage 824 may comprise a hard disk drive, a floppy (registered trademark) disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0056] In at least one embodiment, FIG. 8 shows a system including interconnected hardware devices or "chips", while in other embodiments, FIG. 8 may show an exemplary system on chip ("SoC"). In at least one embodiment, the devices shown in FIG. cc may be interconnected by a proprietary interconnect, a standard interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 800 may be interconnected using a Compute Express Link (CXL) interconnect.
[0057] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the system of FIG. 8 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations of the neural networks described herein, the functionality and / or architecture of the neural networks, or the use cases of the neural networks.
[0058] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0059] FIG. 9 is a block diagram showing an electronic device 900 for utilizing a processor 910, according to at least one embodiment. In at least one embodiment, the electronic device 900 may be, for example, without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0060] In at least one embodiment, system 900 may include, without limitation, a processor 910 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 910 is coupled using a bus or interface such as a 1°C bus, a System Management Bus (“SMBus”), a Low Pin Count (“LPC”) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advance Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, FIG. 9 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 9 may show an exemplary System on Chip (“SoC”). In at least one embodiment, the devices shown in FIG. 9 may be interconnected using proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 9 may be interconnected using a Compute Express Link (CXL) interconnect.
[0061] In at least one embodiment, FIG. 9 shows a display 924, a touch screen 925, a touch pad 930, a Near Field Communications unit (NFC) 945, a sensor hub 940, a thermal sensor 946, an Express Chipset (EC) 935, a Trusted Platform Module (TPM) 938, a BIOS / firmware / flash memory (BIOS, FW flash) 922, a DSP 960, a drive 920 such as a Solid State Disk (SSD) or a Hard Disk Drive (HDD), a wireless local area network unit (WLAN) 950, a Bluetooth unit 952, a Wireless Wide Area Network unit (WWAN) 956, a Global Positioning System (GPS) 955, a camera such as a USB3.0 camera (USB3.0 camera) 954, and / or a Low Power Double Data Rate (LPDDR) memory unit (LPDDR3) 915 implemented, for example, in accordance with the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0062] In at least one embodiment, other components may be communicatively coupled to the processor 910 via the components as described above. In at least one embodiment, an accelerometer 941, an ambient light sensor ("ALS") 942, a compass 943, and a gyroscope 944 may be communicatively coupled to the sensor hub 940. In at least one embodiment, a thermal sensor 939, a fan 937, a keyboard 946, and a touch pad 930 may be communicatively coupled to the EC 935. In at least one embodiment, a speaker 963, headphones 964, and a microphone ("mic") 965 may be communicatively coupled to an audio unit ("audio codec and class D amplifier") 962, and this audio unit may be communicatively coupled to the DSP 960. In at least one embodiment, the audio unit 964 may include, for example and without limitation, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, a SIM card ("SIM") 957 may be communicatively coupled to the WWAN unit 956. In at least one embodiment, components such as the WLAN unit 950 and the Bluetooth unit 952, as well as the WWAN 956, may be implemented in a next generation form factor ("NGFF": Next Generation Form Factor).
[0063] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in FIG. 9 of the system diagram for inference or prediction operations, at least partially based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0064] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, where each of the two or more versions of the image is synthetically and independently generated.
[0065] FIG. 10 shows a computer system 1000 according to at least one embodiment. In at least one embodiment, the computer system 1000 is configured to implement the various processes and methods described throughout this disclosure.
[0066] In at least one embodiment, computer system 1000 includes at least one central processing unit ("CPU") 1002 connected to a communication bus 1010 implemented using any suitable protocol, such as, without limitation, Peripheral Component Interconnect ("PCI"), Peripheral Component Interconnect Express ("PCI-Express"), Accelerated Graphics Port ("AGP"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1000 includes a main memory 1004 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in main memory 1004, which may take the form of random access memory ("RAM"). In at least one embodiment, a network interface subsystem ("network interface") 1022 provides an interface to other computing devices and networks for receiving data from and transmitting data to other systems from computer system 1000.
[0067] In at least one embodiment, computer system 1000 includes, without limitation in at least one embodiment, an input device 1008, a parallel processing system 1012, and a display device 1006 that may be implemented using a conventional cathode ray tube (CRT), liquid crystal display (LCD), light emitting diode (LED), plasma display, or other suitable display technology. In at least one embodiment, user input is received from an input device 1008 such as a keyboard, mouse, touch pad, microphone, etc. In at least one embodiment, each of the above modules may be located on a single semiconductor platform to form a processing system.
[0068] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in system FIG. 10 for inference or prediction operations based at least in part on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0069] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently.
[0070] FIG. 11 shows a computer system 1100 according to at least one embodiment. In at least one embodiment, the computer system 1100 includes, without limitation, a computer 1110 and a USB stick 1120. In at least one embodiment, the computer 1110 may include, without limitation, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, the computer 1110 includes, without limitation, a server, a cloud instance, a laptop, and a desktop computer.
[0071] In at least one embodiment, the USB stick 1120 includes, without limitation, a processing unit 1130, a USB interface 1140, and USB interface logic 1150. In at least one embodiment, the processing unit 1130 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1130 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1130 includes an application-specific integrated circuit (``ASIC'') optimized to perform any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing core 1130 is a tensor processing unit (``TPC'') optimized to perform machine learning inference operations. In at least one embodiment, the processing core 1130 is a vision processing unit (``VPU'') optimized to perform machine vision and machine learning inference operations.
[0072] In at least one embodiment, the USB interface 1140 can be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1140 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1140 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 1150 can include any amount and type of logic that enables the processing unit 1130 to interface with a device (e.g., computer 1110) via the USB connector 1140.
[0073] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in FIG. 11 of the system for inference or prediction operations, at least partially based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0074] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks, at least partially based on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0075] FIG. 12A shows an exemplary architecture in which a plurality of GPUs 1210-1213 are communicatively coupled to a plurality of multi-core processors 1205-1206 via high-speed links 1240-1243 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 1240-1243 support a communication throughput of 4 GB / second, 30 GB / second, 80 GB / second, or more. Various interconnect protocols may be used, including but not limited to PCIe 4.0 or 5.0, and NVLink 2.0.
[0076] Furthermore, in one embodiment, two or more of the GPUs 1210-1213 are interconnected via high-speed links 1229-1230, which may be implemented using the same or different protocols / links as those used for the high-speed links 1240-1243. Similarly, two or more of the multi-core processors 1205-1206 may be connected via high-speed link 1228, which can be a symmetric multi-processor (SMP) bus operating at 20 GB / second, 30 GB / second, 120 GB / second, or more. Alternatively, all communication between the various system components shown in FIG. 12A may be realized using the same protocol / link (e.g., via a common interconnect fabric).
[0077] In one embodiment, each of the multi-core processors 1205-1206 is communicatively coupled to the processor memories 1201-1202 via the memory interconnects 1226-1227, respectively, and each of the GPUs 1210-1213 is communicatively coupled to the GPU memories 1220-1223 via the GPU memory interconnects 1250-1253, respectively. The memory interconnects 1226-1227 and 1250-1253 may utilize the same or different memory access technologies. By way of example and not limitation, the processor memories 1201-1202 and the GPU memories 1220-1223 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics double data rate SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories such as 3D XPoint or Nano-Ram. In one embodiment, some portions of the processor memories 1201-1202 may be volatile memories and other portions may be non-volatile memories (e.g., using a two-level memory (2LM) hierarchy).
[0078] As described below, the various processors 1205-1206 and GPUs 1210-1213 may each be physically coupled to specific memories 1201-1202, 1220-1223, but an integrated memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. For example, each of the processor memories 1201-1202 may each have a 64 GB system memory address space, and each of the GPU memories 1220-1223 may each have a 32 GB system memory address space (in this example, a total of 256 GB of addressable memory is obtained).
[0079] FIG. 12B shows further details of the interconnection between a multi-core processor 1207 and a graphics acceleration module 1246 according to one exemplary embodiment. The graphics acceleration module 1246 may include one or more GPU chips integrated on a line card coupled to the processor 1207 via a high-speed link 1240. Alternatively, the graphics acceleration module 1246 may be integrated in the same package or chip as the processor 1207.
[0080] In at least one embodiment, the processor 1207 shown includes a plurality of cores 1260A - 1260D, each core having a translation lookaside buffer 1261A - 1261D and one or more caches 1262A - 1262D. In at least one embodiment, the cores 1260A - 1260D may include various other components (not shown) for executing instructions and processing data. The caches 1262A - 1262D may comprise level 1 (L1) and level 2 (L2) caches. Further, one or more shared caches 1256 may be included in the caches 1262A - 1262D and may be shared by a set of the cores 1260A - 1260D. For example, one embodiment of the processor 1207 includes 24 cores, each core having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more of the L2 and L3 caches are shared by two adjacent cores. The processor 1207 and the graphics acceleration module 1246 are connected to a system memory 1214, which may include the processor memories 1201 - 1202 of FIG. 12A.
[0081] For the data and instructions stored in the various caches 1262A - 1262D, 1256, and the system memory 1214, coherence is maintained through inter - core communication via the coherence bus 1264. For example, each cache may have associated cache coherence logic / circuitry to communicate via the coherence bus 1264 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented via the coherence bus 1264 to monitor cache accesses.
[0082] In one embodiment, a proxy circuit 1225 communicatively couples the graphics acceleration module 1246 to the coherence bus 1264 so that the graphics acceleration module 1246 can participate in the cache coherence protocol as a peer of cores 1260A - 1260D. In particular, interface 1235 provides a connection to the proxy circuit 1225 via a high - speed link 1240 (e.g., a PCIe bus, NVLink, etc.), and interface 1237 connects the graphics acceleration module 1246 to the link 1240.
[0083] In one implementation form, the accelerator integration circuit 1236 provides services for cache management, memory access, content management, and interrupt management instead of the plurality of graphics processing engines 1231 to 1232 of the graphics acceleration module 1246 and N. Each of the graphics processing engines 1231 to 1232 and N may include a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1231, 1232, and N may include different types of graphics processing engines such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine within the GPU. In at least one embodiment, the graphics acceleration module 1246 may be a GPU having a plurality of graphics processing engines 1231 to 1232, or the graphics processing engines 1231 to 1232 may be individual GPUs integrated in a common package, line card, or chip.
[0084] In one embodiment, the accelerator integration circuit 1236 includes a memory management unit (MMU) 1239 for performing various memory management functions, such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation), and a memory access protocol for accessing the system memory 1214. The MMU 1239 can also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In one implementation, the cache 1238 stores commands and data so that the graphics processing engines 1231-1232 can access it efficiently. In one embodiment, the data stored in the cache 1238 and the graphics memories 1233-1234, M are kept coherent with the core caches 1262A-1262D, 1256, and the system memory 1214. As described above, this may be achieved via the proxy circuit 1225 instead of the cache 1238 and the memories 1233-1234, M (e.g., by sending updates regarding cache line modifications / accesses in the processor caches 1262A-1262D, 1256 to the cache 1238 and receiving updates from the cache 1238).
[0085] The set of registers 1245 stores context data for the threads executed by the graphics processing engines 1231-1232, N, and the context management circuit 1248 manages thread contexts. For example, the context management circuit 1248 may perform save and restore operations to save and restore the contexts of various threads during a context switch (e.g., here, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1248 may store the current register values in a designated area of memory (identified, e.g., by a context pointer). Then, when returning to the context, the context management circuit 1248 may restore the register values. In one embodiment, the interrupt management circuit 1247 receives and processes interrupts received from system devices.
[0086] In one implementation, the virtual / effective addresses from the graphics processing engine 1231 are translated by the MMU 1239 into real / physical addresses of the system memory 1214. One example of the accelerator integration circuit 1236 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 1246 and / or other accelerator devices. The graphics accelerator module 1246 may be dedicated to a single application executed on the processor 1207 or may be shared among multiple applications. In one embodiment, there is a virtualized graphics execution environment in which the resources of the graphics processing engines 1231-1232, N are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.
[0087] In at least one embodiment, the accelerator integration circuit 1236 functions as a bridge to the system for the graphics acceleration module 1246 and provides address translation and system memory cache services. Further, the accelerator integration circuit 1236 may provide a virtualization facility for the host processor to manage the graphics processing engines 1231-1232, N virtualization, interrupts, and memory management.
[0088] Since the hardware resources of the graphics processing engines 1231-1232, N are explicitly mapped to the physical address space seen by the host processor 1207, any host processor can directly address these resources using the effective address value. One function of the accelerator integration circuit 1636 is, in one embodiment, to physically separate the graphics processing engines 1231-1232, N so that they appear as independent units to the system.
[0089] In at least one embodiment, each of one or more graphics memories 1233-1234, M is coupled to a respective one of the graphics processing engines 1231-1232, N. The graphics memories 1233-1234, M store instructions and data processed by the respective graphics processing engines 1231-1232, N. The graphics memories 1233-1234, M may be volatile memories such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memories such as 3D XPoint or Nano-Ram.
[0090] In one embodiment, to reduce data traffic through link 1240, data stored in graphics memories 1233-1234, M is made to be the data most frequently used by graphics processing engines 1231-1232, N, and preferably not used (or at least not frequently used) by cores 1260A-1260D. Similarly, the bias mechanism attempts to keep data required by the cores (and thus preferably not required by graphics processing engines 1231-1232, N) in caches 1262A-1262D, 1256 of the cores, and system memory 1214.
[0091] FIG. 12C shows another exemplary embodiment in which the accelerator integration circuit 1236 is integrated within the processor 1207. At least in this embodiment, the graphics processing engines 1231-1232, N communicate directly with the accelerator integration circuit 1236 via the high-speed link 1240 through interfaces 1237 and 1235 (any form of bus or interface protocol may also be utilized in this case). The accelerator integration circuit 1236 may perform the same operations as described with respect to FIG. 12B, but considering its proximity to the coherence bus 1264 and caches 1262A-1262D, 1256, it may potentially operate at a higher throughput. At least one embodiment supports different programming models including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1236 and a programming model controlled by the graphics acceleration module 1246.
[0092] In at least one embodiment, the graphics processing engines 1231-1232, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can centralize other application requirements to the graphics processing engines 1231-1232, N to achieve virtualization within a VM / partition.
[0093] In at least one embodiment, the graphics processing engines 1231-1232, N may be shared by multiple VM / application partitions. In at least one embodiment, the shared model may use a system hypervisor to virtualize the graphics processing engines 1231-1232, N to enable access by each operating system. In a single partition system without a hypervisor, the graphics processing engines 1231-1232, N are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1231-1232, N to provide access to each process or application.
[0094] In at least one embodiment, the graphics acceleration module 1246 or individual graphics processing engines 1231-1232, N use a process handle to select process elements. In at least one embodiment, the process elements are stored in the system memory 1214 and can be addressed using the translation technique from the effective address to the physical address described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the context of the host process with the graphics processing engines 1231-1232, N (i.e., calling system software to add a process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element within the process element link list.
[0095] FIG. 12D shows an exemplary accelerator integration slice 1290. As used herein, a "slice" comprises a designated portion of the processing resources of an accelerator integration circuit 1236. An application effective address space 1282 within system memory 1214 stores a process element 1283. In one embodiment, the process element 1283 is stored in response to a GPU call 1281 from an application 1280 executing on a processor 1207. The process element 1283 accommodates the process state of the corresponding application 1280. A work descriptor (WD) 1284 accommodated by the process element 1283 can be a single job required by the application, or can accommodate a pointer to a queue of jobs. In at least one embodiment, the WD 1284 is a pointer to a job request queue in the application's address space 1282.
[0096] The graphics acceleration module 1246 and / or individual graphics processing engines 1231-1232, N can be shared by all or a subset of the processes within the system. In at least one embodiment, infrastructure for setting the process state and transmitting the WD 1284 to the graphics acceleration module 1246 to initiate a job in a virtualized environment may be included.
[0097] In at least one embodiment, a dedicated process programming model is implementation specific. In this model, a single process owns the graphics acceleration module 1246 or an individual graphics processing engine 1231. Since the graphics acceleration module 1246 is owned by a single process, when the graphics acceleration module 1246 is allocated, the hypervisor initializes the accelerator integration circuit 1236 for the owning partition, and the operating system initializes the accelerator integration circuit 1236 for the owning process.
[0098] During operation, the WD fetch unit 1291 within the accelerator integration slice 1290 fetches the next WD 1284, which includes the display of work to be performed by one or more graphics processing engines of the graphics acceleration module 1246. As shown, the data from the WD 1284 is stored in the register 1245 and may be used by the MMU 1239, the interrupt management circuit 1247, and / or the context management circuit 1248. For example, one embodiment of the MMU 1239 includes a segment / page walk circuit for accessing the segment / page table 1286 within the OS virtual address space 1285. The interrupt management circuit 1247 may process the interrupt event 1292 received from the graphics acceleration module 1246. When executing a graphics operation, the effective address 1293 generated by the graphics processing engines 1231 - 1232, N is translated to a physical address by the MMU 1239.
[0099] In one embodiment, a set of the same registers 1245 is replicated for each of the graphics processing engines 1231 - 1232, N, and / or the graphics acceleration module 1246 and may be initialized by the hypervisor or the operating system. Each of these replicated registers may be included in the accelerator integration slice 1290. Exemplary registers that may be initialized by the hypervisor are shown in Table 1.
Table 1
[0100] Exemplary registers that may be initialized by the operating system are shown in Table 2.
Table 2
[0101] In one embodiment, each WD1284 is specific to a particular graphics acceleration module 1246 and / or graphics processing engines 1231-1232, N. The WD1284 can contain all the information required for the graphics processing engines 1231-1232, N to perform work, or can be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0102] FIG. 12E shows further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor real address space 1298 in which a process element list 1299 is stored. The hypervisor real address space 1298 is accessible via a hypervisor 1296 that virtualizes the graphics acceleration module engine of the operating system 1295.
[0103] In at least one embodiment, the shared programming model enables all or a subset of processes from all or a subset of partitions within the system to use the graphics acceleration module 1246. There are two programming models in which the graphics acceleration module 1246 is shared by multiple processes and partitions, namely time slice sharing and graphics-directed shared.
[0104] In this model, the system hypervisor 1296 owns the graphics acceleration module 1246 and makes its functions available to all operating systems 1295. In order for the graphics acceleration module 1246 to support the virtualization by the system hypervisor 1296, the graphics acceleration module 1246 may comply with the following: 1) The job requests of the application must be autonomous (i.e., there is no need to maintain the state between jobs), or the graphics acceleration module 1246 must provide a mechanism for saving and restoring the context. 2) The job requests of the application are guaranteed by the graphics acceleration module 1246 to be completed within the specified amount of time, including any translation errors, or the graphics acceleration module 1246 provides a function to preempt the processing of the job. 3) When the graphics acceleration module 1246 operates in the specified shared programming model, fairness must be guaranteed among processes.
[0105] In at least one embodiment, application 1280 needs to make a system call to operating system 1295, along with the type of graphics acceleration module 1246, work descriptor (WD), authority mask register (AMR) value, and context save / restore area pointer (CSRP). In at least one embodiment, the type of graphics acceleration module 1246 describes the acceleration function targeted by the system call. In at least one embodiment, the type of graphics acceleration module 1246 may be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 1246 and can be in the form of a command for graphics acceleration module 1246, a virtual address pointer pointing to a user-defined structure, a virtual address pointer pointing to a command queue, or any other data structure for describing the work performed by graphics acceleration module 1246. In one embodiment, the AMR value is the AMR state for use by the current process. In at least one embodiment, the value passed to the operating system is the same as the application that sets the AMR. If the implementation of accelerator integration circuit 1236 and graphics acceleration module 1246 does not support the user authority mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value and then pass the AMR to the hypervisor call. Hypervisor 1296 may optionally apply the current authority mask override register (AMOR) value and then place the AMR in process element 1283. In at least one embodiment, the CSRP is one of registers 1245 that holds the virtual address of an area within the virtual address space 1282 of the application for graphics acceleration module 1246 to save and restore the context state. This pointer is optional if there is no need to save any state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.
[0106] Upon receiving a system call, the operating system 1295 may verify that the application 1280 is registered and has been granted the right to use the graphics acceleration module 1246. Next, the operating system 1295 calls the hypervisor 1296 with the information shown in Table 3.
Table 3
[0107] Upon receiving a hypervisor call, the hypervisor 1296 verifies that the operating system 1295 is registered and has been granted the right to use the graphics acceleration module 1246. Next, the hypervisor 1296 inserts the process element 1283 into the process element link list of the corresponding type of graphics acceleration module 1246. The process element may include the information shown in Table 4.
Table 4
[0108] In at least one embodiment, the hypervisor initializes the registers 1245 of the plurality of accelerator integration slices 1290.
[0109] As shown in FIG. 12F, in at least one embodiment, an integrated memory is used that is addressable via a common virtual memory address space used to access physical processors 1201-1202 and GPU memories 1220-1223. In this implementation, operations executed by GPUs 1210-1213 utilize the same virtual / effective memory address space as accessing physical processors 1201-1202, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to physical processor 1201, a second portion is allocated to a second physical processor 1202, and a third portion is allocated to GPU memory 1220, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thereby distributed across each of physical processors 1201-1202 and GPU memories 1220-1223 such that any processor or GPU can access any physical memory with virtual addresses mapped to physical memory.
[0110] In one embodiment, bias / coherence management circuits 1294A-1294E in one or more of MMUs 1239A-1239E ensure cache coherence between the caches of one or more host processors (e.g., 1205) and the caches of GPUs 1210-1213 and implement a bias technique to indicate the physical memory in which a particular type of data should be stored. Multiple instances of bias / coherence management circuits 1294A-1294E are shown in FIG. 12F, but the bias / coherence circuit may be implemented within the MMU of one or more host processors 1205 and / or within accelerator integration circuit 1236.
[0111] One embodiment enables the mapping of the GPUs' memories 1220 - 1223 as part of the system memory and makes them accessible using the Shared Virtual Memory (SVM) technique, without incurring a performance degradation associated with full system cache coherence. In at least one embodiment, the GPUs' memories 1220 - 1223 are accessible as system memory without cumbersome cache coherence overhead, providing a beneficial operating environment for GPU offloading. This configuration enables the host processor 1205 software to set operands and access computation results without the overhead of conventional I / O DMA data copies. Such conventional copies require driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access the GPUs' memories 1220 - 1223 without cache coherence overhead may be essential to the execution time of offloaded computations. For example, in the presence of significant streaming write memory traffic, the cache coherence overhead may significantly reduce the effective write bandwidth seen by the GPUs 1210 - 1213. In at least one embodiment, the efficiency of operand setting, access to results, and GPU computation may help in determining the effectiveness of GPU offloading.
[0112] In at least one embodiment, the selection of the GPU bias and the host processor bias is determined by a bias tracker data structure. For example, a bias table may be used, and this table may have a page granularity structure that includes one or two bits per memory page with a GPU (i.e., it may be controlled at the granularity of the memory page). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPUs with memory 1220 - 1223 in a state where a bias cache (for example, to cache frequently used / recently used entries of the bias table) is present or not present in GPUs 1210 - 1213. Alternatively, the entire bias table may be maintained within the GPU.
[0113] In at least one embodiment, an entry of the bias table associated with each access to GPUs with memory 1220 - 1223 is accessed prior to the actual access to the GPU memory, resulting in the following operations. First, local requests from GPUs 1210 - 1213 that find their pages within the GPU bias are transferred directly to the corresponding GPUs with memory 1220 - 1223. Local requests from GPUs that find their pages within the host bias are transferred to the processor 1205 (for example, via a high - speed link as described above). In one embodiment, requests from the processor 1205 that find the requested page within the host processor bias complete the request in the same manner as a normal memory read. Alternatively, requests directed to a GPU - biased page may be transferred to GPUs 1210 - 1213. In at least one embodiment, then, if the GPU is not currently using the page, it may migrate the page to the host processor bias. In at least one embodiment, the bias state of a page can be changed by either a software - based mechanism, a software - based mechanism with hardware assistance, or, for a limited set of cases, simply a hardware - based mechanism.
[0114] One mechanism for changing the bias state utilizes an API call (e.g., OpenCL), where this API call calls the GPU's device driver, and this device driver sends a message to the GPU (or adds a command descriptor to a queue) to change the bias state, and for some transitions, guides the GPU to perform a cache flushing operation on the host. In at least one embodiment, the cache flushing operation is used for the transition from the host processor 1205's bias to the GPU bias, but not for the opposite transition.
[0115] In one embodiment, cache coherence is maintained by the host processor 1205 temporarily rendering GPU-biased pages that cannot be cached. To access these pages, the processor 1205 may request access from the GPU 1210, and the GPU 1210 may either immediately grant access or not. Thus, to reduce communication between the processor 1205 and the GPU 1210, it is beneficial to make GPU-biased pages be requested by the GPU but not by the host processor 1205, or vice versa.
[0116] Inference and / or training logic 615 is used to execute one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B.
[0117] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, where each of the two or more versions of the image is synthetically generated independently.
[0118] FIG. 13 shows an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general purpose processor cores.
[0119] FIG. 13 is a block diagram showing an exemplary system-on-chip integrated circuit 1300 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1300 includes one or more application processors 1305 (e.g., CPUs), at least one graphics processor 1310, and may further include an image processor 1315 and / or a video processor 1320, any of which may be modular IP cores. In at least one embodiment, integrated circuit 1300 includes peripheral devices or bus logic including a USB controller 1325, a UART controller 1330, an SPI / SDIO controller 1335, and an I 2 S / I 2 2C controller 1340. In at least one embodiment, integrated circuit 1300 can include a display device 1345 coupled to one or more of a high-definition multimedia interface (HDMI™) controller 1350 and a mobile industry processor interface (MIPI) display interface 1355. In at least one embodiment, storage may be provided by a flash memory subsystem 1360 including a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1365 to access SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits further include an embedded security engine 1370.
[0120] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the integrated circuit 1300 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the use cases of the neural networks.
[0121] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0122] FIGS. 14A - 14B illustrate an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores, according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general - purpose processor cores.
[0123] FIG. 14A and FIG. 14B are block diagrams showing exemplary graphics processors for use within a SoC according to the embodiments described herein. FIG. 14A shows an exemplary graphics processor 1410 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. FIG. 14B shows a further exemplary graphics processor 1440 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the graphics processor 1410 of FIG. 14A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1440 of FIG. 14B is a high-performance graphics processor core. In at least one embodiment, each of the graphics processors 1410, 1440 can be a variant of the graphics processor 1310 of FIG. 13.
[0124] In at least one embodiment, the graphics processor 1410 includes a vertex processor 1405 and one or more fragment processors 1415A - 1415N (e.g., 1415A, 1415B, 1415C, 1415D - 1415N - 1, and 1415N). In at least one embodiment, the graphics processor 1410 can execute different shader programs via separate logic, whereby the vertex processor 1405 is optimized to execute operations for vertex shader programs, while the one or more fragment processors 1415A - 1415N execute fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 1405 executes the vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 1415A - 1415N use the primitives and vertex data generated by the vertex processor 1405 to generate a frame buffer to be displayed on a display device. In at least one embodiment, the fragment processors 1415A - 1415N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform operations similar to pixel shader programs provided in the Direct 3D API.
[0125] In at least one embodiment, the graphics processor 1410 further includes one or more memory management units (MMUs) 1420A - 1420B, caches 1425A - 1425B, and circuit interconnects 1430A - 1430B. In at least one embodiment, one or more MMUs 1420A - 1420B provide a virtual to physical address mapping for the graphics processor 1410, including the vertex processor 1405 and / or the fragment processors 1415A - 1415N, and they may reference vertex or image / text data stored in memory in addition to vertex or image / text data stored in one or more caches 1425A - 1425B. In at least one embodiment, one or more MMUs 1420A - 1420B may be synchronized with one or more other MMUs in the system, including one or more MMUs associated with one or more of the application processors 1305, image processors 1315, and / or video processors 1320 of FIG. 13, such that each processor 1305 - 1320 can participate in a shared or integrated virtual memory system. In at least one embodiment, one or more circuit interconnects 1430A - 1430B enable the graphics processor 1410 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.
[0126] In at least one embodiment, the graphics processor 1440 includes one or more memory management units (MMUs) 1420A, 1420B, caches 1425A, 1425B, and circuit interconnects 1430A, 1430B of the graphics processor 1410 of FIG. 14A. In at least one embodiment, the graphics processor 1440 provides a unified shader core architecture in which a single core or type of core can execute all types of programmable shader code, including shader program code for implementing a vertex shader, a fragment shader, and / or a compute shader. The graphics processor 1440 includes one or more shader cores 1455A to 1455N (e.g., 1455A, 1455B, 1455C, 1455D, 1455E, 1455F, ~1455N-1, 1455N). In at least one embodiment, the number of shader cores may vary. In at least one embodiment, the graphics processor 1440 includes an inter-core task manager 1445 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1455A to 1455N, and a tiling unit 1458 for accelerating tiling operations for tile-based rendering in which the rendering operation of a scene is subdivided in the image space, for example, to utilize local space coherence within the scene or to optimize the use of internal caches.
[0127] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in integrated circuits 14A and / or 14B for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of neural networks described herein, or use cases of neural networks. Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently.
[0128] FIGS. 15A-15B show further exemplary graphics processor logic according to embodiments described herein. FIG. 15A shows a graphics core 1500, which in at least one embodiment may be included in the graphics processor 1310 of FIG. 13 and in at least one embodiment may be integrated shader cores 1455A-1455N as in FIG. 14B. FIG. 15B shows a highly parallel general-purpose graphics processing unit 1530 suitable for introduction into a multi-chip module in at least one embodiment.
[0129] In at least one embodiment, the graphics core 1500 includes a shared instruction cache 1502, a texture unit 1518, and a cache / shared memory 1520, which are common to the execution resources within the graphics core 1500. In at least one embodiment, the graphics core 1500 can include a plurality of slices 1501A - 1501N, or per-core partitions, and the graphics processor can include a plurality of instances of the graphics core 1500. The slices 1501A - 1501N can include support logic that includes local instruction caches 1504A - 1504N, thread schedulers 1506A - 1506N, thread dispatchers 1508A - 1508N, and sets of registers 1510A - 1510N. In at least one embodiment, the slices 1501A - 1501N can include a set of additional functional units (AFU1512A - 1512N), floating point units (FPU1514A - 1514N), integer arithmetic logic units (ALU1516 - 1516N), address calculation units (ACU1513A - 1513N), double precision floating point units (DPFPU1515A - 1515N), and matrix processing units (MPU1517A - 1517N).
[0130] In at least one embodiment, FPU1514A to 1514N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, and DPFPU1515A to 1515N perform double-precision (64-bit) floating-point operations. In at least one embodiment, ALU1516A to 1516N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured to perform mixed-precision operations. In at least one embodiment, MPU1517A to 1517N can also be configured to perform mixed-precision matrix operations including half-precision floating-point and 8-bit integer operations. In at least one embodiment, MPU1517A to 1517N can perform various matrix operations for accelerating machine learning application frameworks, including enabling support for accelerating general matrix-matrix multiplication (GEMM). In at least one embodiment, AFU1512A to 1512N can perform additional logical operations not supported by a floating-point unit or an integer unit, including trigonometric operations (e.g., sine, cosine, etc.).
[0131] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B. In at least one embodiment, inference and / or training logic 615 may be used in graphics core 1500 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0132] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0133] FIG. 15B shows a general-purpose processing unit (GPGPU) 1530, which in at least one embodiment can be configured to perform high-parallel computing operations by an array of graphics processing units. In at least one embodiment, GPGPU 1530 can be directly linked to other instances of GPGPU 1530 to generate a plurality of GPU clusters to improve the training speed of a deep neural network. In at least one embodiment, GPGPU 1530 includes a host interface 1532 to enable connection with a host processor. In at least one embodiment, host interface 1532 is a PCI Express interface. In at least one embodiment, host interface 1532 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, GPGPU 1530 receives commands from a host processor and uses a global scheduler 1534 to distribute the execution threads associated with these commands to a set of compute clusters 1536A - 1536H. In at least one embodiment, compute clusters 1536A - 1536H share a cache memory 1538. In at least one embodiment, cache memory 1538 can act as a high-level cache for cache memory within compute clusters 1536A - 1536H.
[0134] In at least one embodiment, the GPGPU 1530 includes memories 1544A-1544B coupled to compute clusters 1536A-1536H via a set of memory controllers 1542A-1542B. In at least one embodiment, the memories 1544A-1544B can include various types of memory devices including dynamic random access memory (DRAM), such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory, or graphics random access memory.
[0135] In at least one embodiment, each of the compute clusters 1536A-1536H includes a set of graphics cores, such as the graphics core 1500 of FIG. 15A, and the set of graphics cores can include multiple types of integer and floating point logic units capable of performing compute operations with various precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating point units in each of the compute clusters 1536A-1536H can be configured to perform 16-bit or 32-bit floating point operations, while another subset of the floating point units can be configured to perform 64-bit floating point operations.
[0136] In at least one embodiment, multiple instances of GPGPU1530 can be configured to operate as a compute cluster. In at least one embodiment, the communication used by compute clusters 1536A - 1536H for synchronization and data exchange varies across embodiments. In at least one embodiment, multiple instances of GPGPU1530 communicate via host interface 1532. In at least one embodiment, GPGPU1530 includes I / O hub 1539, which couples GPGPU1530 to GPU link 1540 that enables a direct connection to other instances of GPGPU1530. In at least one embodiment, GPU link 1540 is coupled to a dedicated GPU - to - GPU bridge that enables communication and synchronization between multiple instances of GPGPU1530. In at least one embodiment, GPU link 1540 is coupled to a high - speed interconnect for transmitting and receiving data to / from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU1530 are located in separate data processing systems and communicate via a network device accessible via host interface 1532. In at least one embodiment, GPU link 1540 can be configured to enable a connection to the host processor in addition to, or instead of, host interface 1532.
[0137] In at least one embodiment, the GPGPU 1530 can be configured to train a neural network. In at least one embodiment, the GPGPU 1530 can be used within an inference platform. In at least one embodiment where the GPGPU 1530 is used for inference, the GPGPU may include fewer compute clusters 1536A - 1536H than when the GPGPU is used for training a neural network. In at least one embodiment, the memory technology associated with memories 1544A - 1544B may be different for the inference configuration and the training configuration, and a high bandwidth memory technology is applied to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 1530 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can support one or more 8-bit integer dot product instructions, which may be used during the inference operation of a deployed neural network.
[0138] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the GPGPU 1530 for inference or prediction operations, based at least in part on the training operations of the neural network described herein, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network.
[0139] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently.
[0140] FIG. 16 is a block diagram showing a computing system 1600 according to at least one embodiment. In at least one embodiment, the computing system 1600 includes a processing subsystem 1601 having one or more processors 1602 and a system memory 1604 that communicate via an interconnect path that may include a memory hub 1605. In at least one embodiment, the memory hub 1605 may be a separate component within a chipset component or may be integrated within one or more processors 1602. In at least one embodiment, the memory hub 1605 is coupled to an I / O subsystem 1611 via a communication link 1606. In at least one embodiment, the I / O subsystem 1611 includes an I / O hub 1607 that enables the computing system 1600 to receive input from one or more input devices 1608. In at least one embodiment, the I / O hub 1607 can enable a display controller, which may be included in one or more processors 1602, to provide output to one or more display devices 1610A. In at least one embodiment, one or more display devices 1610A coupled to the I / O hub 1607 can include local, internal, or embedded display devices.
[0141] In at least one embodiment, the processing subsystem 1601 includes one or more parallel processors 1612 coupled to the memory hub 1605 via a bus or other communication link 1613. In at least one embodiment, the communication link 1613 can be one of any number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or can be a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 1612 include a computationally intensive parallel or vector processing system that can include a number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 1612 form a graphics processing subsystem that can output pixels to one of one or more display devices 1610A coupled via the I / O hub 1607. In at least one embodiment, the one or more parallel processors 1612 can also include a display controller and display interface (not shown) that enables direct connection to one or more display devices 1610B.
[0142] In at least one embodiment, the system storage unit 1614 can be connected to the I / O hub 1607 to provide a storage mechanism for the computing system 1600. In at least one embodiment, an I / O switch 1616 can be used to provide an interface mechanism for enabling communication between the I / O hub 1607 and other components such as a network adapter 1618 and / or a wireless network adapter 1619 that may be integrated into the platform, as well as various other devices that can be added via one or more add-in devices 1620. In at least one embodiment, the network adapter 1618 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1619 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radios.
[0143] In at least one embodiment, the computing system 1600 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 1607. In at least one embodiment, the communication paths interconnecting the various components of FIG. 16 may be implemented using any suitable protocol such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express), or other buses or point-to-point communication interfaces such as NV-Link high-speed interconnects, or other interconnect protocols.
[0144] In at least one embodiment, one or more parallel processors 1612 incorporate circuitry optimized for graphics and video processing, including, for example, a video output circuit, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1612 incorporate circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 1600 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1612, memory hub 1605, processor 1602, and I / O hub 1607 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 1600 may be integrated in a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 1600 can be integrated into a multi-chip module (MCM), and this module can be interconnected with other multi-chip modules to form a modular computing system.
[0145] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of FIG. 1600 for inference or prediction operations, based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0146] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0147] Processor FIG. 17A shows a parallel processor 1700 according to at least one embodiment. In at least one embodiment, the various components of parallel processor 1700 may be implemented using one or more integrated circuit devices such as programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 1700 is a variation of one or more parallel processors 1612 shown in FIG. 16 according to an exemplary embodiment.
[0148] In at least one embodiment, parallel processor 1700 includes a parallel processing unit 1702. In at least one embodiment, parallel processing unit 1702 includes an I / O unit 1704 that enables communication with other devices including other instances of parallel processing unit 1702. In at least one embodiment, I / O unit 1704 may be directly connected to other devices. In at least one embodiment, I / O unit 1704 is connected to other devices through the use of a hub or switch interface such as memory hub 1605. In at least one embodiment, the connection between memory hub 1605 and I / O unit 1704 forms a communication link 1613. In at least one embodiment, I / O unit 1704 is connected to a host interface 1706 and a memory crossbar 1716, where host interface 1706 receives commands targeted for execution of processing operations and memory crossbar 1716 receives commands targeted for execution of memory operations.
[0149] In at least one embodiment, when the host interface 1706 receives a command buffer via the I / O unit 1704, the host interface 1706 can direct a work operation for executing these commands to the front end 1708. In at least one embodiment, the front end 1708 is coupled to a scheduler 1710, and this scheduler is configured to distribute commands or other work items to the processing cluster array 1712. In at least one embodiment, the scheduler 1710 ensures that the processing cluster array 1712 is properly configured and in an effective state before tasks are distributed to the processing cluster array 1712. In at least one embodiment, the scheduler 1710 is implemented via firmware logic running on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 1710 can be configured to execute complex scheduling and work distribution operations at a coarse granularity and a fine granularity, enabling rapid preemption of threads running on the processing array 1712 and context switching. In at least one embodiment, the host software can prove the scheduling workload on the processing array 1712 via one of a plurality of graphics processing doorbells. In at least one embodiment, then, the workload can be automatically distributed across the entire processing array 1712 by the scheduler 1710 logic within the microcontroller including the scheduler 1710.
[0150] In at least one embodiment, the processing cluster array 1712 can include up to "N" processing clusters (e.g., cluster 1714A, cluster 1714B to cluster 1714N). In at least one embodiment, each of the clusters 1714A to 1714N of the processing cluster array 1712 can execute a large number of simultaneous threads. In at least one embodiment, the scheduler 1710 can use various scheduling and / or workload distribution algorithms to distribute work to the clusters 1714A to 1714N of the processing cluster array 1712, and these algorithms may vary according to the workload generated for each type of program or calculation. In at least one embodiment, the scheduling may be dynamically handled by the scheduler 1710, or may be partially assisted by the compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 1712. In at least one embodiment, the different clusters 1714A to 1714N of the processing cluster array 1712 can be distributed to process different types of programs or execute different types of calculations.
[0151] In at least one embodiment, the processing cluster array 1712 can be configured to execute various types of parallel processing operations. In at least one embodiment, the processing cluster array 1712 is configured to execute general-purpose parallel computing operations. For example, in at least one embodiment, the processing cluster array 1712 can include logic for performing processing tasks including filtering of video and / or audio data, execution of modeling operations including physical operations, and execution of data conversion.
[0152] In at least one embodiment, the processing cluster array 1712 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 1712 can include texture sampling logic for performing texture operations, as well as additional logic for supporting the performance of such graphics processing operations, including but not limited to mosiac logic and other vertex processing logic. In at least one embodiment, the processing cluster array 1712 can be configured to execute graphics processing related shader programs, such as but not limited to vertex shaders, mosiac shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 1702 can transfer data from the system memory through the I / O unit 1704 for processing. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 1722) during processing and then written back to the system memory.
[0153] In at least one embodiment, when graphics processing is performed using the parallel processing unit 1702, the scheduler 1710 can be configured to divide the processing workload into tasks of approximately equal size so as to more effectively distribute the graphics processing operations among the plurality of clusters 1714A - 1714N of the processing cluster array 1712. In at least one embodiment, a portion of the processing cluster array 1712 can be configured to perform different types of processing. For example, in at least one embodiment, for generating and displaying a rendered image, the first portion may be configured to perform vertex shading and topology generation, the second portion may be configured to perform mosaic and geometry shading, and the third portion may be configured to perform pixel shading or other screen space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 1714A - 1714N can be stored in a buffer so that the intermediate data can be transmitted among the clusters 1714A - 1714N for further processing.
[0154] In at least one embodiment, the processing cluster array 1712 can receive processing tasks to be executed via a scheduler 1710, and the scheduler 1710 receives commands defining the processing tasks from a front end 1708. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands (e.g., which program to execute) defining how the data is to be processed. In at least one embodiment, the scheduler 1710 may be configured to fetch an index corresponding to a task, or may receive an index from the front end 1708. In at least one embodiment, the front end 1708 can be configured to ensure that the processing cluster array 1712 is configured in an active state before a workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is started.
[0155] In at least one embodiment, each of one or more instances of the parallel processing unit 1702 can be coupled to a parallel processor memory 1722. In at least one embodiment, the parallel processor memory 1722 can be accessed via a memory crossbar 1716, and the memory crossbar 1716 can receive memory requests from the processing cluster array 1712 as well as the I / O unit 1704. In at least one embodiment, the memory crossbar 1716 can access the parallel processor memory 1722 via a memory interface 1718. In at least one embodiment, the memory interface 1718 can include a plurality of partitioning units (e.g., partitioning unit 1720A, partitioning unit 1720B - partitioning unit 1720N), and each of these units can be coupled to a portion (e.g., a memory unit) of the parallel processor memory 1722. In at least one embodiment, the number of partitioning units 1720A - 1720N is configured to be equal to the number of memory units, such that the first partitioning unit 1720A has a corresponding first memory unit 1724A, the second partitioning unit 1720B has a corresponding memory unit 1724B, and the Nth partitioning unit 1720N has a corresponding Nth memory unit 1724N. In at least one embodiment, the number of partitioning units 1720A - 1720N may not be equal to the number of memory devices.
[0156] In at least one embodiment, the memory units 1724A - 1724N can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory. In at least one embodiment, the memory units 1724A - 1724N may also include, but are not limited to, 3D stacked memory including high bandwidth memory (HBM). In at least one embodiment, to efficiently use the available bandwidth of the parallel processor memory 1722, a render target, such as a frame buffer or texture map, may be stored across the memory units 1724A - 1724N so that the partition units 1720A - 1720N can write portions of each render target in parallel. In at least one embodiment, the local instance of the parallel processor memory 1722 may be excluded to be advantageous for an integrated memory design that combines system memory and local cache memory.
[0157] In at least one embodiment, any one of clusters 1714A - 1714N of the processing cluster array 1712 can process data to be written to any one of the memory units 1724A - 1724N within the parallel processor memory 1722. In at least one embodiment, the memory crossbar 1716 can be configured to transfer the output of each of clusters 1714A - 1714N to any of the partition units 1720A - 1720N that can perform further processing operations on the output, or to another cluster 1714A - 1714N. In at least one embodiment, each of clusters 1714A - 1714N can communicate with the memory interface 1718 through the memory crossbar 1716 to read from or write to various external memory devices. In at least one embodiment, the memory crossbar 1716 has a connection to the memory interface 1718 for communicating with the I / O unit 1704, as well as a connection to a local instance of the parallel processor memory 1722, enabling processing units within different processing clusters 1714A - 1714N to communicate with the system memory or other memory not local to the parallel processing unit 1702. In at least one embodiment, the memory crossbar 1716 can use virtual channels to separate traffic streams between the clusters 1714A - 1714N and the partition units 1720A - 1720N.
[0158] In at least one embodiment, multiple instances of the parallel processing unit 1702 may be provided on a single add-in card or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 1702 may be configured to interoperate even if they have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 1702 can include floating-point units with higher precision compared to other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 1702 or the parallel processor 1700 can be implemented in various configurations and form factors including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.
[0159] FIG. 17B is a block diagram of a partition unit 1720 according to at least one embodiment. In at least one embodiment, the partition unit 1720 is an instance of one of the partition units 1720A - 1720N of the partition unit of FIG. 17A. In at least one embodiment, the partition unit 1720 includes an L2 cache 1721, a frame buffer interface 1725, and a raster operations unit ( "ROP") 1726. The L2 cache 1721 is a read / write cache configured to perform load and store operations received from the memory crossbar 1716 and the ROP 1726. In at least one embodiment, read misses and urgent write-back requests are output by the L2 cache 1721 to the frame buffer interface 1725 to be processed. In at least one embodiment, updates are also sent to the frame via the frame buffer interface 1725 to be processed. In at least one embodiment, the frame buffer interface 1725 interfaces with one of the memory units of the parallel processor memory, such as the memory units 1724A - 1724N (e.g., within the parallel processor memory 1722) of FIG. 17.
[0160] In at least one embodiment, the ROP 1726 is a processing unit that performs raster operations such as stencil, z-test, and blending. In at least one embodiment, the ROP 1726 then outputs the processed graphics data stored in the graphics memory. In at least one embodiment, the ROP 1726 includes compression logic for compressing depth or color data written to the memory and decompressing depth or color data read from the memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a plurality of compression algorithms. The compression logic executed by the ROP 1726 can be changed based on the statistical characteristics of the data to be compressed. For example, in at least one embodiment, delta color compression is performed tile-by-tile on depth and color data.
[0161] In at least one embodiment, ROP1726 is included within each processing cluster (e.g., clusters 1714A-1714N of FIG. 17A), rather than within partition unit 1720. In at least one embodiment, read and write requests for pixel data, rather than pixel fragment data, are transmitted via memory crossbar 1716. In at least one embodiment, the processed graphics data may be displayed on a display device, such as one of the one or more display devices 1610 of FIG. 16, routed so as to be further processed by processor 1602, or routed so as to be further processed by one of the processing entities within parallel processor 1700 of FIG. 17A.
[0162] FIG. 17C is a block diagram of processing cluster 1714 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of the processing clusters 1714A-1714N of FIG. 17A. In at least one embodiment, one or more of processing clusters 1714 may be configured to execute a number of threads in parallel, where a "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issue techniques are used to support parallel execution of a number of threads without providing a plurality of independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) techniques are used to support parallel execution of a number of threads that are globally synchronized using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0163] In at least one embodiment, the operation of processing cluster 1714 can be controlled via a pipeline manager 1732 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, pipeline manager 1732 receives instructions from scheduler 1710 of FIG. 17A and manages the execution of these instructions via graphics multiprocessor 1734 and / or texture unit 1736. In at least one embodiment, graphics multiprocessor 1734 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within processing cluster 1714. In at least one embodiment, one or more instances of graphics multiprocessor 1734 can be included within processing cluster 1714. In at least one embodiment, graphics multiprocessor 1734 can process data, and data crossbar 1740 may be used to distribute the processed data to one of a plurality of possible destinations including other shader units. In at least one embodiment, pipeline manager 1732 can facilitate the distribution of processed data by specifying the destination of the processed data that will be distributed through data crossbar 1740.
[0164] In at least one embodiment, each graphics multiprocessor 1734 within processing cluster 1714 can include the same set of function execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner such that new instructions can be issued before the previous instruction has completed. In at least one embodiment, the function execution logic supports various operations including integer and floating point arithmetic, comparison operations, boolean operations, bit shifts, and calculations of various algebraic functions. In at least one embodiment, different operations can be executed by leveraging the hardware of the same function units, and any combination of function units may exist.
[0165] In at least one embodiment, the instructions sent to processing cluster 1714 configure a thread. In at least one embodiment, a set of threads being executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within the thread group can be assigned to a different processing engine within graphics multiprocessor 1734. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within graphics multiprocessor 1734. In at least one embodiment, if the thread group includes fewer threads than the number of processing engines, one or more processing engines may be idle during the cycles in which the thread group is being processed. In at least one embodiment, the thread group may also include more threads than the number of processing engines within graphics multiprocessor 1734. In at least one embodiment, if the thread group includes more threads than the processing engines within graphics multiprocessor 1734, processing can be executed over consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 1734.
[0166] In at least one embodiment, the graphics multi-processor 1734 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multi-processor 1734 can forego the internal cache and use the cache memory (e.g., L1 cache 1748) within the processing cluster 1714. In at least one embodiment, each graphics multi-processor 1734 can also access the L2 cache within a partition unit (e.g., partition units 1720A - 1720N of FIG. 17A), which can be shared among all processing clusters 1714 and used to transfer data between threads. In at least one embodiment, the graphics multi-processor 1734 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 1702 can be used as global memory. In at least one embodiment, the processing cluster 1714 includes multiple instances of the graphics multi-processor 1734 that can share common instructions and data, which may be stored in the L1 cache 1748.
[0167] In at least one embodiment, each processing cluster 1714 may include a memory management unit (“MMU”) 1745 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of MMU 1745 may be within the memory interface 1718 of FIG. 17A. In at least one embodiment, MMU 1745 includes a set of page table entries (PTEs) used to map virtual addresses to the physical addresses of tiles and optionally cache line indices. In at least one embodiment, MMU 1745 may include a translation lookaside buffer (TLB) or cache, which may be within the graphics multiprocessor 1734 or L1 cache, or within the processing cluster 1714. In at least one embodiment, the physical addresses are processed to locally distribute surface data access, enabling efficient interleaving of requests among partition units. In at least one embodiment, a cache line index may be used to determine whether a cache line request is a hit or a miss.
[0168] In at least one embodiment, each graphics multi-processor 1734 is coupled to a texture unit 1736 such that the processing cluster 1714 may be configured to perform texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from the L1 cache within the graphics multi-processor 1734 and, if necessary, fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multi-processor 1734 outputs processed tasks to a data crossbar 1740 to provide the processed tasks to another processing cluster 1714 for further processing or stores the processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 1716. In at least one embodiment, a pre-ROP 1742 (pre-raster operation unit) is configured to receive data from the graphics multi-processor 1734 and direct the data to a ROP unit, which may be located within a partitioning unit (e.g., partitioning units 1720A - 170N of FIG. 17A) as described herein. In at least one embodiment, the pre-ROP 1742 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.
[0169] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the graphics processing cluster 1714 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of neural networks, or use cases of neural networks described herein.
[0170] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently.
[0171] FIG. 17D shows a graphics multiprocessor 1734 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 1734 is coupled to a pipeline manager 1732 of the processing cluster 1714. In at least one embodiment, the graphics multiprocessor 1734 has an execution pipeline including, but not limited to, an instruction cache 1752, an instruction unit 1754, an address mapping unit 1756, a register file 1758, one or more general-purpose graphics processing unit (GPGPU) cores 1762, and one or more load / store units 1766. The GPGPU cores 1762 and the load / store units 1766 are coupled to a cache memory 1772 and a shared memory 1770 via a memory and cache interconnect 1768.
[0172] In at least one embodiment, the instruction cache 1752 receives a stream of instructions to be executed from the pipeline manager 1732. In at least one embodiment, the instructions are cached in the instruction cache 1752 and dispatched to be executed by the instruction unit 1754. In at least one embodiment, the instruction unit 1754 can dispatch instructions as thread groups (e.g., warps), and each thread group is assigned to different execution units within the GPGPU core 1762. In at least one embodiment, instructions can access any of the local address space, shared address space, or global address space by specifying an address within the unified address space. In at least one embodiment, instructions can access any of the local, shared, or global address spaces by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 1756 can be used to translate an address in the unified address space to an individual memory address accessible by the load / store unit 1766.
[0173] In at least one embodiment, the register file 1758 provides a set of registers to the functional units of the graphics multiprocessor 1734. In at least one embodiment, the register file 1758 provides temporary storage for operands connected to the data paths of the functional units (e.g., GPGPU core 1762, load / store unit 1766) of the graphics multiprocessor 1734. In at least one embodiment, the register file 1758 is divided among the respective functional units such that each functional unit is allocated a dedicated portion of the register file 1758. In at least one embodiment, the register file 1758 is divided among different warps being executed by the graphics multiprocessor 1734.
[0174] In at least one embodiment, each GPGPU core 1762 can include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions of the graphics multiprocessor 1734. The GPGPU cores 1762 may have the same architecture or different architectures. In at least one embodiment, a first portion of the GPGPU core 1762 includes a single-precision FPU and an integer ALU, and a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU can implement the IEEE 754-2008 standard for floating-point operations or enable variable-precision floating-point operations. In at least one embodiment, the graphics multiprocessor 1734 can further include one or more fixed-function units or special-function units for executing specific functions such as rectangle copy or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores can also include fixed or special-function logic.
[0175] In at least one embodiment, the GPGPU core 1762 includes SIMD logic capable of executing a single instruction on multiple data sets. In at least one embodiment, the GPGPU core 1762 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for the GPGPU core may be generated during compilation by a shader compiler or may be automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logical unit.
[0176] In at least one embodiment, the memory and cache interconnect 1768 is an interconnect network that connects each functional unit of the graphics multiprocessor 1734 to the register file 1758 and the shared memory 1770. In at least one embodiment, the memory and cache interconnect 1768 is a crossbar interconnect that enables the load / store unit 1766 to implement load and store operations between the shared memory 1770 and the register file 1758. In at least one embodiment, the register file 1758 can operate at the same frequency as the GPGPU core 1762, and thus, data transfer between the GPGPU core 1762 and the register file 1758 is very low latency. In at least one embodiment, the shared memory 1770 can be used to enable communication between threads executed by functional units within the graphics multiprocessor 1734. In at least one embodiment, the cache memory 1772 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 1736. In at least one embodiment, the shared memory 1770 can also be used as a program management cache. In at least one embodiment, threads executing on the GPGPU core 1762 can store data programmatically in the shared memory in addition to automatically cached data stored in the cache memory 1772.
[0177] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated as a core on the same package or chip and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). In at least one embodiment, regardless of the method of connecting the GPU, the processor core may distribute work to the GPU in the form of a sequence of commands / instructions included in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0178] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the graphics multiprocessor 1734 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0179] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0180] FIG. 18 shows a multi-GPU computing system 1800 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 1800 can include a processor 1802 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 1806A-D via a host interface switch 1804. In at least one embodiment, the host interface switch 1804 is a PCI Express switch device that couples the processor 1802 to a PCI Express bus, via which the processor 1802 can communicate with the GPGPUs 1806A-D. The GPGPUs 1806A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 1816. In at least one embodiment, the GPU-to-GPU link 1816 is connected to each of the GPGPUs 1806A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU link 1816 enables direct communication between each of the GPGPUs 1806A-D without requiring communication via the host interface bus 1804 to which the processor 1802 is connected. In at least one embodiment, when there is GPU-to-GPU traffic destined for the P2P GPU link 1816, the host interface bus 1804 is kept available to access system memory or to communicate with other instances of the multi-GPU computing system 1800, for example, via one or more network devices. In at least one embodiment, the GPGPUs 1806A-D are connected to the processor 1802 via the host interface switch 1804, and in at least one embodiment, the processor 1802 includes direct support for the P2P GPU link 1816 and can be directly connected to the GPGPUs 1806A-D.
[0181] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the multi-GPU computing system 1800 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions, and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0182] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0183] FIG. 19 is a block diagram of a graphics processor 1900 according to at least one embodiment. In at least one embodiment, the graphics processor 1900 includes a ring interconnect 1902, a pipeline front end 1904, a media engine 1937, and graphics cores 1980A-1980N. In at least one embodiment, the ring interconnect 1902 couples the graphics processor 1900 to other graphics processors or other processing units including one or more general-purpose processor cores. In at least one embodiment, the graphics processor 1900 is one of a number of processors integrated within a multi-core processing system.
[0184] In at least one embodiment, the graphics processor 1900 receives a batch of commands via the ring interconnect 1902. In at least one embodiment, incoming commands are interpreted by the command streamer 1903 of the pipeline front end 1904. In at least one embodiment, the graphics processor 1900 includes scalable execution logic for performing 3D geometry processing and media processing via the graphics cores 1980A - 1980N. In at least one embodiment, for 3D geometry processing commands, the command streamer 1903 supplies the commands to the geometry pipeline 1936. In at least one embodiment, for at least some media processing commands, the command streamer 1903 supplies the commands to the video front end 1934, and the video front end 1934 is coupled to the media engine 1937. In at least one embodiment, the media engine 1937 includes a Video Quality Engine (VQE) 1930 for post - processing of video and images, and a multi - format encode / decode (MFX) 1933 engine that provides hardware - accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 1936 and the media engine 1937 each generate execution threads for the thread execution resources provided by at least one graphics core 1980A.
[0185] In at least one embodiment, the graphics processor 1900 includes a scalable thread execution resource characterized by modular cores 1980A - 1980N (also sometimes referred to as core slices), and each of the modular cores 1980A - 1980N has a plurality of sub - cores 1950A - 1950N, 1960A - 1960N (also sometimes referred to as core sub - slices). In at least one embodiment, the graphics processor 1900 can have any number of graphics cores 1980A - 1980N. In at least one embodiment, the graphics processor 1900 includes a graphics core 1980A having at least a first sub - core 1950A and a second sub - core 1960A. In at least one embodiment, the graphics processor 1900 is a low - power processor having a single sub - core (e.g., 1950A). In at least one embodiment, the graphics processor 1900 includes a plurality of graphics cores 1980A - 1980N, each of which includes a set of first sub - cores 1950A - 1950N and a set of second sub - cores 1960A - 1960N. In at least one embodiment, each sub - core of the first sub - cores 1950A - 1950N includes at least an execution unit 1952A - 1952N and a first set of media / texture samplers 1954A - 1954N. In at least one embodiment, each sub - core of the second sub - cores 1960A - 1960N includes at least an execution unit 1962A - 1962N and a second set of samplers 1964A - 1964N. In at least one embodiment, each sub - core 1950A - 1950N, 1960A - 1960N shares a set of shared resources 1970A - 1970N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.
[0186] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the graphics processor 1900 for inference or prediction operations, at least partially based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0187] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks, at least partially based on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0188] FIG. 20 is a block diagram showing the microarchitecture of a processor 2000 that may include a logic circuit for executing instructions according to at least one embodiment. In at least one embodiment, the processor 2000 may execute instructions including x86 instructions, AMR instructions, special instructions for application specific integrated circuits (ASICs), etc. In at least one embodiment, the processor 2000 may include registers for storing packed data, such as 64-bit wide MMXTM registers in a microprocessor enabled with MMX technology by Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers available in both integer and floating point formats may operate on packed data elements with single instruction multiple data (''SIMD'') and streaming SIMD extensions (''SSE'') instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or more (collectively referred to as ''SSEx'') technologies may hold operands of such packed data. In at least one embodiment, the processor 2000 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0189] In at least one embodiment, the processor 2000 includes an in-order front end (“front end”) 2001 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front end 2001 may include several units. In at least one embodiment, an instruction prefetcher 2026 fetches instructions from memory and supplies the instructions to an instruction decoder 2028, and the instruction decoder decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2028 decodes the received instruction into one or more operations called “microinstructions” or “micro-operations” (also called “micro-ops” or “uops”) that the machine can execute. In at least one embodiment, the instruction decoder 2028 parses the instruction into an opcode and corresponding data, as well as a control field, such that these are used by the microarchitecture and the operations according to at least one embodiment may be executed. In at least one embodiment, a trace cache 2030 may assemble the decoded uops into a program-order sequence or trace in a uop queue 2034 for execution. In at least one embodiment, when the trace cache 2030 encounters a complex instruction, a microcode ROM 2032 provides the uops necessary for the completion of the operation.
[0190] In at least one embodiment, there are instructions that can be translated into a single micro-op, and there are also instructions that require several micro-ops to complete all operations. In at least one embodiment, if more than five micro-ops are required to complete an instruction, the instruction decoder 2028 may access the microcode ROM 2032 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-ops so that it can be processed in the instruction decoder 2028. In at least one embodiment, if a large number of micro-ops are required to complete an operation, the instruction may be stored in the microcode ROM 2032. In at least one embodiment, the trace cache 2030 determines the correct micro-instruction pointer for reading the microcode sequence by referring to an entry-point programmable logic array ("PLA") to complete one or more instructions from the microcode ROM 2032 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2032 finishes sequencing the micro-ops for an instruction, the front end 2001 of the machine may resume fetching micro-ops from the trace cache 2030.
[0191] In at least one embodiment, an out-of-order execution engine (“out-of-order engine”) 2003 may prepare instructions for execution. In at least one embodiment, out-of-order execution logic has a number of buffers to smooth the flow of instructions and change their order, optimizing performance when instructions are scheduled to flow down a pipeline and be executed. In at least one embodiment, the out-of-order execution engine 2003 includes, without limitation, an allocator / register renamer 2040, a memory uop queue 2042, an integer / floating point uop queue 2044, a memory scheduler 2046, a fast scheduler 2002, a slow / general purpose floating point scheduler (“slow / general purpose FP scheduler”) 2004, and a simple floating point scheduler (“simple FP scheduler”) 2006. In at least one embodiment, the fast scheduler 2002, the slow / general purpose floating point scheduler 2004, and the simple floating point scheduler 2006 are also collectively referred to herein as “uop schedulers 2002, 2004, 2006”. In at least one embodiment, the allocator / register renamer 2040 allocates the machine buffers and resources required by each uop for execution. In at least one embodiment, the allocator / register renamer 2040 changes the name of the logical register upon entry into the register file. In at least one embodiment, the allocator / register renamer 2040 also distributes the entry of each uop to one of two uop queues, namely the memory uop queue 2042 for memory operations and the integer / floating point uop queue 2044 for non-memory operations, in front of the memory scheduler 2046 and the uop schedulers 2002, 2004, 2006. In at least one embodiment, the uop schedulers 2002, 2004, 2006 determine when uops are ready for execution based on the availability of the sources of their dependent input register operands and the execution resources required by the uop to complete their operations.In at least one embodiment, the high-speed scheduler 2002 of at least one embodiment may schedule every half of the main clock cycle, and the low-speed / general-purpose floating-point scheduler 2004 and the simple floating-point scheduler 2006 may schedule once per clock cycle of the main processor. In at least one embodiment, the uop schedulers 2002, 2004, 2006 arbitrate dispatch ports to schedule uops for execution.
[0192] In at least one embodiment, the execution block 2011 includes, without limitation, the integer register file / bypass network 2008, the floating-point register file / bypass network (referred to herein as the "FP register file / bypass network") 2010, address generation units ("AGUs") 2012 and 2014, high-speed arithmetic logic units ("high-speed ALUs") 2016 and 2018, low-speed arithmetic logic units ("low-speed ALUs") 2020, floating-point ALUs ("FPs") 2022, and floating-point move units ("FP moves") 2024. In at least one embodiment, the integer register file / bypass network 2008 and the floating-point register file / bypass network 2010 are also referred to herein as the "register files 2008, 2010". In at least one embodiment, the AGUs 2012 and 2014, the high-speed ALUs 2016 and 2018, the low-speed ALUs 2020, the floating-point ALUs 2022, and the floating-point move units 2024 are also referred to herein as the "execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024". In at least one embodiment, the execution block b11 may include, without limitation, any number and type of register files, bypass networks, address generation units, and execution units (including zero) in any combination.
[0193] In at least one embodiment, register files 2008, 2010 may be disposed between uop schedulers 2002, 2004, 2006 and execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024. In at least one embodiment, integer register file / bypass network 2008 performs integer operations. In at least one embodiment, floating point register file / bypass network 2010 performs floating point operations. In at least one embodiment, each of register files 2008, 2010 may include, without limitation, a bypass network that may bypass or transfer recently completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, register files 2008, 2010 may communicate data with each other. In at least one embodiment, integer register file / bypass network 2008 may include, without limitation, two separate register files, namely, one register file for lower 32-bit data and a second register file for upper 32-bit data. In at least one embodiment, since floating point instructions typically have operands with a width of 64 to 128 bits, floating point register file / bypass network 2010 may include, without limitation, 128-bit wide entries.
[0194] In at least one embodiment, the execution units 2012, 2014, 2016, 2018, 2020, 2022, 2024 may execute instructions. In at least one embodiment, the register files 2008, 2010 store operand values of integer and floating-point data that the microinstructions need to execute. In at least one embodiment, the processor 2000 may include any number and combination of execution units 2012, 2014, 2016, 2018, 2020, 2022, 2024 without limitation. In at least one embodiment, the floating-point ALU 2022 and the floating-point shift unit 2024 may execute other operations including floating-point, MMX, SIMD, AVX, and SEE, or special machine learning instructions. In at least one embodiment, the floating-point ALU 2022 includes, without limitation, a floating-point divider in 64-bit chunks and may execute division, square root, and other micro-ops. In at least one embodiment, instructions containing floating-point values may be handled by the floating-point hardware. In at least one embodiment, ALU operations may be passed to the fast ALUs 2016, 2018. In at least one embodiment, the fast ALUs 2016, 2018 may execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, since the slow ALU 2020 may include, without limitation, integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branch processing, most complex integer operations proceed to the slow ALU 2020. In at least one embodiment, memory load / store operations may be performed by the AGUs 2012, 2014. In at least one embodiment, the fast ALU 2016, the fast ALU 2018, and the slow ALU 2020 may execute integer operations on 64-bit data operands. In at least one embodiment, the fast ALU 2016, the fast ALU 2018, and the slow ALU 2020 may be implemented to support various data bit sizes including 16, 32, 128, 256, etc. In at least one embodiment, the floating-point ALU 2022 and the floating-point shift unit 2024 may be implemented to support a wide range of operands having various bit widths.In at least one embodiment, the floating point ALU 2022 and the floating point shift unit 2024 can operate on 128-bit wide packed data operands in conjunction with single-instruction, multiple-data (SIMD) and multimedia instructions.
[0195] In at least one embodiment, the uop schedulers 2002, 2004, 2006 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, since uops may be scheduled and executed speculatively in the processor 2000, the processor 2000 may also include logic for handling memory misses. In at least one embodiment, when a data load misses in the data cache, there may be ongoing dependent operations in the pipeline that have passed through a scheduler with temporarily inaccurate data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use inaccurate data. In at least one embodiment, dependent operations may need to be replayed and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0196] In at least one embodiment, the term "register" may refer to a storage location of an on-board processor that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be accessible from outside the processor (from the perspective of a programmer). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuits within a processor using any number of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, combinations of dedicated physical registers and physically registers dynamically allocated, and the like. In at least one embodiment, an integer register stores 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packed data.
[0197] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of inference and / or training logic 615 may be incorporated into execution block 2011 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs shown in execution block 2011. Additionally, weight parameters may be stored in on-chip or off-chip memories and / or registers (shown or not shown) that make up the ALU of execution block 2011 for performing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0198] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0199] FIG. 21 shows a deep learning application processor 2100 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2100 uses instructions that cause the deep learning application processor 2100 to perform some or all of the processes and techniques described throughout this disclosure when executed by the deep learning application processor 2100. In at least one embodiment, the deep learning application processor 2100 is an application specific integrated circuit (ASIC). In at least one embodiment, the application processor 2100 performs matrix multiplication operations that are “hard-wired” to hardware as a result of executing one or both of the instructions. In at least one embodiment, the deep learning application processor 2100 includes, without limitation, processing clusters 2110(1) to 2110(12), inter-chip links (“ICL”) 2120(1) to 2120(12), inter-chip controllers (“ICC”) 2130(1) to 2130(2), memory controllers (“Mem Ctrlrs”) 2142(1) to 2142(4), high bandwidth memory physical layers (“HBM PHY”) 2144(1) to 2144(4), a management-controller central processing unit (“management-controller CPU”) 2150, a peripheral component interconnect express controller and direct memory access block (“PCIe controller and DMA”) 2170, and a 16-lane peripheral component interconnect express port (“PCI Expressx16”) 2180.
[0200] In at least one embodiment, the processing cluster 2110 may perform deep learning operations including inference or prediction operations based on weight parameters calculated using one or more training techniques including the techniques described herein. In at least one embodiment, each processing cluster 2110 may include any number and type of processors, without limitation. In at least one embodiment, the deep learning application processor 2100 may include any number and type of processing clusters 2100. In at least one embodiment, the inter-chip link 2120 is bidirectional. In at least one embodiment, the inter-chip link 2120 and the inter-chip controller 2130 enable the plurality of deep learning application processors 2100 to exchange information including activation information obtained as a result of executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2100 may include any number and type (including zero) of ICLs 2120 and ICCs 2130.
[0201] In at least one embodiment, the HBM2 2140 provides a total of 32 gigabytes (GB) of memory. The HBM2 2140(i) is associated with both the memory controller 2142(i) and the HBM PHY 2144(i). In at least one embodiment, any number of HBM2s 2140 may provide any type and total amount of high-bandwidth memory and may be associated with any number and type (including zero) of memory controllers 2142 and HBM PHYs 2144. In at least one embodiment, the SPI, I2C, GPIO 2160, the PCIe controller and DMA 2170, and / or the PCIe 2180 may be replaced with any number and type of blocks enabling any number and type of communication standards in any technically feasible manner.
[0202] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the deep learning application processor 2100 is used to train a machine learning model, such as a neural network, to predict or infer information provided to the deep learning application processor 2100. In at least one embodiment, the deep learning application processor 2100 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 2100 itself. In at least one embodiment, the processor 2100 may be used to execute one or more neural network use cases described herein.
[0203] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently.
[0204] FIG. 22 is a block diagram of a neuromorphic processor 2200 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2200 receives one or more inputs from a source external to the neuromorphic processor 2200. In at least one embodiment, these inputs may be sent to one or more neurons 2202 within the neuromorphic processor 2200. In at least one embodiment, the neurons 2202 and their components may be implemented using circuitry or logic that includes one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2200 may include thousands or millions of instances of neurons 2202, without limitation, although any suitable number of neurons 2202 may be used. In at least one embodiment, each instance of a neuron 2202 may include a neuron input 2204 and a neuron output 2206. In at least one embodiment, the neuron 2202 may generate an output, which may be sent to the inputs of other instances of the neuron 2202. For example, in at least one embodiment, the neuron input 2204 and the neuron output 2206 may be interconnected via a synapse 2208.
[0205] In at least one embodiment, neuron 2202 and synapse 2208 may be interconnected such that the neuromorphic processor 2200 processes or analyzes the information received by the neuromorphic processor 2200. In at least one embodiment, neuron 2202 may transmit an output pulse (or "fire" or "spike") when the input received via neuron input 2204 exceeds a threshold. In at least one embodiment, neuron 2202 may sum or integrate the signals received at neuron input 2204. For example, in at least one embodiment, neuron 2202 may be implemented as a leaky integrate-and-fire neuron, where when the sum (referred to as the "membrane potential") exceeds a threshold, neuron 2202 may use a transfer function such as a sigmoid function or a threshold function to generate an output (or "fire"). In at least one embodiment, the leaky integrate-and-fire neuron may sum the signals received at neuron input 2204 to form a membrane potential, and may also apply a decay factor (or leak) to reduce the membrane potential. In at least one embodiment, the leaky integrate-and-fire neuron may fire if a plurality of input signals are received at neuron input 2204 quickly enough such that they exceed the threshold (i.e., before the decay of the membrane potential is too great to prevent firing). In at least one embodiment, neuron 2202 may be implemented using circuitry or logic that receives an input, integrates the input to form a membrane potential, and decays the membrane potential. In at least one embodiment, the input may be averaged, or any other suitable transfer function may be used. Further, in at least one embodiment, neuron 2202 may include, without limitation, comparator circuitry or logic that generates an output spike at neuron output 2206 when the result of applying a transfer function to the neuron input 2204 exceeds a threshold. In at least one embodiment, when neuron 2202 fires, it may ignore the previously received input information, for example, by resetting the membrane potential to 0 or some other suitable default value.In at least one embodiment, when the membrane potential is reset to 0, neuron 2202 may resume normal operation after a suitable period (or refractory period).
[0206] In at least one embodiment, neurons 2202 may be interconnected through synapses 2208. In at least one embodiment, synapses 2208 may be operative to transmit signals from the output of a first neuron 2202 to the input of a second neuron 2202. In at least one embodiment, neurons 2202 may transmit information through more than one instance of synapses 2208. In at least one embodiment, one or more instances of neuron outputs 2206 may be connected through an instance of synapses 2208 to an instance of neuron inputs 2204 of the same neuron 2202. In at least one embodiment, an instance of neuron 2202 that generates an output that will be transmitted through an instance of synapses 2208 may be referred to as a “presynaptic neuron” with respect to that instance of synapses 2208. In at least one embodiment, an instance of neuron 2202 that receives an input that will be transmitted through an instance of synapses 2208 may be referred to as a “postsynaptic neuron” with respect to that instance of synapses 2208. In at least one embodiment, an instance of neuron 2202 may receive inputs from one or more instances of synapses 2208 and may also transmit outputs through one or more instances of synapses 2208, so a single instance of neuron 2202 may thus be both a “presynaptic neuron” and a “postsynaptic neuron” with respect to various instances of synapses 2208.
[0207] In at least one embodiment, neurons 2202 may be organized into one or more layers. Each instance of neuron 2202 may have one neuron output 2206 that can fan out to one or more neuron inputs 2204 through one or more synapses 2208. In at least one embodiment, the neuron output 2206 of neurons 2202 in the first layer 2210 may be connected to the neuron inputs 2204 of neurons 2202 in the second layer 2212. In at least one embodiment, layer 2210 may be referred to as a "feed-forward" layer. In at least one embodiment, each instance of neuron 2202 in an instance of the first layer 2210 may fan out to each instance of neuron 2202 in the second layer 2212. In at least one embodiment, the first layer 2210 may be referred to as a "fully connected feed-forward layer". In at least one embodiment, each instance of neuron 2202 in an instance of the second layer 2212 may fan out to fewer instances of neuron 2202 in the third layer 2214 than all instances of neuron 2202 in the third layer 2214. In at least one embodiment, the second layer 2212 may be referred to as a "sparsely connected feed-forward layer". In at least one embodiment, neurons 2202 in the second layer 2212 may fan out to neurons 2202 in a plurality of other layers, including neurons 2202 in the (same) second layer 2212. In at least one embodiment, the second layer 2212 may be referred to as a "recurrent layer". In at least one embodiment, neuromorphic processor 2200 may include any suitable combination of recurrent layers and feed-forward layers, including both sparsely connected feed-forward layers and fully connected feed-forward layers, without limitation.
[0208] In at least one embodiment, the neuromorphic processor 2200 may include, without limitation, a reconfigurable interconnect architecture for connecting synapses 2208 to neurons 2202, or dedicated hard-wired interconnects. In at least one embodiment, the neuromorphic processor 2200 may include, without limitation, circuitry or logic that enables synapses to be distributed to different neurons 2202 as needed, based on the neural network topology and the fan-in / fan-out of the neurons. For example, in at least one embodiment, the synapses 2208 may be connected to the neurons 2202 using an interconnect fabric such as a network-on-chip or using dedicated connections. In at least one embodiment, the synapse interconnects and their components may be implemented using circuitry or logic.
[0209] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0210] FIG. 23 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, the system 2300 includes one or more processors 2302 and one or more graphics processors 2308, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a number of processors 2302 or processor cores 2307. In at least one embodiment, the system 2300 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile device, a portable device, or an embedded device.
[0211] In at least one embodiment, system 2300 may include, or be incorporated in, a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable gaming console, or an online gaming console. In at least one embodiment, system 2300 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, processing system 2300 may also include, be coupled to, or be integrated within wearable devices such as smartwatch wearable devices, smart eyewear devices, augmented reality devices, or virtual reality devices. In at least one embodiment, processing system 2300 is a television or set-top box device having one or more processors 2302 and a graphical interface generated by one or more graphics processors 2308.
[0212] In at least one embodiment, each of the one or more processors 2302 includes one or more processor cores 2307 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 2307 is configured to process a particular instruction set 2309. In at least one embodiment, instruction set 2309 may facilitate computing via a complex instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction word (VLIW). In at least one embodiment, the processor cores 2307 may each process different instruction sets 2309, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, processor cores 2307 may also include other processing devices such as a digital signal processor (DSP).
[0213] In at least one embodiment, the processor 2302 includes a cache memory 2304. In at least one embodiment, the processor 2302 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of the processor 2302. In at least one embodiment, the processor 2302 may also use an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), and this cache may be shared among the processor cores 2307 using known cache coherence techniques. In at least one embodiment, a register file 2306 is further included in the processor 2302, and this register file may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, the register file 2306 may include general-purpose registers or other registers.
[0214] In at least one embodiment, one or more processors 2302 are coupled to one or more interface buses 2310 to transmit communication signals such as address, data, or control signals between the processor 2302 and other components within the system 2300. In at least one embodiment, the interface bus 2310 can be a processor bus such as a version of a Direct Media Interface (DMI) bus in one embodiment. In at least one embodiment, the interface 2310 is not limited to the DMI bus and may include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 2302 includes an integrated memory controller 2316 and a platform controller hub 2330. In at least one embodiment, the memory controller 2316 facilitates communication between the memory device and other components of the system 2300, while the platform controller hub (PCH) 2330 provides connections to I / O devices via a local I / O bus.
[0215] In at least one embodiment, the memory device 2320 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or any other memory device having suitable performance to serve as a process memory. In at least one embodiment, the memory device 2320 operates as a system memory for the system 2300 and can store data 2322 and instructions 2321 for use when one or more processors 2302 execute an application or process. In at least one embodiment, the memory controller 2316 is also coupled to an optional external graphics processor 2312, which may communicate with one or more graphics processors 2308 within the processor 2302 to perform graphics and media operations. In at least one embodiment, the display device 2311 can be connected to the processor 2302. In at least one embodiment, the display device 2311 can include one or more of an internal display device such as a mobile electronic device or a laptop device, or an external display device attached via a display interface (e.g., a display port, etc.). In at least one embodiment, the display device 2311 can include a head-mounted display (HMD) such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.
[0216] In at least one embodiment, the platform controller hub 2330 enables peripheral devices to be connected to the memory device 2320 and the processor 2302 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 2346, a network controller 2334, a firmware interface 2328, a wireless transceiver 2326, a touch sensor 2325, and a data storage device 2324 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 2324 can be connected via a storage interface (e.g., SATA) or a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 2325 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 2326 can be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2328 enables communication with system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 2334 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 2310. In at least one embodiment, the audio controller 2346 is a multi-channel high-definition audio controller. In at least one embodiment, the system 2300 includes an optional legacy I / O controller 2340 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.In at least one embodiment, the platform controller hub 2330 can also be connected to a connection input device of one or more universal serial bus (USB) controllers 2342, such as a combination of a keyboard and mouse 2343, a camera 2344, or other USB input devices.
[0217] In at least one embodiment, instances of the memory controller 2316 and the platform controller hub 2330 may be integrated into a separate external graphics processor, such as the external graphics processor 2312. In at least one embodiment, the platform controller hub 2330 and / or the memory controller 2316 may be external to one or more processors 2302. For example, in at least one embodiment, the system 2300 can include an external memory controller 2316 and a platform controller hub 2330, which may be configured as a memory controller hub and a peripheral device controller hub within a system chipset that communicates with the processor 2302.
[0218] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the graphics processor 2300. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in the graphics processor 2312. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIGS. 6A or 6B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that constitute the ALU of the graphics processor 2300 for performing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0219] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently.
[0220] FIG. 24 is a block diagram of a processor 2400 having at least one embodiment of one or more processor cores 2402A - 2402N, an integrated memory controller 2414, and an integrated graphics processor 2408. In at least one embodiment, the processor 2400 can include a lesser number of additional cores including the additional core 2402N represented by the dashed square. In at least one embodiment, each of the processor cores 2402A - 2402N includes one or more internal cache units 2404A - 2404N. In at least one embodiment, each processor core can also access one or more shared cache units 2406.
[0221] In at least one embodiment, the internal cache units 2404A - 2404N and the shared cache unit 2406 represent a cache memory hierarchy within the processor 2400. In at least one embodiment, the cache memory units 2404A - 2404N can include at least one level of instruction and data cache within each processor core, and one or more levels of shared intermediate-level cache such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence among the various cache units 2406 and 2404A - 2404N.
[0222] In at least one embodiment, the processor 2400 may also include a set of one or more bus controller units 2416 and a system agent core 2410. In at least one embodiment, the one or more bus controller units 2416 manage a set of peripheral buses such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 2410 provides management functions for various processor components. In at least one embodiment, the system agent core 2410 includes one or more integrated memory controllers 2414 for managing access to various external memory devices (not shown).
[0223] In at least one embodiment, one or more of the processor cores 2402A - 2402N include support for simultaneous multithreading. In at least one embodiment, the system agent core 2410 includes components for coordinating and operating the cores 2402A - 2402N during multithreaded processing. In at least one embodiment, the system agent core 2410 may further include a power control unit (PCU) that includes logic and components for adjusting the power state of one or more of the processor cores 2402A - 2402N and the graphics processor 2408.
[0224] In at least one embodiment, the processor 2400 further includes a graphics processor 2408 for performing graphics processing operations. In at least one embodiment, the graphics processor 2408 is coupled to a shared cache unit 2406 and a system agent core 2410 including one or more integrated memory controllers 2414. In at least one embodiment, the system agent core 2410 also includes a display controller 2411 for causing the output of the graphics processor to be output to one or more coupled displays. In at least one embodiment, the display controller 2411 may also be a separate module coupled to the graphics processor 2408 via at least one interconnect, or may be integrated within the graphics processor 2408.
[0225] In at least one embodiment, a ring-based interconnect unit 2412 is used to couple the internal components of the processor 2400. In at least one embodiment, alternative interconnect units such as point-to-point interconnects, switch interconnects, or other techniques may be used. In at least one embodiment, the graphics processor 2408 is coupled to the ring interconnect 2412 via an I / O link 2413.
[0226] In at least one embodiment, the I / O link 2413 represents at least one of a variety of I / O interconnects including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 2418 such as an eDRAM module. In at least one embodiment, each of the processor cores 2402A - 2402N and the graphics processor 2408 use the embedded memory module 2418 as a shared last-level cache.
[0227] In at least one embodiment, the processor cores 2402A - 2402N are of the same type executing a common instruction set architecture. In at least one embodiment, the processor cores 2402A - 2402N are heterogeneous from the perspective of the instruction set architecture (ISA), where one or more of the processor cores 2402A - 2402N execute a common instruction set, but one or more of the other cores among the processor cores 2402A - 2402N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, the processor cores 2402A - 2402N are heterogeneous from the perspective of the micro - architecture, where one or more cores with relatively high power consumption are coupled with one or more cores with lower power consumption. In at least one embodiment, the processor 2400 can be implemented on one or more chips or as a SoC integrated circuit.
[0228] To perform inference and / or training operations related to one or more embodiments, the inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the processor 2400. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in the graphics processor 2312, the graphics cores 2402A - 2402N, or other components of FIG. 24. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIGS. 6A or 6B. In at least one embodiment, the weight parameters may be stored in on - chip or off - chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 2400 for executing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0229] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0230] FIG. 25 is a block diagram of the hardware logic of a graphics processor core 2500 according to at least one embodiment described herein. In at least one embodiment, the graphics processor core 2500 is included within a graphics core array. In at least one embodiment, the graphics processor core 2500, sometimes referred to as a core slice, can be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 2500 is an example of one graphics core slice, and a graphics processor as described herein can include multiple graphics core slices based on a target power and performance envelope. In at least one embodiment, each graphics core 2500 can include a fixed function block 2530 coupled to a plurality of sub-cores 2501A-2501F, also referred to as sub-slices, that include modular blocks of general purpose and fixed function logic.
[0231] In at least one embodiment, the fixed function block 2530 includes, for example, a geometry / fixed function pipeline 2536 that can be shared by all sub-cores in the graphics processor 2500 in a low performance and / or low power graphics processor implementation. In at least one embodiment, the geometry / fixed function pipeline 2536 includes a 3D fixed function pipeline, a video front end unit, a thread spawner and thread dispatcher, and an integrated return buffer manager that manages an integrated return buffer.
[0232] In at least one embodiment, the fixed function block 2530 also includes a graphics SoC interface 2537, a graphics microcontroller 2538, and a media pipeline 2539. In at least one embodiment, the fixed graphics SoC interface 2537 provides an interface between the graphics core 2500 and other processor cores within the system-on-chip integrated circuit. In at least one embodiment, the graphics microcontroller 2538 is a programmable sub-processor that can be configured to manage various functions of the graphics processor 2500, including thread dispatch, scheduling, and preemption. In at least one embodiment, the media pipeline 2539 includes logic for facilitating the decoding, encoding, preprocessing, and / or postprocessing of multimedia data, including image and video data. In at least one embodiment, the media pipeline 2539 implements media operations via requests to compute logic or sampling logic within sub-cores 2501-2501F.
[0233] In at least one embodiment, the SoC interface 2537 enables the graphics core 2500 to communicate with other components within the SoC, including a general-purpose application processor core (e.g., a CPU), and / or memory hierarchy elements such as a shared last-level cache memory, system RAM, and / or an embedded on-chip or on-package DRAM. In at least one embodiment, the SoC interface 2537 also enables communication with fixed-function devices within the SoC, such as a camera imaging pipeline, enables and / or implements the use of global memory atomics that can be shared between the graphics core 2500 and a CPU within the SoC. In at least one embodiment, the SoC interface 2537 also implements power management control for the graphics core 2500 and can enable an interface between the clock domain of the graphics core 2500 and other clock domains within the SoC. In at least one embodiment, the SoC interface 2537 enables receipt of command buffers from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, the commands and instructions can be dispatched to the media pipeline 2539 when a media operation is to be executed, or to a geometry and fixed-function pipeline (e.g., geometry and fixed-function pipeline 2536, geometry and fixed-function pipeline 2514) when a graphics processing operation is to be executed.
[0234] In at least one embodiment, the graphics microcontroller 2538 can be configured to perform various scheduling and management tasks for the graphics core 2500. In at least one embodiment, the graphics microcontroller 2538 can execute graphics for various graphics parallel engines within the execution unit (EU) arrays 2502A - 2502F, 2504A - 2504F in the sub - cores 2501A - 2501F and / or calculate workload scheduling. In at least one embodiment, the host software running on the CPU core of the SoC including the graphics core 2500 can submit a workload to one of a plurality of graphics processor doorbells that call the scheduling operation for the appropriate graphics engine. In at least one embodiment, the scheduling operation includes determining which workload to execute next, submitting the workload to the command streamer, pre - empting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, the graphics microcontroller 2538 can also provide the graphics core 2500 with the ability to save and restore registers within the graphics core 2500 across the transition to the low - power state independent of the operating system and / or the graphics driver software on the system to facilitate the low - power or idle state of the graphics core 2500.
[0235] In at least one embodiment, the graphics core 2500 may have up to N modular sub-cores, more or fewer than the illustrated sub-cores 2501A - 2501F. For each set of N sub-cores, in at least one embodiment, the graphics core 2500 may also include shared function logic 2510, shared and / or cache memory 2512, geometry / fixed function pipeline 2514, and additional fixed function logic 2516 for accelerating various graphics and computing processing operations. In at least one embodiment, the shared function logic 2510 may include logic units (e.g., samplers, math, and / or inter-thread communication logic) that can be shared by each of the N sub-cores within the graphics core 2500. In at least one embodiment, the fixed shared and / or cache memory 2512 can serve as the last-level cache for the N sub-cores 2501A - 2501F within the graphics core 2500 and can also act as shared memory accessible by multiple sub-cores. In at least one embodiment, the geometry / fixed function pipeline 2514 may be included instead of the geometry / fixed function pipeline 2536 within the fixed function block 2530 and may include the same or similar logic units.
[0236] In at least one embodiment, the graphics core 2500 includes additional fixed function logic 2516 that can include various fixed function acceleration logic for use by the graphics core 2500. In at least one embodiment, the additional fixed function logic 2516 includes an additional geometry pipeline for use in position-limited shading. In position-limited shading, there are at least two geometry pipelines in the geometry / fixed function pipelines 2516, 2536, the full geometry pipeline, and a culling pipeline that can be included within the additional fixed function logic 2516. In at least one embodiment, the culling pipeline is a reduced version of the full geometry pipeline. In at least one embodiment, the full pipeline and the culling pipeline can execute different instances of an application, and each instance has a separate context. In at least one embodiment, position-limited shading can hide long culling runs of discarded triangles and can potentially complete shading faster in some instances. For example, in at least one embodiment, the culling pipeline fetches and shades the vertex position attributes without performing pixel rasterization and rendering to the frame buffer, so the culling pipeline logic within the additional fixed function logic 2516 can execute the position shader in parallel with the main application and generally produce critical results faster than the full pipeline. In at least one embodiment, the culling pipeline can use the generated critical results to compute visibility information for all triangles, regardless of whether those triangles are culled. In at least one embodiment, the full pipeline (which may be called the replay pipeline in this instance) can consume the visibility information to skip culled triangles in order to shade only the visible triangles that are ultimately passed to the rasterization phase.
[0237] In at least one embodiment, the additional fixed function logic 2516 can also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for implementations that include optimizations for machine learning training or inference.
[0238] In at least one embodiment, within each of the graphics sub-cores 2501A-2501F, a set of execution resources may be included for performing graphics operations, media operations, and compute operations in response to requests from a graphics pipeline, a media pipeline, or a shader program. In at least one embodiment, the graphics sub-cores 2501A-2501F include a plurality of EU arrays 2502A-2502F, 2504A-2504F, thread dispatch and inter-thread communication (TD / IC) logic 2503A-2503F, 3D (e.g., texture) samplers 2505A-2505F, media samplers 2506A-2506F, shader processors 2507A-2507F, and shared local memory (SLM). The EU arrays 2502A-2502F, 2504A-2504F each include a plurality of execution units that are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations in the service of graphics operations, media operations, or compute operations including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 2503A-2503F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between the threads executing on the execution units of the sub-core. In at least one embodiment, the 3D samplers 2505A-2505F can read texture or other 3D graphics-related data into memory. In at least one embodiment, the 3D sampler can read texture data in different ways based on the configured sample state associated with a given texture and the texture format. In at least one embodiment, the media samplers 2506A-2506F can perform similar read operations based on the type and format associated with the media data.In at least one embodiment, each of the graphics sub-cores 2501A - 2501F can alternately include an integrated sampler for 3D and media. In at least one embodiment, threads executing on execution units within each of the sub-cores 2501A - 2501F can utilize the shared local memories 2508A - 2508F within each sub-core to enable threads executing within a thread group to execute using a common pool of on-chip memory.
[0239] Inference and / or training logic 615 is used to perform inference and / or training operations related to one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the graphics processor 2510. For example, in at least one embodiment, the training and / or inference techniques described herein may be implemented using one or more of the ALUs implemented in the graphics processor 2312, the graphics microcontroller 2538, the geometry and fixed function pipelines 2514 and 2536, or other logic in FIG. 24. Additionally, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIGS. 6A or 6B. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 2500 to execute one or more of the machine learning algorithms, neural network architecture use cases, or training techniques described herein.
[0240] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0241] FIGS. 26A and 26B illustrate thread execution logic 2600 including an array of processing elements of a graphics processor core, according to at least one embodiment. FIG. 26A illustrates at least one embodiment in which thread execution logic 2600 is used. FIG. 26B is a diagram illustrating exemplary inner details of an execution unit, according to at least one embodiment.
[0242] As shown in FIG. 26A, in at least one embodiment, the thread execution logic 2600 includes a shader processor 2602, a thread dispatcher 2604, an instruction cache 2606, a scalable execution unit array including a plurality of execution units 2608A - 2608N, a sampler 2610, a data cache 2612, and a data port 2614. In at least one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any of execution units 2608A, 2608B, 2608C, 2608D - 2608N - 1, and 2608N) based on, for example, the computational requirements of the workload. In at least one embodiment, the scalable execution units are interconnected via an interconnect fabric linked to each of the execution units. In at least one embodiment, the thread execution logic 2600 includes one or more connections to memory such as system memory or cache memory via the instruction cache 2606, the data port 2614, the sampler 2610, and one or more of the execution units 2608A - 2608N. In at least one embodiment, each execution unit (e.g., 2608A) is a stand - alone programmable general - purpose computing unit that can execute a plurality of simultaneous hardware threads while processing a plurality of data elements in parallel for each thread. In at least one embodiment, the array of execution units 2608A - 2608N is scalable to include any number of individual execution units.
[0243] In at least one embodiment, execution units 2608A - 2608N are primarily used to execute shader programs. In at least one embodiment, shader processor 2602 can process various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 2604. In at least one embodiment, thread dispatcher 2604 arbitrates thread start requests from the graphics and media pipeline and includes logic for instantiating the requested threads on one or more of execution units 2608A - 2608N. For example, in at least one embodiment, the geometry pipeline can dispatch a vertex shader, a tessellation shader, or a geometry shader to thread execution logic for processing. In at least one embodiment, thread dispatcher 2604 can also process runtime thread spawning requests from the executing shader program.
[0244] In at least one embodiment, execution units 2608A - 2608N support an instruction set that includes native support for many standard 3D graphics shader instructions, such that shader programs from graphics libraries (e.g., Direct3D and OpenGL) are executed with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders). In at least one embodiment, each of execution units 2608A - 2608N, which includes one or more arithmetic - logic units (ALUs), can issue multiple single - instruction multiple - data (SIMD) executions, enabling an efficient execution environment despite high memory - access latency through multithreaded operation. In at least one embodiment, each hardware thread within each execution unit has a dedicated high - bandwidth register file and associated independent thread state. In at least one embodiment, execution is issued multiple times per clock to a pipeline that can perform integer arithmetic, single - precision and double - precision floating - point arithmetic, SIMD branch performance, logical operations, transcendental operations, and various other operations. In at least one embodiment, while waiting for data from memory or one of the shared functions, the dependent logic within execution units 2608A - 2608N puts the waiting threads to sleep until the requested data is returned. In at least one embodiment, while the waiting threads are asleep, the hardware resources may be dedicated to the processing of other threads. For example, in at least one embodiment, during the latency associated with vertex - shader operation, the execution unit can execute a different type of shader program, including a pixel shader, a fragment shader, or a different vertex shader.
[0245] In at least one embodiment, each execution unit of execution units 2608A - 2608N operates on an array of data elements. In at least one embodiment, the number of data elements is the "execution size" or the number of channels for an instruction. In at least one embodiment, an execution channel is a logical unit of execution related to access, masking, and flow control within an instruction for data elements. In at least one embodiment, the number of channels may be independent of the number of physical arithmetic logic units (ALUs) or floating - point units (FPUs) for a particular graphics processor. In at least one embodiment, execution units 2608A - 2608N may support integer and floating - point data types.
[0246] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements may be stored in a register as a packed data type, and the execution unit processes various elements based on the data size of the elements. For example, in at least one embodiment, when operating on a 256 - bit wide vector, the 256 bits of the vector are stored in a register, and the execution unit operates on the vector as four separate 64 - bit packed data elements (data elements of quad - word (QW) size), eight separate 32 - bit packed data elements (data elements of double - word (DW) size), sixteen separate 16 - bit packed data elements (data elements of word (W) size), or thirty - two separate 8 - bit data elements (data elements of byte (B) size). However, in at least one embodiment, different vector widths and register sizes are conceivable.
[0247] In at least one embodiment, one or more execution units can be combined to form fused execution units 2609A - 2609N having thread control logic (2607A - 2607N) common to the fused EUs. In at least one embodiment, multiple EUs can be fused to form an EU group. In at least one embodiment, each EU in the fused EU group can be configured to execute separate SIMD hardware threads. The number of EUs in the fused EU group can vary according to different embodiments. In at least one embodiment, various SIMD widths including, but not limited to, SIMD8, SIMD16, and SIMD32 can be executed for each EU. In at least one embodiment, each fused graphics execution unit 2609A - 2609N includes at least two execution units. For example, in at least one embodiment, fused execution unit 2609A includes a first EU 2608A, a second EU 2608B, and thread control logic 2607A common to the first EU 2608A and the second EU 2608A. In at least one embodiment, thread control logic 2607A controls the threads executed in fused graphics execution unit 2609A so that each EU within fused execution units 2609A - 2609N can be executed using a common instruction pointer register.
[0248] In at least one embodiment, one or more internal instruction caches (e.g., 2606) are included in thread execution logic 2600 to cache thread instructions for the execution units. In at least one embodiment, one or more data caches (e.g., 2612) are included to cache thread data during thread execution. In at least one embodiment, sampler 2610 is included to perform texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, sampler 2610 includes special texture or media sampling functions and processes texture or media data during sampling before providing the sampled data to the execution units.
[0249] During execution, in at least one embodiment, the graphics and media pipeline sends thread start requests to thread execution logic 2600 via thread spawning and dispatch logic. In at least one embodiment, when a group of geometric objects is processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 2602 is called to further compute output information and write the results to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader computes the values of various vertex attributes that will be interpolated between the rasterized objects. In at least one embodiment, the pixel processor logic within shader processor 2602 then executes a pixel shader program or fragment shader program with an application programming interface (API). In at least one embodiment, to execute the shader program, shader processor 2602 dispatches threads to execution units (e.g., 2608A) via thread dispatcher 2604. In at least one embodiment, shader processor 2602 uses the texture sampling logic of sampler 2610 to access texture data of a texture map stored in memory. In at least one embodiment, through arithmetic operations on the texture data and input geometry data, the pixel color data for each geometry fragment is computed or one or more pixels are discarded so that they are not further processed.
[0250] In at least one embodiment, data port 2614 provides a memory access mechanism for thread execution logic 2600 to output processed data to memory so that it can be further processed in the graphics processor output pipeline. In at least one embodiment, data port 2614 includes or is coupled to one or more cache memories (e.g., data cache 2612) to cache data for memory access through the data port.
[0251] As shown in FIG. 26B, in at least one embodiment, graphics execution unit 2608 can include instruction fetch unit 2637, general register file array (GRF) 2624, architecture register file array (ARF) 2626, thread arbiter 2622, dispatch unit 2630, branch unit 2632, a set of SIMD floating point units (FPUs) 2634, and in at least one embodiment, a set of dedicated integer SIMD ALUs 2635. In at least one embodiment, GRF 2624 and ARF 2626 include a set of general register files and architecture register files associated with each simultaneous hardware thread, and this hardware thread may be active in graphics execution unit 2608. In at least one embodiment, the per-thread architecture state is maintained in ARF 2626, and the data used during thread execution is stored in GRF 2624. In at least one embodiment, the execution state of each thread, including the instruction pointer for each thread, can be held in the thread-specific registers of ARF 2626.
[0252] In at least one embodiment, the graphics execution unit 2608 has an architecture that is a combination of simultaneous multi-threading (SMT) and fine-grained interleaved multi-threading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, where the resources of the execution unit are divided across the logic used to execute multiple simultaneous threads.
[0253] In at least one embodiment, the graphics execution unit 2608 can co-issue a plurality of instructions, which may each be different instructions. In at least one embodiment, the thread arbitration device 2622 of the graphics execution unit thread 2608 can enable an instruction to be dispatched for execution to one of the transmission unit 2630, the branch unit 2642, or the SIMD FPU 2634. In at least one embodiment, each execution thread can access 128 general-purpose registers in the GRF 2624, where each register can store 32 bytes that are accessible as a vector of SIMD8 elements of 32-bit data elements. In at least one embodiment, each execution unit thread can access 4 kilobytes in the GRF 2624, but the embodiments are not so limited, and more or fewer resources may be provided in other embodiments. In at least one embodiment, up to 7 threads can be executed simultaneously, but the number of threads per execution unit can also vary depending on the embodiment. In at least one embodiment where 7 threads can access 4 kilobytes, the GRF 2624 can store a total of 28 kilobytes. In at least one embodiment, a flexible addressing mode can enable multiple registers to be addressed together to construct wider registers or represent a strided rectangular block data structure.
[0254] In at least one embodiment, memory operations, sampler operations, and other long-latency system communications are dispatched via "send" instructions executed by the message delivery sending unit 2630. In at least one embodiment, branch instructions are dispatched to a dedicated branch unit 2632 to facilitate SIMD divergence and eventual convergence.
[0255] In at least one embodiment, the graphics execution unit 2608 includes one or more SIMD floating-point units (FPUs) 2634 for performing floating-point operations. In at least one embodiment, the FPU 2634 also supports integer calculations. In at least one embodiment, the FPU 2634 can perform up to M 32-bit floating-point (or integer) operations in SIMD, or up to 2M 16-bit integer operations, or 16-bit floating-point operations in SIMD. In at least one embodiment, at least one of the FPUs provides extended mathematical functions to support high-throughput transcendental mathematical functions and double-precision 64-bit floating-point. In at least one embodiment, a set of 8-bit integer SIMD ALUs 2635 may also be present and may be specifically optimized to perform operations related to machine learning calculations.
[0256] In at least one embodiment, an array of multiple instances of the graphics execution unit 2608 may be instantiated in a graphics sub-core group (e.g., a sub-slice). In at least one embodiment, the execution unit 2608 can execute instructions across multiple execution channels. In at least one embodiment, each thread executed by the graphics execution unit 2608 is executed on a different channel.
[0257] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the execution logic 2600. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIGS. 6A or 6B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that make up the ALU of the execution logic 2600 for performing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0258] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0259] FIG. 27 shows a parallel processing unit (PPU) 2700 according to at least one embodiment. In at least one embodiment, the PPU 2700 is composed of machine-readable code that causes the PPU 2700 to execute some or all of the processes and techniques described throughout this disclosure when executed by the PPU 2700. In at least one embodiment, the PPU 2700 is a multi-threaded processor, which is implemented on one or more integrated circuit devices and utilizes multi-threading as a latency hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) in parallel across multiple threads. In at least one embodiment, a thread refers to an execution thread and is an instantiation of a set of instructions configured to be executed by the PPU 2700. In at least one embodiment, the PPU 2700 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing 3D graphics data in order to generate 2D image data that can be displayed on a display device such as a liquid crystal display (LCD) device. In at least one embodiment, the PPU 2700 is utilized to perform calculations such as linear algebra operations and machine learning operations. FIG. 27 shows an exemplary parallel processor for illustrative purposes only and should be construed as a non-limiting example of a processor architecture contemplated within the scope of this disclosure, and it should be construed that any suitable processor may be utilized to add to and / or replace the same processor.
[0260] In at least one embodiment, one or more PPU2700s are configured to accelerate applications in high performance computing ("HPC"), data centers, and machine learning. In at least one embodiment, PPU2700 is configured to accelerate deep learning systems and applications including the following non-limiting examples: autonomous vehicle platforms, deep learning, high-precision audio, images, text recognition systems, intelligent video analysis, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, etc.
[0261] In at least one embodiment, PPU2700 includes, without limitation, an input / output ("I / O") unit 2706, a front-end unit 2710, a scheduler unit 2712, a work distribution unit 2714, a hub 2716, a crossbar ("Xbar") 2720, one or more general processing clusters ("GPC") 2718, and one or more partition units ("memory partition units") 2722. In at least one embodiment, PPU2700 is connected to a host processor or another PPU2700 via one or more high-speed GPU interconnects ("GPU interconnects") 2708. In at least one embodiment, PPU2700 is connected to a host processor or other peripheral devices via interconnect 2702. In at least one embodiment, PPU2700 is connected to a local memory having one or more memory devices ("memory"). In at least one embodiment, memory device 2704 includes, without limitation, one or more dynamic random access memory ("DRAM") devices. In at least one embodiment, one or more DRAM devices may be configured as, and / or be configurable as, a high bandwidth memory ("HBM") subsystem in which multiple DRAM dies are stacked within each device.
[0262] In at least one embodiment, the high-speed GPU interconnect 2708 may refer to a wired-based multi-lane communication link that is used by the system for scaling and that includes one or more PPU 2700s in combination with one or more central processing units (“CPUs”), and supports cache coherence and CPU mastering between the PPU 2700s and the CPUs. In at least one embodiment, data and / or commands are transmitted to / from another unit of the PPU 2700, such as one or more copy engines, video encoders, video decoders, power management units, and other components not explicitly shown in FIG. 27, via the high-speed GPU interconnect 2708 and through the hub 2716.
[0263] In at least one embodiment, the I / O unit 2706 is configured to communicate (e.g., commands, data) with a host processor (not shown in FIG. 27) via the system bus 2702. In at least one embodiment, the I / O unit 2706 communicates with the host processor directly via the system bus 2702 or via one or more intermediate devices such as one or more memory bridges. In at least one embodiment, the I / O unit 2706 may communicate with one or more other processors such as one or more of the PPU 2700s via the system bus 2702. In at least one embodiment, the I / O unit 2706 implements a peripheral component interconnect express (“PCIe”) interface to enable communication via a PCIe bus. In at least one embodiment, the I / O unit 2706 implements an interface for communicating with external devices.
[0264] In at least one embodiment, the I / O unit 2706 decodes packets received via the system bus 2702. In at least one embodiment, at least some of the packets represent commands configured to cause the PPU 2700 to perform various operations. In at least one embodiment, the I / O unit 2706 transmits the decoded commands to various other units of the PPU 2700 specified by the commands. In at least one embodiment, the commands are transmitted to the front-end unit 2710 and / or to the hub 2716 or to one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown in FIG. 27) of the PPU 2700. In at least one embodiment, the I / O unit 2706 is configured to route communications between various logical units of the PPU 2700.
[0265] In at least one embodiment, a program executed by the host processor encodes a command stream in a buffer that enables the PPU 2700 to receive and process a workload. In at least one embodiment, the workload includes instructions and data to be processed by these instructions. In at least one embodiment, the buffer is an area in memory accessible (e.g., writable / readable) by both the host processor and the PPU 2700, and the host interface unit may be configured to access the buffer in system memory connected to the system bus 2702 via a memory request transmitted via the system bus 2702 by the I / O unit 2706. In at least one embodiment, the host processor writes the command stream to the buffer and then transmits a pointer indicating the start point of the command stream to the PPU 2700, whereby the front-end unit 2710 receives a pointer indicating one or more command streams, manages the one or more command streams, reads commands from the command stream, and transfers the commands to various units of the PPU 2700.
[0266] In at least one embodiment, the front - end unit 2710 is coupled to a scheduler unit 2712 that configures various GPCs 2718 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 2712 is configured to track state information related to various tasks managed by the scheduler unit 2712, where the state information may indicate which GPC 2718 a task is assigned to, whether the task is active or inactive, the priority level associated with the task, and the like. In at least one embodiment, the scheduler unit 2712 manages the execution of multiple tasks in one or more of the GPCs 2718.
[0267] In at least one embodiment, the scheduler unit 2712 is coupled to a work distribution unit 2714 configured to dispatch tasks for execution on the GPCs 2718. In at least one embodiment, the work distribution unit 2714 tracks the number of scheduled tasks received from the scheduler unit 2712, and the work distribution unit 2714 manages a pending task pool and an active task pool for each of the GPCs 2718. In at least one embodiment, the pending task pool comprises a number of slots (e.g., 32 slots) for tasks assigned to be processed by a particular GPC 2718, and the active task pool comprises a number of slots (e.g., 4 slots) for tasks being actively processed by the GPC 2718, such that when one of the GPCs 2718 completes execution of a task, that task is removed from the active task pool of the GPC 2718 and one of the other tasks from the pending task pool is selected and scheduled for execution on the GPC 2718. In at least one embodiment, when an active task is idle on a GPC 2718, such as while waiting for data dependencies to be resolved, the active task is removed from the GPC 2718 and returned to the pending task pool, during which time another task from the pending task pool is selected and scheduled for execution on the GPC 2718.
[0268] In at least one embodiment, the work distribution unit 2714 communicates with one or more GPCs 2718 via an X-bar 2720. In at least one embodiment, the X-bar 2720 is an interconnect network that couples many of the units of the PPU 2700 to other units of the PPU 2700 and can be configured to couple the work distribution unit 2714 to a particular GPC 2718. In at least one embodiment, one or more other units of the PPU 2700 may also be connected to the X-bar 2720 via a hub 2716.
[0269] In at least one embodiment, the task is managed by a scheduler unit 2712 and dispatched to one of the GPCs 2718 by a work distribution unit 2714. The GPC 2718 is configured to process the task and generate a result. In at least one embodiment, the result may be consumed by other tasks within the GPC 2718, routed to a different GPC 2718 via the X-bar 2720, or stored in the memory 2704. In at least one embodiment, the result can be written to the memory 2704 via a partition unit 2722, which implements a memory interface for reading and writing data to / from the memory 2704. In at least one embodiment, the result can be sent to another PPU 2704 or CPU via the high-speed GPU interconnect 2708. In at least one embodiment, the PPU 2700 includes, without limitation, U partition units 2722 equal in number to the number of separate individual memory devices 2704 coupled to the PPU 2700. In at least one embodiment, the partition unit 2722 is further described in more detail below in conjunction with FIG. 29.
[0270] In at least one embodiment, the host processor executes a driver kernel that implements an application programming interface (API) that enables scheduling of operations for one or more applications executing on the host processor to be executed on the PPU 2700. In at least one embodiment, multiple compute applications are executed simultaneously by the PPU 2700, and the PPU 2700 provides isolation, quality of service (“QoS”), and independent address spaces for the multiple compute applications. In at least one embodiment, an application generates instructions (e.g., in the form of API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 2700, and the driver kernel outputs the tasks to one or more streams being processed by the PPU 2700. In at least one embodiment, each task comprises one or more groups of related threads, which may be referred to as warps. In at least one embodiment, a warp comprises multiple related threads (e.g., 32 threads) that can be executed in parallel. In at least one embodiment, cooperating threads may refer to multiple threads that include instructions for executing a task and exchange data via shared memory. In at least one embodiment, threads and cooperating threads are described in further detail in conjunction with FIG. 29, in accordance with at least one embodiment.
[0271] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, a deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the PPU 2700. In at least one embodiment, the PPU 2700 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the PPU 2700. In at least one embodiment, the PPU 2700 may be used to perform one or more neural network use cases described herein.
[0272] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0273] FIG. 28 shows a general purpose processing cluster (``GPC'') 2800 according to at least one embodiment. In at least one embodiment, GPC 2800 is GPC 2718 of FIG. 27. In at least one embodiment, each GPC 2800 includes, without limitation, several hardware units for processing tasks, and each GPC 2800 includes, without limitation, a pipeline manager 2802, a pre-raster operations unit (``PROP'') 2804, a raster engine 2808, a work distribution crossbar (``WDX'') 2816, a memory management unit (``MMU'') 2818, one or more data processing clusters (``DPC'': Data Processing Clusters) 2806, and any suitable combination of parts.
[0274] In at least one embodiment, the operation of GPC2800 is controlled by pipeline manager 2802. In at least one embodiment, pipeline manager 2802 manages the configuration of one or more DPCs 2806 to process tasks assigned to GPC2800. In at least one embodiment, pipeline manager 2802 configures at least one of one or more DPCs 2806 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, DPC2806 is configured to execute a vertex shader program on programmable streaming multi-processor ("SM") 2814. In at least one embodiment, pipeline manager 2802 is configured to route packets received from a work distribution unit to appropriate logical units within GPC2800, and some packets may be routed to fixed function hardware unit PROP2804 and / or raster engine 2808, and other packets may be routed to DPC2806 to be processed by primitive engine 2812 or SM2814. In at least one embodiment, pipeline manager 2802 configures at least one of DPCs 2806 to implement a neural network model and / or a computing pipeline.
[0275] In at least one embodiment, the PROP unit 2804 is configured to route data generated by the raster engine 2808 and the DPC 2806 to the raster operation (ROP) unit of the partition unit 2722, which was described in more detail above in conjunction with FIG. 27. In at least one embodiment, the PROP unit 2804 is configured to perform optimizations for color blending, organize pixel data, perform address translation, and perform other operations. In at least one embodiment, the raster engine 2808 includes, without limitation, several fixed-function hardware units configured to perform various raster operations in at least one embodiment, and the raster engine 2808 includes, without limitation, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile combination engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives the transformed vertices, generates a plane equation associated with the geometric primitive defined by the vertices, and the plane equation is sent to the coarse raster engine to generate coverage information for the primitive (e.g., the x, y coverage mask of the tile), and the output of the coarse raster engine is sent to the culling engine, where fragments associated with primitives that fail the z-test are culled and sent to the clipping engine, where fragments outside the frustum are clipped. In at least one embodiment, the fragments that have passed clipping and culling are passed to the fine raster engine to generate attributes for the pixel fragments based on the plane equation generated by the setup engine. In at least one embodiment, the output of the raster engine 2808 includes fragments that will be processed by any suitable entity, such as by a fragment shader implemented within the DPC 2806.
[0276] In at least one embodiment, each DPC2806 included in GPC2800 includes, without limitation, an M-Pipe Controller (MPC) 2810, a primitive engine 2812, one or more SMs 2814, and any suitable combination thereof. In at least one embodiment, MPC2810 controls the operation of DPC2806 to route packets received from pipeline manager 2802 to appropriate units within DPC2806. In at least one embodiment, packets associated with vertices are routed to a primitive engine 2812 configured to fetch vertex attributes associated with the vertices from memory, whereas, in contrast, packets associated with shader programs may be sent to SM2814.
[0277] In at least one embodiment, SM2814 includes, without limitation, a programmable streaming processor configured to process tasks represented by a number of threads. In at least one embodiment, SM2814 is multi-threaded and configured to execute multiple threads (e.g., 32 threads) from a particular group of threads simultaneously, implementing a single instruction multiple data (SIMD) architecture, where each thread within a group of threads (warp) is configured to process a different data set based on the same instruction set. In at least one embodiment, all threads within a thread group execute the same instruction. In at least one embodiment, SM2814 implements a single instruction multiple thread (SIMT) architecture, where each thread of a thread group is configured to process a different data set based on the same set of instructions, but individual threads within a thread group are allowed to diverge during execution. In at least one embodiment, the program counter, call stack, and execution state are maintained per warp, enabling simultaneous processing between warps and serial execution within a warp when threads within a warp diverge. In another embodiment, the program counter, call stack, and execution state are maintained per individual thread, enabling equal simultaneous processing between all threads, within a warp, and between warps. In at least one embodiment, the execution state is maintained per individual thread, and threads executing the same instruction may be converged and executed in parallel for greater efficiency. At least one embodiment of SM2814 is described in further detail below.
[0278] In at least one embodiment, the MMU 2818 provides an interface between the GPC 2800 and a memory partition unit (e.g., partition unit 2722 of FIG. 27), and the MMU 2818 provides virtual address to physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, the MMU 2818 provides one or more translation lookaside buffers ("TLBs") for performing translation from virtual addresses to physical addresses of memory.
[0279] In order to perform inference and / or training operations related to one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the GPC 2800. In at least one embodiment, the GPC 2800 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system, or by the GPC 2800. In at least one embodiment, the GPC 2800 may be used to execute one or more use cases of the neural networks described herein.
[0280] In order to perform inference and / or training operations related to one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically and independently generated.
[0281] FIG. 29 shows a memory partition unit 2900 of a parallel processing unit (PPU) according to at least one embodiment. In at least one embodiment, the partition unit 2900 includes, without limitation, a raster operation (ROP) unit 2902, a level 2 (L2) cache 2904, a memory interface 2906, and any suitable combination thereof. In at least one embodiment, the memory interface 2906 is coupled to the memory. In at least one embodiment, the memory interface 2906 may implement a 32, 64, 128, 1024-bit data bus, or a similar implementation, for high-speed data transfer. In at least one embodiment, the PPU incorporates U memory interfaces 2906 per pair of the partition unit 2900 with one memory interface 2906 per pair, where each pair of the partition unit 2900 is connected to a corresponding memory device. For example, in at least one embodiment, the PPU may be connected to a maximum of Y memory devices, such as a high-bandwidth memory stack, or a graphics double data rate, version 5, synchronous dynamic random access memory (GDDR5 SDRAM).
[0282] In at least one embodiment, the memory interface 2906 implements a second generation high bandwidth memory (“HBM2”) memory interface, and Y is equal to half of U. In at least one embodiment, the HBM2 memory stack is positioned in the same physical package as the PPU, realizing substantial power and area savings compared to conventional GDDR5 SDRAM systems. In at least one embodiment, each HBM2 stack includes, without limitation, four memory dies, Y is equal to 4, and each HBM2 stack includes a total of eight channels of two 128-bit channels per die and a data bus width of 1024 bits. In at least one embodiment, the memory supports a single-bit error correction two-bit error detection (“SECDED”) error correction code (“ECC”) to protect data. In at least one embodiment, the ECC provides higher reliability for compute applications vulnerable to data corruption.
[0283] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partition unit 2900 supports integrated memory to provide a single integrated virtual address space for the central processing unit (“CPU”) and PPU memory, enabling sharing of data between virtual memory systems. In at least one embodiment, the frequency of access by the PPU to memory located in other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the pages more frequently. In at least one embodiment, the high-speed GPU interconnect 2708 supports an address translation service to enable the PPU to directly access the CPU's page table, realizing full access by the PPU to CPU memory.
[0284] In at least one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In at least one embodiment, the copy engine can generate a page error for an address not mapped to a page table, and then the memory partition unit 2900 maps the address to the page table in response to the page error, and then the copy engine executes the transfer. In at least one embodiment, the memory is pinned (e.g., made non-pageable) for multiple operations of the copy engine among multiple processors, reducing substantially available memory. In at least one embodiment, in the case of a hardware page error, the address can be passed to the copy engine regardless of whether the memory page is resident, and the copy process is transparent.
[0285] According to at least one embodiment, data from the memory 2704 of FIG. 27 or other system memory is fetched by the memory partition unit 2900 and stored in the L2 cache 2904, which is located on-chip and shared among various GPCs. In at least one embodiment, each memory partition unit 2900 includes, without limitation, at least a portion of the L2 cache associated with the corresponding memory device. In at least one embodiment, lower-level caches are implemented in various units within the GPC. In at least one embodiment, each of the SM2814s may implement a level 1 ("L1") cache, where the L1 cache is private memory dedicated to a particular SM2814, and data from the L2 cache 2904 is fetched and stored in each of the L1 caches for processing by the functional units of the SM2814. In at least one embodiment, the L2 cache 2904 is coupled to the memory interface 2906 and the X bar 2720.
[0286] In at least one embodiment, the ROP unit 2902 performs graphics raster operations related to pixel colors, such as color compression and pixel blending. In at least one embodiment, the ROP unit 2902, in conjunction with the raster engine 2808, implements a depth test to receive the depth of the sample location associated with the pixel fragment from the culling engine of the raster engine 2808. In at least one embodiment, the depth is tested against the corresponding depth in the depth buffer of the sample location associated with the fragment. In at least one embodiment, when the fragment passes the depth test of the sample location, the ROP unit 2902 updates the depth buffer and sends the result of the depth test to the raster engine 2808. It will be appreciated that the number of partition units 2900 may be different from the number of GPCs, and thus each ROP unit 2902 may be coupled to each of the GPCs in at least one embodiment. In at least one embodiment, the ROP unit 2902 tracks packets received from different GPCs and determines that the results generated by the ROP unit 2902 are routed through the X-bar 2720.
[0287] FIG. 30 shows a streaming multi-processor ( "SM") 3000 according to at least one embodiment. In at least one embodiment, SM3000 is the SM2814 of FIG. 28. In at least one embodiment, SM3000 includes, without limitation, an instruction cache 3002, one or more scheduler units 3004, a register file 3008, one or more processing cores ( "cores") 3010, one or more special function units ( "SFU": special function unit) 3012, one or more load / store units ( "LSU" load / store unit) 3014, an interconnect network 3016, a shared memory / level 1 ( "L1") cache 3018, and any suitable combination thereof. In at least one embodiment, the work distribution unit dispatches tasks for execution in the general-purpose processing cluster ( "GPC") of the parallel processing unit ( "PPU"), and each task is distributed to a specific data processing cluster ( "DPC") within the GPC. When the task is related to a shader program, the task is distributed to one of the SM3000. In at least one embodiment, the scheduler unit 3004 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks assigned to the SM3000. In at least one embodiment, the scheduler unit 3004 schedules thread blocks so that they can be executed as warps of parallel threads, where each thread block is distributed to at least one warp. In at least one embodiment, each warp executes threads. In at least one embodiment, the scheduler unit 3004 manages multiple different thread blocks, distributes warps to different thread blocks, and then dispatches instructions from multiple different synchronization groups to various functional units (e.g., processing core 3010, SFU 3012, and LSU 3014) during each clock cycle.
[0288] In at least one embodiment, a cooperative group refers to a programming model for organizing a group of communicating threads, which enables a developer to express the granularity at which threads communicate and allows for a richer and more efficient representation of parallel decomposition. In at least one embodiment, the cooperative launch API supports synchronization between thread blocks so that a parallel algorithm can be executed. In at least one embodiment, applications of conventional programming models provide a single simple construct for synchronizing cooperating threads, namely a barrier (e.g., the syncthreads() function) across all threads of a thread block. However, in at least one embodiment, a programmer may define thread groups smaller than the granularity of a thread block, synchronize within the defined groups, and enable higher performance, design flexibility, and software reuse in the form of a collective functional interface across the entire set of groups. In at least one embodiment, the cooperative group allows a programmer to explicitly define a group of threads at the granularity of sub-blocks (i.e., the same size as a single thread) and at the granularity of multi-blocks, and to perform collective operations such as synchronization on the threads within the cooperative group. In at least one embodiment, the programming model supports clean composition across software boundaries, thereby enabling libraries and utility functions to synchronize safely within their local contexts without the need to make assumptions about convergence. In at least one embodiment, the primitives of the cooperative group enable a new pattern of cooperative parallelism that includes producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire grid of thread blocks without limiting them.
[0289] In at least one embodiment, the dispatch unit 3006 is configured to send instructions to one or more of the functional units, and the scheduler unit 3004 includes two dispatch units 3006 without limitation that enable two different instructions from the same warp to be dispatched during each clock cycle. In at least one embodiment, each scheduler unit 3004 includes a single dispatch unit 3006 or an additional dispatch unit 3006.
[0290] In at least one embodiment, each SM3000 includes, without limitation, a register file 3008 that provides a set of registers to the functional units of the SM3000 in at least one embodiment. In at least one embodiment, the register file 3008 is divided among the functional units such that each functional unit is allocated a dedicated portion of the register file 3008. In at least one embodiment, the register file 3008 is divided among different warps being executed by the SM3000, and the register file 3008 provides temporary storage for operands connected to the data paths of the functional units. In at least one embodiment, each SM3000 includes, without limitation, a plurality of L processing cores 3010. In at least one embodiment, each SM3000 includes, without limitation, a large number (e.g., 128 or more) of individual processing cores 3010 in at least one embodiment. In at least one embodiment, each processing core 3010 includes, without limitation, a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit that includes a floating-point arithmetic logic unit and an integer arithmetic logic unit without limitation. In at least one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point operations. In at least one embodiment, the processing core 3010 includes, without limitation, 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.
[0291] The tensor core is configured to perform matrix operations according to at least one embodiment. In at least one embodiment, one or more tensor cores are included in the processing core 3010. In at least one embodiment, the tensor core is configured to perform deep learning matrix operations such as convolution operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4×4 matrix and performs a matrix multiply and accumulate operation D = A×B + C, where A, B, C, and D are 4×4 matrices.
[0292] In at least one embodiment, the input matrices A and B for matrix multiplication are 16-bit floating-point matrices, and the sum matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor core operates on 16-bit floating-point input data with a 32-bit floating-point sum. In at least one embodiment, the 16-bit floating-point multiplication uses 64 operations, resulting in a full-precision product, which is then added using 32-bit floating-point addition with other intermediate products of a 4×4×4 matrix multiplication. Using the tensor core, in at least one embodiment, much larger two-dimensional or even higher-dimensional matrix operations constructed from these small elements are performed. In at least one embodiment, an API such as the CUDA9 C++ API exposes special matrix load operations, matrix multiply and accumulate operations, and matrix store operations to efficiently use the tensor core from a CUDA-C++ program. In at least one embodiment, at the CUDA level, the warp-level interface assumes a matrix of size 16×16 across all 32 threads of a warp.
[0293] In at least one embodiment, each SM3000 includes, without limitation, M SFU3012s that execute special functions (such as attribute evaluation, reciprocal square root, etc.). In at least one embodiment, the SFU3012s include, without limitation, tree traversal units configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFU3012s include, without limitation, texture units configured to perform filtering operations on texture maps. In at least one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texels) from memory and a sample texture map and generate sampled texture values for use in a shader program executed by the SM3000. In at least one embodiment, the texture map is stored in the shared memory / L1 cache 3018. In at least one embodiment, the texture unit implements texture operations such as filtering operations using mip maps (e.g., texture maps with different levels of detail) according to at least one embodiment. In at least one embodiment, each SM3000 includes, without limitation, two texture units.
[0294] In at least one embodiment, each SM3000 includes, without limitation, N LSU3014s that implement load and store operations between the shared memory / L1 cache 3018 and the register file 3008. In at least one embodiment, each SM3000 includes, without limitation, an interconnect network 3016 that connects each functional unit to the register file 3008 and connects the LSU3014 to the register file 3008 and the memory locations of the shared memory / L1 cache 3018. In at least one embodiment, the interconnect network 3016 is a crossbar, and this crossbar may be configured to connect any functional unit to any register in the register file 3008 and connect the LSU3014 to the register file 3008 and the memory locations of the shared memory / L1 cache 3018.
[0295] In at least one embodiment, the shared memory / L1 cache 3018 is, in at least one embodiment, an array of on-chip memory that enables data storage and communication between the SM3000 and the primitive engine and between the threads of the SM3000. In at least one embodiment, the shared memory / L1 cache 3018, without limitation, has a storage capacity of 128 KB and is on the path from the SM3000 to the partition unit. In at least one embodiment, the shared memory / L1 cache 3018 is, in at least one embodiment, used to cache reads and writes. In at least one embodiment, one or more of the shared memory / L1 cache 3018, the L2 cache, and the memory is auxiliary storage.
[0296] In at least one embodiment, by combining a data cache and a shared memory function into a single memory block, the performance for both types of memory access is improved. In at least one embodiment, the capacity can be used as, or is available as, a cache by programs that do not use the shared memory. Thus, when the shared memory is configured to use half of the capacity, texture and load / store operations can use the remaining capacity. According to at least one embodiment, by integrating into the shared memory / L1 cache 3018, the shared memory / L1 cache 3018 can function as a high-throughput pipe for streaming data, while at the same time providing high-bandwidth and low-latency access to frequently reused data. In at least one embodiment, when configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function graphics processing unit is bypassed to create a much simpler programming model. In the configuration of general-purpose parallel computing, the work distribution unit directly assigns and distributes thread blocks to the DPC in at least one embodiment. In at least one embodiment, the threads within a block execute the same program using unique thread IDs in the calculation so that each thread surely generates a unique result, use the SM3000 to execute the program and perform the calculation, use the shared memory / L1 cache 3018 to communicate between threads, and use the LSU3014 to read and write to the global memory through the shared memory / L1 cache 3018 and the memory partition unit. In at least one embodiment, when configured for general-purpose parallel computing, the SM3000 writes commands that the scheduler unit 3004 can use to launch new work on the DCP.
[0297] In at least one embodiment, the PPU is included in or coupled to a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., a wireless portable device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head-mounted display, a portable electronic device, etc. In at least one embodiment, the PPU is embodied on a single semiconductor substrate. In at least one embodiment, the PPU is included in a system-on-chip ("SoC") together with one or more other devices such as an additional PPU, memory, a reduced instruction set computer ("RISC") CPU, a memory management unit ("MMU"), a digital-to-analog converter ("DAC").
[0298] In at least one embodiment, the PPU may be included in a graphics card that includes one or more memory devices. The graphics card may be configured to interface with a PCIe slot on the motherboard of a desktop computer. In at least one embodiment, the PPU may be an integrated graphics processing unit ("iGPU") included in the chipset of the motherboard.
[0299] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the SM3000. In at least one embodiment, the SM3000 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system, or by the SM3000. In at least one embodiment, the SM3000 may be used to execute one or more use cases of the neural networks described herein.
[0300] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently.
[0301] In at least one embodiment, a single semiconductor platform may refer to a single monolithic semiconductor-based integrated circuit or chip. In at least one embodiment, a multi-chip module with improved connectivity may be used that simulates on-chip operations and provides significant improvements over using conventional central processing units (“CPUs”) and bus implementations. In at least one embodiment, various modules may also be arranged separately or in various combinations of the semiconductor platform, depending on the user's desires.
[0302] In at least one embodiment, a computer program in the form of machine-readable executable code or a computer control logic algorithm is stored in main memory 1004 and / or secondary storage. When the computer program is executed by one or more processors, it enables system 1000 to perform various functions according to at least one embodiment. In at least one embodiment, memory 1004, storage, and / or any other storage are possible examples of computer-readable media. In at least one embodiment, secondary storage may refer to any suitable storage device or system, such as a hard disk drive and / or a removable storage drive, representing a floppy (registered trademark) disk drive, a magnetic tape drive, a compact disk drive, a digital versatile disk (DVD) drive, a recording device, a universal serial bus (USB) flash memory, etc. In at least one embodiment, the architectures and / or functions of various previous figures are implemented in the context of CPU 1002, a parallel processing system 1012, an integrated circuit capable of realizing at least a part of the functions of both CPUs 1002, a parallel processing system 1012, a chipset (e.g., a group of integrated circuits designed and sold to function as units for performing related functions), and any suitable combination of integrated circuits.
[0303] In at least one embodiment, the architectures and / or functions of various previous figures are implemented in the context of general purpose computer systems, circuit board systems, game console systems dedicated to entertainment purposes, and application specific systems, among others. In at least one embodiment, computer system 1000 may take the form of a desktop computer, laptop computer, tablet computer, server, supercomputer, smartphone (e.g., wireless portable device), personal digital assistant (“PDA”), digital camera, vehicle, head-mounted display, portable electronic device, mobile phone device, television, workstation, game console, embedded system, and / or any other type of logic.
[0304] In at least one embodiment, parallel processing system 1012 includes, without limitation, a plurality of parallel processing units (“PPU”) 1014 and associated memory 1016. In at least one embodiment, PPU 1014 is connected to a host processor or other peripheral device via interconnect 1018 and switch 1020 or multiplexer. In at least one embodiment, parallel processing system 1012 distributes computational tasks across PPU 1014, which may be parallelizable, for example, as part of distributing computational tasks across multiple graphics processing unit (“GPU”) thread blocks. In at least one embodiment, the memory is shared and accessible across some or all of PPU 1014 (e.g., for read and / or write access), but such shared memory may result in performance disadvantages over the use of local memory and registers resident in PPU 1014. In at least one embodiment, the operation of PPU 1014 is synchronized by the use of commands such as _syncthreads(), and all threads in a block (e.g., executed across multiple PPU 1014) reach a certain execution point of the code before proceeding.
[0305] Virtual computing platform Examples are disclosed for a virtual computing platform for advanced computing, such as image inference and image processing. Referring to FIG. 31, it is an example data flow diagram of a process 3100 for generating and introducing a pipeline for image processing and inference according to at least one example. In at least one example, the process 3100 may be introduced for use with one or more facilities 3102, such as a medical facility, a hospital, a healthcare institution, a clinic, a research or diagnostic laboratory, etc., along with an imaging device, a processing device, a genomics device, a gene sequencing device, a radiation device, and / or other types of devices. In at least one example, the process 3100 may be introduced to perform genomics analysis and inference on sequencing data. Examples of genomic analysis that can be performed using the systems and processes described herein include, without limitation, variant calling, mutation detection, and quantification of gene expression. The process 3100 may be executed within the training system 3104 and / or within the deployment system 3106. In at least one example, the training system 3104 is used to train, deploy, and implement a machine learning model (e.g., a neural network, an object detection algorithm, a computer vision algorithm, etc.) for use in the deployment system 3106. In at least one example, the deployment system 3106 is configured to offload processing and computing resources across a distributed computing environment to reduce infrastructure requirements at the facility 3102. In at least one example, the deployment system 3106 may provide a streamlined platform for selecting, customizing, and implementing virtual appliances for use with an imaging device (e.g., MRI, CT scan, X-ray, ultrasound, etc.) or a sequencing device at the facility 3102. In at least one example, the virtual appliance may include a software-defined application for performing one or more processing operations on imaging data generated by an imaging device, a sequencing device, a radiation device, and / or other types of devices.In at least one embodiment, one or more applications in the pipeline may use or call services of the onboarding system 3106 (such as inference, virtualization, computing, AI, etc.) during the execution of the application.
[0306] In at least one embodiment, some of the applications used in the advanced processing and inference pipeline may use a machine learning model or other AI to perform one or more processing steps. In at least one embodiment, the machine learning model may be trained at the facility 3102 using data 3108 (such as imaging data) generated at the facility 3102 and stored in one or more image archive and communication system (PACS) servers at the facility 3102, may be trained using imaging or sequencing data 3108 from one or more other facilities (such as different hospitals, research institutes, clinics, etc.), or may be a combination thereof. In at least one embodiment, the training system 3104 may be used to provide applications, services, and / or other resources for generating a practical and deployable machine learning model for the onboarding system 3106.
[0307] In at least one embodiment, the model registry 3124 may be backed up by an object storage that can support version management and object metadata. In at least one embodiment, the object storage may be accessible, for example, from within a cloud platform, via a compatibility application programming interface (API) of cloud storage (e.g., cloud 3226 of FIG. 32). In at least one embodiment, a machine learning model within the model registry 3124 may be uploaded, listed, modified, or deleted by a system developer or partner interacting with the API. In at least one embodiment, the API may provide access to a way for a user with appropriate credentials to associate a model with an application, thereby enabling the model to be executed as part of running a containerized instance of the application.
[0308] In at least one embodiment, the training pipeline 3204 (FIG. 32) may include a situation where the facility 3102 is training its own machine learning model or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 3108 generated by an imaging device, a sequencing device, and / or other types of devices may be received. In at least one embodiment, when the imaging data 3108 is received, AI-assisted annotation 3110 may be used to assist in generating an annotation corresponding to the imaging data 3108 that will be used as ground truth data for the machine learning model. In at least one embodiment, the AI-assisted annotation 3110 may include one or more machine learning models (e.g., a convolutional neural network (CNN)), which may be trained to generate an annotation corresponding to a specific type of imaging data 3108 (e.g., from a specific device) and / or a specific type of abnormality within the imaging data 3108. In at least one embodiment, the AI-assisted annotation 3110 may then be used directly to generate ground truth data or may be adjusted or fine-tuned using an annotation tool (e.g., by researchers, clinicians, physicians, scientists, etc.). In at least one embodiment, in some examples, labeled clinical data 3112 (e.g., annotations provided by clinicians, physicians, scientists, technicians, etc.) may be used as ground truth data for training the machine learning model. In at least one embodiment, the AI-assisted annotation 3110, the labeled clinical data 3112, or a combination thereof may be used as ground truth data for training the machine learning model. In at least one embodiment, the trained machine learning model may be referred to as the output model 3116 and may be used by the introduction system 3106 described herein.
[0309] In at least one example, the training pipeline 3204 (FIG. 32) may include a situation where the facility 3102 requires a machine learning model to execute one or more processing tasks for one or more applications within the onboarding system 3106, but the facility 3102 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for such purposes). In at least one example, an existing machine learning model may be selected from the model registry 3124. In at least one example, the model registry 3124 may include machine learning models trained to perform various different inference tasks on imaging data. In at least one example, the machine learning models of the model registry 3124 may be trained on imaging data from a facility different from the facility 3102 (e.g., a facility in a remote location). In at least one example, the machine learning model may be trained on imaging data from one location, two locations, or any number of locations. In at least one example, when trained on imaging data from a particular location, the training may be performed at that location or at least in a manner that protects the confidentiality of the imaging data or restricts the transfer of the imaging data outside the facility (e.g., in accordance with HIPPA regulations, privacy regulations). In at least one example, when a model is trained or partially trained at one location, the machine learning model may be added to the model registry 3124. In at least one example, the machine learning model may then be retrained or updated at any number of other facilities, and the retrained or updated model may be made available in the model registry 3124. In at least one example, the machine learning model may then be selected from the model registry 3124, may be referred to as the output model 3116, and may be used in the onboarding system 3106 to execute one or more processing tasks for one or more applications of the onboarding system.
[0310] In at least one example, the training pipeline 3204 (FIG. 32) may include a facility 3102 where the scenario requires a machine learning model to perform one or more processing tasks for one or more applications within the introduction system 3106, but the facility 3102 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for such purposes). In at least one example, the machine learning model selected from the model registry 3124 may not be fine-tuned or optimized for the imaging data 3108 generated at the facility 3102 because there may be differences in the population, genetic variations, robustness of the training data used to train the machine learning model, diversity of anomalies in the training data, and / or other issues associated with the training data. In at least one example, AI-assisted annotation 3110 may be used to assist in generating annotations corresponding to the imaging data 3108 that will be used as ground truth data for retraining or updating the machine learning model. In at least one example, labeled clinical data 3112 (e.g., annotations provided by clinicians, physicians, scientists, technicians, etc.) may be used as ground truth data for training the machine learning model. In at least one example, retraining or updating the machine learning model may be referred to as model training 3114. In at least one example, model training 3114, such as AI-assisted annotation 3110, labeled clinical data 3112, or a combination thereof, may be used as ground truth data for retraining or updating the machine learning model. In at least one example, the trained machine learning model may be referred to as an output model 3116 and may be used by the deployment system 106 as described herein.
[0311] In at least one embodiment, the introduction system 3106 may include software 3118, services 3120, hardware 3122, and / or other components, features, and functions. In at least one embodiment, the introduction system 3106 may include a software “stack” such that the software 3118 may be built on top of the services 3120, may use the services 3120 to perform some or all of the processing tasks, and the services 3120 and software 3118 may be built on top of the hardware 3122 and may use the hardware 3122 to perform the processing, storage, and / or other computing tasks of the introduction system 3106. In at least one embodiment, the software 3118 may include any number of different containers, where each container may perform an instantiation of an application. In at least one embodiment, each application may perform one or more processing tasks of an advanced processing and inference pipeline (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, there may be any number of containers capable of performing data processing tasks on the imaging data 3108 (or other types of data such as those described herein) generated by a device, for each type of device such as an imaging device (e.g., CT, MRI, X-ray, ultrasound, sonography, echocardiogram, etc.), a sequencing device, a radiation device, a genomics device, etc.In at least one embodiment, the advanced processing and inference pipeline re-converts the output (e.g., to DICOM (Digital Imaging and Communications in Medicine) data, radiology information system (RIS) data, clinical information system (CIS) data, remote procedure call (RPC) data, data substantially compliant with a Representational State Transfer (REST) interface, data substantially compliant with a file-based interface, and / or raw data, etc., types of data that can be used) through the pipeline for storage and display at facility 3102, and in addition to the containers that receive and compose the imaging data used by each container and / or used by facility 3102, it may be defined based on the selection of different containers desired or required to process the imaging data 3108. In at least one embodiment, a combination of containers within software 3118 (e.g., that constitutes the pipeline) may be referred to as a virtual appliance (described in more detail herein), and the virtual appliance may utilize service 3120 and hardware 3122 to execute some or all of the processing tasks of the application instantiated in the container.
[0312] In at least one embodiment, the data processing pipeline may receive input data (e.g., imaging data 3108) in DICOM, RIS, CIS, REST-compliant, RPC, raw, and / or other formats in response to an inference request (e.g., a request from a user of the introduction system 3106 such as a clinician, physician, radiologist, etc.). In at least one embodiment, the input data may represent one or more images, videos, and / or other data representations generated by one or more imaging devices, sequencing devices, radiation devices, genomics devices, and / or other types of devices. In at least one embodiment, the data may undergo preprocessing as part of the data processing pipeline and be prepared so that it can be processed by one or more applications. In at least one embodiment, postprocessing may be performed on the output of one or more inference tasks or other processing tasks of the pipeline so that output data is prepared for the next application and / or so that output data is prepared for transmission and / or use by a user (e.g., in response to an inference request). In at least one embodiment, the inference task may be performed by one or more machine learning models such as a trained or introduced neural network, and this model may include the output model 3116 of the training system 3104.
[0313] In at least one embodiment, tasks of a data processing pipeline may be encapsulated in containers, each of which represents a separate fully functional instantiation of an application and a virtualized computing environment in which a machine learning model can be referenced. In at least one embodiment, a container or application may be issued to a private (e.g., restricted access) zone of a container registry (described in more detail herein), and a trained or introduced model may be stored in a model registry 3124 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., an image of a container) may be available in a container registry, and when selected by a user from the container registry for introduction into a pipeline, the image may be used to generate a container for instantiating the application so that it can be used on the user's system.
[0314] In at least one example, a developer (such as a software developer, clinician, physician, etc.) may develop, publish, and store an application (such as in a container) to perform image processing and / or inference on supplied data. In at least one example, the development, publishing, and / or storage may be performed using a software development kit (SDK) associated with the system (such as to ensure the developed application and / or container conforms to or is compatible with the system). In at least one example, the developed application may be tested locally (such as at a first facility, for data from the first facility) using an SDK that can support at least a portion of service 3120 as a system (such as system 3200 of FIG. 32). In at least one example, because a DICOM object can contain anywhere from one to hundreds of images or other types of data and there are variations in the data, the developer may be responsible for managing the extraction and preparation of the input DICOM data (such as configuring the setup for the application, building preprocessing into the application, etc.). In at least one example, once the application is verified by system 3200 (such as for accuracy, safety, patient privacy, etc.), it is made available in a container registry for selection and / or implementation by a user (such as a hospital, clinic, research institute, healthcare provider, etc.), and one or more processing tasks may be performed on data at the user's facility (such as a second facility).
[0315] In at least one embodiment, the developer may then share the application or container over a network so that it can be accessed and used by a user of the system (e.g., system 3200 of FIG. 32). In at least one embodiment, the completed and verified application or container may be stored in a container registry, and the associated machine learning model may be stored in model registry 3124. In at least one embodiment, a requesting entity (e.g., a user of a medical facility) making an inference or image processing request may browse the container registry and / or model registry 3124 to search for applications, containers, datasets, machine learning models, etc., select a desired combination of elements for inclusion in a data processing pipeline, and send an imaging processing request. In at least one embodiment, the request may include the input data (and in some instances, associated patient data) required to execute the request, and / or may include the selection of the application and / or machine learning model to be executed when processing the request. In at least one embodiment, the request is then passed to one or more components of the ingress system 3106 (e.g., the cloud) to execute the processing of the data processing pipeline. In at least one embodiment, the processing by the ingress system 3106 may include referring to elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 3124. In at least one embodiment, when a result is generated by the pipeline, the result may be returned to and viewed by the user (e.g., locally, viewed in a viewing application suite running on an in-house workstation or terminal). In at least one embodiment, a radiologist may receive results from a data processing pipeline including any number of applications and / or containers, where the results may include anomaly detection in X-rays, CT scans, MRIs, etc.
[0316] In at least one embodiment, service 3120 may be utilized to assist in processing or executing applications or containers in a pipeline. In at least one embodiment, service 3120 may include computing services, artificial intelligence (AI) services, visualization services, and / or other types of services. In at least one embodiment, service 3120 may provide common functionality to one or more applications of software 3118, whereby the functionality may be abstracted with respect to services that can be called or utilized by the applications. In at least one embodiment, the functionality provided by service 3120 may be executed dynamically and more efficiently, and at the same time, may scale well by enabling applications to process data in parallel (e.g., using parallel computing platform 3230 (FIG. 32)). Instead of requiring each application sharing the same functionality provided by service 3120 to have its own instance of service 3120, service 3120 may be shared among various applications. In at least one embodiment, the service may include an inference server or engine that may be used, as a non-limiting example, to perform detection or segmentation tasks. In at least one embodiment, a model training service capable of providing functionality for training and / or retraining a machine learning model may be included. In at least one embodiment, a data augmentation service capable of providing extraction, resizing, scaling, and / or other augmentation of GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compliant, RPC, raw, etc.) may be further included. In at least one embodiment, a visualization service capable of adding image rendering effects such as ray tracing, rasterization, noise removal, sharpening, etc. may be used to add a sense of realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual device service capable of realizing beamforming, segmentation, inference, imaging, and / or support for other applications within a pipeline of virtual devices may be included.
[0317] In at least one embodiment, when service 3120 includes an AI service (e.g., an inference service), one or more machine learning models associated with an application for anomaly detection (e.g., tumors, abnormal growths, scarring, etc.) may be executed by calling the inference service (e.g., an inference server) (as an API call) to execute the machine learning model, or its processing, as part of the application execution. In at least one embodiment, when another application includes one or more machine learning models for a segmentation task, the application may call the inference service to execute a machine learning model for executing one or more of the processing operations associated with the segmentation task. In at least one embodiment, software 3118 implementing an advanced processing and inference pipeline including a segmentation applic...
Claims
1. A processor comprising one or more circuits for training one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently, the two or more versions of the image being synthetically generated by a renderer, the two or more versions corresponding to an initial resolution and at least one output resolution, a processor.
2. The processor according to claim 1, wherein the one or more neural networks are trained to perform real-time upsampling of an input image at the initial resolution to one or more images at the at least one output resolution using only synthetically generated training data.
3. The processor according to claim 2, wherein the one or more circuits are further for injecting one or more rendering artifacts into the synthetically generated training data during training of the one or more neural networks.
4. The processor according to claim 1, wherein the renderer is modified to be deterministic and the two or more versions include a pixel-matched version of the image.
5. The processor according to claim 1, wherein the one or more circuits are further for generating a reference image using some samples per pixel, reconstructed using a filter that uses a determined dither offset.
6. A system comprising one or more processors for training one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being synthetically generated independently, A system in which two or more versions of the image are synthetically generated by a renderer, and the two or more versions correspond to an initial resolution and at least one output resolution. **Claim 7** The system according to claim 6, wherein the one or more neural networks are trained to perform real-time upsampling of an input image at the initial resolution to one or more images at the at least one output resolution using only synthetically generated training data. **Claim 8** The system according to claim 7, wherein the one or more processors are further for injecting one or more rendering artifacts into the synthetically generated training data during training of the one or more neural networks. **Claim 9** The system according to claim 6, wherein the renderer is modified to be deterministic and the two or more versions include a pixel-matched version of the image. **Claim 10** The system according to claim 6, wherein the one or more circuits are further for generating a reference image using several per-pixel samples reconstructed using a filter that uses a determined dither offset. **Claim 11** A method comprising training one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being independently synthetically generated, the two or more versions of the image being synthetically generated by a renderer, and the two or more versions corresponding to an initial resolution and at least one output resolution. **Claim 12** Further comprising training the one or more neural networks to perform real-time upsampling of an input image at the initial resolution to one or more images at the at least one output resolution using only synthetically generated training data, the method according to claim 11.
13. Injecting one or more rendering artifacts into the synthetically generated training data during training of the one or more neural networks Further comprising the method according to claim 12.
14. The method according to claim 11, wherein the renderer is modified to be deterministic and the two or more versions include a pixel-matched version of the image.
15. The method according to claim 11, wherein the one or more circuits are further for generating a reference image using some samples per pixel, reconstructed using a filter that uses a determined dither offset.
16. A machine-readable medium storing a set of instructions that, when executed by one or more processors, cause the one or more processors to at least Train one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being independently synthetically generated, The two or more versions of the image are synthetically generated by a renderer, and the two or more versions correspond to an initial resolution and at least one output resolution, a machine-readable medium.
17. The machine-readable medium of claim 16, wherein the one or more neural networks are trained to perform real-time upsampling of an input image at an initial resolution to one or more images at the at least one output resolution using only synthetically generated training data.
18. The machine-readable medium of claim 17, wherein the one or more circuits are further for injecting one or more rendering artifacts into the synthetically generated training data during training of the one or more neural networks.
19. The machine-readable medium of claim 16, wherein the renderer is modified to be deterministic and the two or more versions of the image include pixel-matched versions.
20. The machine-readable medium of claim 16, wherein the one or more circuits are further for generating a reference image using several samples per pixel, reconstructed using a filter that uses a determined dither offset.
21. One or more processors for training one or more neural networks based at least in part on two or more versions of an image, each of the two or more versions of the image being independently synthetically generated, one or more processors; A memory for storing network parameters for the one or more neural networks comprising A network training system, wherein the two or more versions of the image are synthetically generated by a renderer, and the two or more versions correspond to an initial resolution and at least one output resolution.
22. The network training system according to claim 21, wherein the one or more neural networks are trained to perform real-time upsampling of an input image at the initial resolution to one or more images at the at least one output resolution using only synthetically generated training data. **Claim 23** The network training system according to claim 22, wherein the one or more circuits are further for injecting one or more rendering artifacts into the synthetically generated training data during training of the one or more neural networks. **Claim 24** The network training system according to claim 21, wherein the renderer is modified to be deterministic and the two or more versions of the image include pixel-matched versions. **Claim 25** The network training system according to claim 21, wherein the one or more circuits are further for generating a reference image using several samples per pixel, reconstructed using a filter that uses a determined jitter offset. **Claim 26** The processor according to claim 1, wherein the one or more circuits generate an image based at least in part on an incomplete version of the image and an estimated complete version of the image using the one or more neural networks.
Citation Information
Patent Citations
Wavelength split type optical exchange channel
JP1988030092A
Systems and methods for generating and transmitting image sequences based on sampled color information
WO2020068140A1
Graphics processing chip with machine-learning based shader
WO2020172043A2