Image blending using one or more neural networks

A deep learning-based system uses neural networks to predict blending coefficients for upscaling images, addressing resource-intensity and artifact issues in high-resolution content production, achieving high-quality, artifact-free images at higher frame rates.

JP7763714B2Active Publication Date: 2025-11-04NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022077353
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-12-20
Filing Date
2022-05-10
Publication Date
2025-11-04
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

Image and video content production at higher resolutions is resource-intensive, and determining optimal blending weights for temporal smoothing between frames can result in noisy or artifact-prone images.

Method used

A deep learning-based system uses neural networks to predict blending coefficients for upscaling images, incorporating motion vectors and historical frame data to reduce artifacts and improve temporal stability.

Benefits of technology

The system achieves high-quality, artifact-free images at multiple times the original resolution with reduced computational resources, enabling real-time rendering at higher frame rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007763714000005
    Figure 0007763714000005
  • Figure 0007763714000006
    Figure 0007763714000006
  • Figure 0007763714000007
    Figure 0007763714000007
Patent Text Reader

Abstract

To provide a device, a system, and a technique for reconstructing one or more images.SOLUTION: In at least one embodiment, one or more circuits are to use one or more neural networks to adjust one or more pixel blending weights.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] At least one embodiment relates to processing resources used to perform and facilitate artificial intelligence. For example, at least one embodiment relates to a processor or computing system used to train neural networks according to various novel techniques described herein. Summary of the Invention [Problem to be solved by the invention]

[0002] Image and video content is increasingly being produced at higher resolutions and displayed on higher-quality displays. Techniques for producing higher-quality content are often very resource-intensive, especially at modern frame rates, which can be problematic for devices with limited resource capacity. Blending data from current and previous frames in a sequence can help improve the quality of this content by providing some temporal smoothing and accumulation of pixel data between frames, but determining optimal blending weights can be difficult, and inappropriate blending weights can produce images that are too noisy or have artifacts such as ghosting or temporal instability. [Means for solving the problem]

[0003] Various embodiments according to the present disclosure will now be described with reference to the drawings. [Brief explanation of the drawings]

[0004] [Figure 1] FIG. 1 illustrates an exemplary temporal upsampling pipeline, according to at least one embodiment. [Figure 2] FIG. 1 illustrates a system for blending color values ​​of images in a sequence, according to at least one embodiment. [Figure 3A] FIG. 1 illustrates components of an architecture including a refinement network, according to at least one embodiment. [Figure 3B] FIG. 1 illustrates components of an architecture including a refinement network, according to at least one embodiment. [Figure 3C] FIG. 1 illustrates components of an architecture including a refinement network, according to at least one embodiment. [Figure 4A] FIG. 1 illustrates a process for generating images in a sequence, according to at least one embodiment. [Figure 4B] FIG. 1 illustrates a process for generating images in a sequence, according to at least one embodiment. [Figure 5] FIG. 1 illustrates components of a system for providing image content, according to at least one embodiment. [Figure 6A] FIG. 1 illustrates inference and / or training logic, according to at least one embodiment. [Figure 6B] FIG. 1 illustrates inference and / or training logic, according to at least one embodiment. [Figure 7] FIG. 1 illustrates an exemplary data center system, according to at least one embodiment. [Figure 8] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 9] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 10] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 11] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 12A] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 12B] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 12C]FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 12D] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 12E] FIG. 1 illustrates a shared programming model, according to at least one embodiment. [Figure 12F] FIG. 1 illustrates a shared programming model, according to at least one embodiment. [Figure 13] FIG. 1 illustrates an exemplary integrated circuit and associated graphics processor, according to at least one embodiment. [Figure 14A] FIG. 1 illustrates an exemplary integrated circuit and associated graphics processor, according to at least one embodiment. [Figure 14B] FIG. 1 illustrates an exemplary integrated circuit and associated graphics processor, according to at least one embodiment. [Figure 15A] FIG. 10 illustrates additional exemplary graphics processor logic, according to at least one embodiment. [Figure 15B] FIG. 10 illustrates additional exemplary graphics processor logic, according to at least one embodiment. [Figure 16] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 17A] FIG. 1 illustrates a parallel processor, according to at least one embodiment. [Figure 17B] FIG. 1 illustrates a partition unit, according to at least one embodiment. [Figure 17C] FIG. 1 illustrates a processing cluster, according to at least one embodiment. [Figure 17D] FIG. 1 illustrates a graphics multiprocessor according to at least one embodiment. [Figure 18] FIG. 1 illustrates a multi-graphics processing unit (GPU) system, according to at least one embodiment. [Figure 19] FIG. 1 illustrates a graphics processor according to at least one embodiment. [Figure 20] FIG. 1 illustrates a micro-architecture of a processor, according to at least one embodiment. [Figure 21] FIG. 1 illustrates a deep learning application processor, according to at least one embodiment. [Figure 22] FIG. 1 illustrates an exemplary neuromorphic processor, according to at least one embodiment. [Figure 23] FIG. 1 illustrates at least a portion of a graphics processor, according to at least one embodiment. [Figure 24] FIG. 1 illustrates at least a portion of a graphics processor, according to at least one embodiment. [Figure 25] FIG. 1 illustrates at least a portion of a graphics processor core, according to at least one embodiment. [Figure 26A] FIG. 1 illustrates at least a portion of a graphics processor core, according to at least one embodiment. [Figure 26B] FIG. 1 illustrates at least a portion of a graphics processor core, according to at least one embodiment. [Figure 27] FIG. 1 illustrates a parallel processing unit (“PPU”), according to at least one embodiment. [Figure 28] FIG. 1 illustrates a general purpose processing cluster (“GPC”), according to at least one embodiment. [Figure 29] FIG. 1 illustrates a memory partition unit of a parallel processing unit (“PPU”) according to at least one implementation. [Figure 30] FIG. 1 illustrates a streaming multiprocessor, according to at least one embodiment. [Figure 31] FIG. 1 is an exemplary data flow diagram of an advanced computing pipeline, according to at least one embodiment. [Figure 32] FIG. 1 is a system diagram of an exemplary system for training, calibrating, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment. [Figure 33A] FIG. 1 is a data flow diagram of a process for training a machine learning model, according to at least one embodiment. [Figure 33B] FIG. 1 is an exemplary diagram of a client-server architecture for enhancing an annotation tool with pre-trained annotation models, according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0005] In at least one embodiment, an upscaling system 100 such as that shown in FIG. 1 can be used to increase the resolution of one or more images, such as images or video frames in a sequence or video stream, as part of a deep learning-based super-sampling or super-resolution process. In at least one embodiment, this may include a representation of one or more objects in a scene, such as a scene of live gameplay. In at least one embodiment, a rendering engine 102 or renderer can output images of one or more objects at a first resolution that will be upscaled to one or more higher output resolutions. In at least one embodiment, real-time temporal image reconstruction can be performed at a higher resolution than the resolution at which the image is generated by the rendering engine. In at least one embodiment, the temporal aspect of this process can include blending color values ​​of corresponding points between a current frame 106 and at least one previous or historical frame 124 in the sequence. In at least one embodiment, to ensure that this blending is performed for corresponding points on objects in these frames, this previous history color data can be warped based on detected motion between the history frame and the current frame, as may be indicated by a set of motion vectors output from the rendering engine or otherwise determined. In at least one embodiment, such warping can ensure that points, such as various image feature points, are tracked over time and corresponding color values ​​are used in blending, which can help reduce the presence of artifacts such as noise or flicker during playback. In at least one embodiment, and as discussed in more detail elsewhere herein, the super-sampling algorithm can utilize a neural network that predicts blending coefficients to determine how much to weight the color values ​​of a current pixel in the current frame and corresponding history pixels from previous warped history frames.In at least one embodiment, such an algorithm can also utilize a filtering kernel to generate a new, higher-resolution output image from a set of inputs. In at least one embodiment, the output image quality of such a network can depend, at least in part, on the information available in this input, which may include information such as the current luma frame, historical luma, training history, and a color variation mask or motion vector difference buffer. In at least one embodiment, an application can render an aliased 1 sample-per-pixel (spp) image at 1080p (Full HD) resolution, and the algorithm can reconstruct an anti-aliased 2160p (4k) image from this input image and any such side information sequences provided by the application. In at least one embodiment, such a process can be extended to other resolutions with other upscaling ratios, including the case of pure anti-aliasing where the input and output resolutions are equal.

[0006] In at least one embodiment, content such as video game content or animations can be generated using a renderer 102, rendering engine, or other such content generation means of such a system. In at least one embodiment, the renderer 102 can receive an input of one or more frames of a sequence and can generate images or video frames using stored content 104 (e.g., maps and graphic assets) modified at least in part based on the input. In at least one embodiment, this renderer 102 can be part of a rendering pipeline, such as one that can utilize rendering software such as Unreal Engine 4 from Epic Games, Inc., which can provide features such as deferred shading, global illumination, light transparency processing, post-processing, and graphics processing unit ("GPU") particle simulation using vector fields. In at least one embodiment, the amount of processing required for this complex rendering of full high-resolution images can make it difficult to render these video frames to meet current frame rates, such as at least 60 frames per second (fps). In at least one embodiment, renderer 102 may instead be used to generate rendered images 106 at a resolution lower than one or more final output resolutions in order to meet timing requirements, reduce processing resource requirements, etc. In at least one embodiment, this lower resolution rendered image 106 may be processed using upscaler 108 to generate an upscaled image 110 that represents the content of the lower resolution rendered image 106 at a resolution equal to (or at least close to) the target output resolution.

[0007] In at least one embodiment, an upscaler system 108 (which may take the form of a service, system, module, or device) may be used to upscale individual frames of a video or animation sequence. In at least one embodiment, the amount of upscaling to be performed may depend on the initial resolution of the rendered image and the target resolution of the display, such as ranging from 1080p to 4k resolution. In at least one embodiment, additional processing, which may include anti-aliasing and temporal smoothing, may be performed as part of the upsampling process. In at least one embodiment, appropriate reconstruction filters may be utilized, such as filters such as anisotropic Gaussian filters or dynamic filter networks (DFNs). In at least one embodiment, the upsampling process may take into account sub-pixel jitter, which may be applied on a frame-by-frame basis.

[0008] In at least one embodiment, deep learning can be used to infer these upsampled video frames of the sequence. In at least one embodiment, temporal reconstruction can be used to provide a combination of anti-aliasing and super-resolution. In at least one embodiment, information from the corresponding sequence of video frames can be used to infer a higher quality upsampled image. In at least one embodiment, one or more heuristics based on prior knowledge of the rendering pipeline can be used without requiring learning from data. In at least one embodiment, this can include jitter-aware upsampling and accumulating samples at the upsampled resolution. In at least one embodiment, this jitter offset data can be provided along with the current input video frame and the previous inferred frame as input to an upscaler 108 that includes at least one neural network for inferring a higher quality upsampled image 110 that would be generated by the upsampling algorithm alone. In at least one embodiment, this upsampling necessarily shifts the jitter offset 122 and the samples per frame so that they are aligned with a history buffer that may be at a higher resolution.

[0009] In at least one embodiment, the upscaled image 110 can be provided as input to a neural network 112 to determine one or more blending coefficients or weights. In at least one embodiment, the neural network 112 also receives as input, along with the upscaled image 110, a previous high-resolution image in the sequence, which has been warped and provided to the neural network 112. In at least one embodiment, the neural network 112 can also receive other input features, such as those related to spatial and temporal variations as discussed herein. In at least one embodiment, deep learning is used to reconstruct images for real-time rendering at a resolution multiple (e.g., 2-9) times higher than the actual rendering resolution. In at least one embodiment, the reconstructed image quality from such a process is comparable to or even exceeds original resolution rendering, at least in terms of detail, temporal stability, and lack of common artifacts such as ghosting or lag. In at least one embodiment, the neural network can also determine at least some filtering to be applied when reconstructing or blending the current image with the previous image. In at least one embodiment, this information may then be provided along with the upscaled image 110 to a blending component 114 to be blended with at least one previous image in the sequence. In at least one embodiment, jitter offset data 122 may also be provided as an input to the blending component 114. In at least one embodiment, this blending of the current image with previous (or historical) images in the sequence may aid in temporal convergence to a nice, sharp, high-resolution output image 116, which may then be provided for presentation via a display 120 or other such presentation mechanism.In at least one embodiment, a copy of this high-resolution output image 116 may also be stored in a history buffer 118 or other such storage location for blending with later-generated images in the sequence. In at least one embodiment, such a process may leverage deep learning to reconstruct images for real-time rendering at a resolution multiple times (e.g., 2x, 4x, or 8x) higher than the actual rendering resolution, with reconstructed image quality at least comparable to the original resolution rendering in terms of detail, temporal stability, and lack of common artifacts such as ghosting or lag. In at least one embodiment, tensor cores may be used to accelerate the reconstruction rate, making this rendering process much more sample-efficient and enabling significant increases in frames per second for a variety of applications by using techniques such as those presented herein.

[0010] In at least one embodiment, using buffered information in such a system may include a component 200 such as that shown in FIG. 2. In at least one embodiment, three primary input sources are utilized, including a color buffer 202, a motion vector buffer 204, and a depth buffer 206. In at least one embodiment, a preprocessor 208, such as may include one or more processes running on one or more processors on one or more computing devices, may receive as input color information for the current frame as generated by a rendering engine or application, and the output of a warper 210, such as a warp function or application running on one or more processors of one or more devices. In at least one embodiment, this warper 210 receives as input motion vector information for the current frame as stored in the motion vector buffer 204. In at least one embodiment, the warper 210 may receive this data directly from an application or renderer and may not utilize a dedicated buffer. In at least one embodiment, this temporal process also provides as input to warper 210 high-resolution color data from the previous image in the sequence, as stored in history buffer 214. In at least one embodiment, as previously described, information for each final output image may also be stored in history buffer 214 for use in generating subsequent images or frames in the sequence. In at least one embodiment, warper 210 may utilize this motion vectors and depth data to warp pixel data or color data of particular features in the previous image to corresponding pixel locations in the current image frame, effectively using these motion vectors to map pixel data or color data of particular features in the two images to corresponding pixel locations of features, thus allowing comparison and blending of color values ​​of similar features. In at least one embodiment, preprocessor 208 may perform any associated processing on the current color data from color buffer 202 or the warped previous color data from warper 210.In at least one embodiment, this data, after any pre-processing, is provided as input to a neural network or other deep learning (DL) based generator 212 that can analyze this data to determine pixel-specific weights for each pixel location in the image to be generated. In at least one embodiment, this generated data is sent to a post-processor 216, which may include one or more processes running on one or more processors of one or more computing devices that can output a final high-resolution color image 218. In at least one embodiment, this post-processor can also output information that is stored in a high-resolution color and history buffer 214 for use in generating subsequent images in the current sequence.

[0011] In at least one embodiment, generating a frame using such a technique may involve an application providing a low-resolution jittered input image and associated jitter values, low-resolution backward motion vectors for each individual input image pixel, and other quantities such as exposure values ​​and a depth buffer to a reconstruction algorithm. In at least one embodiment, these low-resolution input (backward) motion vectors may be used to warp a previous frame output image to align with the geometry in the current time stamp. In at least one embodiment, this low-resolution current frame image is upsampled to the resolution of the output image 218 using an upsampling algorithm. In at least one embodiment, a neural network 212 is used to infer a weight value w for each output pixel (at the output resolution). In at least one embodiment, a high-resolution output image of the current frame may be created as follows: Output = w * (upsampled current frame input image) + (1-w)*(previous warped output image)

[0012] In at least one embodiment, and in this type of temporal image reconstruction algorithm, a significant factor in the resulting image quality (IQ) can be attributed to the weighting factor w described above. In at least one embodiment, w meets various criteria, including favoring the current input image when a region in the output image becomes unoccluded due to the motion of objects in the scene being rendered, or weighting the color values ​​of the current image more heavily, such as when w=1.0. In at least one embodiment, when a region in the output image is visible (and similarly shaded) in previous frames, an optimal weighting factor can result in appropriate blending between these previous output images and the current input image. In at least one embodiment, this blending can favor historical data more, such as when the value of w approaches zero because more frames have made this region visible.

[0013] In at least one embodiment, the network can base this prediction weight at least in part on the current frame input image and the warped previous frame output image. In at least one embodiment, whenever the upsampled current image has significantly different values ​​from the warped previous frame output image and therefore will look very different when displayed, the neural network can predict a high value of weighting factor w that gives more importance to the upsampled current frame input image. In at least one embodiment, when the current image has similar values ​​to the warped previous frame output image and therefore will look very similar when displayed, the neural network can predict a low value of weighting factor w that gives more importance to the warped previous frame output image.

[0014] In at least one embodiment, motion vector difference information can be used as an additional modality or input, as discussed herein. In at least one embodiment, an additional buffer, such as a motion buffer 220, can be used as another source of input in such a system 200. In at least one embodiment, other buffers, such as at least one motion data buffer or depth data buffer, as discussed herein, can also be utilized. In at least one embodiment, this motion buffer 220, also referred to herein as a history motion buffer, can store new or additional motion vector data that can persist between frames. In at least one embodiment, current motion vectors from the motion vector buffer 204 can be stored in one or more forms, such as those that may correspond to a transformation process, to be used for subsequent frames. In at least one embodiment, this motion vector information can be provided as an additional input to the warper 210. Here, in at least one embodiment, the warping function of the warper 210 can warp not only the high-resolution color history data from the buffer 214, but also this previous motion vector data from the motion buffer 220. In at least one embodiment, the temporal calculation means 222 can perform calculations as described with respect to the associated equations above, in which warped motion vector data from the warper 210 is processed with current motion vector data from the motion vector buffer 204 to determine a difference or difference region. In at least one embodiment, this calculation can include determining the difference, then a norm, and applying the associated function as previously described. In at least one embodiment, this temporal calculation can then be provided as an external input to the pre-processor 208 and then passed to the generator 212 for use in determining more accurate pixel-specific weightings as discussed herein, enabling this DL-based network to produce higher quality results.

[0015] In at least one embodiment, neural networks used in temporal upsampling systems such as those shown in FIG. 2 can benefit from the use of lower resolutions by one or more neural networks, such as generator 212, thereby providing significant reductions in processing costs, memory, bandwidth, and other computational or data storage / transmission resources. However, in at least one embodiment, such resolution reductions may also reduce the quality of the output at the final higher output resolution. In at least one embodiment, a network architecture can be utilized to provide parameters for spatiotemporal accumulation and upsampling by utilizing a two-stage prediction process. In at least one embodiment, the first stage can operate at a lower resolution, such as for U-Net architecture 306 as shown in FIG. 3. In at least one embodiment, inputs 302 from a rendering engine or other such source can be passed to a downsample and folding module 304 or process, which can reduce the resolution of these inputs 302 for processing by one or more neural networks, in this case, convolutional neural networks (CNNs) such as U-Net 306. In at least one embodiment, the output of U-NET 306 can be passed to an unfolding and upsampling module 308, which can upsample the output to a higher resolution, such as that which may correspond to the resolution of input 302 or the target resolution. In at least one embodiment, those higher resolution results can then be passed to a refinement network 310, which also receives input 302, which may be at this same higher resolution. In at least one embodiment, refinement network 310 can then improve the predictions from U-Net 306 at the higher per-pixel resolution, such as that of input 302.In at least one embodiment, such an approach can also result in modifying the computational structure of at least some network layers of U-Net 306 from a dense representation to a sparse representation or a representation utilizing a reduced-precision numerical format. In at least one embodiment, the resolution of the data in FIG. 3A corresponds to the thickness of the corresponding arrow of the data flow, with thicker arrows representing data at a higher resolution and thinner arrows, such as those flowing into or out of U-Net 306, at a lower, downsampled resolution.

[0016] In at least one embodiment, a CNN, such as U-Net, is a deep but costly network capable of analyzing a large range of input features. In at least one embodiment, the network may be evaluated at a resolution lower than the full output resolution due, at least in part, to this high cost. In at least one embodiment, all inputs may be provided to the network at this lower resolution, and the output may also be at this lower resolution. In at least one embodiment, these lower-resolution outputs from the network may be upscaled to a higher resolution, such as the output resolution. To prevent the network's lower operating resolution from limiting the quality of the results, a second stage of the prediction process may take these upscaled full-resolution outputs and the network's inputs as inputs and generate more refined predictions. In at least one embodiment, the refinement network 310 may be a shallow convolutional neural network (CNN) operating at full resolution, or at least at a higher resolution than U-Net 306. In at least one embodiment, the upsampling system can then reconstruct the image at significantly higher quality, comparable to U-Net 306 operating at full resolution, but at much lower cost and in some cases with much lower latency.

[0017] In at least one embodiment, an exemplary architecture 336 of a refinement network is shown in configuration 330 of FIG. 3B . In at least one embodiment, an alternative refinement network architecture 362 can be utilized that includes similar components but also includes skip connections from the U-Net output 334 to the refinement output 366. In at least one embodiment, these skip connections can be utilized between convolutional layers, which can help improve image quality. In at least one embodiment, the number of convolutional layers and the corresponding kernel size and number of features can vary within different refinement networks, as this can depend at least in part on the available computational budget. In at least one embodiment, the refinement network can consist of four convolutional layers, with concatenated network inputs processed in two separate branches. In at least one embodiment, the first branch processes these inputs through a 3×3 convolution 340 combined with an activation function, and then through a 1×1 convolution 344 + activation function. In at least one embodiment, the second branch computes a trained 3x3 convolution 342 whose output is added to the output of the previous convolution with a skip connection by an addition module or process 346. In at least one embodiment, this sum is processed by another 1x1 convolution 348 to produce an output tensor that can be added to the set of U-Net outputs, such as through a final skip connection that is added using another addition module or process 364 as shown in FIG. 3C. In at least one embodiment, one or more 3x3 convolutions may be replaced with 1x1 convolutions to further reduce execution time costs. In at least one embodiment, the logic of any of the components shown in FIGS. 3A, 3B, and 3C, such as the execution of the convolutions, may be implemented in one or more processors or systems of similar or different types, as illustrated or suggested in this disclosure.

[0018] In at least one embodiment, the refinement network is a small, shallow refinement network that operates pixel-by-pixel at the full output resolution. In at least one embodiment, such a refinement network combines the upsampled results from the U-Net with the full-resolution image data. In at least one embodiment, such a refinement network can refine primarily based on the residuals from the U-Net. In at least one embodiment, the refinement network can make small adjustments to the upsampled output from the deep U-Net, thereby improving aspects of the upsampling algorithm used to upsample these lower-resolution blending weights. In at least one embodiment, such an approach can generate high-frequency or high-resolution output data by looking at the full-resolution input image in the spatial domain and utilizing deep features coming from the U-Net to provide context for what needs to be done with the image data. In at least one embodiment, the top branch of the stacked convolutions in the refinement network can perform most of this processing. In at least one embodiment, regular full-resolution input data can be concatenated with the upscaled features from this U-Net. In at least one embodiment, the lower branches are then mapped to the re-downsampled resolution, with downward arrows representing the corresponding set of residuals being determined. In at least one embodiment, the refinement network can then compare the low-resolution feature set and the high-resolution feature set and the residuals between them. In at least one embodiment, this residual data can then be used to determine a small offset to apply to the high-resolution image space based at least in part on these differences.In at least one embodiment, this high-resolution per-pixel data is input along with these lower-resolution features from the U-Net, and the output from the U-Net is deep-learned features that lack the high-frequency content of the higher-resolution raw image data. In at least one embodiment, a refinement network can generate refinements to the data used to determine pixel weights for one or more output images. In at least one embodiment, such a refinement network can generate other data refinements and potentially relevant refinements discussed in more detail elsewhere herein, including, among other such options, the number of channels (e.g., 5) of learned history (which can be averaged per channel after network processing and warped by the RGB history of the next frame), the number of channels (e.g., 3) for filter parameters, such as anisotropic Gaussian filter parameters, and at least one channel for adaptive jitter-aware blending of data from at least one history frame and the current frame.

[0019] In at least one embodiment, at least some layers of the refinement network 310 and / or the U-Net 306 may have their weights and / or activations represented in a sparse manner. In at least one embodiment, some zero-valued elements may not need to be stored as explicit values. In at least one embodiment, such a sparse representation may allow for a savings in the number of mathematical operations required to compute these layers. In at least one embodiment, the sparse activations and / or weights may also be stored such that they use less storage than corresponding dense representations and reduce memory bandwidth used by transferring these activations and / or weights. In at least one embodiment, certain sparse representation formats may also allow for the use of certain deep learning hardware that is sparse and utilizes a reduced number of mathematical operations, resulting in higher throughput and / or efficiency. In at least one embodiment, at least some layers of the refinement network 310 and / or the U-Net 306 may have their weights and / or activations represented in a reduced-precision numerical format. In at least one embodiment, using such a reduced-precision format allows weights and activations to be stored in a smaller space, reducing the bandwidth required for their transmission. In at least one embodiment, reduced-precision activations and weights can be used by certain deep learning hardware to increase processing and transmission throughput and / or efficiency. In at least one embodiment, at least some layers of the refinement network 310 and / or U-Net 306 can be represented using both a sparse representation and a reduced-numeric-precision format. In at least one embodiment, such a process can generate parameter estimates at quarter resolution, which can then be upscaled to the target resolution. In at least one embodiment, image quality comparable to running the primary network at full resolution can be achieved, but at approximately 20% of the runtime cost.In at least one embodiment, the use of a refinement network allows the network to produce more accurate predictions and significantly sharper images.

[0020] In at least one embodiment, a process 400 for generating an image may be utilized, as shown in FIG. 4A . In at least one embodiment, image data for a current image and at least one previous image in a sequence may be obtained 402, such as image data received from a rendering engine or application and previous image data stored in a buffer or repository. In at least one embodiment, this image data may include data such as per-pixel color data, motion vector data, and depth data, among other such options. In at least one embodiment, this image data may be downsampled to a lower resolution to reduce the computational, storage, or other resource costs of processing this data. In at least one embodiment, any logic for this or another such process may be executed on any processor or system, or combination thereof, illustrated or described in this disclosure. In at least one embodiment, this lower resolution image data may be processed using a deep neural network, such as a U-Net or other CNN, to infer a set of blending weights or other such image parameters 406. In at least one embodiment, these blending weights may be upsampled 408 to the original resolution of the image data before it was downsampled. In at least one embodiment, these upsampled blending weights may be provided as input to a refinement network along with the current and previous image data at the original (non-downsampled) resolution 410. In at least one embodiment, a set of blending weights (or other image parameters) that are adjusted based at least in part on favorable features in the image data at the original higher resolution may be received from the refinement network 412. In at least one embodiment, one or more output images may be generated 414 by blending, at least in part, pixel-specific color values ​​from the current image and at least one previous image according to these adjusted blending weights.

[0021] In at least one embodiment, a process 450 for generating an image may be performed, as shown in FIG. 4B. In at least one embodiment, one or more pixel blending weights may be generated 452. In at least one embodiment, these one or more pixel blending weights may be adjusted 454, such as by using a refinement network that takes higher resolution image data as input. In at least one embodiment, these adjusted pixel blending weights, such as those that may be generated by a spatiotemporal super-sampling system, may be used to generate one or more images 456.

[0022] In at least one embodiment, the amount of resources required to determine blending weights can be reduced by utilizing fewer than all available color channels. In at least one embodiment, this can include utilizing only the luma channel, which represents luminance in an image, as distinct from its chrominance. In at least one embodiment, a single channel, such as the luma channel, can be used instead of using full RGB (red-green-blue) or other color values. In at least one embodiment, luma information can be determined for both the current frame and a previous reconstructed or historical frame. In at least one embodiment, this can allow for a significant reduction in the information to be processed from six channels of information (current RGB and previous RGB) to two channels (current luma and previous luma). In at least one embodiment, only the luma information of the current and previous frames (and a variance mask, if utilized) is provided as input to the neural network. In at least one embodiment, the neural network does not infer color values, but rather can utilize a single channel of information to infer a filter to be applied to color or pixel values. In at least one embodiment, the network effectively determines the extent to which previous pixel data can be reused in the reconstruction of the current frame, and to this end, a single color data channel is sufficient for both the current and previous frames. In at least one embodiment, luma is utilized because the human eye is more sensitive to luma (i.e., changes in luminance) than to changes in color. In at least one embodiment, reducing the dimensionality of this technique also helps increase generalization. As noted, in at least one embodiment, a variance mask can also be generated using only luma values, where the luma values ​​of pixels from historical frames are compared against the corresponding luma mean and standard deviation values ​​of the current frame.

[0023] In at least one embodiment, one or more color filters may be applied to at least the current frame prior to blending. In at least one embodiment, an upsampling process such as that discussed with respect to FIG. 2 may be utilized to upsample an image being rendered at a first resolution to an image at a higher resolution, such as a target output resolution. In at least one embodiment, this upsampling may result in a somewhat blocky or jagged image, which may be partially addressed through an anti-aliasing process. In at least one embodiment, blending may be improved by first applying a color filter to remove these jagged edges in the upsampled image. In at least one embodiment, applying this filtering to the upsampled image may help increase the effective resolution of the image. In at least one embodiment, a filter such as a Gaussian filter may be applied. In at least one embodiment, this may be a parameterized anisotropic Gaussian filter. In at least one embodiment, an upsampling process with a large ratio (e.g., 9x) may result in an image that appears relatively jagged or low-resolution, since there will generally be 3-pixel by 3-pixel blocks that share the same color value. In at least one embodiment, this blockiness can be significantly reduced by applying a filter to remove these colors. In at least one embodiment, generating significantly more pixel data may result in the need for significantly more output channels in the output layer of the network, such as 25 channels in the output layer of the network for a 5x5 filter. In at least one embodiment, for very high resolutions, this additional data may cause the network to run very slowly. In at least one embodiment, instead of explicitly predicting all 25 numbers for the filter, a parameterized model can be predicted for the filter.In at least one embodiment, this can involve inferring only a small set of numbers (e.g., 3) for this parameterized model, which can be applied to these 25 relevant pixels. In at least one embodiment, these three numbers can later be instantiated into 25 numbers by interpreting these three numbers as parameters of an anisotropic Gaussian kernel or other parameterized dynamic filter kernel. In at least one embodiment, such an approach can be used with algorithms that do not utilize neural networks or machine learning.

[0024] In at least one embodiment, it may be desirable to further reduce the processing, memory, and other resources utilized in such a process. In at least one embodiment, images and inputs provided to the neural network can first be downsampled to allow the neural network to operate at a lower resolution. In at least one embodiment, because the neural network determines filter or blending weights to be applied to image locations rather than individual pixel color values, the network can be operated at a lower resolution to reduce resource requirements and latency while maintaining high image reconstruction quality. In at least one embodiment, this can include downsampling at least the current image and previous reconstructed images to half resolution or another reduced resolution. In at least one embodiment, the neural network can be trained at full resolution or a reduced resolution, but can run at the reduced resolution during inference. In at least one embodiment, this can serve to decouple the display resolution from the resolution at which the neural network operates. In at least one embodiment, a downsampling filter can be applied, such as may include 2x2 downsampling using a regular box filter, although other filters and ratios can be utilized. In at least one embodiment, the blending weights determined at this lower resolution can be applied to a higher resolution image or set of images to reconstruct an image at a target output resolution. In at least one embodiment, the output of this neural network can be upsampled before applying the blending and filtering. In at least one embodiment, the blending weight outputs are smooth, so that differences in resolution do not significantly affect the quality when applying those blending weights.

[0025] In at least one embodiment, a sub-pixel jitter offset can be determined and utilized as described above. In at least one embodiment, an upsampling system such as that described with respect to FIG. 2 can be used to upscale individual frames of the sequence. In at least one embodiment, this can include jitter-aware upsampling and accumulating samples at the upsampled resolution. In at least one embodiment, this previous process data can be provided along with the current input frame and the previous inferred frame as input to an upsampler system including at least one neural network for inferring an upsampled output image. In at least one embodiment, the upsampling process can be performed for each individual pixel of the lower resolution rendered image. In at least one embodiment, as a result of the upsampling process, color information from that pixel can be applied to a corresponding pixel region in the larger-sized upscaled image. In at least one embodiment, this pixel in the upscaled image can be segmented (or mapped) into multiple individual pixels. In at least one embodiment, the upsampling can be a 4x upsampling, where each pixel of the input image is segmented into four higher resolution pixels.

[0026] In at least one embodiment, color is determined for a pixel having a central pixel location. In at least one embodiment, the lower resolution image to be rendered then has a single color value reported for this pixel centered at this central pixel location. However, in at least one embodiment, jittering may be performed between frames or images in a sequence, where the center point of the color determination is slightly shifted to another point within this pixel. In at least one embodiment, this may correspond to a sample point offset of a sub-pixel offset from the pixel center. In at least one embodiment, a pixel analysis area (e.g., a 3x3 pixel analysis area) may still be used to determine color information for a given pixel, but the location of this 3x3 pixel analysis is shifted slightly based on the jitter location used to center the pixel analysis area. In at least one embodiment, these jitter locations may vary between frames in a sequence, either randomly or according to a determined pattern or sequence. In at least one embodiment, the color data determined for this pixel may include data such as saturation, RGB (red-green-blue) color values, or a luminance (e.g., luma) value that represents the brightness of the pixel rather than its final color value, with the luma typically being combined with the saturation color value to generate the final pixel value.

[0027] In at least one embodiment, a process for generating a sequence of images can be performed, where an image (or video frame) is rendered at a first resolution. In at least one embodiment, the image can be part of any suitable type of content, such as may be related to video, gaming, virtual reality (VR), augmented reality (AR), or other such application or content type. In at least one embodiment, the resolution can be derived from the rendering engine or provide desired performance. In at least one embodiment, the rendered image can be upsampled to a second, higher resolution using jitter-aware upsampling, where a sub-pixel offset is determined for the image, which can differ from a previous offset for one or more previous images in the sequence. In at least one embodiment, the upsampled image can be provided to a neural network to determine one or more blending weights to be used to blend the upscaled image with the previous image.

[0028] In at least one embodiment, image quality and stability can be further improved by including additional computed features as inputs to both the U-Net and the refinement network. In one embodiment, both the U-Net and the refinement network receive the current frame's luma, warping history luma, network-trained memory, and motion vector difference information. In yet another embodiment, a temporal color variance mask is computed as an additional feature to improve temporal stability.

[0029] In at least one embodiment, the network can utilize a larger set of input features to provide high image quality, as otherwise limited input information can result in problems such as temporal instability, ghosting artifacts, and poor anti-aliasing. In at least one embodiment, this can include information such as spatial gradients. In at least one embodiment, spatial gradients are calculated based at least in part on differences between neighboring spatial pixels, while temporal gradients are calculated based at least in part on differences in color values ​​between the current frame and the warping history or warping history frames. In at least one embodiment, these spatial gradients can include horizontal or vertical gradients, or other such gradients determined spatially with respect to a pixel of interest in the current image or frame. In at least one embodiment, these spatial gradients can represent differences between neighboring horizontal and / or vertical pixels in the input image. In at least one embodiment, these spatial gradients can represent this difference information using full color (e.g., RGB), a single color channel, grayscale, or other such color values. In at least one embodiment, this spatial gradient information can be used in addition to the temporal gradient information to enable improved blending and color determination of pixels of interest within the current frame.

[0030] In at least one embodiment, a temporal color gradient can be calculated for a pixel's current color value, which can be a color in any relevant color space, such as full color (e.g., RGB), a single color channel, or grayscale. In at least one embodiment, this can be compared against a warping history buffer. In at least one embodiment, these temporal color gradients can also represent the difference between the current frame and the warped history frame, and the motion vector for that pixel location in the current frame is applied to the location from the history frame to ensure that the color value of the corresponding point or feature is used in blending. In at least one embodiment, motion vector differences, such as those that may take the form of the absolute value of the difference between the motion vector components in the current frame relative to the warped history frame, can also be used. In at least one embodiment, using temporal color gradients can help the network better detect and manage rapid color changes or occlusions at various locations, thereby better preventing the occurrence of high-frequency temporal phenomena such as flicker. In at least one embodiment, these temporal color gradients can also be determined based on two or more previous history frames, such as those that can be determined for the previous two, three, or more images or frames in a sequence.

[0031] In at least one embodiment, this delta value between horizontally or vertically neighboring pixel values ​​within a row or column can be analyzed in a full color space, a single-channel color space, or grayscale, such as may be performed in luma space only as discussed elsewhere herein. In at least one embodiment, the use of such spatial gradients can allow for significantly improved identification of edges and various types of texture gradients, allowing the network to produce higher quality image output.

[0032] Additional features can be provided to the neural network as part of a rich combination of input features for purposes such as temporal super-sampling. In at least one embodiment, these additional features can include current frame luma values ​​and warped history frame luma values, as well as current color values ​​and warped color values. In at least one embodiment, learned history data can be provided, which may include any history data about the video or image sequence that can assist in providing higher quality image output. In at least one embodiment, the luma channel of color data, to which human vision is particularly sensitive, can be provided as an input to the network corresponding to a different color space and written to a history buffer that can be accessed to upsample subsequent images or video frames in the sequence. In at least one embodiment, stored RGB data has all three channels warped together, and then the luma history is provided as an input to the network. In at least one embodiment, warping can be performed using motion vectors provided by the rendering engine and then stored as the warped and blurred color history of the luma channel of the color information.

[0033] In at least one embodiment, depth difference information may also be provided. In at least one embodiment, this may take the form of a log of the absolute value of the difference between depth values ​​in the current frame relative to the warped history frame. In at least one embodiment, this depth data may be provided by a rendering engine or application, and warping may be applied to history pixel locations as performed on color data. In at least one embodiment, the depth information may help ensure that appropriate color values ​​are blended, such as color values ​​of the same object at a given pixel location. In at least one embodiment, this may be managed similarly to how motion vector differences are calculated and managed, except now depth data is stored and warped. In at least one embodiment, it may be beneficial to utilize both warped history depth data and warped history motion data, as they can provide very different, and often complementary, information about what is happening in a given scene. In at least one embodiment, this may include capturing instances of depth differences when no motion vector differences actually exist, and vice versa. In at least one embodiment, using both data can help the network regarding the semantics of an object or scene, which can help the network make better overall decisions.

[0034] In at least one embodiment, a color variance mask may also be provided and utilized. In at least one embodiment, the mask may represent the difference between the mean value of a pixel neighborhood in the current frame and the corresponding warping history value, divided by the standard deviation of the pixel neighborhood in the current frame multiplied by a scalar value. In at least one embodiment, the mask may be generated using full-color, single-channel, or grayscale color data. In at least one embodiment, the mask may represent the color variation between a center pixel in a neighborhood and a set of neighboring pixels within the neighborhood. In at least one embodiment, using less than full-color color data may significantly reduce the amount of computation and storage required to calculate and utilize the mask.

[0035] In at least one embodiment, advantages can be obtained by using variations or combinations of values. In at least one embodiment, this can include using numerically reduced-precision versions of one or more of the motion vector differences, chromatic variance, and / or luma variance. In at least one embodiment, these lower-precision versions still provide a significant (or at least significant) improvement in output quality, but do not require as much resource overhead. In at least one embodiment, a user can have the option to select between precision or performance for these or other such values. In at least one embodiment, sums of differences, such as sums of motion vector differences, depth differences, and chromatic variance, may also be used.

[0036] In at least one embodiment, by using features that cover both spatial and temporal variations per pixel, the network is able to better understand when an object or portion of an object is occluded and / or unoccluded, and how lighting and shading change over time. In at least one embodiment, the network can then make better decisions regarding how much data to accumulate and where to accumulate data across frames, versus when to use only the upscaled data for the current frame. In at least one embodiment, this results in a more stable image with less temporal artifacts such as flickering, ghosting, and flashing. In at least one embodiment, such an approach can also provide higher upscaling ratios, increasing application performance, such as improving the performance of video games that utilize such upscaling.

[0037] In at least one embodiment, at least some of these values ​​may be available with less precision, such as by using a low-precision floating-point format. In at least one embodiment, this may include storing some data in FP32, some in FP16, and some in FP8. In at least one embodiment, the precision of various values ​​may be reduced where appropriate, which may be configurable by a user, application, or other such entity or source. In at least one embodiment, the FP8 format may include one sign bit, five exponent bits, and a two-bit mantissa. In at least one embodiment, different internal representations may also be used that utilize the same bit width but may utilize different bit formats, for example, possibly trading fewer mantissa bits for a different number of exponent bits. In at least one embodiment, this may represent a trade-off regarding the range of precision logarithmic value representation. In at least one embodiment, reducing the precision of certain data values ​​may help offset the additional overhead and possible hit to performance that may be experienced when utilizing additional features as discussed herein, which may require more processing, memory, and other such factors or resources. In at least one embodiment, feature inputs can be combined to attempt to reduce the amount of data to be processed, and therefore reduce the overall impact on performance, while utilizing this additional data to help improve the quality of the output image data. In at least one embodiment, this may include summing feature inputs, such as motion vectors and color variance data. In at least one embodiment, other combinations of features can be combined, such as through summation or multiplication.

[0038] In at least one embodiment, the client device 502 can generate this content for a session, such as a gaming session or a video viewing session, using components of a content application 504 on the client device 502 and data stored locally on the client device as shown in Figure 5. In at least one embodiment, a content application 524 (e.g., a gaming or streaming media application) executing on a content server 520 can initiate a session associated with at least the client device 502, utilizing user data stored in a session manager and user database 534, such that content 532, as needed for this type of content or platform, can be determined by a content manager 526, rendered using a rendering engine 528, and transmitted to the client device 502 using an appropriate transmission manager 522 for transmission by download, streaming, or another such transmission channel. In at least one embodiment, a client device 502 receiving this content can provide this content to a corresponding content application 504, which may additionally or alternatively include a rendering engine 510 for rendering at least a portion of this content for presentation via the client device 502, such as presenting video content through a display 506, presenting audio such as voice and music through at least one audio playback device 508, such as speakers or headphones.In at least one embodiment, at least a portion of this content may already be stored, rendered, or accessible on the client device 502 so that at least that portion of the content does not need to be transmitted over the network 540, such as when the content may have been previously downloaded or stored locally on a hard drive or optical disc. In at least one embodiment, this content may be transferred to the client device 502 from the server 520 or content database 534 using a transmission mechanism such as data streaming. In at least one embodiment, at least a portion of this content may be obtained or streamed from another source, such as a third-party content service 550, which may also include a content application 552 for generating or providing the content. In at least one embodiment, portions of this functionality may be performed using multiple computing devices or multiple processors within one or more computing devices, such as those that may include a combination of a CPU and a GPU.

[0039] In at least one embodiment, the content application 524 includes a content manager 526 that can determine or analyze content before the content is sent to the client device 502. In at least one embodiment, the content manager 526 can also include or cooperate with other components that can generate, modify, or enhance the content to be provided. In at least one embodiment, this can include a rendering engine 528 for rendering content, such as aliased content, at a first resolution. In at least one embodiment, an upsampling or scaling component 530 can generate at least one additional version of the image at a different, higher or lower resolution and can perform at least some processing, such as anti-aliasing. In at least one embodiment, a blending component 532, which may include at least one neural network, can perform blending of one or more of the images against one or more previous images, as discussed herein. In at least one embodiment, the content manager 526 can then select an image or video frame of the appropriate resolution to send to the client device 502. In at least one embodiment, content application 504 on client device 502 may also include components such as rendering engine 510, upsampling module 512, and blending module 514, such that any or all of this functionality may additionally or alternatively be performed on client device 502. In at least one embodiment, content application 552 on third-party content serving system 550 may also include such functionality.In at least one embodiment, the location where at least a portion of this functionality is performed may be configurable or may depend on factors such as the type of client device 502 or the availability of a network connection with adequate bandwidth, among other such factors. In at least one embodiment, the upsampling module 530 or the blending module 532 may include one or more neural networks to perform or support this functionality, and those neural networks (or at least the network parameters of those networks) may be provided by the content server 520 or a third-party system 550. In at least one embodiment, a system for content generation may include any suitable combination of hardware and software in one or more locations. In at least one embodiment, the generated image or video content in one or more resolutions may also be provided to or made available to other client devices 560, such as for downloading or streaming from a media source that stores a copy of the image or video content. In at least one embodiment, this may include transmitting images of game content for a multiplayer game where different client devices are capable of displaying that content at different resolutions, including one or more super resolutions.

[0040] Logic of inference and training Figure 6A illustrates inference and / or training logic 615 used to perform inference and / or training operations for one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and / or 6B.

[0041] In at least one embodiment, the inference and / or training logic 615 may include, without limitation, code and / or data storage 601 for storing forward and / or output weights, and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network that are trained and / or used to infer in one or more embodiments. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 601 for storing graph code or other software for controlling the timing and / or sequence of logic loaded with weights and / or other parameter information, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weights or other parameter information into a processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, code and / or data storage 601 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 601 may be included with other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache, or system memory.

[0042] In at least one embodiment, any portion of code and / or data storage 601 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 601 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 601 is internal or external to a processor, or whether it is comprised of DRAM, SRAM, flash, or some other type of storage, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in neural network inference and / or training, or any combination of these factors.

[0043] In at least one embodiment, the inference and / or training logic 615 may include, without limitation, code and / or data storage 605 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used to infer in accordance with one or more aspects of the embodiment. In at least one embodiment, the code and / or data storage 605 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more aspects of the embodiment while backpropagating input / output data and / or weight parameters during training and / or inference using one or more aspects of the embodiment. In at least one embodiment, training logic 615 may include or be coupled to code and / or data storage 605 for storing graph code or other software for timing and / or sequencing control, and code and / or data storage 605 may be loaded with weights and / or other parameter information to configure logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weights or other parameter information into processor ALUs based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 605 may be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 caches or system memory. In at least one embodiment, any portion of code and / or data storage 605 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 605 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether code and / or data storage 605 is internal or external to the processor, for example, or whether it is comprised of DRAM, SRAM, flash, or some other type of storage, may depend on the storage available on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inferencing and / or training of the neural network, or any combination of these factors.

[0044] In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be separate storage structures. In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be the same storage structure. In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 601 and code and / or data storage 605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0045] In at least one embodiment, the inference and / or training logic 615 may include one or more arithmetic logic units (“ALUs”) 610, including, without limitation, integer and / or floating point units, for performing logical and / or arithmetic operations based at least in part on or indicated by the training and / or inference code (e.g., graph code), the results of which may generate activations (e.g., output values ​​from layers or neurons in a neural network) stored in activation storage 620, which are functions of input / output and / or weight parameter data stored in code and / or data storage 601 and / or code and / or data storage 605. In at least one embodiment, the activations stored in activation storage 620 are generated according to linear algebra and / or matrix-based calculations performed by ALU 610 in response to executing instructions or other code, where weight values ​​stored in code and / or data storage 605 and / or code and / or data storage 601 are used as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 605, or code and / or data storage 601, or in another storage, on-chip or off-chip.

[0046] In at least one embodiment, ALU 610 is included within one or more processors or other hardware logic devices or circuits, while in other embodiments, ALU 610 may be external to the processors or other hardware logic devices or circuits that use them (e.g., a coprocessor). In at least one embodiment, ALU 610 may be included within an execution unit of a processor or may otherwise be included within an ALU bank accessible by execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 601, code and / or data storage 605, and activation storage 620 may be in the same processor or other hardware logic devices or circuits, while in other embodiments, they may be in different processors or other hardware logic devices or circuits, or some combination of the same processor or other hardware logic devices or circuits and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 620 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry, and may be fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic.

[0047] In at least one embodiment, activation storage 620 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 620 may be completely or partially internal to or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether activation storage 620 is internal or external to a processor, or whether it is comprised of DRAM, SRAM, flash, or some other type of storage, for example, may depend on available on-chip versus off-chip storage, latency requirements of the training and / or inference functions being performed, batch sizes of data used in inference and / or training of the neural network, or any combination of these factors. In at least one embodiment, the inference and / or training logic 615 shown in Figure 6A may be used in conjunction with an application specific integrated circuit ("ASIC") such as Google's Tensorflow® processing unit, Graphcore™'s inference processing unit (IPU), or Intel Corp's Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 615 shown in Figure 6A may be used in conjunction with other hardware such as central processing unit ("CPU") hardware, graphics processing unit ("GPU") hardware, or field programmable gate arrays ("FPGAs").

[0048] FIG. 6B illustrates inference and / or training logic 615, according to at least one embodiment. In at least one embodiment, inference and / or training logic 615 may include, without limitation, hardware logic in which computational resources are dedicated to, or otherwise used only in conjunction with, weight values ​​or other information corresponding to one or more layers of neurons in a neural network. In at least one embodiment, inference and / or training logic 615 illustrated in FIG. 6B may be used in conjunction with an application-specific integrated circuit (ASIC), such as Google's Tensorflow® processing unit, Graphcore™'s inference processing unit (IPU), or Intel Corp.'s Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, inference and / or training logic 615 illustrated in FIG. 6B may be used in conjunction with other hardware, such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field-programmable gate arrays (FPGAs). In at least one embodiment, inference and / or training logic 615 includes, without limitation, code and / or data storage 601 and code and / or data storage 605, which may be used to store code (e.g., graph code), weight and / or bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment shown in FIG. 6B , code and / or data storage 601 and code and / or data storage 605 are each associated with dedicated computational resources, such as computation hardware 602 and computation hardware 606, respectively. In at least one embodiment, computation hardware 602 and computation hardware 606 each include one or more ALUs that perform mathematical functions, such as linear algebraic functions, solely on the information stored in code and / or data storage 601 and code and / or data storage 605, respectively, with the results stored in activation storage 620.

[0049] In at least one embodiment, each of the code and / or data storage 601 and 605 and corresponding computational hardware 602 and 606 corresponds to a different layer of a neural network, such that activations resulting from one "storage / computation pair 601 / 602" of code and / or data storage 601 and computational hardware 602 are provided as input to the next "storage / computation pair 605 / 606" of code and / or data storage 605 and computational hardware 606 to reflect the conceptual organization of the neural network. In at least one embodiment, storage / computation pairs 601 / 602 and 605 / 606 may correspond to two or more layers of the neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 615 after or in parallel with storage / computation pairs 601 / 602 and 605 / 606.

[0050] Data Center 7 illustrates an exemplary data center 700 in which at least one embodiment may be used. In at least one embodiment, the data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0051] 7, in at least one embodiment, a data center infrastructure layer 710 may include a resource orchestrator 712, grouped computing resources 714, and node computing resources (“node CRs”) 716(1) through 716(N), where “N” represents any positive integer. In at least one embodiment, node CRs 716(1) through 716(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (solid-state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules. In at least one embodiment, one or more of the nodes CR 716(1)-716(N) may be a server having one or more of the computing resources described above.

[0052] In at least one embodiment, grouped computing resources 714 may include separate groups of node CRs housed within one or more racks (not shown), or multiple racks housed in a data center at various geographic locations (also not shown). Separate groups of node CRs within grouped computing resources 714 may include grouped compute resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide compute resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power supply modules, cooling modules, and network switches in any combination.

[0053] In at least one embodiment, resource orchestrator 712 may configure or otherwise control one or more nodes CR 716(1)-716(N) and / or grouped computing resources 714. In at least one embodiment, resource orchestrator 712 may include a software design infrastructure (“SDI”) management entity for data center 700. In at least one embodiment, resource orchestrator may include hardware, software, or some combination thereof.

[0054] 7 , framework layer 720 includes a job scheduler 722, a configuration manager 724, a resource manager 726, and a distributed file system 728. In at least one embodiment, framework layer 720 may include frameworks to support software 732 in software layer 730 and / or one or more applications 742 in application layer 740. In at least one embodiment, software 732 or applications 742 may each include web-based service software or applications, such as those offered by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 720 may be a type of free and open-source software web application framework, such as, but not limited to, Apache Spark™ (hereinafter “Spark”), which can use distributed file system 728 for large-scale data processing (e.g., “big data”). In at least one embodiment, job scheduler 722 may include a Spark driver to facilitate scheduling of workloads supported by various tiers of data center 700. In at least one embodiment, configuration manager 724 may be capable of configuring different tiers, such as software tier 730, as well as framework tier 720, which includes Spark and distributed file system 728 to support large-scale data processing. In at least one embodiment, resource manager 726 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support distributed file system 728 and job scheduler 722. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 714 in data center infrastructure tier 710.In at least one embodiment, resource manager 726 may manage these mappings or allocated computing resources in conjunction with resource orchestrator 712.

[0055] In at least one embodiment, software 732 included in software layer 730 may include software used by nodes CR 716(1)-716(N), grouped computing resources 714, and / or at least a portion of distributed file system 728 of framework layer 720. The one or more types of software may include, but are not limited to, internet web page searching software, email virus scanning software, database software, and streaming video content software.

[0056] In at least one embodiment, the applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the nodes CRs 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 728 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive compute, and machine learning applications including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0057] In at least one embodiment, any of configuration manager 724, resource manager 726, and resource orchestrator 712 may implement any number and types of self-corrective actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-corrective actions may enable data center operators of data center 700 to avoid determining potentially faulty configurations and eliminate underutilized and / or underperforming portions of the data center.

[0058] In at least one embodiment, data center 700 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, machine learning models may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 700. In at least one embodiment, trained machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 700 by using weight parameters calculated by one or more techniques described herein.

[0059] In at least one embodiment, the data center may use a CPU, application specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the resources described above. Additionally, one or more of the software and / or hardware resources described above may be configured as a service to enable a user to train or perform inference on information, such as image recognition, speech recognition, or other artificial intelligence services.

[0060] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of Figure 7 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0061] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0062] Computer Systems 8 is a block diagram illustrating an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof 800 formed with a processor that may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, computer system 800 may include components such as, without limitation, a processor 802 for using an execution unit that includes logic for executing algorithms for processing data in accordance with the present disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 800 may include a processor such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, computer system 800 may run a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux), embedded software, and / or graphical user interfaces may also be used.

[0063] Embodiments may be used in other devices, such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, embedded applications may include a microcontroller, a digital signal processor ("DSP"), a system-on-chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system capable of executing one or more instructions according to at least one embodiment.

[0064] In at least one embodiment, computer system 800 may include, without limitation, a processor 802, which may include one or more execution units 808 for performing machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, computer system 800 is a single-processor desktop or server system, while in other embodiments, computer system 800 may be a multiprocessor system. In at least one embodiment, processor 802 may include, without limitation, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 802 may be coupled to a processor bus 810, which may transmit digital signals between processor 802 and other components within computer system 800.

[0065] In at least one embodiment, processor 802 may include, without limitation, level 1 ("L1") internal cache memory ("cache") 804. In at least one embodiment, processor 802 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory may be external to processor 802. Other embodiments may include a combination of both internal and external cache, depending on the particular implementation and needs. In at least one embodiment, register file 806 may store different types of data in various registers, including, without limitation, integer registers, floating-point registers, status registers, and an instruction pointer register.

[0066] In at least one embodiment, processor 802 also includes an execution unit 808, including, without limitation, logic for performing integer and floating-point operations. In at least one embodiment, processor 802 may also include microcode (“u-code”) read-only memory (“ROM”) that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 808 may include logic for a packed instruction set 809. In at least one embodiment, including packed instruction set 809, along with associated circuitry for executing the instructions, in the instruction set of general-purpose processor 802 allows operations used by many multimedia applications to be performed using packed data in general-purpose processor 802. In one or more embodiments, many multimedia applications can be accelerated and run more efficiently by performing operations on packed data using the full width of the processor's data bus, thereby eliminating the need to transfer smaller units of data between the processor's data bus to perform one or more operations on one data element at a time.

[0067] In at least one embodiment, execution unit 808 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 800 may include, without limitation, memory 820. In at least one embodiment, memory 820 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other memory device. In at least one embodiment, memory 820 may store instructions 819 and / or data 821 represented by data signals that may be executed by processor 802.

[0068] In at least one embodiment, a system logic chip may be coupled to processor bus 810 and memory 820. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (“MCH”) 816, and processor 802 may communicate with MCH 816 via processor bus 810. In at least one embodiment, MCH 816 may provide a high-bandwidth memory path 818 to memory 820 for storing instructions and data, and for storing graphics commands, data, and textures. In at least one embodiment, MCH 816 may route data signals between processor 802, memory 820, and other components of computer system 800, and may bridge data signals between processor bus 810, memory 820, and system I / O 822. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 816 may be coupled to memory 820 via a high-bandwidth memory path 818, and the graphics / video card 812 may be coupled to the MCH 816 via an accelerated graphics port (“AGP”) interconnect 814.

[0069] In at least one embodiment, computer system 800 may use system I / O 822, a proprietary hub interface bus, to couple MCH 816 to I / O controller hub (“ICH”) 830. In at least one embodiment, ICH 830 may provide direct connectivity to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 820, a chipset, and processor 802. Examples may include, without limitation, an audio controller 829, a firmware hub (“flash BIOS”) 828, a wireless transceiver 826, data storage 824, a legacy I / O controller 823 including a user input and keyboard interface 825, a serial expansion port 827 such as a Universal Serial Bus (“USB”), and a network controller 834. Data storage 824 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0070] In at least one embodiment, Figure 8 illustrates a system including interconnected hardware devices or "chips," while in other embodiments, Figure 8 may illustrate an exemplary system-on-a-chip ("SoC"). In at least one embodiment, the devices illustrated in Figure cc may be interconnected using a proprietary interconnect, a standard interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 800 may be interconnected using a compute express link (CXL) interconnect.

[0071] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of Figure 8 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0072] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0073] 9 is a block diagram illustrating an electronic device 900 for utilizing a processor 910, according to at least one embodiment. In at least one embodiment, electronic device 900 may be, for example, without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0074] In at least one embodiment, system 900 may include a processor 910 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices, including, without limitation, a 10G bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advance Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, and 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus, or other bus or interface. In at least one embodiment, FIG. 9 illustrates a system including interconnected hardware devices or "chips," while in other embodiments, FIG. 9 may illustrate an exemplary system-on-a-chip ("SoC"). In at least one embodiment, the devices illustrated in FIG. 9 may be interconnected with a proprietary interconnect, a standard interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 9 may be interconnected using a Compute Express Link (CXL) interconnect.

[0075] In at least one embodiment, FIG. 9 illustrates a display 924, a touch screen 925, a touch pad 930, a Near Field Communications unit ("NFC") 945, a sensor hub 940, a thermal sensor 946, an Express Chipset ("EC") 935, a Trusted Platform Module ("TPM") 938, a BIOS / firmware / flash memory ("BIOS,FW flash") 922, a DSP 960, a drive 920, such as a Solid State Disk ("SSD") or a Hard Disk Drive ("HDD"), a wireless local area network unit ("WLAN") 950, a Bluetooth unit 952, a Wireless Wide Area Network unit ("WWAN") 956, a Global Positioning System (GPS), a Bluetooth module 952, a Bluetooth 954 module 956, a Bluetooth module 954, a Bluetooth module 956, a Bluetooth module 958, a Bluetooth module 958, a Bluetooth module 959, a Bluetooth module 960 ... The memory may include a USB 3.0 System unit 955, a camera such as a USB 3.0 camera ("USB 3.0 Camera") 954, and / or a Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3") 915, for example, implemented to the LPDDR3 standard. Each of these components may be implemented in any suitable manner.

[0076] In at least one embodiment, other components may be communicatively coupled to processor 910 via the components described above. In at least one embodiment, an accelerometer 941, an ambient light sensor (“ALS”) 942, a compass 943, and a gyroscope 944 may be communicatively coupled to sensor hub 940. In at least one embodiment, a thermal sensor 939, a fan 937, a keyboard 946, and a touchpad 930 may be communicatively coupled to EC 935. In at least one embodiment, a speaker 963, headphones 964, and a microphone (“mic”) 965 may be communicatively coupled to an audio unit (“audio codec and class D amplifier”) 962, which may be communicatively coupled to DSP 960. In at least one embodiment, audio unit 964 may include, for example, without limitation, an audio coder / decoder (“codec”) and a class D amplifier. In at least one embodiment, a SIM card (“SIM”) 957 may be communicatively coupled to the WWAN unit 956. In at least one embodiment, components such as the WLAN unit 950 and Bluetooth unit 952, as well as the WWAN unit 956, may be implemented in a Next Generation Form Factor (“NGFF”).

[0077] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of Figure 9 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0078] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0079] 10 illustrates a computer system 1000 according to at least one embodiment. In at least one embodiment, the computer system 1000 is configured to implement the various processes and methods described throughout this disclosure.

[0080] In at least one embodiment, computer system 1000 includes at least one central processing unit ("CPU") 1002 connected to a communication bus 1010 implemented using any suitable protocol, such as, without limitation, PCI (Peripheral Component Interconnect), Peripheral Component Interconnect Express ("PCI-Express"), AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1000 includes main memory 1004 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data is stored in main memory 1004, which may be in the form of random access memory ("RAM"). In at least one embodiment, network interface subsystem (“network interface”) 1022 provides an interface with other computing devices and networks to receive data from other systems and transmit data from computer system 1000 to other systems.

[0081] In at least one embodiment, computer system 1000 includes, without limitation, an input device 1008, a parallel processing system 1012, and a display device 1006, which may be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light emitting diode ("LED"), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from input device 1008, such as a keyboard, mouse, touch pad, microphone, or the like. In at least one embodiment, each of the above modules may be located on a single semiconductor platform to form a processing system.

[0082] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of Figure 10 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0083] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0084] 11 illustrates a computer system 1100 according to at least one embodiment. In at least one embodiment, computer system 1100 may include, without limitation, a computer 1110 and a USB stick 1120. In at least one embodiment, computer 1110 may include, without limitation, any number and type of processor (not shown) and memory (not shown). In at least one embodiment, computer 1110 includes, without limitation, a server, a cloud instance, a laptop, and a desktop computer.

[0085] In at least one embodiment, USB stick 1120 includes, without limitation, a processing unit 1130, a USB interface 1140, and USB interface logic 1150. In at least one embodiment, processing unit 1130 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 1130 may include, without limitation, any number and types of processing cores (not shown). In at least one embodiment, processing core 1130 comprises an application specific integrated circuit ("ASIC") optimized to perform any quantity and type of operations related to machine learning. For example, in at least one embodiment, processing core 1130 is a tensor processing unit ("TPC") optimized to perform machine vision and machine learning inference operations. In at least one embodiment, processing core 1130 is a vision processing unit ("VPU") optimized to perform machine vision and machine learning inference operations.

[0086] In at least one embodiment, USB interface 1140 may be any type of USB connector or socket. For example, in at least one embodiment, USB interface 1140 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, USB interface 1140 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 1150 may include any amount and type of logic that enables processing unit 1130 to interface with a device (e.g., computer 1110) via USB connector 1140.

[0087] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of Figure 11 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0088] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0089] 12A illustrates an exemplary architecture in which multiple GPUs 1210-1213 are communicatively coupled to multiple multi-core processors 1205-1206 via high-speed links 1240-1243 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 1240-1243 support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or more. Various interconnect protocols may be used, including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0.

[0090] Additionally, in one embodiment, two or more of GPUs 1210-1213 may be interconnected via high-speed links 1229-1230, which may be implemented using the same or different protocol / links as used for high-speed links 1240-1243. Similarly, two or more of multi-core processors 1205-1206 may be connected via high-speed link 1228, which may be a symmetric multi-processor (SMP) bus operating at 20 GB / s, 30 GB / s, 120 GB / s, or more. Alternatively, all communications between the various system components shown in FIG. 12A may be achieved using the same protocol / links (e.g., via a common interconnect fabric).

[0091] In one embodiment, each multi-core processor 1205-1206 is communicatively coupled to processor memory 1201-1202 via memory interconnects 1226-1227, respectively, and each GPU 1210-1213 is communicatively coupled to GPU memory 1220-1223 via GPU memory interconnects 1250-1253, respectively. Memory interconnects 1226-1227 and 1250-1253 may utilize the same or different memory access technologies. By way of example, and not limitation, processor memory 1201-1202 and GPU memory 1220-1223 may be volatile memory such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or non-volatile memory such as 3D XPoint or Nano-RAM. In one embodiment, some portions of processor memory 1201-1202 may be volatile memory and other portions may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0092] As described below, the various processors 1205-1206 and GPUs 1210-1213 may each be physically coupled to a particular memory 1201-1202, 1220-1223, and a unified memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. For example, the processor memories 1201-1202 may each have 64 GB of system memory address space, and the GPU memories 1220-1223 may each have 32 GB of system memory address space (resulting in this example in a total of 256 GB of addressable memory).

[0093] 12B shows further details of the interconnection between multi-core processor 1207 and graphics acceleration module 1246 according to one example embodiment. Graphics acceleration module 1246 may include one or more GPU chips integrated on a line card that is coupled to processor 1207 via high-speed link 1240. Alternatively, graphics acceleration module 1246 may be integrated in the same package or chip as processor 1207.

[0094] In at least one embodiment, the illustrated processor 1207 includes multiple cores 1260A-1260D, each having a translation lookaside buffer 1261A-1261D and one or more caches 1262A-1262D. In at least one embodiment, the cores 1260A-1260D may include various other components (not shown) for executing instructions and processing data. The caches 1262A-1262D may comprise level 1 (L1) and level 2 (L2) caches. Additionally, one or more shared caches 1256 may be included in the caches 1262A-1262D and shared by the set of cores 1260A-1260D. For example, one embodiment of the processor 1207 includes 24 cores, each with its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more L2 and L3 caches are shared by two adjacent cores. The processor 1207 and graphics acceleration module 1246 are connected to system memory 1214, which may include processor memory 1201-1202 of FIG. 12A.

[0095] Coherence is maintained for data and instructions stored in the various caches 1262A-1262D, 1256 and system memory 1214 through inter-core communication via coherence bus 1264. For example, each cache may have associated cache coherence logic / circuitry for communicating via coherence bus 1264 in response to detecting a read or write to a particular cache line. In one embodiment, a cache snooping protocol is implemented via coherence bus 1264 to monitor cache accesses.

[0096] In one embodiment, proxy circuit 1225 communicatively couples graphics acceleration module 1246 to coherence bus 1264 to enable graphics acceleration module 1246 to participate in cache coherence protocols as a peer of cores 1260A-1260D. In particular, interface 1235 provides a connection to proxy circuit 1225 over high-speed link 1240 (e.g., PCIe bus, NVLink, etc.), and interface 1237 connects graphics acceleration module 1246 to link 1240.

[0097] In one embodiment, the accelerator integrated circuit 1236 provides cache management, memory access, content management, and interrupt management services on behalf of the multiple graphics processing engines 1231, 1232, N of the graphics acceleration module 1246. The graphics processing engines 1231, 1232, N may each comprise a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1231, 1232, N may comprise different types of graphics processing engines within a GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module 1246 may be a GPU having multiple graphics processing engines 1231, 1232, N, or the graphics processing engines 1231, 1232, N may be individual GPUs integrated into a common package, line card, or chip.

[0098] In one embodiment, accelerator integrated circuitry 1236 includes a memory management unit (MMU) 1239 for performing various memory management functions, such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation), and a memory access protocol for accessing system memory 1214. MMU 1239 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In one embodiment, cache 1238 may store commands and data for efficient access by graphics processing engines 1231-1232, N. In one embodiment, data stored in cache 1238 and graphics memory 1233-1234, M is kept coherent with core caches 1262A-1262D, 1256, and system memory 1214. As described above, this may be accomplished via proxy circuit 1225 on behalf of cache 1238 and memory 1233-1234, M (e.g., sending updates regarding modifications / accesses of cache lines in processor caches 1262A-1262D, 1256 to cache 1238 and receiving updates from cache 1238).

[0099] A set of registers 1245 stores context data for threads executed by graphics processing engines 1231-1232, N, and context management circuit 1248 manages thread contexts. For example, context management circuit 1248 may perform save and restore operations to save and restore the context of various threads during a context switch (e.g., where a first thread is saved and a second thread is saved so that the second thread can be executed by the graphics processing engine). For example, during a context switch, context management circuit 1248 may store current register values ​​in a designated area of ​​memory (e.g., identified by a context pointer). Then, when returning to the context, context management circuit 1248 may restore the register values. In one embodiment, interrupt management circuit 1247 receives and processes interrupts received from system devices.

[0100] In one embodiment, virtual / effective addresses from the graphics processing engine 1231 are translated to real / physical addresses in the system memory 1214 by the MMU 1239. One embodiment of the accelerator integration circuit 1236 supports multiple (e.g., four, eight, or sixteen) graphics accelerator modules 1246 and / or other accelerator devices. The graphics accelerator modules 1246 may be dedicated to a single application running on the processor 1207 or may be shared among multiple applications. In one embodiment, a virtualized graphics execution environment exists in which the resources of the graphics processing engines 1231-1232, N, are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources may be subdivided into "slices," which are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.

[0101] In at least one embodiment, the accelerator integration circuitry 1236 acts as a bridge to the system for the graphics acceleration module 1246, providing address translation and system memory caching services. Additionally, the accelerator integration circuitry 1236 may provide a virtualization facility for the host processor to manage virtualization, interrupts, and memory management for the graphics processing engines 1231-1232, N.

[0102] The hardware resources of the graphics processing engines 1231-1232, N are explicitly mapped into the real address space seen by the host processor 1207, so that any host processor can directly address these resources using effective address values. One function of the accelerator integrated circuit 1236 is to physically separate the graphics processing engines 1231-1232, N so that they appear as independent units to the system.

[0103] In at least one embodiment, one or more graphics memories 1233-1234, M are respectively coupled to each of the graphics processing engines 1231-1232, N. The graphics memories 1233-1234, M store instructions and data that are processed by the respective graphics processing engines 1231-1232, N. The graphics memories 1233-1234, M may be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memory such as 3D XPoint or Nano-Ram.

[0104] In one embodiment, to reduce data traffic over link 1240, a biasing technique is used to ensure that the data stored in graphics memory 1233-1234, M is data that will be most frequently used by graphics processing engines 1231-1232, N, and preferably is data that is not used (or at least not frequently used) by cores 1260A-1260D. Similarly, the biasing mechanism attempts to keep data needed by the cores (and therefore preferably not needed by graphics processing engines 1231-1232, N) in the cores' caches 1262A-1262D, 1256 and system memory 1214.

[0105] FIG. 12C illustrates another exemplary embodiment in which the accelerator integration circuitry 1236 is integrated within the processor 1207. In at least this embodiment, the graphics processing engines 1231-1232, N communicate directly with the accelerator integration circuitry 1236 via high-speed link 1240 via interface 1237 and interface 1235 (again, any form of bus or interface protocol can be utilized). The accelerator integration circuitry 1236 may perform the same operations as described with respect to FIG. 12B, but may potentially operate at a higher throughput given its proximity to the coherence bus 1264 and caches 1262A-1262D, 1256. At least one embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuitry 1236 and a programming model controlled by the graphics acceleration module 1246.

[0106] In at least one embodiment, graphics processing engines 1231-1232, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to graphics processing engines 1231-1232, N, achieving virtualization within a VM / partition.

[0107] In at least one embodiment, graphics processing engines 1231-1232,N may be shared by multiple VM / application partitions. In at least one embodiment, the sharing model may use a system hypervisor to virtualize graphics processing engines 1231-1232,N and allow access by each operating system. In a single-partition system without a hypervisor, graphics processing engines 1231-1232,N are owned by the operating system. In at least one embodiment, the operating system may virtualize graphics processing engines 1231-1232,N and provide access to each process or application.

[0108] In at least one embodiment, the graphics acceleration module 1246 or an individual graphics processing engine 1231-1232,N selects a process element using a process handle. In at least one embodiment, the process element is stored in system memory 1214 and is addressable using the effective address to real address translation techniques described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to a host process when registering the host process's context with the graphics processing engine 1231-1232,N (i.e., calling system software to add the process element to the process element linked list). In at least one embodiment, the low-order 16 bits of the process handle may be the offset of the process element within the process element linked list.

[0109] FIG. 12D illustrates an exemplary accelerator integration slice 1290. As used herein, a "slice" comprises a designated portion of the processing resources of the accelerator integration circuitry 1236. An application effective address space 1282 in system memory 1214 stores a process element 1283. In at least one embodiment, the process element 1283 is stored in response to a GPU call 1281 from an application 1280 running on the processor 1207. The process element 1283 contains the process state of the corresponding application 1280. A work descriptor (WD) 1284 contained in the process element 1283 can be a single job requested by the application or may contain a pointer to a queue of jobs. In at least one embodiment, the WD 1284 is a pointer to a job request queue in the application's address space 1282.

[0110] The graphics acceleration module 1246 and / or the individual graphics processing engines 1231-1232, N may be shared by all or a subset of the processes in the system. In at least one embodiment, infrastructure may be included for setting process state and sending WD 1284 to the graphics acceleration module 1246 to start a job in a virtualized environment.

[0111] In at least one embodiment, the dedicated process programming model is implementation specific, in which a single process owns the graphics acceleration module 1246 or an individual graphics processing engine 1231. Because the graphics acceleration module 1246 is owned by a single process, when the graphics acceleration module 1246 is allocated, the hypervisor initializes the accelerator integration circuitry 1236 for the owning partition, and the operating system initializes the accelerator integration circuitry 1236 for the owning process.

[0112] In operation, WD fetch unit 1291 in accelerator integrated slice 1290 fetches the next WD 1284, which contains an indication of work to be performed by one or more graphics processing engines of graphics acceleration module 1246. As shown, data from WD 1284 is stored in register 1245 and may be used by MMU 1239, interrupt management circuit 1247, and / or context management circuit 1248. For example, one embodiment of MMU 1239 includes segment / page walk circuitry for accessing segment / page table 1286 within OS virtual address space 1285. Interrupt management circuit 1247 may process interrupt events 1292 received from graphics acceleration module 1246. When performing graphics operations, effective addresses 1293 generated by graphics processing engines 1231-1232, N are translated into real addresses by MMU 1239.

[0113] In one embodiment, the same set of registers 1245 may be replicated for each graphics processing engine 1231-1232, N, and / or graphics acceleration module 1246 and initialized by the hypervisor or operating system. Each of these replicated registers may be included in the accelerator integration slice 1290. Exemplary registers that may be initialized by the hypervisor are shown in Table 1. [Table 1]

[0114] Exemplary registers that may be initialized by the operating system are shown in Table 2. [Table 2]

[0115] In one embodiment, each WD 1284 is specific to a particular graphics acceleration module 1246 and / or graphics processing engine 1231-1232, N. The WD 1284 contains all the information the graphics processing engine 1231-1232, N needs to do its work, or it can be a pointer to a memory location where the application has set up a command queue for work to be completed.

[0116] 12E shows further details of an exemplary embodiment of the sharing model. This embodiment includes a hypervisor real address space 1298 in which a process element list 1299 is stored. The hypervisor real address space 1298 is accessible through a hypervisor 1296 that virtualizes the graphics acceleration module engine of the operating system 1295.

[0117] In at least one embodiment, a shared programming model allows all or a subset of processes from all or a subset of partitions in the system to use the graphics acceleration module 1246. There are two programming models in which the graphics acceleration module 1246 is shared by multiple processes and partitions: timeslice shared and graphics-directed shared.

[0118] In this model, the system hypervisor 1296 owns the graphics acceleration module 1246 and makes its functionality available to all operating systems 1295. In order for the graphics acceleration module 1246 to support virtualization by the system hypervisor 1296, the graphics acceleration module 1246 may comply with the following: 1) application job requests must be autonomous (i.e., no state needs to be maintained between jobs) or the graphics acceleration module 1246 must provide a mechanism for saving and restoring context; 2) the graphics acceleration module 1246 must guarantee that the application job requests are completed in a specified amount of time, including any translation errors, or the graphics acceleration module 1246 provides the ability to preempt job processing; and 3) the graphics acceleration module 1246 must ensure fairness between processes when operating in a specified shared programming model.

[0119] In at least one embodiment, the application 1280 must make a system call to the operating system 1295 with the graphics acceleration module 1246 type, a work descriptor (WD), an authorization mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the graphics acceleration module 1246 type describes the acceleration function targeted by the system call. In at least one embodiment, the graphics acceleration module 1246 type may be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 1246 and may be in the form of a graphics acceleration module 1246 command, an effective address pointer to a user-defined structure, an effective address pointer to a queue of commands, or any other data structure for describing the work to be performed by the graphics acceleration module 1246. In one embodiment, the AMR value is the AMR state to use for the current process. In at least one embodiment, the value passed to the operating system is the same as the application setting the AMR. If an embodiment of the accelerator integrated circuit 1236 and the graphics acceleration module 1246 does not support a User Permission Mask Override Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR to the hypervisor call. The hypervisor 1296 may optionally apply the current Permission Mask Override Register (AMOR) value before placing the AMR in the process element 1283. In at least one embodiment, the CSRP is one of the registers 1245 that contains the effective address of an area in the application's effective address space 1282 for the graphics acceleration module 1246 to save and restore context state. This pointer is optional if no state needs to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.

[0120] Upon receiving the system call, the operating system 1295 may verify that the application 1280 is registered and authorized to use the graphics acceleration module 1246. The operating system 1295 then calls the hypervisor 1296 with the information shown in Table 3. [Table 3]

[0121] Upon receiving the hypervisor call, the hypervisor 1296 verifies that the operating system 1295 is registered and authorized to use the graphics acceleration module 1246. The hypervisor 1296 then places the process element 1283 into the process element linked list of the corresponding graphics acceleration module 1246 type. The process element may include the information shown in Table 4. [Table 4]

[0122] In at least one embodiment, the hypervisor initializes registers 1245 of multiple accelerator integrated slices 1290.

[0123] As shown in FIG. 12F, at least one embodiment uses unified memory that is addressable via a common virtual memory address space used to access physical processor memories 1201-1202 and GPU memories 1220-1223. In this embodiment, operations performed on GPUs 1210-1213 utilize the same virtual / effective memory address space as those used to access processor memories 1201-1202, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1201, a second portion is allocated to second processor memory 1202, a third portion is allocated to GPU memory 1220, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thereby distributed across each of the processor memories 1201-1202 and GPU memories 1220-1223, allowing either processor or GPU to access either physical memory, with virtual addresses mapped to physical memory.

[0124] In one embodiment, bias / coherence management circuits 1294A-1294E in one or more of the MMUs 1239A-1239E ensure cache coherence between the caches of one or more host processors (e.g., 1205) and the caches of the GPUs 1210-1213 and implement biasing techniques to indicate the physical memory in which certain types of data should be stored. While multiple instances of bias / coherence management circuits 1294A-1294E are shown in FIG. 12F, the bias / coherence circuits may also be implemented within the MMUs of one or more host processors 1205 and / or within the accelerator integration circuit 1236.

[0125] One embodiment allows the GPU-attached memory 1220-1223 to be mapped as part of system memory and accessible using shared virtual memory (SVM) techniques, but without the performance penalty associated with full system cache coherence. In at least one embodiment, the GPU-attached memory 1220-1223 can be accessed as system memory without cumbersome cache coherence overhead, providing a beneficial operating environment for GPU offload. This configuration allows the host processor 1205 software to set up operands and access computation results without the overhead of traditional I / O DMA data copies. These traditional copies require driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, being able to access the GPU-attached memory 1220-1223 without cache coherence overhead can be crucial to the execution time of the offloaded computation. For example, in cases where there is significant streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPUs 1210-1213. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation can be useful in determining the effectiveness of GPU offloading.

[0126] In at least one embodiment, the selection of the GPU bias and the host processor bias is determined by a bias tracker data structure. For example, a bias table may be used, which may be a page-granular structure containing one or two bits per GPU-attached memory page (i.e., controlled at memory page granularity). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPU-attached memories 1220-1223, with or without a bias cache in the GPU 1210-1213 (e.g., for caching frequently / recently used entries of the bias table). Alternatively, the bias table may be maintained entirely within the GPU.

[0127] In at least one embodiment, the bias table entry associated with each access to GPU-biased memory 1220-1223 is accessed prior to the actual access to the GPU memory, resulting in the following actions: First, a local request from a GPU 1210-1213 that finds its page in the GPU bias is forwarded directly to the corresponding GPU memory 1220-1223. A local request from a GPU that finds its page in the host bias is forwarded to the processor 1205 (e.g., via the high-speed link described above). In one embodiment, a request from the processor 1205 that finds the requested page in the host processor bias completes the request similar to a normal memory read. Alternatively, a request directed to a GPU-biased page may be forwarded to the GPU 1210-1213. In at least one embodiment, the GPU may then migrate the page to the host processor bias if it is not currently using the page. In at least one embodiment, the bias state of a page can be changed by either a software-based mechanism, a hardware-assisted software-based mechanism, or for a limited set of cases, solely by a hardware-based mechanism.

[0128] One mechanism for changing the bias state utilizes an API call (e.g., OpenCL) that calls the GPU's device driver, which sends a message (or queues a command descriptor) to the GPU to change the bias state and, for some transitions, directs the GPU to perform a cache flushing operation in the host. In at least one embodiment, a cache flushing operation is used for transitions from host processor 1205 bias to GPU bias, but not for transitions in the opposite direction.

[0129] In one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages uncacheable by the host processor 1205. To access these pages, the processor 1205 may request access from the GPU 1210, which may or may not immediately grant the access. Therefore, to reduce communication between the processor 1205 and the GPU 1210, it is beneficial for GPU-biased pages to be requested by the GPU but not by the host processor 1205, or vice versa.

[0130] To implement one or more embodiments, inference and / or training logic 615 is used. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and / or 6B.

[0131] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0132] 13 illustrates an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general-purpose processor cores.

[0133] 13 is a block diagram illustrating an exemplary system-on-chip integrated circuit 1300 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1300 includes one or more application processors 1305 (e.g., CPUs), at least one graphics processor 1310, and may further include an image processor 1315 and / or a video processor 1320, any of which may be modular IP cores. In at least one embodiment, integrated circuit 1300 includes a USB controller 1325, a UART controller 1330, an SPI / SDIO controller 1335, and an I / O controller 1340. 2 S / I 2 The integrated circuit 1300 may include peripheral or bus logic including a HDMI controller 1340. In at least one embodiment, the integrated circuit 1300 may include a display device 1345 coupled to one or more of a high-definition multimedia interface (HDMI®) controller 1350 and a mobile industry processor interface (MIPI) display interface 1355. In at least one embodiment, storage may be provided by a flash memory subsystem 1360 including a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1365 for accessing an SDRAM or SRAM memory device. In at least one embodiment, some integrated circuits further include an embedded security engine 1370.

[0134] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in integrated circuit 1300 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0135] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0136] 14A-14B illustrate an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0137] 14A-14B are block diagrams illustrating exemplary graphics processors for use within an SoC, according to embodiments described herein. FIG. 14A illustrates an exemplary graphics processor 1410 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores, according to at least one embodiment. FIG. 14B illustrates a further exemplary graphics processor 1440 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the graphics processor 1410 of FIG. 14A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1440 of FIG. 14B is a high-performance graphics processor core. In at least one embodiment, each of the graphics processors 1410, 1440 can be a variation of the graphics processor 1310 of FIG. 13.

[0138] In at least one embodiment, graphics processor 1410 includes a vertex processor 1405 and one or more fragment processors 1415A-1415N (e.g., 1415A, 1415B, 1415C, 1415D-1415N-1, and 1415N). In at least one embodiment, graphics processor 1410 can execute different shader programs through separate logic, such that vertex processor 1405 is optimized to perform operations for vertex shader programs, while one or more fragment processors 1415A-1415N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, vertex processor 1405 executes the vertex processing stage of a 3D graphics pipeline, generating primitive and vertex data. In at least one embodiment, fragment processors 1415A-1415N use the primitive and vertex data generated by vertex processor 1405 to generate a frame buffer that is displayed on a display device. In at least one embodiment, fragment processors 1415A-1415N are optimized to execute fragment shader programs provided in the OpenGL API, which may be used to perform operations similar to pixel shader programs provided in the Direct 3D API.

[0139] In at least one embodiment, the graphics processor 1410 further includes one or more memory management units (MMUs) 1420A-1420B, caches 1425A-1425B, and circuit interconnects 1430A-1430B. In at least one embodiment, the one or more MMUs 1420A-1420B provide virtual-to-physical address mapping for the graphics processor 1410, including the vertex processor 1405 and / or fragment processors 1415A-1415N, which may reference vertex or image / text data stored in memory in addition to vertex or image / text data stored in one or more caches 1425A-1425B. In at least one embodiment, one or more MMUs 1420A-1420B may be synchronized with other MMUs in the system, including one or more MMUs associated with one or more application processors 1305, image processor 1315, and / or video processor 1320 of Figure 13, allowing each processor 1305-1320 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1430A-1430B enable graphics processor 1410 to interface with other IP cores in the SoC via the SoC's internal bus or via a direct connection.

[0140] In at least one embodiment, graphics processor 1440 includes one or more MMUs 1420A-1420B, caches 1425A-1425B, and circuit interconnects 1430A-1430B of graphics processor 1410 of FIG. 14A. In at least one embodiment, graphics processor 1440 includes one or more shader cores 1455A-1455N (e.g., 1455A, 1455B, 1455C, 1455D, 1455E, 1455F-1455N-1, and 1455N) that provide a unified shader core architecture in which a single core, or type, or cores can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, graphics processor 1440 includes an inter-core task manager 1445 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1455A-1455N, and a tiling unit 1458 for accelerating tiling operations for tile-based rendering, where rendering operations of a scene are subdivided in image space, e.g., to exploit local spatial coherence within a scene or to optimize internal cache usage.

[0141] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B . In at least one embodiment, inference and / or training logic 615 may be used in integrated circuits 14A and / or 14B for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein. Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0142] 15A-15B illustrate further exemplary graphics processor logic according to embodiments described herein. FIG. 15A illustrates a graphics core 1500, which, in at least one embodiment, may be included in graphics processor 1310 of FIG. 13, or, in at least one embodiment, may be integrated shader cores 1455A-1455N, as in FIG. 14B. FIG. 15B illustrates a highly parallel, general-purpose graphics processing unit 1530 suitable for incorporation into a multi-chip module in at least one embodiment.

[0143] In at least one embodiment, graphics core 1500 includes a shared instruction cache 1502, a texture unit 1518, and a cache / shared memory 1520, which are common to execution resources within graphics core 1500. In at least one embodiment, graphics core 1500 may include multiple slices 1501A-1501N, or partitions per core, and a graphics processor may include multiple instances of graphics core 1500. Slices 1501A-1501N may include supporting logic, including local instruction caches 1504A-1504N, thread schedulers 1506A-1506N, thread dispatchers 1508A-1508N, and sets of registers 1510A-1510N. In at least one embodiment, slices 1501A-1501N may include a set of additional functional units (AFUs 1512A-1512N), floating point units (FPUs 1514A-1514N), integer arithmetic logic units (ALUs 1516-1516N), address calculation units (ACUs 1513A-1513N), double precision floating point units (DPFPUs 1515A-1515N), and matrix processing units (MPUs 1517A-1517N).

[0144] In at least one embodiment, the FPUs 1514A-1514N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, and the DPFPUs 1515A-1515N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 1516A-1516N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 1517A-1517N can also be configured for mixed-precision matrix operations, including half-precision floating-point and 8-bit integer operations. In at least one embodiment, the MPUs 1517A-1517N can perform various matrix operations to accelerate machine learning application frameworks, including being able to support general matrix-matrix multiplication (GEMM) acceleration. In at least one embodiment, AFUs 1512A-1512N can perform additional logical operations not supported by the floating-point unit or integer unit, including trigonometric operations (e.g., sine, cosine, etc.).

[0145] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with FIGURES 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in graphics core 1500 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0146] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0147] FIG. 15B illustrates a general-purpose processing unit (GPGPU) 1530, which, in at least one embodiment, can be configured to enable highly parallel computational operations by an array of graphics processing units. In at least one embodiment, the GPGPU 1530 can be directly linked to other instances of the GPGPU 1530 to create multiple GPU clusters to improve the training speed of deep neural networks. In at least one embodiment, the GPGPU 1530 includes a host interface 1532 for connecting to a host processor. In at least one embodiment, the host interface 1532 is a PCI Express interface. In at least one embodiment, the host interface 1532 can be a vendor-specific communications interface or fabric. In at least one embodiment, the GPGPU 1530 receives commands from the host processor and, using a global scheduler 1534, distributes execution threads associated with these commands to a set of compute clusters 1536A-1536H. In at least one embodiment, the compute clusters 1536A-1536H share a cache memory 1538. In at least one embodiment, the cache memory 1538 can act as a higher level cache for the cache memories within the compute clusters 1536A-1536H.

[0148] In at least one embodiment, GPGPU 1530 includes memory 1544A-1544B coupled to compute clusters 1536A-1536H via a set of memory controllers 1542A-1542B. In at least one embodiment, memory 1544A-1544B can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.

[0149] In at least one embodiment, compute clusters 1536A-1536H each include a set of graphics cores, such as graphics core 1500 of FIG. 15A, which may include multiple types of integer and floating-point logic units capable of performing computational operations with various precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each of compute clusters 1536A-1536H may be configured to perform 16-bit or 32-bit floating-point operations, while another subset of the floating-point units may be configured to perform 64-bit floating-point operations.

[0150] In at least one embodiment, multiple instances of GPGPU 1530 can be configured to operate as a compute cluster. In at least one embodiment, the communications used by compute clusters 1536A-1536H for synchronization and data exchange vary across embodiments. In at least one embodiment, multiple instances of GPGPU 1530 communicate through host interface 1532. In at least one embodiment, GPGPU 1530 includes I / O hub 1539, which couples GPGPU 1530 to GPU link 1540, which allows direct connection to other instances of GPGPU 1530. In at least one embodiment, GPU link 1540 is coupled to a dedicated GPU-to-GPU bridge, which allows communication and synchronization between multiple instances of GPGPU 1530. In at least one embodiment, GPU link 1540 is coupled to a high-speed interconnect for sending and receiving data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 1530 are located in separate data processing systems and communicate via a network device accessible via host interface 1532. In at least one embodiment, GPU link 1540 can be configured to allow connection to a host processor in addition to, or instead of, host interface 1532.

[0151] In at least one embodiment, the GPGPU 1530 can be configured to train a neural network. In at least one embodiment, the GPGPU 1530 can be used within an inference platform. In at least one embodiment, when the GPGPU 1530 is used for inference, the GPGPU may include fewer compute clusters 1536A-1536H than when the GPGPU is used to train a neural network. In at least one embodiment, the memory technology associated with memories 1544A-1544B may be different between the inference configuration and the training configuration, with higher bandwidth memory technology being devoted to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 1530 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can support one or more 8-bit integer dot product instructions, which may be used during inference operations of a deployed neural network.

[0152] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in GPGPU 1530 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0153] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0154] 16 is a block diagram illustrating a computing system 1600 according to at least one embodiment. In at least one embodiment, computing system 1600 includes a processing subsystem 1601 having one or more processors 1602 and system memory 1604 that communicate via an interconnection path that may include a memory hub 1605. In at least one embodiment, memory hub 1605 may be a separate component within a chipset component or may be integrated within one or more processors 1602. In at least one embodiment, memory hub 1605 is coupled to an I / O subsystem 1611 via communication link 1606. In at least one embodiment, I / O subsystem 1611 includes an I / O hub 1607 that can enable computing system 1600 to receive input from one or more input devices 1608. In at least one embodiment, I / O hub 1607 can enable a display controller, which may be included in one or more processors 1602 and provide output to one or more display devices 1610A. In at least one embodiment, the one or more display devices 1610A coupled to I / O hub 1607 can include local, internal, or embedded display devices.

[0155] In at least one embodiment, processing subsystem 1601 includes one or more parallel processors 1612 coupled to memory hub 1605 via a bus or other communication link 1613. In at least one embodiment, communication link 1613 may be one of any number of standard-based communication link technologies or protocols, such as, but not limited to, PCI Express, or may be a vendor-specific communication interface or fabric. In at least one embodiment, one or more parallel processors 1612 form a computationally intensive parallel or vector processing system that may include multiple processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, one or more parallel processors 1612 form a graphics processing subsystem that can output pixels to one of one or more display devices 1610A coupled via I / O hub 1607. In at least one embodiment, the one or more parallel processors 1612 may also include a display controller and display interface (not shown) that allows for direct connection to one or more display devices 1610B.

[0156] In at least one embodiment, a system storage unit 1614 may be connected to an I / O hub 1607 to provide a storage mechanism for the computing system 1600. In at least one embodiment, an I / O switch 1616 may be used to provide an interface mechanism to enable communication between the I / O hub 1607 and other components, such as a network adapter 1618 and / or a wireless network adapter 1619, which may be integrated into the platform, as well as various other devices that may be added via one or more add-in devices 1620. In at least one embodiment, the network adapter 1618 may be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1619 may include one or more of Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more wireless radios.

[0157] In at least one embodiment, computing system 1600 may include other components not shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to I / O hub 1607. In at least one embodiment, the communication paths interconnecting the various components of FIG. 16 may be implemented using any suitable protocol, such as a Peripheral Component Interconnect (PCI)-based protocol (e.g., PCI-Express), or other bus or point-to-point communication interface, such as an NV-Link high-speed interconnect, or other interconnection protocol.

[0158] In at least one embodiment, one or more parallel processors 1612 incorporate circuitry optimized for graphics and video processing, including, for example, video output circuitry, forming a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1612 incorporate circuitry optimized for general-purpose processing. In at least one embodiment, components of computing system 1600 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1612, memory hub 1605, processor 1602, and I / O hub 1607 may be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, components of computing system 1600 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 1600 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules to form a modular computing system.

[0159] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of Figure 1600 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0160] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0161] Processor 17A illustrates a parallel processor 1700 according to at least one embodiment. In at least one embodiment, various components of parallel processor 1700 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the illustrated parallel processor 1700 is a variation of one or more parallel processors 1612 shown in FIG. 16 according to an example embodiment.

[0162] In at least one embodiment, parallel processor 1700 includes parallel processing units 1702. In at least one embodiment, parallel processing units 1702 include I / O units 1704 that enable communication with other devices, including other instances of parallel processing units 1702. In at least one embodiment, I / O units 1704 may be directly connected to other devices. In at least one embodiment, I / O units 1704 are connected to other devices through the use of a hub or switch interface, such as memory hub 1605. In at least one embodiment, the connection between memory hub 1605 and I / O units 1704 forms communication link 1613. In at least one embodiment, I / O units 1704 are connected to host interface 1706 and memory crossbar 1716, where host interface 1706 receives commands directed to the execution of processing operations and memory crossbar 1716 receives commands directed to the execution of memory operations.

[0163] In at least one embodiment, when host interface 1706 receives command buffers via I / O unit 1704, host interface 1706 can direct work operations to front end 1708 to execute these commands. In at least one embodiment, front end 1708 is coupled to scheduler 1710, which is configured to distribute commands or other work items to processing cluster array 1712. In at least one embodiment, scheduler 1710 ensures that processing cluster array 1712 is properly configured and in a valid state before tasks are distributed to processing cluster array 1712. In at least one embodiment, scheduler 1710 is implemented via firmware logic running on a microcontroller. In at least one embodiment, microcontroller-implemented scheduler 1710 is configurable to perform complex scheduling and work distribution operations at both coarse and fine granularities, enabling rapid preemption and context switching of threads executing on processing array 1712. In at least one embodiment, host software can initiate scheduling workloads across processing array 1712 via one of multiple graphics processing doorbells. In at least one embodiment, the workload can then be automatically distributed across processing cluster array 1712 by scheduler 1710 logic within the microcontroller that includes scheduler 1710.

[0164] In at least one embodiment, processing cluster array 1712 includes up to “N” processing clusters (e.g., cluster 1714A, cluster 1714B through cluster 1714N). In at least one embodiment, each cluster 1714A through 1714N of processing cluster array 1712 is capable of executing a large number of simultaneous threads. In at least one embodiment, scheduler 1710 can allocate work to clusters 1714A through 1714N of processing cluster array 1712 using various scheduling and / or work distribution algorithms, which may vary depending on the workload generated by each program or type of computation. In at least one embodiment, scheduling may be handled dynamically by scheduler 1710 or may be partially assisted by compiler logic during compilation of program logic configured to be executed by processing cluster array 1712. In at least one embodiment, different clusters 1714A through 1714N of processing cluster array 1712 may be allocated to process different types of programs or perform different types of computations.

[0165] In at least one embodiment, processing cluster array 1712 may be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster array 1712 may be configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, processing cluster array 1712 may include logic for performing processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.

[0166] In at least one embodiment, the processing cluster array 1712 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 1712 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic for performing texture operations, as well as mosaic logic and other vertex processing logic. In at least one embodiment, the processing cluster array 1712 may be configured to execute graphics processing related shader programs, such as, but not limited to, vertex shaders, mosaic shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 1702 may transfer data from system memory via the I / O unit 1704 for processing. In at least one embodiment, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 1722) during processing and then written back to system memory.

[0167] In at least one embodiment, when graphics processing is performed using parallel processing unit 1702, scheduler 1710 can be configured to divide the processing workload into roughly equal-sized tasks to better distribute graphics processing operations among multiple clusters 1714A-1714N of processing cluster array 1712. In at least one embodiment, portions of processing cluster array 1712 can be configured to perform different types of processing. For example, in at least one embodiment, to generate and display a rendered image, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform mosaic and geometry shading, and a third portion may be configured to perform pixel shading or other screen space operations. In at least one embodiment, intermediate data generated by one or more of clusters 1714A-1714N may be stored in a buffer so that the intermediate data can be transmitted between clusters 1714A-1714N for further processing.

[0168] In at least one embodiment, the processing cluster array 1712 can receive processing tasks to be performed via a scheduler 1710, which receives commands defining the processing tasks from the front end 1708. In at least one embodiment, a processing task can include an index of the data to be processed, e.g., surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data should be processed (e.g., which program to execute). In at least one embodiment, the scheduler 1710 can be configured to fetch the index corresponding to the task or can receive the index from the front end 1708. In at least one embodiment, the front end 1708 can be configured to ensure that the processing cluster array 1712 is configured to a valid state before a workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is initiated.

[0169] In at least one embodiment, each of one or more instances of parallel processing unit 1702 can be coupled to parallel processor memory 1722. In at least one embodiment, parallel processor memory 1722 can be accessed via memory crossbar 1716, which can receive memory requests from processing cluster array 1712 as well as I / O unit 1704. In at least one embodiment, memory crossbar 1716 can access parallel processor memory 1722 via memory interface 1718. In at least one embodiment, memory interface 1718 can include multiple partition units (e.g., partition unit 1720A, partition unit 1720B through partition unit 1720N), each of which can be coupled to a portion (e.g., a memory unit) of parallel processor memory 1722. In at least one embodiment, the number of partition units 1720A-1720N is configured to be equal to the number of memory units, such that a first partition unit 1720A has a corresponding first memory unit 1724A, a second partition unit 1720B has a corresponding memory unit 1724B, and an Nth partition unit 1720N has a corresponding Nth memory unit 1724N. In at least one embodiment, the number of partition units 1720A-1720N does not have to be equal to the number of memory devices.

[0170] In at least one embodiment, the memory units 1724A-1724N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, the memory units 1724A-1724N may also include 3D stacked memory, including, but not limited to, high-bandwidth memory (HBM). In at least one embodiment, to efficiently use the available bandwidth of the parallel processor memory 1722, render targets, such as frame buffers or texture maps, may be stored across the memory units 1724A-1724N, allowing the partition units 1720A-1720N to write portions of each render target in parallel. In at least one embodiment, local instances of the parallel processor memory 1722 may be omitted in favor of a unified memory design that uses a combination of system memory and local cache memory.

[0171] In at least one embodiment, any one of the clusters 1714A-1714N in the processing cluster array 1712 can process data that is to be written to any one of the memory units 1724A-1724N in the parallel processor memory 1722. In at least one embodiment, the memory crossbar 1716 can be configured to forward the output of each cluster 1714A-1714N to any partition unit 1720A-1720N or to another cluster 1714A-1714N that can perform further processing operations on the output. In at least one embodiment, each cluster 1714A-1714N can communicate with a memory interface 1718 through the memory crossbar 1716 to read from or write to various external memory devices. In at least one embodiment, memory crossbar 1716 has connections to memory interface 1718 for communicating with I / O unit 1704, as well as connections to local instances of parallel processor memory 1722, allowing processing units in different processing clusters 1714A-1714N to communicate with system memory or other memory not local to parallel processing unit 1702. In at least one embodiment, memory crossbar 1716 can use virtual channels to separate traffic streams between clusters 1714A-1714N and partition units 1720A-1720N.

[0172] In at least one embodiment, multiple instances of parallel processing unit 1702 may be provided on a single add-in card, or multiple add-in cards may be interconnected. In at least one embodiment, different instances of parallel processing unit 1702 may be configured to interoperate even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of parallel processing unit 1702 may include higher precision floating-point units than other instances. In at least one embodiment, systems incorporating one or more instances of parallel processing unit 1702 or parallel processor 1700 may be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.

[0173] FIG. 17B is a block diagram of a partition unit 1720 according to at least one embodiment. In at least one embodiment, partition unit 1720 is an instance of one of partition units 1720A-1720N of FIG. 17A. In at least one embodiment, partition unit 1720 includes an L2 cache 1721, a frame buffer interface 1725, and a raster operation unit (“ROP”) 1726. L2 cache 1721 is a read / write cache configured to execute load and store operations received from memory crossbar 1716 and ROP 1726. In at least one embodiment, read misses and urgent writeback requests are output by L2 cache 1721 to frame buffer interface 1725 for processing. In at least one embodiment, updates are also sent to the frame buffer via frame buffer interface 1725 for processing. In at least one embodiment, frame buffer interface 1725 interfaces with one of the memory units of a parallel processor memory, such as memory units 1724A-1724N (eg, in parallel processor memory 1722) of FIG.

[0174] In at least one embodiment, ROP1726 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. In at least one embodiment, ROP1726 then outputs the processed graphics data stored in graphics memory. In at least one embodiment, ROP1726 includes compression logic for compressing depth or color data being written to memory and decompressing depth or color data being read from memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a number of compression algorithms. The compression logic performed by ROP1726 can be modified based on statistical characteristics of the data being compressed. For example, in at least one embodiment, delta color compression is performed on the depth and color data on a tile-by-tile basis.

[0175] In at least one embodiment, ROP 1726 is included within each processing cluster (e.g., clusters 1714A-1714N of FIG. 17A ) rather than within partition unit 1720. In at least one embodiment, read and write requests for pixel data, rather than pixel fragment data, are transmitted through memory crossbar 1716. In at least one embodiment, processed graphics data may be displayed on a display device, such as one of one or more display devices 1610 of FIG. 16 , may be routed for further processing by processor 1602, or may be routed for further processing by one of the processing entities in parallel processor 1700 of FIG. 17A .

[0176] FIG. 17C is a block diagram of a processing cluster 1714 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of processing clusters 1714A-1714N of FIG. 17A. In at least one embodiment, one or more of the processing clusters 1714 may be configured to execute multiple threads in parallel, where a "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single-instruction, multiple-data (SIMD) instruction issue techniques are used to support parallel execution of multiple threads without providing multiple independent instruction units. In at least one embodiment, single-instruction, multiple-thread (SIMT) techniques are used to support parallel execution of multiple, generally synchronized threads using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.

[0177] In at least one embodiment, operation of the processing cluster 1714 may be controlled via a pipeline manager 1732, which distributes processing tasks to the SIMT parallel processors. In at least one embodiment, the pipeline manager 1732 receives instructions from the scheduler 1710 of FIG. 17A and manages the execution of those instructions via the graphics multiprocessor 1734 and / or the texture unit 1736. In at least one embodiment, the graphics multiprocessor 1734 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 1714. In at least one embodiment, one or more instances of the graphics multiprocessor 1734 may be included within the processing cluster 1714. In at least one embodiment, the graphics multiprocessor 1734 may process data, and a data crossbar 1740 may be used to distribute the processed data to one of several possible destinations, including other shader units. In at least one embodiment, the pipeline manager 1732 can facilitate distribution of the processed data by specifying destinations for the processed data to be distributed through the data crossbar 1740.

[0178] In at least one embodiment, each graphics multiprocessor 1734 in a processing cluster 1714 may include an identical set of function execution logic (e.g., arithmetic logic units, load-store units, etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, allowing new instructions to be issued before previous instructions complete. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifts, and calculation of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.

[0179] In at least one embodiment, instructions sent to processing cluster 1714 constitute threads. In at least one embodiment, a set of threads executing across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program on different input data. In at least one embodiment, each thread in a thread group can be assigned to a different processing engine in graphics multiprocessor 1734. In at least one embodiment, a thread group may include fewer threads than the number of processing engines in graphics multiprocessor 1734. In at least one embodiment, if a thread group includes fewer threads than the number of processing engines, one or more processing engines may be idle during the cycle in which the thread group is processed. In at least one embodiment, a thread group may also include more threads than the number of processing engines in graphics multiprocessor 1734. In at least one embodiment, if a thread group includes more threads than the number of processing engines in graphics multiprocessor 1734, processing may be performed over consecutive clock cycles. In at least one embodiment, multiple thread groups may execute simultaneously on graphics multiprocessor 1734.

[0180] In at least one embodiment, the graphics multiprocessor 1734 includes internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 1734 can forgo the internal cache and use cache memory (e.g., L1 cache 1748) within the processing cluster 1714. In at least one embodiment, each graphics multiprocessor 1734 can also access an L2 cache within a partition unit (e.g., partition units 1720A-1720N in FIG. 17 ), which may be shared among all processing clusters 1714 and used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 1734 can also access off-chip global memory, which may include one or more of the local parallel processor memories and / or system memories. In at least one embodiment, any memory external to the parallel processing units 1702 may be used as global memory. In at least one embodiment, processing cluster 1714 may include multiple instances of graphics multiprocessor 1734 and share common instructions and data, which may be stored in L1 cache 1748.

[0181] In at least one embodiment, each processing cluster 1714 may include a memory management unit (“MMU”) 1745 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of MMU 1745 may be in memory interface 1718 of FIG. 17A . In at least one embodiment, MMU 1745 includes a set of page table entries (PTEs) used to map virtual addresses to physical addresses of tiles and optionally cache line indexes. In at least one embodiment, MMU 1745 may include an address translation lookaside buffer (TLB) or cache, which may be in graphics multiprocessor 1734 or an L1 cache, or processing cluster 1714. In at least one embodiment, physical addresses are processed to locally distribute surface data accesses, allowing efficient interleaving of requests across partition units. In at least one embodiment, the cache line index may be used to determine whether a cache line request is a hit or a miss.

[0182] In at least one embodiment, processing cluster 1714 may be configured such that each graphics multiprocessor 1734 is coupled to a texture unit 1736 to perform texture mapping operations, such as determining texture sample locations, reading texture data, and filtering the texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within graphics multiprocessor 1734 and, as needed, fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 1734 outputs processed tasks to data crossbar 1740 to provide the processed tasks to another processing cluster 1714 for further processing, or stores the processed tasks in an L2 cache, local parallel processor memory, or system memory via memory crossbar 1716. In at least one embodiment, a pre-ROP 1742 (pre-raster operation unit) is configured to receive data from the graphics multiprocessor 1734 and direct the data to the ROP units, which may be located within partition units (e.g., partition units 1720A-1720N in FIG. 17A ) as described herein. In at least one embodiment, the pre-ROP 1742 unit can perform color blending optimizations, organize pixel color data, and perform address translation.

[0183] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with FIGURES 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in graphics processing cluster 1714 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0184] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0185] 17D illustrates a graphics multiprocessor 1734 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 1734 couples with a pipeline manager 1732 of a processing cluster 1714. In at least one embodiment, the graphics multiprocessor 1734 has an execution pipeline including, but not limited to, an instruction cache 1752, an instruction unit 1754, an address mapping unit 1756, a register file 1758, one or more general-purpose graphics processing unit (GPGPU) cores 1762, and one or more load / store units 1766. The GPGPU cores 1762 and the load / store units 1766 are coupled to a cache memory 1772 and a shared memory 1770 via a memory and cache interconnect 1768.

[0186] In at least one embodiment, instruction cache 1752 receives a stream of instructions to execute from pipeline manager 1732. In at least one embodiment, instructions are cached in instruction cache 1752 and dispatched for execution by instruction unit 1754. In at least one embodiment, instruction unit 1754 can dispatch instructions as thread groups (e.g., warps), with each thread group assigned to a different execution unit within GPGPU core 1762. In at least one embodiment, instructions can access either local, shared, or global address spaces by specifying addresses in the unified address space. In at least one embodiment, address mapping unit 1756 can be used to translate addresses in the unified address space into individual memory addresses accessible by load / store unit 1766.

[0187] In at least one embodiment, register file 1758 provides a set of registers to the functional units of graphics multiprocessor 1734. In at least one embodiment, register file 1758 provides temporary storage for operands connected to the data paths of the functional units (e.g., GPGPU core 1762, load / store unit 1766) of graphics multiprocessor 1734. In at least one embodiment, register file 1758 is partitioned among each of the functional units, such that each functional unit is allocated a dedicated portion of register file 1758. In one embodiment, register file 1758 is partitioned among the different warps being executed by graphics multiprocessor 1734.

[0188] In at least one embodiment, GPGPU cores 1762 may each include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions for graphics multiprocessor 1734. GPGPU cores 1762 may have similar or different architectures. In at least one embodiment, a first portion of GPGPU core 1762 includes a single-precision FPU and an integer ALU, and a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may perform IEEE 754-2008 standard floating-point operations or may enable variable-precision floating-point operations. In at least one embodiment, graphics multiprocessor 1734 may further include one or more fixed-function or special-function units for performing specific functions, such as rectangular copy or pixel-blending operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed or special-function logic.

[0189] In at least one embodiment, GPGPU core 1762 includes SIMD logic capable of executing a single instruction on multiple data sets. In at least one embodiment, GPGPU core 1762 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for the GPGPU core may be generated at compile time by a shader compiler or may be generated automatically when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model may execute via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations may execute in parallel via a single SIMD8 logical unit.

[0190] In at least one embodiment, memory and cache interconnect 1768 is an interconnect network connecting each functional unit of graphics multiprocessor 1734 to register file 1758 and shared memory 1770. In at least one embodiment, memory and cache interconnect 1768 is a crossbar interconnect that allows load / store unit 1766 to perform load and store operations between shared memory 1770 and register file 1758. In at least one embodiment, register file 1758 can operate at the same frequency as GPGPU cores 1762, and therefore data transfers between GPGPU cores 1762 and register file 1758 have very low latency. In at least one embodiment, shared memory 1770 can be used to enable communication between threads executing in functional units within graphics multiprocessor 1734. In at least one embodiment, cache memory 1772 can be used, for example, as a data cache to cache texture data communicated between functional units and texture unit 1736. In at least one embodiment, shared memory 1770 can also be used as a program-managed cache. In at least one embodiment, threads running on GPGPU cores 1762 can programmatically store data in the shared memory in addition to automatically caching data stored in cache memory 1772.

[0191] In at least one embodiment, a parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated into the same package or chip as the core or may be communicatively coupled to the core via an internal (i.e., internal to the package or chip) processor bus / interconnect. In at least one embodiment, regardless of how the GPU is connected, the processor core may allocate work to such GPU in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0192] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with FIGURES 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in graphics multiprocessor 1734 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0193] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0194] FIG. 18 illustrates a multi-GPU computing system 1800, according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 1800 may include a processor 1802 coupled to multiple general-purpose graphics processing units (GPGPUs) 1806A-D via a host interface switch 1804. In at least one embodiment, the host interface switch 1804 is a PCI Express switch device that couples the processor 1802 to a PCI Express bus, via which the processor 1802 can communicate with the GPGPUs 1806A-D. The GPGPUs 1806A-D may be interconnected via a set of high-speed point-to-point GPU-to-GPU links 1816. In at least one embodiment, the GPU-to-GPU links 1816 are connected to each of the GPGPUs 1806A-D via dedicated GPU links. In at least one embodiment, the P2P GPU link 1816 allows direct communication between each of the GPGPUs 1806A-D without requiring communication through the host interface bus 1804 to which the processor 1802 is connected. In at least one embodiment, when there is GPU-to-GPU traffic directed to the P2P GPU link 1816, the host interface bus 1804 remains available for access to system memory or for communication with other instances of the multi-GPU computing system 1800, for example, via one or more network devices. In at least one embodiment, the GPGPUs 1806A-D are connected to the processor 1802 through the host interface switch 1804, and in at least one embodiment, the processor 1802 includes direct support for the P2P GPU link 1816 and can be directly connected to the GPGPUs 1806A-D.

[0195] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with Figures 6A and 6B. In at least one embodiment, inference and / or training logic 615 may be used in multi-GPU computing system 1800 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0196] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0197] 19 is a block diagram of a graphics processor 1900 according to at least one embodiment. In at least one embodiment, the graphics processor 1900 includes a ring interconnect 1902, a pipeline front end 1904, a media engine 1937, and graphics cores 1980A-1980N. In at least one embodiment, the ring interconnect 1902 couples the graphics processor 1900 to other processing units, including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, the graphics processor 1900 is one of multiple processors integrated within a multi-core processing system.

[0198] In at least one embodiment, graphics processor 1900 receives batches of commands via ring interconnect 1902. In at least one embodiment, the incoming commands are interpreted by command streamer 1903 of pipeline front end 1904. In at least one embodiment, graphics processor 1900 includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 1980A-1980N. In at least one embodiment, for 3D geometry processing commands, command streamer 1903 supplies the commands to geometry pipeline 1936. In at least one embodiment, for at least some media processing commands, command streamer 1903 supplies the commands to video front end 1934, which is coupled to media engine 1937. In at least one embodiment, the media engine 1937 includes a Video Quality Engine (VQE) 1930 for video and image post-processing and a Multi-Format Encode / Decode (MFX) 1933 engine that provides hardware-accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 1936 and the media engine 1937 each spawn execution threads for thread execution resources provided by at least one graphics core 1980A.

[0199] In at least one embodiment, graphics processor 1900 includes scalable thread execution resources characterized by modular cores 1980A-1980N (sometimes referred to as core slices), each of which has multiple sub-cores 1950A-1950N, 1960A-1960N (sometimes referred to as core sub-slices). In at least one embodiment, graphics processor 1900 can have any number of graphics cores 1980A-1980N. In at least one embodiment, graphics processor 1900 includes graphics core 1980A having at least a first sub-core 1950A and a second sub-core 1960A. In at least one embodiment, graphics processor 1900 is a low-power processor having a single sub-core (e.g., 1950A). In at least one embodiment, graphics processor 1900 includes multiple graphics cores 1980A-1980N, each including a set of first sub-cores 1950A-1950N and a set of second sub-cores 1960A-1960N. In at least one embodiment, each of first sub-cores 1950A-1950N includes at least a first set of execution units 1952A-1952N and media / texture samplers 1954A-1954N. In at least one embodiment, each of second sub-cores 1960A-1960N includes at least a second set of execution units 1962A-1962N and samplers 1964A-1964N. In at least one embodiment, each sub-core 1950A-1950N, 1960A-1960N shares a set of shared resources 1970A-1970N. In at least one embodiment, the shared resources include shared cache memory and pixel operating logic.

[0200] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with FIGURES 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in graphics processor 1900 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0201] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0202] FIG. 20 is a block diagram illustrating the micro-architecture of a processor 2000 that may include logic circuits for executing instructions, according to at least one embodiment. In at least one embodiment, the processor 2000 may execute instructions, including x86 instructions, AMR instructions, special instructions for application-specific integrated circuits (ASICs), and the like. In at least one embodiment, the processor 2000 may include registers for storing packed data, such as 64-bit wide MMX™ registers in microprocessors enabled with MMX technology by Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers, available in both integer and floating-point formats, may operate on packed data elements with Single Instruction Multiple Data (“SIMD”) and Streaming SIMD Extensions (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers associated with SSE2, SSE3, SSE4, AVX, or higher (collectively referred to as “SSEx”) technology may hold such packed data operands. In at least one embodiment, the processor 2000 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.

[0203] In at least one embodiment, processor 2000 includes an in-order front end ("front end") 2001 that fetches instructions to be executed and prepares the instructions for later use in the processor pipeline. In at least one embodiment, front end 2001 may include several units. In at least one embodiment, an instruction prefetcher 2026 fetches instructions from memory and provides the instructions to an instruction decoder 2028, which decodes or interprets the instructions. For example, in at least one embodiment, instruction decoder 2028 decodes received instructions into one or more operations, called "microinstructions" or "micro-operations" (also called "micro-ops" or "uops"), that the machine can execute. In at least one embodiment, instruction decoder 2028 parses instructions into opcodes and corresponding data and control fields that may be used by the micro-architecture to perform operations in accordance with at least one embodiment. In at least one embodiment, trace cache 2030 may assemble the decoded uops into program-order sequences, or traces, in uop queue 2034 for execution. In at least one embodiment, when trace cache 2030 encounters a complex instruction, microcode ROM 2032 provides the uops necessary to complete the operation.

[0204] In at least one embodiment, some instructions can be converted into a single micro-op, while other instructions require several micro-ops to complete the entire operation. In at least one embodiment, if an instruction requires more than four micro-ops to complete, the instruction decoder 2028 may access the microcode ROM 2032 to execute the instruction. In at least one embodiment, the instruction may be decoded into a smaller number of micro-ops for processing in the instruction decoder 2028. In at least one embodiment, if an operation requires a large number of micro-ops to complete, the instruction may be stored in the microcode ROM 2032. In at least one embodiment, the trace cache 2030 references an entry point programmable logic array (“PLA”) to determine the correct microinstruction pointer to read the microcode sequence from to complete one or more instructions from the microcode ROM 2032, in accordance with at least one embodiment. In at least one embodiment, after the microcode ROM 2032 has finished sequencing micro-ops for an instruction, the machine front end 2001 may resume fetching micro-ops from the trace cache 2030.

[0205] In at least one embodiment, an out-of-order execution engine ("out-of-order engine") 2003 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth and reorder the flow of instructions to optimize performance as instructions are scheduled for execution down the pipeline. In at least one embodiment, the out-of-order execution engine 2003 includes, without limitation, an allocator / register renamer 2040, a memory uop queue 2042, an integer / floating point uop queue 2044, a memory scheduler 2046, a fast scheduler 2002, a slow / general purpose floating point scheduler ("slow / general purpose FP scheduler") 2004, and a simple floating point scheduler ("simple FP scheduler") 2006. In at least one embodiment, the fast scheduler 2002, the slow / general purpose floating point scheduler 2004, and the simple floating point scheduler 2006 are also collectively referred to herein as "uop schedulers 2002, 2004, 2006." In at least one embodiment, the allocator / register renamer 2040 allocates the machine buffers and resources required by each uop to execute. In at least one embodiment, the allocator / register renamer 2040 renames the logical registers upon entry into the register file. In at least one embodiment, allocator / register renamer 2040 also allocates each uop's entry to one of two uop queues: memory uop queue 2042 for memory operations and integer / floating point uop queue 2044 for non-memory operations, before memory scheduler 2046 and uop schedulers 2002, 2004, 2006. In at least one embodiment, uop schedulers 2002, 2004, 2006 determine when uops are ready to execute based on the readiness of the sources of their dependent input register operands and the availability of the execution resources required by the uop to complete their operations.In at least one embodiment, the fast scheduler 2002 may schedule every half of the main clock cycle, and the slow / general purpose floating point scheduler 2004 and simple floating point scheduler 2006 may schedule once per main processor clock cycle. In at least one embodiment, the uop schedulers 2002, 2004, 2006 arbitrate for dispatch ports to schedule uops for execution.

[0206] In at least one embodiment, execution block 2011 includes, without limitation, integer register file / bypass network 2008, floating point register file / bypass network (“FP register file / bypass network”) 2010, address generation units (“AGUs”) 2012 and 2014, fast arithmetic logic units (ALUs) (“fast ALUs”) 2016 and 2018, slower arithmetic logic unit (“slower ALU”) 2020, floating point ALU (“FP”) 2022, and floating point move unit (“FP move”) 2024. In at least one embodiment, integer register file / bypass network 2008 and floating point register file / bypass network 2010 are also referred to herein as “register files 2008, 2010.” In at least one embodiment, AGUs 2012 and 2014, fast ALUs 2016 and 2018, slow ALU 2020, floating-point ALU 2022, and floating-point move unit 2024 are also referred to herein as "execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024." In at least one embodiment, execution block b11 may include any number and type of register files (including zero), bypass networks, address generation units, and execution units, in any combination, without limitation.

[0207] In at least one embodiment, register files 2008, 2010 may be located between uop schedulers 2002, 2004, 2006 and execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024. In at least one embodiment, integer register file / bypass network 2008 performs integer operations. In at least one embodiment, floating point register file / bypass network 2010 performs floating point operations. In at least one embodiment, each of register files 2008, 2010 may include, without limitation, a bypass network that may bypass or forward recently completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, register files 2008, 2010 may communicate data with each other. In at least one embodiment, integer register file / bypass network 2008 may include, without limitation, two separate register files: one register file for lower 32-bit data and a second register file for higher 32-bit data. In at least one embodiment, floating-point instructions typically have operands that are 64 to 128 bits wide, and therefore floating-point register file / bypass network 2010 may include, without limitation, 128-bit wide entries.

[0208] In at least one embodiment, execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024 may execute instructions. In at least one embodiment, register files 2008 and 2010 store integer and floating-point data operand values ​​required by microinstructions to execute. In at least one embodiment, processor 2000 may include any number and combination of execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024, without limitation. In at least one embodiment, floating-point ALU 2022 and floating-point move unit 2024 may execute floating-point, MMX, SIMD, AVX, and SEE, or other operations, including special machine learning instructions. In at least one embodiment, the floating-point ALU 2022 may include, without limitation, a 64-bit floating-point divider to perform division, square root, and remaining micro-ops. In at least one embodiment, instructions involving floating-point values ​​may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to the fast ALUs 2016, 2018. In at least one embodiment, the fast ALUs 2016, 2018 may perform high-speed operations with an effective latency of half a clock cycle. In at least one embodiment, the slow ALU 2020 may include, without limitation, integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branching, with most complex integer operations proceeding to the slow ALU 2020. In at least one embodiment, memory load / store operations may be performed by the ALUs 2012, 2014. In at least one embodiment, fast ALU 2016, fast ALU 2018, and slow ALU 2020 may perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 2016, fast ALU 2018, and slow ALU 2020 may be implemented to support various data bit sizes, including 16, 32, 128, 256, etc. In at least one embodiment, floating-point ALU 2022 and floating-point move unit 2024 may be implemented to support wide operands having various bit widths.In at least one embodiment, the floating-point ALU 2022 and floating-point move unit 2024 may operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.

[0209] In at least one embodiment, the uop schedulers 2002, 2004, 2006 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, because uops may be speculatively scheduled and executed in the processor 2000, the processor 2000 may also include logic to handle memory misses. In at least one embodiment, if a data load misses in the data cache, there may be dependent operations in progress in the pipeline past the scheduler that have temporarily incorrect data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use the incorrect data. In at least one embodiment, the dependent operations may need to be replayed, and the independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of a processor may also be designed to capture instruction sequences for text string comparison operations.

[0210] In at least one embodiment, the term "register" may refer to an on-board processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be available externally to the processor (from a programmer's perspective). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuitry within the processor using any number of different techniques, such as dedicated physical registers, dynamically allocated physical registers using register renaming, or a combination of dedicated and dynamically allocated physical registers. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packed data.

[0211] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B . In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the execution block 2011 and other memory or registers, shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs shown in the execution block 2011. Furthermore, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the execution block 2011 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0212] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0213] 21 illustrates a deep learning application processor 2100 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2100 uses instructions that, when executed by the deep learning application processor 2100, cause the deep learning application processor 2100 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 2100 is an application specific integrated circuit (ASIC). In at least one embodiment, the application processor 2100 performs a matrix multiplication operation, both "hardwired" in hardware, as a result of executing one or more instructions or both. In at least one embodiment, deep learning application processor 2100 includes, without limitation, processing clusters 2110(1)-2110(12), inter-chip links (“ICLs”) 2120(1)-2120(12), inter-chip controllers (“ICCs”) 2130(1)-2130(2), memory controllers (“Mem Ctrlr”) 2142(1)-2142(4), high-bandwidth memory physical layers (“HBM PHYs”) 2144(1)-2144(4), management-controller central processing unit (“management-controller CPU”) 2150, peripheral component interconnect express controller and direct memory access block (“PCIe controller and DMA”) 2170, and a 16-lane peripheral component interconnect express port (“PCI Express x16”) 2180.

[0214] In at least one embodiment, the processing clusters 2110 may perform deep learning operations, including inference or prediction operations, based on weight parameters calculated using one or more training techniques, including the techniques described herein. In at least one embodiment, each processing cluster 2110 may include any number and types of processors, without limitation. In at least one embodiment, the deep learning application processor 2100 may include any number and types of processing clusters 2100. In at least one embodiment, the inter-chip link 2120 is bidirectional. In at least one embodiment, the inter-chip link 2120 and the inter-chip controller 2130 enable multiple deep learning application processors 2100 to exchange information, including activation information resulting from running one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2100 may include any number and types (including zero) of ICLs 2120 and ICCs 2130.

[0215] In at least one embodiment, the HBM2 2140 provides a total of 32 Gigabytes (GB) of memory. Each HBM2 2140(i) is associated with both a memory controller 2142(i) and an HBM PHY 2144(i). In at least one embodiment, any number of HBM2 2140s may provide any type and total amount of high-bandwidth memory and may be associated with any number and types of memory controllers 2142 and HBM PHYs 2144 (including zero). In at least one embodiment, the SPI, I2C, GPIO 2160, PCIe controller and DMA 2170, and / or PCIe 2180 may be replaced with any number and types of blocks enabling any number and types of communication standards in any technically feasible manner.

[0216] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B. In at least one embodiment, the deep learning application processor 2100 is used to train a machine learning model, such as a neural network, to predict or infer information provided to the deep learning application processor 2100. In at least one embodiment, the deep learning application processor 2100 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 2100. In at least one embodiment, the processor 2100 may be used to perform one or more neural network use cases described herein.

[0217] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0218] FIG. 22 is a block diagram of a neuromorphic processor 2200, according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2200 receives one or more inputs from sources external to the neuromorphic processor 2200. In at least one embodiment, these inputs may be sent to one or more neurons 2202 within the neuromorphic processor 2200. In at least one embodiment, the neurons 2202 and their components may be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2200 may include, without limitation, thousands or millions of instances of neurons 2202, although any suitable number of neurons 2202 may be used. In at least one embodiment, each instance of a neuron 2202 may include a neuron input 2204 and a neuron output 2206. In at least one embodiment, a neuron 2202 may generate an output, which may be sent to an input of another instance of a neuron 2202. For example, in at least one embodiment, neuron input 2204 and neuron output 2206 may be interconnected via synapse 2208 .

[0219] In at least one embodiment, neurons 2202 and synapses 2208 may be interconnected such that neuromorphic processor 2200 operates to process or analyze information received by neuromorphic processor 2200. In at least one embodiment, neuron 2202 may send an output pulse (or "fire" or "spike") when input received via neuron input 2204 exceeds a threshold. In at least one embodiment, neuron 2202 may sum or integrate signals received at neuron input 2204. For example, in at least one embodiment, neuron 2202 may be implemented as a leaky integrate-and-fire neuron, where if the sum (referred to as the "membrane potential") exceeds a threshold, neuron 2202 may generate an output (or "fire") using a transfer function such as a sigmoid function or a threshold function. In at least one embodiment, a leaky integrate-and-fire neuron may sum signals received at neuron input 2204 into a membrane potential and may apply a decay factor (or leakage) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-fire neuron may fire if multiple input signals are received at neuron input 2204 quickly enough to exceed a threshold (i.e., before the membrane potential decays too little to cause firing). In at least one embodiment, neuron 2202 may be implemented using circuitry or logic that receives inputs, integrates the inputs into a membrane potential, and decays the membrane potential. In at least one embodiment, the inputs may be averaged, or any other suitable transfer function may be used. Further, in at least one embodiment, neuron 2202 may include, without limitation, comparator circuitry or logic that generates an output spike at neuron 2206 when the result of applying the transfer function to neuron 2204 exceeds a threshold. In at least one embodiment, neuron 2202 may ignore previously received input information when firing, for example, by resetting the membrane potential to 0 or another suitable default value.In at least one embodiment, once the membrane potential is reset to zero, neuron 2202 may resume normal operation after a suitable period (or refractory period).

[0220] In at least one embodiment, neurons 2202 may be interconnected through synapses 2208. In at least one embodiment, synapses 2208 may operate to transmit a signal from an output of a first neuron 2202 to an input of a second neuron 2202. In at least one embodiment, neurons 2202 may transmit information through two or more instances of synapses 2208. In at least one embodiment, one or more instances of neuron outputs 2206 may be connected to instances of neuron inputs 2204 of the same neuron 2202 through instances of synapses 2208. In at least one embodiment, an instance of neuron 2202 that generates an output to be transmitted through an instance of synapse 2208 may be referred to as a “pre-synaptic neuron” with respect to that instance of synapse 2208. In at least one embodiment, an instance of neuron 2202 that receives an input to be transmitted through an instance of synapse 2208 may be referred to as a “post-synaptic neuron” with respect to that instance of synapse 2208. In at least one embodiment, an instance of neuron 2202 may receive input from one or more instances of synapse 2208 and may send output through one or more instances of synapse 2208, so that a single instance of neuron 2202 may therefore be both a "pre-synaptic neuron" and a "post-synaptic neuron" with respect to various instances of synapse 2208.

[0221] In at least one embodiment, neurons 2202 may be organized into one or more layers. Each instance of a neuron 2202 may have one neuron output 2206 that can fan out to one or more neuron inputs 2204 through one or more synapses 2208. In at least one embodiment, a neuron output 2206 of a neuron 2202 in a first layer 2210 may be connected to a neuron input 2204 of a neuron 2202 in a second layer 2212. In at least one embodiment, layer 2210 may be referred to as a "feed-forward" layer. In at least one embodiment, each instance of a neuron 2202 in an instance of a first layer 2210 may fan out to each instance of a neuron 2202 in a second layer 2212. In at least one embodiment, first layer 2210 may be referred to as a "fully connected feed-forward layer." In at least one embodiment, each instance of neuron 2202 in the second layer 2212 may fan out to fewer than all instances of neuron 2202 in the third layer 2214. In at least one embodiment, the second layer 2212 may be referred to as a "sparsely connected feed-forward layer." In at least one embodiment, the neurons 2202 in the second layer 2212 may fan out to neurons 2202 in multiple other layers, including neurons 2202 in the same second layer 2212. In at least one embodiment, the second layer 2212 may be referred to as a "recurrent layer." In at least one embodiment, the neuromorphic processor 2200 may include any suitable combination of recurrent and feed-forward layers, including, without limitation, both sparsely connected feed-forward layers and fully connected feed-forward layers.

[0222] In at least one embodiment, neuromorphic processor 2200 may include, without limitation, a reconfigurable interconnect architecture or dedicated hardwired interconnects for connecting synapses 2208 to neurons 2202. In at least one embodiment, neuromorphic processor 2200 may include, without limitation, circuitry or logic that allows synapses to be allocated to different neurons 2202 as needed based on the neural network topology and the fan-in / fan-out of the neurons. For example, in at least one embodiment, synapses 2208 may be connected to neurons 2202 using an interconnect fabric, such as a network-on-chip, or using dedicated connections. In at least one embodiment, synaptic interconnects and their components may be implemented using circuitry or logic.

[0223] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0224] 23 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 2300 includes one or more processors 2302 and one or more graphics processors 2308 and may be a single-processor desktop system, a multi-processor workstation system, or a server system having multiple processors 2302 or processor cores 2307. In at least one embodiment, system 2300 is a processing platform integrated into a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.

[0225] In at least one embodiment, system 2300 may include or be incorporated into a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a handheld game console, or an online game console. In at least one embodiment, system 2300 is a mobile phone, a smart phone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing system 2300 may also include, be coupled to, or be integrated into a wearable device, such as a smart watch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 2300 is a television or set-top box device having one or more processors 2302 and a graphical interface generated by one or more graphics processors 2308.

[0226] In at least one embodiment, the one or more processors 2302 each include one or more processor cores 2307 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 2307 is configured to process a particular instruction set 2309. In at least one embodiment, the instruction set 2309 may facilitate computing via complex instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction word (VLIW). In at least one embodiment, each processor core 2307 may process a different instruction set 2309, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, the processor cores 2307 may also include other processing devices, such as a digital signal processor (DSP).

[0227] In at least one embodiment, processor 2302 includes cache memory 2304. In at least one embodiment, processor 2302 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory is shared among various components of processor 2302. In at least one embodiment, processor 2302 also uses an external cache (e.g., a level 3 (L3) cache or last level cache (LLC)) (not shown), which may be shared among processor cores 2307 using known cache coherence techniques. In at least one embodiment, processor 2302 further includes register file 2306, which may include different types of registers (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, register file 2306 may include general-purpose registers or other registers.

[0228] In at least one embodiment, the one or more processors 2302 are coupled to one or more interface buses 2310 to transmit communication signals, such as address, data, or control signals, between the processors 2302 and other components in the system 2300. In at least one embodiment, the interface bus 2310 may be a processor bus, such as, in one embodiment, a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface 2310 is not limited to a DMI bus, but may include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 2302 includes an integrated memory controller 2316 and a platform controller hub 2330. In at least one embodiment, the memory controller 2316 facilitates communication between memory devices and other components of the system 2300, while the platform controller hub (PCH) 2330 provides connectivity to I / O devices via a local I / O bus.

[0229] In at least one embodiment, memory device 2320 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or any other memory device with performance suitable for serving as process memory. In at least one embodiment, memory device 2320 may operate as system memory for system 2300, storing data 2322 and instructions 2321 for use by one or more processors 2302 when executing applications or processes. In at least one embodiment, memory controller 2316 also couples to an optional external graphics processor 2312, which may communicate with one or more graphics processors 2308 within processor 2302 to perform graphics and media operations. In at least one embodiment, a display device 2311 may be connected to processor 2302. In at least one embodiment, display device 2311 may include one or more of an internal display device, such as a mobile electronic device or laptop device, or an external display device attached via a display interface (e.g., display port, etc.). In at least one embodiment, display device 2311 may include a head-mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) or augmented reality (AR) applications.

[0230] In at least one embodiment, platform controller hub 2330 allows peripheral devices to connect to memory device 2320 and processor 2302 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 2346, a network controller 2334, a firmware interface 2328, a wireless transceiver 2326, a touch sensor 2325, and a data storage device 2324 (e.g., hard disk drive, flash memory, etc.). In at least one embodiment, data storage device 2324 can be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, touch sensor 2325 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, wireless transceiver 2326 may be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, firmware interface 2328 enables communication with system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, network controller 2334 may enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) couples to interface bus 2310. In at least one embodiment, audio controller 2346 is a multi-channel high-definition audio controller. In at least one embodiment, system 2300 includes an optional legacy I / O controller 2340 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.In at least one embodiment, platform controller hub 2330 can also connect to one or more universal serial bus (USB) controller 2342 connected input devices, such as a keyboard and mouse 2343 combination, a camera 2344, or other USB input devices.

[0231] In at least one embodiment, instances of memory controller 2316 and platform controller hub 2330 may be integrated into a separate external graphics processor, such as external graphics processor 2312. In at least one embodiment, platform controller hub 2330 and / or memory controller 2316 may be external to one or more processors 2302. For example, in at least one embodiment, system 2300 may include external memory controller 2316 and platform controller hub 2330, which may be configured as a memory controller hub and a peripheral controller hub within a system chipset that communicates with processor 2302.

[0232] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B . In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into graphics processor 2300. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in graphics processor 2312. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that shown in FIG. 6A or FIG. 6B . In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of graphics processor 2300 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0233] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0234] 24 is a block diagram of a processor 2400 having one or more processor cores 2402A-2402N, an integrated memory controller 2414, and an integrated graphics processor 2408, according to at least one embodiment. In at least one embodiment, the processor 2400 may include a number of additional cores, including additional core 2402N, represented by a dashed box. In at least one embodiment, each of the processor cores 2402A-2402N includes one or more internal cache units 2404A-2404N. In at least one embodiment, each processor core also has access to one or more shared cache units 2406.

[0235] In at least one embodiment, internal cache units 2404A-2404N and shared cache unit 2406 represent a cache memory hierarchy within processor 2400. In at least one embodiment, cache memory units 2404A-2404N may include at least one level of instruction and data cache within each processor core, as well as one or more levels of shared mid-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherency between the various cache units 2406 and 2404A-2404N.

[0236] In at least one embodiment, processor 2400 may also include a set of one or more bus controller units 2416 and a system agent core 2410. In at least one embodiment, one or more bus controller units 2416 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, system agent core 2410 provides management functions for various processor components. In at least one embodiment, system agent core 2410 includes one or more integrated memory controllers 2414 for managing access to various external memory devices (not shown).

[0237] In at least one embodiment, one or more of the processor cores 2402A-2402N include support for simultaneous multithreading. In at least one embodiment, the system agent core 2410 includes components for coordinating and operating the cores 2402A-2402N during multithreaded processing. In at least one embodiment, the system agent core 2410 may further include a power control unit (PCU), which includes logic and components for coordinating the power state of one or more of the processor cores 2402A-2402N and the graphics processor 2408.

[0238] In at least one embodiment, processor 2400 further includes a graphics processor 2408 for performing graphics processing operations. In at least one embodiment, graphics processor 2408 is coupled to a shared cache unit 2406 and a system agent core 2410 that includes one or more integrated memory controllers 2414. In at least one embodiment, system agent core 2410 also includes a display controller 2411 for directing output of the graphics processor to one or more coupled displays. In at least one embodiment, display controller 2411 may also be a separate module coupled to graphics processor 2408 via at least one interconnect or may be integrated within graphics processor 2408.

[0239] In at least one embodiment, a ring-based interconnect unit 2412 is used to couple the internal components of processor 2400. In at least one embodiment, alternative interconnect units such as a point-to-point interconnect, a switched interconnect, or other techniques may be used. In at least one embodiment, graphics processor 2408 is coupled to ring interconnect 2412 via I / O link 2413.

[0240] In at least one embodiment, I / O link 2413 represents at least one of a variety of I / O interconnects, including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 2418, such as an eDRAM module. In at least one embodiment, each of processor cores 2402A-2402N and graphics processor 2408 use embedded memory module 2418 as a shared last-level cache.

[0241] In at least one embodiment, processor cores 2402A-2402N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, processor cores 2402A-2402N are heterogeneous in terms of instruction set architecture (ISA), where one or more of processor cores 2402A-2402N execute a common instruction set, while one or more other of processor cores 2402A-2402N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, processor cores 2402A-2402N are heterogeneous in terms of microarchitecture, where one or more cores with relatively higher power consumption are combined with one or more cores with lower power consumption. In at least one embodiment, processor 2400 may be implemented on one or more chips or may be implemented as an SoC integrated circuit.

[0242] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into processor 2400. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in graphics processor 2312, graphics cores 2402A-2402N, or other components of FIG. 24. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that shown in FIG. 6A or 6B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of graphics processor 2400 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0243] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0244] FIG. 25 is a block diagram of hardware logic for a graphics processor core 2500 according to at least one embodiment described herein. In at least one embodiment, the graphics processor core 2500 is included within a graphics core array. In at least one embodiment, the graphics processor core 2500, sometimes referred to as a core slice, may be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 2500 is an example of a single graphics core slice; the graphics processors described herein may include multiple graphics core slices based on a target power and performance envelope. In at least one embodiment, each graphics core 2500 may include a fixed function block 2530 coupled to multiple sub-cores 2501A-2501F, also referred to as sub-slices, that include modular blocks of general-purpose and fixed-function logic.

[0245] In at least one embodiment, fixed function block 2530 includes a geometry / fixed function pipeline 2536 that may be shared by all sub-cores within graphics processor 2500, e.g., in lower performance and / or lower power graphics processor embodiments. In at least one embodiment, geometry / fixed function pipeline 2536 includes a 3D fixed function pipeline, a video front end unit, a thread spawner and thread dispatcher, and a unified return buffer manager that manages a unified return buffer.

[0246] In at least one embodiment, fixed function block 2530 also includes graphics SoC interface 2537, graphics microcontroller 2538, and media pipeline 2539. In at least one embodiment, fixed graphics SoC interface 2537 provides an interface between graphics core 2500 and other processor cores within the system-on-chip integrated circuit. In at least one embodiment, graphics microcontroller 2538 is a programmable sub-processor configurable to manage various functions of graphics processor 2500, including thread dispatch, scheduling, and preemption. In at least one embodiment, media pipeline 2539 includes logic to facilitate decoding, encoding, pre-processing, and / or post-processing of multimedia data, including image and video data. In at least one embodiment, media pipeline 2539 performs media operations via requests to compute logic or sampling logic within sub-cores 2501-2501F.

[0247] In at least one embodiment, SoC interface 2537 enables graphics core 2500 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as shared last-level cache memory, system RAM, and / or embedded on-chip or on-package DRAM. In at least one embodiment, SoC interface 2537 also enables communication with fixed-function devices within the SoC, such as a camera imaging pipeline, and enables and / or implements the use of global memory atomics that can be shared between graphics core 2500 and a CPU within the SoC. In at least one embodiment, SoC interface 2537 can also implement power management control for graphics core 2500 and provide an interface between the graphics core 2500 clock domain and other clock domains within the SoC. In at least one embodiment, SoC interface 2537 can receive command buffers from a command streamer and global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores in the graphics processor. In at least one embodiment, the commands and instructions can be dispatched to a media pipeline 2539 when media operations are performed, or to a geometry and fixed function pipeline (e.g., geometry and fixed function pipeline 2536, geometry and fixed function pipeline 2514) when graphics processing operations are performed.

[0248] In at least one embodiment, graphics microcontroller 2538 can be configured to perform various scheduling and management tasks for graphics core 2500. In at least one embodiment, graphics microcontroller 2538 can execute graphics and / or compute workload scheduling on various graphics parallel engines in execution unit (EU) arrays 2502A-2502F, 2504A-2504F within sub-cores 2501A-2501F. In at least one embodiment, host software running on a CPU core of an SoC including graphics core 2500 can submit a workload to one of multiple graphics processor doorbells, which invokes scheduling operations on the appropriate graphics engine. In at least one embodiment, scheduling operations include determining which workload to run next, submitting the workload to a command streamer, preempting existing workloads running on the engines, managing the progress of the workload, and notifying host software when the workload is complete. In at least one embodiment, graphics microcontroller 2538 can also facilitate low power or idle states of graphics core 2500, providing graphics core 2500 with the ability to save and restore registers within graphics core 2500 across low power state transitions, independent of the operating system and / or graphics driver software on the system.

[0249] In at least one embodiment, graphics core 2500 may have up to N modular sub-cores, more or less than the illustrated sub-cores 2501A-2501F. For each set of N sub-cores, in at least one embodiment, graphics core 2500 may also include shared function logic 2510, shared and / or cache memory 2512, geometry / fixed function pipeline 2514, and additional fixed function logic 2516 for accelerating various graphics and compute processing operations. In at least one embodiment, shared function logic 2510 may include logic units (e.g., sampler, math, and / or inter-thread communication logic) that can be shared by each of the N sub-cores in graphics core 2500. In at least one embodiment, fixed, shared, and / or cache memory 2512 may be a last-level cache for N sub-cores 2501A-2501F in graphics core 2500 and may also serve as a shared memory accessible to multiple sub-cores. In at least one embodiment, geometry / fixed function pipeline 2514 may be included in place of geometry / fixed function pipeline 2536 in fixed function block 2530 and may include the same or similar logical units.

[0250] In at least one embodiment, graphics core 2500 includes additional fixed-function logic 2516, which can include various fixed-function acceleration logic for use by graphics core 2500. In at least one embodiment, additional fixed-function logic 2516 includes an additional geometry pipeline for use with position-only shading. For position-only shading, there are at least two geometry pipelines: a full geometry pipeline in geometry / fixed-function pipelines 2516, 2536, and a cull pipeline, which is an additional geometry pipeline that may be included in additional fixed-function logic 2516. In at least one embodiment, the cull pipeline is a scaled-down version of the full geometry pipeline. In at least one embodiment, the full pipeline and the cull pipeline can run different instances of an application, each with a separate context. In at least one embodiment, position-only shading can hide long cull runs of truncated triangles, allowing shading to complete earlier in some instances. For example, in at least one embodiment, because the cull pipeline fetches and shades vertex position attributes without rasterizing and rendering pixels to the frame buffer, the cull pipeline logic in the additional fixed-function logic 2516 can execute position shaders in parallel with the main application, producing critical results overall faster than the full pipeline. In at least one embodiment, the cull pipeline can use the produced critical results to compute visibility information for all triangles, regardless of whether they are culled. In at least one embodiment, the full pipeline (which may be referred to in this instance as the replay pipeline) can consume the visibility information and shade only the visible triangles, skipping over the culled triangles, which are ultimately passed to the rasterization phase.

[0251] In at least one embodiment, additional fixed function logic 2516 may also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for embodiments involving optimization of machine learning training or inference.

[0252] In at least one embodiment, each graphics sub-core 2501A-2501F includes a set of execution resources that may be used to perform graphics operations, media operations, and compute operations in response to requests from a graphics pipeline, a media pipeline, or a shader program. In at least one embodiment, the graphics sub-cores 2501A-2501F include a plurality of EU arrays 2502A-2502F, 2504A-2504F, thread dispatch and inter-thread communication (TD / IC) logic 2503A-2503F, 3D (e.g., texture) samplers 2505A-2505F, media samplers 2506A-2506F, shader processors 2507A-2507F, and shared local memory (SLM) 2508A-2508F. Each of the EU arrays 2502A-2502F, 2504A-2504F includes multiple execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations in service of graphics, media, or compute operations, including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 2503A-2503F performs local thread dispatch and thread control operations for the execution units within a sub-core and facilitates communication between threads running on the sub-core's execution units. In at least one embodiment, the 3D samplers 2505A-2505F can read textures or other 3D graphics-related data into memory. In at least one embodiment, the 3D samplers can read texture data differently based on the configured sample state and texture format associated with a given texture. In at least one embodiment, media samplers 2506A-2506F can perform similar read operations based on the type and format associated with the media data.In at least one embodiment, each graphics sub-core 2501A-2501F can alternatively include an integrated 3D and media sampler. In at least one embodiment, threads executing on execution units within each sub-core 2501A-2501F can utilize shared local memory 2508A-2508F within each sub-core to allow threads executing within a thread group to execute using a common pool of on-chip memory.

[0253] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B. In at least one embodiment, some or all of inference and / or training logic 615 may be incorporated into graphics processor 2510. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the graphics processor 2312, graphics microcontroller 2538, geometry and fixed function pipelines 2514 and 2536, or ALUs embodied in other logic of FIG. 24. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that shown in FIG. 6A or FIG. 6B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of graphics processor 2500 for executing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0254] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0255] 26A-26B illustrate thread execution logic 2600 including an array of processing elements of a graphics processor core, according to at least one embodiment. Figure 26A illustrates at least one embodiment in which thread execution logic 2600 is used. Figure 26B illustrates exemplary internal details of an execution unit, according to at least one embodiment.

[0256] 26A , in at least one embodiment, thread execution logic 2600 includes a shader processor 2602, a thread dispatcher 2604, an instruction cache 2606, a scalable execution unit array including multiple execution units 2608A-2608N, a sampler 2610, a data cache 2612, and a data port 2614. In at least one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any of execution units 2608A, 2608B, 2608C, 2608D-2608N-1 and 2608N), for example, based on the computational requirements of a workload. In at least one embodiment, the scalable execution units are interconnected via an interconnect fabric that links to each of the execution units. In at least one embodiment, the thread execution logic 2600 includes one or more connections to memory, such as system memory or cache memory, via one or more of the instruction cache 2606, the data port 2614, the sampler 2610, and the execution units 2608A-2608N. In at least one embodiment, each execution unit (e.g., 2608A) is a standalone, programmable, general-purpose computational unit capable of executing multiple simultaneous hardware threads, processing multiple data elements in parallel per thread. In at least one embodiment, the array of execution units 2608A-2608N is scalable to include any number of individual execution units.

[0257] In at least one embodiment, the execution units 2608A-2608N are primarily used to execute shader programs. In at least one embodiment, the shader processor 2602 processes various shader programs and can dispatch execution threads associated with the shader programs via the thread dispatcher 2604. In at least one embodiment, the thread dispatcher 2604 includes logic for arbitrating thread initiation requests from the graphics and media pipelines and instantiating the requested threads on one or more of the execution units 2608A-2608N. For example, in at least one embodiment, the geometry pipeline can dispatch a vertex shader, mosaic shader, or geometry shader to thread execution logic for processing. In at least one embodiment, the thread dispatcher 2604 can also handle run-time thread spawning requests from executing shader programs.

[0258] In at least one embodiment, the execution units 2608A-2608N support an instruction set that includes native support for many standard 3D graphics shader instructions, allowing shader programs from graphics libraries (e.g., Direct3D and OpenGL) to execute with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., compute and media shaders). In at least one embodiment, each execution unit 2608A-2608N, including one or more arithmetic logic units (ALUs), can issue multiple single instruction, multiple data (SIMD) executions, allowing for multithreaded operation and an efficient execution environment despite high memory access latency. In at least one embodiment, each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. In at least one embodiment, execution is issued multiple times per clock to a pipeline capable of performing integer operations, single- and double-precision floating-point operations, SIMD branch performance, logical operations, transcendental operations, and various other operations. In at least one embodiment, while waiting for data from memory or one of the shared functions, subordinate logic within the execution units 2608A-2608N puts the waiting thread to sleep until the requested data is returned. In at least one embodiment, while the waiting thread is asleep, hardware resources may be dedicated to processing other threads. For example, in at least one embodiment, during a delay associated with a vertex shader operation, the execution unit may execute another type of shader program, including a pixel shader, a fragment shader, or a different vertex shader.

[0259] In at least one embodiment, each of execution units 2608A-2608N operates on an array of data elements. In at least one embodiment, the number of data elements is the "execution size," or the number of channels for an instruction. In at least one embodiment, an execution channel is a logical unit of execution related to data element access, masking, and flow control within an instruction. In at least one embodiment, the number of channels may be independent of the number of physical arithmetic logic units (ALUs) or floating-point units (FPUs) for a particular graphics processor. In at least one embodiment, execution units 2608A-2608N may support integer and floating-point data types.

[0260] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements may be stored in registers as packed data types, and the execution unit processes the various elements based on the data size of the elements. For example, in at least one embodiment, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four separate 64-bit packed data elements (quad-word (QW) sized data elements), eight separate 32-bit packed data elements (double-word (DW) sized data elements), sixteen separate 16-bit packed data elements (word (W) sized data elements), or thirty-two separate 8-bit data elements (byte (B) sized data elements). However, in at least one embodiment, different vector widths and register sizes are contemplated.

[0261] In at least one embodiment, one or more execution units can be combined into fused execution units 2609A-2609N with thread control logic (2607A-2607N) common to the fused EU. In at least one embodiment, multiple EUs can be fused into an EU group. In at least one embodiment, each EU in a fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in a fused EU group can vary depending on the embodiment. In at least one embodiment, various SIMD widths can be executed per EU, including, but not limited to, SIMD8, SIMD16, and SIMD32. In at least one embodiment, each fused graphics execution unit 2609A-2609N includes at least two execution units. For example, in at least one embodiment, the fused execution unit 2609A includes a first EU 2608A, a second EU 2608B, and thread control logic 2607A common to the first EU 2608A and the second EU 2608B. In at least one embodiment, thread control logic 2607A controls threads executing in fused graphics execution unit 2609A, allowing each EU in fused execution units 2609A-2609N to execute using a common instruction pointer register.

[0262] In at least one embodiment, one or more internal instruction caches (e.g., 2606) are included in the thread execution logic 2600 for caching thread instructions for the execution units. In at least one embodiment, one or more data caches (e.g., 2612) are included for caching thread data during thread execution. In at least one embodiment, a sampler 2610 is included for performing texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, the sampler 2610 includes special texture or media sampling functionality to process texture or media data during the sampling process before providing the sampled data to the execution units.

[0263] During execution, in at least one embodiment, the graphics and media pipeline sends thread start requests to the thread execution logic 2600 via thread spawning and dispatch logic. In at least one embodiment, once a group of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 2602 is invoked to further compute output information and write the results to an output surface (e.g., a color buffer, a depth buffer, a stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader calculates values ​​for various vertex attributes that are to be interpolated between the rasterized objects. In at least one embodiment, the pixel processor logic within the shader processor 2602 then executes a pixel shader program or fragment shader program with an application programming interface (API). In at least one embodiment, to execute shader programs, shader processor 2602 dispatches threads to execution units (e.g., 2608A) via thread dispatcher 2604. In at least one embodiment, shader processor 2602 accesses texture data from texture maps stored in memory using texture sampling logic in sampler 2610. In at least one embodiment, pixel color data for each geometry fragment is computed by arithmetic operations on the texture data and input geometry data, or one or more pixels are truncated so that they are not further processed.

[0264] In at least one embodiment, data port 2614 provides a memory access mechanism for thread execution logic 2600 to output processed data to memory for further processing in the graphics processor output pipeline. In at least one embodiment, data port 2614 includes or is coupled to one or more cache memories (e.g., data cache 2612) to cache data for memory accesses via the data port.

[0265] As shown in FIG. 26B , in at least one embodiment, graphics execution unit 2608 may include an instruction fetch unit 2637, a general register file array (GRF) 2624, an architectural register file array (ARF) 2626, a thread arbiter 2622, a send unit 2630, a branch unit 2632, a set of SIMD floating-point units (FPUs) 2634, and, in at least one embodiment, a set of dedicated integer SIMD ALUs 2635. In at least one embodiment, GRF 2624 and ARF 2626 include a general register file and a set of architectural register files associated with each concurrent hardware thread that may be active in graphics execution unit 2608. In at least one embodiment, per-thread architectural state is maintained in ARF 2626, and data used during thread execution is stored in GRF 2624. In at least one embodiment, the execution state of each thread, including the instruction pointer for each thread, may be kept in thread-specific registers in ARF 2626.

[0266] In at least one embodiment, the graphics execution units 2608 have an architecture that is a combination of simultaneous multi-threading (SMT) and fine-grained interleaved multi-threading (IMT). In at least one embodiment, the architecture has a modular organization that can be tuned at design time based on the target number of concurrent threads and the number of registers per execution unit, where the resources of the execution unit are divided across the logic used to execute multiple concurrent threads.

[0267] In at least one embodiment, the graphics execution unit 2608 can jointly issue multiple instructions, which may each be a different instruction. In at least one embodiment, the thread arbiter 2622 of a graphics execution unit thread 2608 can dispatch an instruction to one of the send unit 2630, the branch unit 2642, or the SIMD FPU 2634 for execution. In at least one embodiment, each execution thread can access 128 general-purpose registers in the GRF 2624, where each register can store 32 bytes accessible as a vector of SIMD8 elements of 32-bit data elements. In at least one embodiment, each execution unit thread can access 4 Kbytes in the GRF 2624, although embodiments are not limited in this manner and other embodiments may provide more or less resources. In at least one embodiment, up to seven threads can execute simultaneously, although the number of threads per execution unit can also vary depending on the embodiment. In at least one embodiment, where seven threads can access 4 Kbytes, the GRF 2624 can store a total of 28 Kbytes. In at least one embodiment, flexible addressing modes allow multiple registers to be addressed together to build wider registers or to represent strided rectangular block data structures.

[0268] In at least one embodiment, memory operations, sampler operations, and other long latency system communications are dispatched via "send" instructions executed by message passing send unit 2630. In at least one embodiment, branch instructions are dispatched to a dedicated branch unit 2632 to facilitate SIMD divergence and eventual convergence.

[0269] In at least one embodiment, the graphics execution unit(s) 2608 includes one or more SIMD floating-point units (FPUs) 2634 for performing floating-point operations. In at least one embodiment, the FPUs 2634 also support integer calculations. In at least one embodiment, the FPUs 2634 can SIMD up to M 32-bit floating-point (or integer) operations or up to 2M 16-bit integer or 16-bit floating-point operations. In at least one embodiment, at least one of the FPUs provides extended mathematical functionality to support high-throughput transcendental functions and double-precision 64-bit floating-point. In at least one embodiment, a set of 8-bit integer SIMD ALUs 2635 is also present and may be specifically optimized to perform operations related to machine learning calculations.

[0270] In at least one embodiment, an array of multiple instances of graphics execution unit 2608 may be instantiated in a graphics sub-core group (e.g., a sub-slice). In at least one embodiment, execution unit 2608 can execute instructions across multiple execution channels. In at least one embodiment, each thread executing on a graphics execution unit 2608 executes on a different channel.

[0271] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B . In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the execution logic 2600. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that shown in FIG. 6A or FIG. 6B . In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the execution logic 2600 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0272] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0273] FIG. 27 illustrates a parallel processing unit (“PPU”) 2700 according to at least one embodiment. In at least one embodiment, the PPU 2700 comprises machine-readable code that, when executed by the PPU 2700, causes the PPU 2700 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the PPU 2700 is a multi-threaded processor that is embodied in one or more integrated circuit devices and utilizes multi-threading as a latency-hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) in parallel on multiple threads. In at least one embodiment, a thread refers to a thread of execution, which is an instantiation of a set of instructions configured to be executed by the PPU 2700. In at least one embodiment, PPU 2700 is a graphics processing unit ("GPU") configured to implement a graphics rendering pipeline for processing three-dimensional ("3D") graphics data to generate two-dimensional ("2D") image data for display on a display device, such as a liquid crystal display ("LCD") device. In at least one embodiment, PPU 2700 is utilized to perform computations such as linear algebra and machine learning operations. It should be understood that FIG. 27 depicts an exemplary parallel processor for illustrative purposes only and should be construed as a non-limiting example of processor architectures contemplated within the scope of the present disclosure, and that any suitable processor may be utilized in addition to and / or to replace the same.

[0274] In at least one embodiment, one or more PPUs 2700 are configured to accelerate High Performance Computing ("HPC"), data center, and machine learning applications. In at least one embodiment, the PPUs 2700 are configured to accelerate deep learning systems and applications, including, but not limited to, the following: autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analytics, molecular simulation, drug discovery, disease diagnostics, weather forecasting, big data analytics, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations.

[0275] In at least one embodiment, PPU 2700 includes, without limitation, an input / output ("I / O") unit 2706, a front-end unit 2710, a scheduler unit 2712, a work distribution unit 2714, a hub 2716, a crossbar ("Xbar") 2720, one or more general processing clusters ("GPCs") 2718, and one or more partition units ("memory partition units") 2722. In at least one embodiment, PPU 2700 is connected to a host processor or other PPUs 2700 via one or more high-speed GPU interconnects ("GPU interconnects") 2708. In at least one embodiment, PPU 2700 is connected to a host processor or other peripheral devices via interconnect 2702. In at least one embodiment, PPU 2700 is connected to local memory comprising one or more memory devices ("memory") 2704. In at least one embodiment, memory device 2704 includes, without limitation, one or more dynamic random access memory ("DRAM") devices. In at least one embodiment, one or more DRAM devices may be configured and / or configurable as a high bandwidth memory ("HBM") subsystem with multiple DRAM dies stacked within each device.

[0276] In at least one embodiment, high-speed GPU interconnect 2708 may refer to a wire-based, multi-lane communication link used by a system to scale, including one or more PPUs 2700 in combination with one or more central processing units (“CPUs”), supporting cache coherence between the PPUs 2700 and the CPUs, and CPU mastering. In at least one embodiment, data and / or commands are transmitted by high-speed GPU interconnect 2708 through hub 2716 to and from other units of PPU 2700, such as one or more copy engines, video encoders, video decoders, power management units, and other components that may not be explicitly shown in FIG. 27 .

[0277] In at least one embodiment, I / O unit 2706 is configured to receive and send communications (e.g., commands, data) from a host processor (not shown in FIG. 27 ) via system bus 2702. In at least one embodiment, I / O unit 2706 communicates with the host processor directly via system bus 2702 or through one or more intermediate devices, such as memory bridges. In at least one embodiment, I / O unit 2706 may communicate with one or more other processors, such as one or more of PPUs 2700, via system bus 2702. In at least one embodiment, I / O unit 2706 implements a Peripheral Component Interconnect Express (“PCIe”) interface to enable communication over a PCIe bus. In at least one embodiment, I / O unit 2706 implements an interface for communicating with external devices.

[0278] In at least one embodiment, I / O unit 2706 decodes packets received via system bus 2702. In at least one embodiment, at least some of the packets represent commands configured to cause PPU 2700 to perform various operations. In at least one embodiment, I / O unit 2706 transmits the decoded commands to various other units of PPU 2700 specified by the commands. In at least one embodiment, the commands are transmitted to front end unit 2710 and / or to hub 2716 or other units of PPU 2700, such as one or more copy engines, video encoders, video decoders, or power management units (not explicitly shown in FIG. 27 ). In at least one embodiment, I / O unit 2706 is configured to route communications between various logical units of PPU 2700.

[0279] In at least one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload to the PPU 2700 for processing. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in memory accessible (e.g., writeable / readable) by both the host processor and the PPU 2700, and the host interface unit may be configured to access the buffer in the system memory connected to the system bus 2702 via memory requests sent by the I / O unit 2706 over the system bus 2702. In at least one embodiment, the host processor writes the command stream to the buffer and then sends a pointer to the start of the command stream to the PPU 2700, whereupon the front end unit 2710 receives the pointer to one or more command streams and manages the one or more command streams, reading commands from the command streams and forwarding the commands to various units of the PPU 2700.

[0280] In at least one embodiment, front end unit 2710 is coupled to a scheduler unit 2712 that configures various GPCs 2718 to process tasks defined by one or more command streams. In at least one embodiment, scheduler unit 2712 is configured to track state information related to the various tasks managed by scheduler unit 2712, where the state information may indicate which GPC 2718 a task is assigned to, whether the task is active or inactive, the priority level associated with the task, etc. In at least one embodiment, scheduler unit 2712 manages the execution of multiple tasks on one or more of GPCs 2718.

[0281] In at least one embodiment, scheduler unit 2712 is coupled to a work distribution unit 2714 configured to dispatch tasks for execution on GPCs 2718. In at least one embodiment, work distribution unit 2714 tracks the number of scheduled tasks received from scheduler unit 2712, and work distribution unit 2714 manages a pending task pool and an active task pool for each of GPCs 2718. In at least one embodiment, the pending task pool comprises a number of slots (e.g., 32 slots) containing tasks assigned to be processed by a particular GPC 2718, and the active task pool comprises a number of slots (e.g., 4 slots) for tasks being actively processed by a GPC 2718, such that when one of GPCs 2718 completes execution of a task, the task is removed from the active task pool of that GPC 2718, and one of the other tasks from the pending task pool is selected and scheduled to run on the GPC 2718. In at least one embodiment, when an active task is idle on the GPC2718, such as while waiting for a data dependency to be resolved, the active task is removed from the GPC2718 and returned to the pending task pool, while another task from the pending task pool is selected and scheduled to run on the GPC2718.

[0282] In at least one embodiment, work distribution unit 2714 communicates with one or more GPCs 2718 via X-bar 2720. In at least one embodiment, X-bar 2720 is an interconnection network coupling many of the units of PPU 2700 to other units of PPU 2700 and can be configured to couple work distribution unit 2714 to a particular GPC 2718. In at least one embodiment, one or more other units of PPU 2700 may also be connected to X-bar 2720 via hub 2716.

[0283] In at least one embodiment, tasks are managed by scheduler unit 2712 and dispatched by work distribution unit 2714 to one of GPCs 2718. GPC 2718 is configured to process the task and generate a result. In at least one embodiment, the result may be consumed by other tasks within GPC 2718, routed to a different GPC 2718 via Xbar 2720, or stored in memory 2704. In at least one embodiment, the result may be written to memory 2704 via partition unit 2722, which implements a memory interface for reading and writing data to / from memory 2704. In at least one embodiment, the result may be sent to another PPU 2704 or a CPU via high-speed GPU interconnect 2708. In at least one embodiment, PPU 2700 includes U partition units 2722, equal to, but not limited to, the number of separate individual memory devices 2704 coupled to PPU 2700. In at least one embodiment, partition unit 2722 is described in further detail below in conjunction with FIG.

[0284] In at least one embodiment, the host processor executes a driver kernel that implements an application programming interface (API) that allows one or more applications running on the host processor to schedule operations for execution on the PPU 2700. In at least one embodiment, multiple compute applications are executed simultaneously by the PPU 2700, which provides isolation, quality of service (QoS), and independent address spaces for the multiple compute applications. In at least one embodiment, an application generates instructions (e.g., in the form of API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 2700, and the driver kernel outputs the tasks to one or more streams that are processed by the PPU 2700. In at least one embodiment, each task comprises one or more groups of related threads, which may be referred to as warps. In at least one embodiment, a warp comprises multiple related threads (e.g., 32 threads) that can execute in parallel. In at least one embodiment, cooperating threads may refer to multiple threads that contain instructions for performing tasks and exchange data via shared memory. In at least one embodiment, threads and cooperating threads are described in further detail in accordance with at least one embodiment in conjunction with FIG. 29.

[0285] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the PPU 2700. In at least one embodiment, the PPU 2700 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the PPU 2700. In at least one embodiment, the PPU 2700 may be used to perform one or more neural network use cases described herein.

[0286] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0287] FIG. 28 illustrates a general purpose processing cluster (“GPC”) 2800 according to at least one embodiment. In at least one embodiment, the GPC 2800 is the GPC 2718 of FIG. 27. In at least one embodiment, each GPC 2800 includes several hardware units for processing tasks, including, without limitation, a pipeline manager 2802, a pre-raster operations unit (“PROP”) 2804, a raster engine 2808, a work distribution crossbar (“WDX”) 2816, a memory management unit (“MMU”) 2818, one or more data processing clusters (“DPC”) 2806, and any suitable combination of parts.

[0288] In at least one embodiment, the operation of the GPC 2800 is controlled by a pipeline manager 2802. In at least one embodiment, the pipeline manager 2802 manages the configuration of one or more DPCs 2806 to process tasks allocated to the GPC 2800. In at least one embodiment, the pipeline manager 2802 configures at least one of the one or more DPCs 2806 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, the DPC 2806 is configured to execute vertex shader programs on a programmable streaming multi-processor (“SM”) 2814. In at least one embodiment, pipeline manager 2802 is configured to route packets received from the work distribution unit to the appropriate logical unit within GPC 2800, where some packets may be routed to a fixed function hardware unit of PROP 2804 and / or Raster Engine 2808, and other packets may be routed to DPC 2806 for processing by Primitive Engine 2812 or SM 2814. In at least one embodiment, pipeline manager 2802 configures at least one of DPC 2806 to implement a neural network model and / or computing pipeline.

[0289] In at least one embodiment, the PROP unit 2804 is configured to route data generated by the raster engine 2808 and the DPC 2806 to a raster operation (ROP) unit of the partition unit 2722, which is described in more detail above in conjunction with FIG. 27. In at least one embodiment, the PROP unit 2804 is configured to perform color blending optimization, organize pixel data, perform address translation, and perform other operations. In at least one embodiment, the raster engine 2808 includes, without limitation, several fixed-function hardware units configured to perform various raster operations, which in at least one embodiment include, without limitation, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile coalescing engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives the transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices, which are sent to a coarse raster engine to generate coverage information for the primitives (e.g., x,y coverage masks for tiles), and the output of the coarse raster engine is sent to a culling engine, which culls fragments associated with primitives that fail a z-test, and to a clipping engine, which clips fragments that are outside the view frustum. In at least one embodiment, fragments that pass clipping and culling are passed to a fine raster engine, which generates attributes for the pixel fragments based on the plane equations generated by the setup engine. In at least one embodiment, the output of the raster engine 2808 includes fragments that may be processed by any suitable entity, such as by a fragment shader implemented in the DPC 2806.

[0290] In at least one embodiment, each DPC 2806 included in GPC 2800 includes, without limitation, an M-Pipe Controller (“MPC”) 2810, a Primitive Engine 2812, one or more SMs 2814, and any suitable combination thereof. In at least one embodiment, MPC 2810 controls the operation of DPC 2806, routing packets received from pipeline manager 2802 to the appropriate unit within DPC 2806. In at least one embodiment, packets associated with vertices are routed to primitive engine 2812, which is configured to fetch vertex attributes associated with the vertices from memory; in contrast, packets associated with shader programs may be sent to SM 2814.

[0291] In at least one embodiment, SM2814 includes, without limitation, a programmable streaming processor configured to process tasks represented by several threads. In at least one embodiment, SM2814 is multithreaded and configured to simultaneously execute multiple threads (e.g., 32 threads) from a particular group of threads and implements a single instruction, multiple data (SIMD) architecture, where each thread in a group of threads (warp) is configured to process different data sets based on the same set of instructions. In at least one embodiment, all threads in a thread group execute the same instructions. In at least one embodiment, SM2814 implements a single instruction, multiple thread (SIMT) architecture, where each thread in a thread group is configured to process different data sets based on the same set of instructions, but individual threads within a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained per warp to enable concurrent processing between warps and serial execution within a warp when threads within a warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, allowing equal concurrency across all threads, within a warp, and across warps. In at least one embodiment, execution state is maintained for each individual thread, and threads executing the same instructions may converge and execute in parallel for more efficiency. At least one embodiment of SM2814 is described in further detail below.

[0292] In at least one embodiment, MMU 2818 provides an interface between GPC 2800 and a memory partition unit (e.g., partition unit 2722 of FIG. 27), where MMU 2818 provides virtual-to-physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, MMU 2818 provides one or more translation lookaside buffers ("TLBs") for performing translations from virtual addresses to physical addresses in memory.

[0293] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B. In at least one embodiment, a deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to GPC2800. In at least one embodiment, GPC2800 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by GPC2800. In at least one embodiment, GPC2800 may be used to perform one or more neural network use cases described herein.

[0294] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0295] FIG. 29 illustrates a memory partition unit 2900 of a parallel processing unit (“PPU”) according to at least one embodiment. In at least one embodiment, the memory partition unit 2900 includes, without limitation, a raster operation (“ROP”) unit 2902, a level 2 (“L2”) cache 2904, a memory interface 2906, and any suitable combination thereof. In at least one embodiment, the memory interface 2906 is coupled to memory. In at least one embodiment, the memory interface 2906 may implement a 32, 64, 128, 1024-bit data bus, or similar implementation, for high-speed data transfer. In at least one embodiment, the PPU incorporates U memory interfaces 2906, one memory interface 2906 per pair of partition units 2900, where each pair of partition units 2900 is connected to a corresponding memory device. For example, in at least one embodiment, the PPU may be connected to up to Y memory devices, such as a high-bandwidth memory stack, or Graphics Double Data Rate, Version 5, Synchronous Dynamic Random Access Memory ("GDDR5 SDRAM").

[0296] In at least one embodiment, memory interface 2906 implements a high bandwidth memory second generation (“HBM2”) memory interface, where Y is equal to one-half of U. In at least one embodiment, the HBM2 memory stack is located in the same physical package as the PPU, achieving substantial power and area savings over conventional GDDR5 SDRAM systems. In at least one embodiment, each HBM2 stack includes, without limitation, four memory dies, where Y is equal to four, and each HBM2 stack includes eight channels, two 128-bit channels per die, and a 1024-bit data bus width. In at least one embodiment, the memory supports Single-Error Correcting Double-Error Detecting (“SECDED”) error correcting code (“ECC”) to protect data. In at least one embodiment, ECC provides higher reliability for computing applications that are susceptible to data corruption.

[0297] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partition unit 2900 supports unified memory, providing a single, unified virtual address space for the central processing unit ("CPU") and PPU memory, enabling data sharing between virtual memory systems. In at least one embodiment, the memory partition unit 2900 tracks how frequently the PPU accesses memory located on other processors and ensures that memory pages are moved to the physical memory of PPUs that access the pages more frequently. In at least one embodiment, the high-speed GPU interconnect 2708 supports address translation services, allowing the PPU to directly access the CPU's page tables, giving the PPU full access to CPU memory.

[0298] In at least one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In at least one embodiment, the copy engine can generate a page fault for an address not mapped in a page table, and the memory partition unit 2900 then responds to the page fault by mapping the address to a page table, after which the copy engine performs the transfer. In at least one embodiment, memory is pinned (e.g., page movement is disabled) for multiple copy engine operations between multiple processors, effectively reducing available memory. In at least one embodiment, if there is a hardware page fault, the address can be passed to the copy engine regardless of whether the memory page is resident, and the copy process is transparent.

[0299] According to at least one embodiment, data from memory 2704 of FIG. 27 or other system memory is fetched by memory partition unit 2900 and stored in L2 cache 2904, which is located on-chip and shared among various GPCs. In at least one embodiment, each memory partition unit 2900 includes, without limitation, at least a portion of the L2 cache associated with a corresponding memory device. In at least one embodiment, lower level caches are implemented in various units within a GPC. In at least one embodiment, each of SMs 2814 may implement a level 1 (“L1”) cache, where the L1 cache is private memory dedicated to that particular SM2814, and data from L2 cache 2904 is fetched and stored in each of the L1 caches for processing by the functional units of that SM2814. In at least one embodiment, L2 cache 2904 is coupled to memory interface 2906 and X-bar 2720.

[0300] In at least one embodiment, the ROP unit 2902 performs graphics raster operations related to pixel color, such as color compression and pixel blending. In at least one embodiment, the ROP unit 2902 performs a depth test in conjunction with the raster engine 2808 to receive depths of sample locations associated with pixel fragments from a culling engine of the raster engine 2808. In at least one embodiment, the depth is tested against the corresponding depth in a depth buffer for the sample location associated with the fragment. In at least one embodiment, if the fragment passes the depth test for the sample location, the ROP unit 2902 updates the depth buffer and sends the results of the depth test to the raster engine 2808. It will be appreciated that the number of partition units 2900 may differ from the number of GPCs, and thus, in at least one embodiment, each ROP unit 2902 may be coupled to a respective GPC. In at least one embodiment, ROP unit 2902 tracks packets received from different GPCs and determines to which one to route the results produced by ROP unit 2902 through X-bar 2720.

[0301] Figure 30 illustrates a streaming multiprocessor ("SM") 3000, according to at least one embodiment. In at least one embodiment, the SM 3000 is the SM2814 of Figure 28. In at least one embodiment, the SM 3000 includes, without limitation, an instruction cache 3002, one or more scheduler units 3004, a register file 3008, one or more processing cores ("cores") 3010, one or more special function units ("SFUs") 3012, one or more load / store units ("LSUs") 3014, an interconnect network 3016, a shared memory / level 1 ("L1") cache 3018, and / or any suitable combination thereof. In at least one embodiment, the work distribution unit dispatches tasks for execution to general purpose processing clusters (“GPCs”) of a parallel processing unit (“PPU”), with each task being allocated to a particular data processing cluster (“DPC”) within the GPC, and, if the task is associated with a shader program, the task being allocated to one of the SMs 3000. In at least one embodiment, the scheduler unit 3004 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks assigned to the SMs 3000. In at least one embodiment, the scheduler unit 3004 schedules the thread blocks for execution as warps of parallel threads, with each thread block being allocated to at least one warp. In at least one embodiment, each warp executes a thread. In at least one embodiment, scheduler unit 3004 manages multiple different thread blocks, allocates warps to the different thread blocks, and then dispatches instructions from multiple different interlocking groups to various functional units (e.g., processing cores 3010, SFUs 3012, and LSUs 3014) during each clock cycle.

[0302] In at least one embodiment, a coordination group refers to a programming model for organizing groups of communicating threads, which allows developers to express the granularity at which threads communicate, enabling richer and more efficient expression of parallel decompositions. In at least one embodiment, a coordination launch API supports synchronization between thread blocks to enable parallel algorithms to be executed. In at least one embodiment, applications of traditional programming models provide a single, simple construct for synchronizing coordinated threads: a barrier across all threads in a thread block (e.g., the syncthreads() function). However, in at least one embodiment, programmers may define thread groups at a smaller granularity than a thread block and synchronize within the defined group, enabling higher performance, design flexibility, and software reuse in the form of a collective, group-wide functional interface. In at least one embodiment, coordination groups allow programmers to explicitly define groups of threads at the granularity of a sub-block (i.e., as large as a single thread) and multi-block, and perform collective operations, such as synchronization, on the threads within the coordination group. In at least one embodiment, the programming model supports clean composition across software boundaries, allowing libraries and utility functions to safely synchronize within their local context without having to make assumptions about convergence. In at least one embodiment, the interlocking group primitive enables new patterns of interlocking parallelism, including but not limited to producer-consumer parallelism, opportunistic parallelism, and global synchronization across a grid of thread blocks.

[0303] In at least one embodiment, the dispatch unit 3006 is configured to send instructions to one or more of the functional units, and the scheduler unit 3004 includes, without limitation, two dispatch units 3006 that enable two different instructions from the same warp to be dispatched during each clock cycle. In at least one embodiment, each scheduler unit 3004 includes a single dispatch unit 3006 or additional dispatch units 3006.

[0304] In at least one embodiment, each SM3000 includes, without limitation, a register file 3008 that provides a set of registers to the functional units of the SM3000. In at least one embodiment, the register file 3008 is divided among each of the functional units such that each functional unit is allocated a dedicated portion of the register file 3008. In at least one embodiment, the register file 3008 is divided among the different warps being executed by the SM3000, and the register file 3008 provides temporary storage for operands connected to the data paths of the functional units. In at least one embodiment, each SM3000 includes, without limitation, a plurality of L processing cores 3010. In at least one embodiment, each SM3000 includes, without limitation, a large number (e.g., 128 or more) of individual processing cores 3010. In at least one embodiment, each processing core 3010 includes, without limitation, fully pipelined, single-precision, double-precision, and / or mixed-precision processing units, including, without limitation, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In at least one embodiment, processing core 3010 includes, without limitation, 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0305] The tensor cores are configured to perform matrix operations according to at least one embodiment. In at least one embodiment, one or more tensor cores are included in processing core 3010. In at least one embodiment, the tensor cores are configured to perform deep learning matrix operations, such as convolution operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4×4 matrix and performs a matrix multiply and accumulate operation D=A×B+C, where A, B, C, and D are 4×4 matrices.

[0306] In at least one embodiment, inputs A and B of a matrix multiplication are 16-bit floating-point matrices, and sum matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor cores operate on 16-bit floating-point input data with a 32-bit floating-point sum. In at least one embodiment, the 16-bit floating-point multiplication uses 64 operations, resulting in a full-precision product, which is then added using 32-bit floating-point addition with other intermediate products of a 4x4x4 matrix multiplication. Tensor cores are used, in at least one embodiment, to perform much larger two-dimensional or even higher-dimensional matrix operations that are built from these smaller elements. In at least one embodiment, an API such as the CUDA9 C++ API exposes specialized matrix load, matrix multiply-and-accumulate, and matrix store operations to efficiently use the tensor cores from CUDA-C++ programs. In at least one embodiment, at the CUDA level, the warp level interface assumes a matrix of size 16x16 across all 32 threads of a warp.

[0307] In at least one embodiment, each SM 3000 includes M SFUs 3012 that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In at least one embodiment, the SFUs 3012 include, without limitation, a tree traversal unit configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFUs 3012 include, without limitation, a texture unit configured to perform filtering operations on texture maps. In at least one embodiment, the texture unit is configured to load texture maps (e.g., 2D arrays of texels) from memory and sample texture maps to generate sampled texture values ​​for use by shader programs executed by the SM 3000. In at least one embodiment, the texture maps are stored in shared memory / level 1 cache 3018. In at least one embodiment, the texture unit performs texture operations such as filtering operations using mip maps (e.g., texture maps with different levels of detail), according to at least one embodiment. In at least one embodiment, each SM 3000 includes, without limitation, two texture units.

[0308] Each SM 3000, in at least one embodiment, includes, without limitation, N LSUs 3014 that perform load and store operations between a shared memory / L1 cache 3018 and a register file 3008. Each SM 3000, in at least one embodiment, includes, without limitation, an interconnection network 3016 that connects each of the functional units to the register file 3008 and that connects the LSUs 3014 to the register file 3008, and the shared memory / L1 cache 3018. In at least one embodiment, the interconnection network 3016 is a crossbar that may be configured to connect any of the functional units to any of the registers in the register file 3008 and to connect the LSUs 3014 to memory locations in the register file 3008 and the shared memory / L1 cache 3018.

[0309] In at least one embodiment, shared memory / L1 cache 3018 is an array of on-chip memory that, in at least one embodiment, enables data storage and communication between SM3000 and the primitive engine and between threads of SM3000. In at least one embodiment, shared memory / L1 cache 3018 has, without limitation, 128 KB of storage capacity and is in the path from SM3000 to the partition unit. In at least one embodiment, shared memory / L1 cache 3018 is, in at least one embodiment, used to cache reads and writes. In at least one embodiment, one or more of shared memory / L1 cache 3018, L2 cache, and memory are auxiliary storage.

[0310] In at least one embodiment, combining data cache and shared memory functionality into a single memory block improves performance for both types of memory access. In at least one embodiment, capacity can be used or made available as a cache by programs that do not use shared memory, such that if the shared memory is configured to use half of the capacity, texture and load / store operations can use the remaining capacity. According to at least one embodiment, integration within shared memory / L1 cache 3018 allows shared memory / L1 cache 3018 to act as a high-throughput conduit for streaming data while simultaneously providing high-bandwidth, low-latency access to frequently reused data. In at least one embodiment, when configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function graphics processing unit is bypassed, creating a much simpler programming model. In a general-purpose parallel computing configuration, the work distribution unit, in at least one embodiment, assigns and distributes thread blocks directly to DPCs. In at least one embodiment, threads within a block run the same program using unique thread IDs in computations to ensure each thread produces unique results, use SM 3000 to execute the program and perform computations, use shared memory / L1 cache 3018 to communicate between threads, and use LSU 3014 to read and write global memory via shared memory / L1 cache 3018 and memory partition unit 3014. In at least one embodiment, when configured for general-purpose parallel computing, SM 3000 writes commands that scheduler unit 3004 can use to launch new work on the DPC.

[0311] In at least one embodiment, the PPU is included in or coupled to a desktop computer, laptop computer, tablet computer, server, supercomputer, smart phone (e.g., a wireless handheld device), personal digital assistant ("PDA"), digital camera, vehicle, head-mounted display, portable electronic device, etc. In at least one embodiment, the PPU is embodied on a single semiconductor substrate. In at least one embodiment, the PPU is included in a system-on-chip ("SoC") with one or more other devices, such as additional PPUs, memory, a reduced instruction set computing ("RISC") CPU, a memory management unit ("MMU"), a digital-to-analog converter ("DAC"), etc.

[0312] In at least one embodiment, the PPU may be included in a graphics card that includes one or more memory devices. The graphics card may be configured to interface with a PCIe slot on a motherboard of a desktop computer. In at least one embodiment, the PPU may be an integrated graphics processing unit ("iGPU") included in a chipset on the motherboard.

[0313] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the SM3000. In at least one embodiment, the SM3000 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the SM3000. In at least one embodiment, the SM3000 may be used to perform one or more neural network use cases described herein.

[0314] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used in conjunction with the components of these figures to adjust one or more pixel blending weights using one or more neural networks.

[0315] At least one embodiment can be described in light of the following paragraphs.

[0316] 1. A processor, comprising: one or more circuits for adjusting one or more pixel blending weights using one or more neural networks; A processor comprising:

[0317] 2. The processor of claim 1, wherein the one or more circuits are further for receiving one or more initial blending weights from a deep neural network that processes image data at a first resolution and for upsampling the initial blending weights to a second resolution.

[0318] 3. The processor of claim 2, wherein the one or more circuits are further for providing the upsampled blending weights, along with the one or more images at the second resolution, to a refinement neural network of the one or more neural networks to adjust the upsampled pixel blending weights based at least in part on image features in the one or more images.

[0319] 4. The processor of claim 3, wherein the one or more circuits are further for downsampling image data from the second resolution before providing the image data of the one or more images to the deep neural network at the first resolution.

[0320] 5. The processor of claim 3, wherein the deep neural network is a U-Net and the refinement network is a shallow convolutional neural network including one or more skip connections between convolutional layers in a sparse CNN architecture.

[0321] 6. The processor of claim 1, wherein the one or more neural networks are further for adjusting one or more additional parameters related to at least one of historical pixel data, filter parameter data, pixel jitter data, filtered pixel data, or warped pixel data.

[0322] 7. A system comprising: one or more processors for adjusting one or more pixel blending weights using one or more neural networks; A system comprising:

[0323] 8. The system of claim 7, wherein the one or more circuits are further for receiving one or more initial blending weights from a deep neural network that processes image data at a first resolution and for upsampling the initial blending weights to a second resolution.

[0324] 9. The system of claim 8, wherein the one or more circuits are further for providing the upsampled blending weights, along with the one or more images at the second resolution, to a refinement neural network of the one or more neural networks to adjust the upsampled blending weights based at least in part on image features in the one or more images.

[0325] 10. The system of claim 9, wherein the one or more circuits are further for downsampling image data from the second resolution before providing the image data for the one or more images to the deep neural network at the first resolution.

[0326] 11. The system of claim 9, wherein the deep neural network is a U-Net and the refinement network is a shallow convolutional neural network.

[0327] 12. The system of claim 9, wherein the refinement network includes one or more skip connections between convolutional layers in a sparse CNN architecture.

[0328] 13. A method comprising: 10. A method comprising: adjusting one or more pixel blending weights using one or more neural networks.

[0329] 14. The method of claim 13, further comprising receiving one or more initial blending weights from a deep neural network that processes image data at a first resolution, and upsampling the initial blending weights to a second resolution.

[0330] 15. The method of claim 14, further comprising providing the upsampled blending weights, along with the one or more images at the second resolution, to a refinement neural network of the one or more neural networks to adjust the upsampled blending weights based at least in part on image features in the one or more images.

[0331] 16. The method of claim 15, further comprising downsampling the image data from the second resolution before providing the image data for the one or more images at the first resolution to the deep neural network.

[0332] 17. The method of claim 15, wherein the deep neural network is a U-Net and the refinement network is a shallow convolutional neural network.

[0333] 18. The method of claim 15, wherein the refinement network includes one or more skip connections between convolutional layers in a sparse CNN architecture.

[0334] 19. A machine-readable medium having stored thereon an instruction set that, when executed by one or more processors, causes the one or more processors to perform at least:

[0335] A machine-readable medium for receiving one or more initial blending weights from a deep neural network that processes image data at a first resolution and upsampling the initial blending weights to a second resolution.

[0336] 20. The instructions, when executed, further cause one or more processors to: ...

Claims

1. 1. A processor, comprising: one or more circuits for generating one or more pixel blending weights used to blend one or more high resolution representations of two or more images using one or more neural networks; the one or more circuits are further for receiving one or more pixel blending weights from a deep neural network that processes image data at a first resolution and upsampling the pixel blending weights to a second resolution; the one or more circuits are further for providing the upsampled pixel blending weights, along with one or more images at the second resolution, to a refinement neural network of the one or more neural networks to adjust the upsampled pixel blending weights based at least in part on image features in the one or more images.

2. 2. The processor of claim 1, wherein the one or more circuits are further for downsampling image data of the one or more images from the second resolution before providing the image data at the first resolution to the deep neural network.

3. 2. The processor of claim 1, wherein the deep neural network is a U-Net and the refinement neural network is a shallow convolutional neural network including one or more skip connections between convolutional layers in a sparse CNN architecture.

4. 2. The processor of claim 1, wherein the one or more neural networks are further for adjusting one or more additional parameters related to at least one of historical pixel data, filter parameter data, pixel jitter data, filtered pixel data, or warped pixel data.

5. 1. A system comprising: one or more processors for using one or more neural networks to generate one or more pixel blending weights used to blend one or more high resolution representations of the two or more images using two or more images; the one or more processors are further for receiving one or more pixel blending weights from a deep neural network that processes image data at a first resolution and upsampling the pixel blending weights to a second resolution; the one or more processors are further for providing the upsampled pixel blending weights, along with one or more images at the second resolution, to a refinement neural network of the one or more neural networks to adjust the upsampled pixel blending weights based at least in part on image features in the one or more images.

6. 6. The system of claim 5, wherein the one or more processors are further for downsampling image data of the one or more images from the second resolution before providing the image data at the first resolution to the deep neural network.

7. 6. The system of claim 5, wherein the deep neural network is a U-Net and the refinement neural network is a shallow convolutional neural network.

8. 6. The system of claim 5, wherein the refinement neural network comprises one or more skip connections between convolutional layers in a sparse CNN architecture.

9. 1. A method comprising: using one or more neural networks to generate one or more pixel blending weights used to blend one or more high resolution representations of the two or more images using two or more images; receiving one or more pixel blending weights from a deep neural network that processes image data at a first resolution; and upsampling the pixel blending weights to a second resolution; and providing the upsampled pixel blending weights, along with one or more images at the second resolution, to a refinement neural network of the one or more neural networks to adjust the upsampled pixel blending weights based at least in part on image features in the one or more images.

10. 10. The method of claim 9, further comprising downsampling the image data of the one or more images from the second resolution before providing the image data at the first resolution to the deep neural network.

11. 10. The method of claim 9, wherein the deep neural network is a U-Net and the refinement neural network is a shallow convolutional neural network.

12. 10. The method of claim 9, wherein the refinement neural network comprises one or more skip connections between convolutional layers in a sparse CNN architecture.

13. 10. A machine-readable medium having stored thereon an instruction set, the instruction set, when executed by one or more processors according to claim 1, causing the one or more processors to perform at least: receiving one or more pixel blending weights from a deep neural network that processes image data at a first resolution; and upsampling the pixel blending weights to a second resolution; The set of instructions, when executed, further causes the one or more processors to: determining one or more second colors based at least in part on one or more depth variations of the one or more pixels; The set of instructions, when executed, further causes the one or more processors to: a refinement neural network of the one or more neural networks configured to provide the upsampled pixel blending weights, along with one or more images at the second resolution, to a refinement neural network of the one or more neural networks to adjust the upsampled pixel blending weights based at least in part on image features in the one or more images.

14. The set of instructions, when executed, further causes the one or more processors to:

14. The machine-readable medium of claim 13, further configured to downsample image data of the one or more images from the second resolution before providing the image data at the first resolution to the deep neural network.

15. 14. The machine-readable medium of claim 13, wherein the deep neural network is a U-Net and the refinement neural network is a shallow convolutional neural network.

16. 14. The machine-readable medium of claim 13, wherein the refinement neural network comprises one or more skip connections between convolutional layers in a sparse CNN architecture.

17. 1. An image reconstruction system comprising: one or more processors for using one or more neural networks to generate one or more pixel blending weights used to blend one or more high resolution representations of the two or more images using two or more images; a memory for storing network parameters of said one or more neural networks; Equipped with the one or more processors are further for receiving one or more pixel blending weights from a deep neural network that processes image data at a first resolution and upsampling the pixel blending weights to a second resolution; the one or more processors are further configured to provide the upsampled pixel blending weights, along with one or more images at the second resolution, to a refinement neural network of the one or more neural networks to adjust the upsampled pixel blending weights based at least in part on image features in the one or more images.

18. 18. The image reconstruction system of claim 17, wherein the one or more processors are further for downsampling image data of the one or more images from a second resolution before providing the image data at the first resolution to the deep neural network.

19. 18. The image reconstruction system of claim 17, wherein the deep neural network is a U-Net and the refinement neural network is a shallow convolutional neural network.

20. 18. The image reconstruction system of claim 17, wherein the refinement neural network comprises one or more skip connections between convolutional layers in a sparse CNN architecture.

Citation Information

Patent Citations

  • Generation of high dynamic range visual media

    US20190096046A1