Generation of Image Blending Weights
A neural network-based approach predicts blending coefficients and applies anisotropic kernels to enhance image quality and stability in high-resolution image and video generation, addressing resource-intensity and artifact issues in existing techniques.
Patent Information
- Application Number
- JP2022077394
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-29
- Filing Date
- 2022-05-10
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-05-10
AI Technical Summary
Existing techniques for generating high-quality image and video content at higher resolutions are resource-intensive and struggle with determining optimal blending weights, leading to noisy or artifact-prone outputs due to ghosting or temporal instability.
A neural network-based approach is used to predict blending coefficients and apply anisotropic kernels for upsampling, aligning with motion vectors and spatial features to reduce checkerboard artifacts and improve image quality.
The method enhances image quality by reducing checkerboard artifacts and improving temporal stability, achieving output resolution multiple times higher than the original with reduced resource consumption.
Smart Images

Figure 0007701307000005 
Figure 0007701307000006 
Figure 0007701307000007
Abstract
Description
Technical Field
[0001] At least one embodiment relates to the processing of resources used to execute and facilitate artificial intelligence. For example, at least one embodiment relates to a processor or computing system used to train a neural network by various novel techniques described herein.
Background Art
[0002] Image and video content are increasingly being generated at higher resolutions and displayed on higher quality displays. Techniques for generating higher quality content are often very resource intensive, especially at modern frame rates, which can be a problem for devices with limited resource capacity. Blending the data of current and previous frames in a sequence can help improve the quality of this content by providing some temporal smoothing and accumulation of pixel data between frames, but it is difficult to determine the optimal blending weights, and the use of inappropriate blending weights can result in images that are too noisy or have artifacts such as ghosting or temporal instability.
[0003] With reference to the drawings, various embodiments according to the present disclosure will be described.
Brief Description of the Drawings
[0004]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 1E
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12A
Figure 12B
Figure 12C
Figure 12D
Figure 12E
Figure 12F
Figure 13
Figure 14A
Figure 14B
Figure 15A
Figure 15B
Figure 16
Figure 17A
Figure 17B
Figure 17C
Figure 17D
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26A
Figure 26B
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33A
Figure 33B
DETAILED DESCRIPTION OF THE INVENTION
[0005] In at least one embodiment, an upscaling process such as a deep learning-based super-sampling or super-resolution process can be used to increase the resolution of one or more images, such as images or video frames within a sequence or video stream. In at least one embodiment, as shown in FIG. 1A, this can include upsampling from a set 102 of lower-resolution pixels to a set 104 of higher-resolution pixels, such as by 4x upsampling as illustrated. In at least one embodiment, this can include the representation of one or more objects within a scene, such as a scene of live game play. In at least one embodiment, a rendering engine can output an image of one or more objects at a first resolution that will be upscaled to one or more higher output resolutions. In at least one embodiment, real-time temporal image reconstruction can be performed at a resolution higher than the resolution at which the image is generated by the rendering engine. In at least one embodiment, the temporal aspect of this process can include blending the color values of corresponding points between the current frame and at least one previous or history frame within the sequence. In at least one embodiment, to ensure that this blending is performed for corresponding points on objects within these frames, this previous history color data can be warped based on the motion detected between this history frame and this current frame such that it can be indicated by a set of motion vectors output from this rendering engine or determined in some other manner. In at least one embodiment, such warping can ensure that points such as feature points of various images are tracked over time and the corresponding color values are used for blending, which can help reduce the presence of artifacts such as noise or flicker during playback.
[0006] In at least one embodiment, and as discussed in more detail elsewhere in this specification, the super-sampling algorithm can utilize a neural network that predicts blending coefficients to determine the amount by which the color values of the current pixel of the current frame and the corresponding historical pixel from the previous warping history frame should be weighted. In at least one embodiment, such an algorithm can also utilize a filtering kernel to generate a new higher-resolution output image from a set of inputs. In at least one embodiment, the output image quality of such a network may at least partially depend on the information available in this input, such as the current luma frame, historical luma, learning history, and color variation mask or motion vector difference buffer. In at least one embodiment, the application may render an aliased 1 sample per pixel (spp) image at a resolution of 1080p (FullHD), and this algorithm may reconstruct an anti-aliased 2160p (4k) image from this input image and any such side information sequence provided by this application. In at least one embodiment, such a process can be extended to other resolutions with other upscaling ratios, including the case of pure anti-aliasing where the input resolution and the output resolution are equal.
[0007] In at least one embodiment, the rendering and / or upsampling process can utilize sample locations 106 within each lower-resolution pixel to determine the color of its pixels, as shown within image frame 100. In at least one embodiment, this sampling location can have a certain amount of random / offset or dither applied between frames to enable capturing of fine or sub-pixel details, as will be discussed in more detail later herein. In at least one embodiment, the color of the image to be rendered can include a set of horizontal coloring strips as shown in FIG. 1A. In at least one embodiment, sample points 406 are shown that fail to sample any of this medium gray that occupies most of this image space of the current image or frame to be rendered, such that the result is that white or dark gray is sampled. In at least one embodiment, the color values provided by that sampling can correspond to a subset of higher-resolution pixels as shown in image 140 of FIG. 1C. In at least one embodiment, the original approach is to apply these colors to all (here four) higher-resolution pixels corresponding to a single lower-resolution pixel, and such an approach loses at least some of this fine sub-pixel detail used to obtain this dither offset. In at least one embodiment, these sampled color values can then be considered only for, or primarily for, those higher-resolution pixels 142, 144 that included one of these sample points. In at least one embodiment, this leaves most of the pixels lacking color values within this current image, shown here by a stippled pattern, which then does not participate in this dither-aware blending of the current input color and the warped previous output color. In at least one embodiment, these color values can be blended with the colors from the previous frame 120, as shown in FIG. 1B, and those colors may have been processed in other ways such as warping, filtering, or as discussed in more detail elsewhere herein.In at least one embodiment, an image 160 such as that shown in FIG. 1D can be provided as a result of blending the color values of the current frame from FIG. 1C with the warping color values of the previous or history frame from FIG. 1B. In at least one embodiment, this blending preserves some of this fine detail represented in FIG. 1C, but this blending results in a kind of checkerboard pattern consisting of pixels 162, 166 where current color information exists and pixels 164, 168 where current color information does not exist. In at least one embodiment, as a result, a lower quality image can be provided, at least because such a checkerboard pattern did not exist in the original image data that was rendered or was to be rendered.
[0008] In at least one embodiment, the spatio-temporal upsampling process can utilize, as part of an image recognition algorithm, a jittered input image and associated jitter values, as well as a set of low-resolution backward motion vectors per input image pixel, and optionally other quantities such as exposure values and depth buffers. In at least one embodiment, these low-resolution input (backward) motion vectors are used to warp the high-resolution output image of the previous frame to be geometrically shaped and positionally aligned with the current time step. In at least one embodiment, a neural network can be utilized that can infer a set of anisotropic reconstruction kernel parameters for upsampling and filtering the current input image, based at least in part on the current input image and the warped output image of the previous frame. In at least one embodiment, this neural network (or a separate neural network) can also infer one or more weighting coefficients for blending this upsampled input image with at least one warped previous output image. In at least one embodiment, this current input image is upsampled according to these predicted kernel parameters, and the high-resolution output image of the current frame is blended from, or composed of, the current input color, the anisotropically kernel-upsampled current input color, and the output color of the warped previous frame.
[0009] In at least one embodiment, the blending process can be described using the following exemplary quantities. g: The gating coefficient predicted by the neural network b: The jitter-aware blending coefficient predicted by the neural network p: Parameters that form the reconstruction kernel predicted by the neural network, such as in the case of an anisotropic Gaussian kernel that defines a covariance matrix j: The current input jitter vector indicating the sampling position within the input pixel c u: The current input sample color at the input pixel coordinates u (integer coordinates) h: The output color of the warped previous frame K(x,p): The value of the reconstruction kernel parameterized by p at offset x. The kernel is centered at each output pixel.
[0010] In at least one embodiment, using these definitions, an operation for determining the output color from the perspective of a particular output pixel can be described. In at least one embodiment, when the input pixel overlaps with that output pixel, the input color is blended into the output pixel. In at least one embodiment, a prediction dither-aware blending coefficient b is used as a blending weight such that it can be given by the following. h JAB =(1 - b)h+bc u If the vector u + j hits this output pixel h Otherwise
[0011] In at least one embodiment, an upsampled color value can be determined at this output pixel. In at least one embodiment, this can be achieved by forming a weighted sum of the input color values, using the predicted reconstruction kernel values as weights such that it can be given by the following. c UP =Σ u {K(u + j,p)c u} / Σ u K(u + j,p)
[0012] In at least one embodiment, using the predicted gating coefficient g as a blending weight such that it can be given by the following, the output color y can be obtained by blending both the dither-aware blending history color h JAB and the upsampled input color c UP . y=(1 - g)h JAB +gc UP
[0013] In at least one embodiment, the first dither-aware blend step using the nearest color has the advantage that details at high frequencies (e.g., up to the Nyquist frequency at the output resolution) can be accumulated in this output image. In at least one embodiment, this is at least partially due to the input color being accumulated in the output pixel that is closest in the dither-aware blend step. In at least one embodiment, anti-aliasing can be improved by using an anisotropic kernel in the upsampling and related second blending step. In at least one embodiment, the anisotropic kernel is (at least conceptually) aligned with the edges in this image, and as a result, the aliased input color is smoothed along the edges. In at least one embodiment, such an approach may not be optimal for upsampling. In at least one embodiment, the situation of 4x upsampling, which is 2 times in the x direction and 2 times in the y direction, can be considered. In at least one embodiment, the underlying geometry may include thin horizontal (or vertical) features that are shaded with a color that is significantly different from the neighboring content, as shown by the stripes of white pixel color in FIG. 1A. In at least one embodiment, the output color of the previous frame may be as shown in FIG. 1B. In at least one embodiment, the current input sample hits this thin white stripe, and thus the input color is significantly different from the warped previous output color data, as shown in FIG. 1C. In at least one embodiment, as a result of blending this input color with the nearest output color, there will be distinct checkerboard artifacts as described above.
[0014] In at least one embodiment, one or more mechanisms can be used to prevent or reduce the impact of this checkerboard pattern from reaching the output color that can be returned to the rendering or upsampling application. In at least one embodiment, the neural network is such that the checkerboard JABTo avoid this, a low dither-aware blending coefficient b can be predicted. In at least one embodiment, the neural network predicts a high gating coefficient g, a checkerboard h JAB image can rely more strongly on the upsampled input color than on the image. In at least one embodiment, any of these mechanisms can result in the output being mainly composed of anisotropic kernel upsampled input colors accumulated over multiple frames. In at least one embodiment, a possible drawback is a situation where the input samples cannot contribute mainly to the colors of these nearest output pixels, which may limit or prevent the reconstruction of high-frequency details. In at least one embodiment, such problems may be further exacerbated when the neural network is operated in the downsampling mode, in which the predicted gating coefficient g can be used for a block of output pixels, so that it is not impossible, but difficult, to accumulate any high-frequency details in this output unless this checkerboard dither-aware blend image is passed to this output.
[0015] In at least one embodiment, the upsampling process can perform blending by upsampling the input color using one or more predicted anisotropic kernels as described above. In at least one embodiment, this upsampling process can take into account the total kernel weight w UP as follows. w UP =Σ u K(u + j, p) c UP =Σ u {K(u + j, p)c u} / wUP
[0016] In at least one embodiment, this dither-aware blending step uses this upsampled color c UPcan be blended with the output color h of the warped previous frame. In at least one embodiment, this blending weight is the product of the predicted blending coefficient from this neural network and a total anisotropic kernel weight such as bw UP and the like. h AJAB =(1 - bw UP )h + bw UP c UP In at least one embodiment, the output can be formed to be given by the following. y=(1 - g)h AJAB + gc UP
[0017] It may be noted that in at least one embodiment, since both of these two blending steps operate on only two inputs h and cUP, they can be combined into a single blending operation. In at least one embodiment, a combined blending coefficient can be formed from the total weight w UP and the DNN predicted weights b and g. w = g+(1 - g)*b*w UP
[0018] In at least one embodiment, the output can then be formed as a single blending operation using this combined coefficient to be given by the following. y=(1 - w)h+wc UP
[0019] In at least one embodiment, the total blending weight given in the above equation is a function of the two DNN predictions g and b, and the kernel weight w UP . In at least one embodiment, these two DNN predictions can be used to balance between reusing the history color and relying on the current input color. In at least one embodiment, the contribution of a color sample to a particular output pixel is the weight of that sample (here w UP) may depend in part on the weights of all other samples. In at least one embodiment, this can be formulated as a temporal recursive execution, where the weight of a new sample w UP reached at a given time step is compared to the total cumulative weight over past time steps given by w hist . In at least one embodiment, the total blending weight formula in this context can be used to approximate the ratio of the current sample weight to the new total sample weight such that it can be given by the following. w = w UP / (w hist + w UP )
[0020] In at least one embodiment, this weighting factor is a function of only two quantities, namely the current sample weight w UP and the history weight w hist . In at least one embodiment, other formulations can also be applied, such as when the total weight is a function of one or more quantities β i predicted by a neural network and the sample weight wUP. w = f(w UP , β i )
[0021] In at least one embodiment, and by way of non - limiting example, the neural network can directly predict the history weight w hist through, for example, an appropriate activation function that ensures it is non - negative such that it can be given by the following. w = w UP / (ReLU(β i ) + w UP )
[0022] In at least one embodiment, such an approach can provide various advantages over other blending techniques. Referring to the above illustration of thin vertical image features, in at least one embodiment, the jitter-aware blending step utilizes anisotropic kernel weights instead of nearest neighbor weights, so the output of such an operation can significantly reduce the checkerboards included. In at least one embodiment, this is possible when the prediction kernel is positionally aligned with the underlying image feature (image feature; image feature point) such that the neural network behaves so when properly trained. In at least one embodiment, image 180 of FIG. 1E shows the output of one such anisotropic jitter-aware blend operation. In at least one embodiment, X locations 182, 186 specify the center locations of the anisotropic kernel or filter. In at least one embodiment, each output pixel can have an associated kernel, but only three are shown here for illustrative purposes. In at least one embodiment, cells with a checkerboard pattern represent the output pixels with the minimum weight associated with them. In at least one embodiment, these prediction kernels are elongated along these image features, so no checkerboard pattern occurs and the final output is similar to that shown in FIG. 1B. In at least one embodiment, by using such an approach, it can be made possible for the upsampling process to be sufficiently prepared to reconstruct a detailed image with significantly reduced checkerboard artifacts.
[0023] In at least one embodiment, such a technique can be advantageously used in a temporal upsampling system, such as a deep learning-based super-sampling or super-resolution system. In at least one embodiment, the components of one such system 200 are shown in FIG. 2, which can be used to perform image reconstruction as discussed with respect to FIG. 1. In at least one embodiment, the components of system 200 can be implemented on one or more processing components of similar or different types, including any of those discussed herein. In at least one embodiment, content such as video game content or animation can be generated using a renderer 202, a rendering engine, or other such content generation means of such a system. In at least one embodiment, the renderer 202 can receive an input of one or more frames of a sequence and generate an image or video frame using stored content 204 (e.g., maps and graphic assets) that is modified at least in part based on that input. In at least one embodiment, this renderer 202 can utilize rendering software such as Unreal Engine 4 by Epic Games, Inc., which can provide functions such as deferred shading, global illumination, light transmission processing, post-processing, and GPU particle simulation using a vector field, and can be part of a rendering pipeline, such as one that can be used to generate a rendered image 206 at a resolution lower than one or more final output resolutions in order to meet timing requirements and reduce processing resource requirements, for example, which may make it difficult to render these video frames to meet a current frame rate, such as at least 60 frames per second (fps).In at least one embodiment, this low-resolution rendering image 206 can be processed using an upscaler 208 to generate an upscaled image 210 that represents the content of the low-resolution rendering image 206 at a resolution equal to (or at least close to) the target output resolution.
[0024] In at least one embodiment, an upscaler system 208 (which can take the form of a service, system, module, or device) can be used to upscale individual frames of a video or animation sequence. In at least one embodiment, the amount of upscaling to be performed can depend on the initial resolution of the rendering image and the target resolution of the display, such as ranging from 1080p to 4k. In at least one embodiment, additional processing, such as may include anti-aliasing and temporal smoothing, can be performed as part of the upsampling process. In at least one embodiment, a suitable reconstruction filter, such as may include a filter such as an anisotropic Gaussian filter or a dynamic filter network (DFN), can be utilized. In at least one embodiment, the upsampling process can take into account subpixel dithering that can be applied on a per-frame basis.
[0025] In at least one embodiment, deep learning can be used to infer these upsampled video frames of the sequence. In at least one embodiment, temporal reconstruction can be used to provide a combination of anti-aliasing and super-resolution. In at least one embodiment, information from the corresponding sequence of video frames can be used to infer a higher quality upsampled image. In at least one embodiment, one or more heuristics based on previous knowledge of the rendering pipeline that do not require learning from data can be used. In at least one embodiment, this can include jitter-aware upsampling and accumulation of samples at the upsampling resolution. In at least one embodiment, this jitter offset data can be provided as input to an upscaler 208 that includes at least one neural network for inferring a higher quality upsampled image 210 generated by the upsampling algorithm alone, along with the current input video frame and the previously inferred frames. In at least one embodiment, this upsampling necessarily shifts the jitter offset 222 and samples per frame such that they are positionally aligned with the history buffer that can be at a higher resolution.
[0026] In at least one embodiment, this upscaled image 210 can be provided as an input to a neural network 212 for determining one or more blending coefficients or blending weights. In at least one embodiment, the neural network 212 also receives, as inputs, this upscaled image 210 along with the warped previous high-resolution images in this sequence that are provided to the neural network 212. In at least one embodiment, the neural network 212 can also receive other input features as being related to spatial and temporal variations as discussed herein. In at least one embodiment, deep learning is used to reconstruct an image for real-time rendering at a resolution that is a multiple (e.g., 2 to 9) times higher than the actual rendering resolution. In at least one embodiment, the reconstructed image quality from such a process is comparable to or even exceeds the rendering at the original resolution, at least with respect to details, temporal stability, and the absence of common artifacts such as ghosting or lag. In at least one embodiment, the neural network can also determine at least some filtering to be applied when reconstructing or blending the current image with previous images. In at least one embodiment, this information can then be provided to a blending component 214 along with this upscaled image 210 to be blended with at least one previous image in this sequence. In at least one embodiment, dither offset data 222 can also be provided as an input to this blending component 214. In at least one embodiment, this blending of the current image with previous (or history) images in the sequence can help to temporally converge to a good, sharp high-resolution output image 216, which can then be provided for presentation via a display 220 or other such presentation mechanism.In at least one embodiment, a copy of this high-resolution output image 216 can also be stored in the history buffer 218 or other such storage location for blending with images generated later in this sequence. In at least one embodiment, such a process utilizes deep learning to have a reconstructed image quality that at least matches the rendering at the original resolution with respect to details, temporal stability, and the absence of common artifacts such as ghosting or lag, and can reconstruct an image for real-time rendering at a resolution that is multiple times (e.g., 2 times, 4 times, or 8 times) higher than the actual rendering resolution. In at least one embodiment, tensor cores can be used to accelerate the reconstruction speed, and by using the techniques presented herein as such, this rendering process can be made much more sample-efficient and can significantly increase the frames per second for various applications.
[0027] In at least one embodiment, using buffered information in a system such as that described with respect to FIG. 2 can include components 300 such as those shown in FIG. 3. In at least one embodiment, three primary input sources are utilized, including a color buffer 302, a motion vector buffer 304, and a depth buffer 306. In at least one embodiment, a pre-processor 308, such as may include one or more processes operating on one or more processors on one or more computing devices, can receive as input the color information of the current frame as generated by a rendering engine or application, and the output of a warper 310, such as a warp function or application that executes on one or more processors of one or more devices. In at least one embodiment, this warper 310 receives as input the motion vector information of the current frame as stored in the motion vector buffer 304, and the depth information of the current frame as stored in the depth buffer 306. In at least one embodiment, the warper 310 may receive this data directly from an application or renderer and may not utilize a dedicated buffer. In at least one embodiment, this temporal process also provides as input to the warper 310 high-resolution color data from previous images in the sequence as stored in the history buffer 314. In at least one embodiment, as described above, the information for each final output image can also be stored in the history buffer 314 for use in generating subsequent images or frames in the sequence. In at least one embodiment, the warper 310 can use this motion vector and depth data to efficiently use these motion vectors to warp the pixel data or color data of specific features of a previous image to corresponding pixel locations within the current image frame, to the corresponding pixel locations of features in these two images, and thus can compare and blend the color values of similar features.In at least one embodiment, the pre-processor 308 can perform any relevant processing on the current color data from the color buffer 302 or the warped previous color data from the warper 310. In at least one embodiment, this data after any pre-processing is provided as input to a neural network or other deep learning (DL)-based generator 312 that can determine pixel-specific weights for each pixel location in the image to be generated by analyzing this data. In at least one embodiment, this generated data is sent to a post-processor 316 that may include one or more processes executed on one or more processors of one or more computing devices that can output a final high-resolution color image 318. In at least one embodiment, this post-processor can also output information stored in the high-resolution color and history buffer 314 for use in generating subsequent images within the current sequence.
[0028] In at least one embodiment, the generation of a frame using such an approach can include the application providing a low-resolution jitter input image and associated jitter values, low-resolution backward motion vectors for each individual input image pixel, and other quantities such as exposure values and depth buffers to a reconstruction algorithm. In at least one embodiment, these low-resolution input (backward) motion vectors can be used to warp the previous frame output image to align with the geometry and position within the current time stamp. In at least one embodiment, this low-resolution current frame image is upsampled to the resolution of the output image 318 using an upsampling algorithm. In at least one embodiment, a neural network 312 is used to infer the weight value w for each output pixel (at the output resolution). In at least one embodiment, the high-resolution output image of the current frame can be created as follows. Output = w * (upsampled current frame input image) + (1 - w) * (warped previous output image)
[0029] In at least one embodiment, also in this type of temporal image reconstruction image algorithm, an important factor in the resulting image quality (IQ) can be due to the above-mentioned weighting coefficient w. In at least one embodiment, w conforms to various criteria, including when the area in the output image is not occluded due to the movement of an object in the scene being rendered, whether this weighting coefficient prefers the current input image, or when w = 1.0, etc., and weights the color values of the current image more heavily. In at least one embodiment, when the area in the output image was visible (and similarly shaded) in a previous frame, the optimal weighting coefficient can result in an appropriate blending between these previous output images and the current input image. In at least one embodiment, this blending can prefer historical data, such as when the value of w approaches zero, since more frames have been made to see this area.
[0030] In at least one embodiment, the network can make this prediction weight at least partially based on the current frame input image and the warped previous frame output image. In at least one embodiment, whenever the upsampled current image has significantly different values from the warped previous frame output image and thus looks very different when displayed, the neural network can predict a high-value weighting coefficient w that gives more importance to the upsampled current frame input image. In at least one embodiment, when the current image has values similar to the warped previous frame output image and thus looks very similar when displayed, the neural network can predict a low-value weighting coefficient w that gives more importance to the warped previous frame output image.
[0031] In at least one embodiment, the motion vector difference information can be used as an additional modality or input as discussed herein. In at least one embodiment, an additional buffer such as motion buffer 320 can be used as another source of input within such a system 300. In at least one embodiment, other buffers can be utilized as may include at least one motion data buffer or depth data buffer as discussed herein. In at least one embodiment, this motion buffer 320, also referred to herein as the history motion buffer, can store new or additional motion vector data that can persist across frames. In at least one embodiment, the current motion vectors from motion vector buffer 304 can be stored in one or more forms such as can be adapted to the conversion process for use in subsequent frames. In at least one embodiment, this motion vector information can be provided as an additional input to warper 310. In at least one embodiment, here, the warping function of warper 310 can warp not only the high-resolution color history data from buffer 314, but also this previous motion vector data from motion buffer 320. In at least one embodiment, the temporal calculation means 322 can perform the calculations as described in relation to the above related equations, in which the warping motion vector data from warper 310 is processed with the current motion vector data from motion vector buffer 304 to determine a difference or difference region. In at least one embodiment, this calculation can include determining a difference and then a norm and applying related functions as described above. In at least one embodiment, this temporal calculation can then be provided as an external input to pre-processor 308 and then passed to generator 312 for use in determining more accurate pixel-specific weighting as discussed herein, enabling this DL-based network to generate higher quality results.
[0032] In at least one embodiment, a process 400 for generating an image, as shown in FIG. 4A, can be executed. In at least one embodiment, data can be received 402 from a renderer (e.g., a rendering engine or application) if the data includes low-resolution dithered image data (or image data at a first, unrendered resolution), the dither values used for this low-resolution image data, and the motion vectors of the current image in the sequence. In at least one embodiment, these motion vectors can be used 404 to warp the historical image data of at least one previous image to align the historical image data with the geometric shape and position of the current frame. In at least one embodiment, the image data of the current frame and the warped previous frame can be provided 406 as an input to a neural network. In at least one embodiment, if anisotropic parameters can be inferred for features such as edges detected in the input image data, pixel-wise weighting coefficients and anisotropic reconstruction kernel parameters (e.g., an elongated anisotropic Gaussian filter) can be received 408 from the neural network. In at least one embodiment, these anisotropic reconstruction kernel parameters, also referred to herein as anisotropic filters, can be used 410 to upsample the image data of the current frame to a higher second output resolution. In at least one embodiment, by using an anisotropic filter, it is possible for the color data of one or more neighboring pixels to be considered in a pixel color blending process along features such as edges that can be associated with checkerboards or other such artifacts. In at least one embodiment, pixel-specific blending can be performed 412, such as by applying an anisotropic filter to the blending weights determined for individual pixels based at least in part on the current input color, the output color of the warped previous frame, and the corresponding anisotropic filter or reconstruction kernel parameters.In at least one embodiment, for a given pixel, all contributions of the anisotropic filter can be combined, and then this combined filter value or weight can be used to modify or refine the initial prediction blending weight of this pixel location received from this neural network. In at least one embodiment, these blending color values can be used to generate an output image 414. In at least one embodiment, this output image can be provided for presentation as part of this image or video sequence. In at least one embodiment, the color data of this output image and these current motion vectors can also be stored in respective buffers, and as a result, this data can be used to determine the per-pixel weights of the next image in this sequence.
[0033] In at least one embodiment, a process 450 for generating one or more images as shown in FIG. 4B can be executed. In at least one embodiment, one or more anisotropic filters can be determined 452 to cause one or more output images to be generated. In at least one embodiment, these anisotropic filters can be inferred by a neural network analyzing image data to cause these one or more images to be generated. In at least one embodiment, these one or more anisotropic filters can be used to determine one or more pixel-specific blending weights 454. In at least one embodiment, one or more output images can be generated 456 using these determined pixel-specific blending weights applied to the color data of the current image and at least one previous image in the sequence.
[0034] In at least one embodiment, the amount of resources required to determine blending weights can be reduced by utilizing fewer available color channels than all. In at least one embodiment, this can include utilizing only the luma channel that represents the luminance in the image, which is separate from its chrominance. In at least one embodiment, instead of using full RGB (red, green, blue) or other color values, a single channel such as the luma channel can also be used. In at least one embodiment, the luma information can be determined for both the current frame and the previous reconstructed frame, or history frame. In at least one embodiment, this can enable a significant reduction of the information to be processed from six information channels (current RGB and previous RGB) to two channels (current luma and previous luma). In at least one embodiment, only the luma information of this current frame and the previous frame (and the variance mask if utilized) is provided as an input to this neural network. In at least one embodiment, this neural network can utilize a single information channel in order to infer not the color values but the filters to be applied to the color values or pixel values. In at least one embodiment, this network effectively determines the extent to which previous pixel data can be reused for the reconstruction of the current frame, and for this purpose, a single color data channel is sufficient for both the current frame and the previous frame. In at least one embodiment, luma is utilized because the human eye is more sensitive to luma (i.e., changes in luminance) than to changes in color. In at least one embodiment, reducing the dimensionality of this approach can also help increase generalization. In at least one embodiment, as mentioned, a variance mask can also be generated using only luma values, and the luma values of the pixels from the history frame are compared against the corresponding luma average and standard deviation values of the current frame.
[0035] In at least one embodiment, one or more color filters can be applied to at least the current frame before blending. In at least one embodiment, an upsampling process as discussed with respect to FIG. 2 can be utilized to upsample an image rendered at a first resolution to an image at a higher resolution such as a target output resolution. In at least one embodiment, as a result of this upsampling, an image with some block noise or jaggedness may be produced, which can be partially addressed through an anti-aliasing process. In at least one embodiment, blending can be improved by first applying a color filter to remove these jagged edges within the upsampled image. In at least one embodiment, applying this filtering to the upsampled image can help increase the effective resolution of this image. In at least one embodiment, a filter such as a Gaussian filter can be applied. In at least one embodiment, this can be a parameterized anisotropic Gaussian filter. In at least one embodiment, since there will generally be 3 pixel × 3 pixel blocks that share the same color value, an upsampling process with a large ratio (e.g., 9 times) can result in an image that appears relatively jagged or of low resolution. In at least one embodiment, this block noise can be significantly reduced by applying a filter to remove these colors. In at least one embodiment, as a result of generating significantly more pixel data, significantly more output channels may be required within the output layer of the network, such as 25 channels within the output layer of the network for a 5×5 filter. In at least one embodiment, for very high resolutions, this additional data can potentially make the execution of this network very slow. In at least one embodiment, instead of explicitly predicting all 25 numbers for this filter, a parameterized model can be predicted for that filter.In at least one embodiment, this can include inferring only a small set (e.g., 3) of numbers for this parameterized model, which can be applied to these 25 relevant pixels. In at least one embodiment, these three numbers can later be instantiated to 25 numbers by interpreting these three numbers as parameters of an anisotropic Gaussian kernel, or other parameterized dynamic filter kernel. In at least one embodiment, such an approach can be used with an algorithm that does not utilize neural networks or machine learning.
[0036] In at least one embodiment, it may be desirable to further reduce the processing, memory, and other resources utilized by such a process. In at least one embodiment, the images and inputs provided to the neural network can first be downsampled to operate the neural network at a lower resolution. In at least one embodiment, since the neural network determines the filters or blending weights to be applied to the locations of the image rather than the individual pixel color values, it is possible to operate this network at a lower resolution to reduce resource requirements and latency while maintaining high image reconstruction quality. In at least one embodiment, this can include downsampling at least the current image and the previous reconstructed image to half the resolution, or another reduced resolution. In at least one embodiment, the neural network can be trained at full resolution or a reduced resolution, but can be executed at a reduced resolution during inference. In at least one embodiment, this can serve to separate the display resolution from the resolution at which the neural network operates. In at least one embodiment, a downsampling filter can be applied, such as may include 2×2 downsampling using a regular box filter, but other filters and ratios can also be utilized. In at least one embodiment, the blending weights determined at this lower resolution can be applied to a higher resolution image or set of images to reconstruct the image at the target output resolution. In at least one embodiment, the output of the neural network can be upsampled before being applied to blending and filtering. In at least one embodiment, the blending weight output is smooth, and as a result, the difference in resolution does not significantly affect the quality when applying those blending weights.
[0037] In at least one embodiment, as described above, sub-pixel dither offsets can be determined and utilized. In at least one embodiment, an upsampling system, such as that described with respect to FIG. 2, can be used to upscale individual frames of a sequence. In at least one embodiment, this can include dither-aware upsampling and accumulation of samples at the upsampling resolution. In at least one embodiment, this previous process data can be provided as input to an upsampler system that includes at least one neural network for inferring an upsampled output image, along with the current input frame and the previously inferred frame. In at least one embodiment, the upsampling process can be performed for each individual pixel of a lower resolution rendering image. In at least one embodiment, as a result of the upsampling process, the color information from that pixel can be applied to a corresponding pixel region within the larger sized upscaled image. In at least one embodiment, this pixel within the upscaled image can be segmented (or mapped) into a plurality of individual pixels. In at least one embodiment, the upsampling can be a 4x upsampling, where each pixel of the input image is segmented into four higher resolution pixels.
[0038] In at least one embodiment, the color is determined for the pixel having the center pixel position. In at least one embodiment, the lower resolution image to be rendered then has a single color value reported for this pixel, centered on this center pixel location. However, in at least one embodiment, dithering can be performed between frames or images in the sequence and the center point of color determination is slightly shifted to another point within this pixel. In at least one embodiment, this can correspond to a sample point offset of just a sub-pixel offset from the pixel center. In at least one embodiment, a pixel analysis region (e.g., a 3×3 pixel analysis region) can still be used to determine the color information of a given pixel, but the location of this 3×3 pixel analysis is slightly shifted based on the dither location used to center that pixel analysis region. In at least one embodiment, these dither locations can vary between frames of the sequence, either randomly or according to a determined pattern or sequence. In at least one embodiment, the color data determined for this pixel can include data such as saturation, RGB (red, green, blue) color values, or a photometric (e.g., luma) value representing the luminance of the pixel rather than its final color value, where luma is typically combined with the saturation color value to produce the final pixel value.
[0039] In at least one embodiment, at least some of these values can be utilized with lower precision, such as using a low-precision floating-point format. In at least one embodiment, this can include storing some data in FP32, some in FP16, and some in FP8. In at least one embodiment, the precision of various values can be reduced when appropriate, as may be configured by a user, application, or other such entity or source. In at least one embodiment, the FP8 format may include one sign bit, five exponent bits, and two mantissa bits. In at least one embodiment, different internal representations that utilize the same bit width but different bit formats can be used, for example, reducing the number of mantissa bits in exchange for a different number of exponent bits in some cases. In at least one embodiment, this can represent a trade-off regarding the range of precision logarithmic values. In at least one embodiment, reducing the precision of specific data values can help offset additional overhead and possible impacts on performance experienced when utilizing additional features as discussed herein, which may require more processing, memory, and other such factors or resources. In at least one embodiment, an attempt can be made to reduce the amount of data to be processed while utilizing this additional data to help improve the quality of the output image data by combining the feature inputs, and thus reduce the overall impact on performance. In at least one embodiment, this can include summing feature inputs such as motion vectors and color dispersion data. In at least one embodiment, other combinations of features can be combined through summation or multiplication, etc.
[0040] In at least one embodiment, a process for generating a sequence of images can be executed, and the image (or video frame) is rendered at a first resolution. In at least one embodiment, this image can be part of any suitable type of content as being related to video, gaming, virtual reality (VR), augmented reality (AR), or other such application or content type. In at least one embodiment, this resolution can be derived from a rendering engine or can be one that provides desired performance. In at least one embodiment, this rendered image can be upsampled to a second, higher resolution using jitter-aware upsampling, and a determined subpixel offset, which can be different from previous offsets of one or more previous images within this sequence, is determined for this image. In at least one embodiment, this upsampled image is provided to a neural network to determine one or more blending weights to be used to blend this upscaled image with previous images.
[0041] In at least one embodiment, the client device 502 can generate this content for a session, such as a gaming session or a video viewing session, using components of the content application 504 on the client device 502 and data stored locally on the client device as shown in FIG. 5. In at least one embodiment, a content application 524 (e.g., a gaming or streaming media application) running on the content server 520 can initiate at least a session associated with the client device 502, assuming it can utilize user data stored in the session manager and user database 534, and if required for this type of content or platform, the content 532 can be determined by the content manager 526, can be rendered using the rendering engine 528, and can be transmitted to the client device 502 using an appropriate transmission manager 522 for download, streaming, or transmission via another such transmission channel. In at least one embodiment, the client device 502 that receives this content can provide this content to the corresponding content application 504, and the content application may include a rendering engine 510 for rendering at least a portion of this content for presentation via the client device 502, such as presentation of video content through the display 506, and presentation of audio, such as voice and music, through at least one audio playback device 508, such as a speaker or headphones, in addition to or alternatively.In at least one embodiment, at least a portion of the content may already be stored, rendered, or accessible on the client device 502 so that transmission via the network 540 is not necessary for at least that portion of the content, such as when the content has been previously downloaded or is stored locally on a hard drive or optical disk. In at least one embodiment, a transmission mechanism such as data streaming can be used to transfer this content from the server 520 or the content database 534 to the client device 502. In at least one embodiment, at least a portion of this content can be obtained or streamed from another source such as a third - party content service 550 which may also include a content application 552 for generating or providing the content. In at least one embodiment, portions of this functionality can be executed using multiple processors within one or more computing devices, such as those that may include a combination of multiple computing devices or a combination of a CPU and a GPU.
[0042] In at least one embodiment, the content application 524 includes a content manager 526 that can determine or analyze the content before this content is sent to the client device 502. In at least one embodiment, the content manager 526 can also include or cooperate with other components that can generate, modify, or enhance the content to be provided. In at least one embodiment, this can include a rendering engine 528 for rendering content, such as aliased content, at a first resolution. In at least one embodiment, an upsampling or scaling component 530 can generate at least one additional version of this image at a different resolution, higher or lower, and can perform at least some processing, such as anti-aliasing. In at least one embodiment, a blending component 532, which may include at least one neural network, can perform blending of one or more of those images with one or more previous images, as discussed herein. In at least one embodiment, the content manager 526 can then select an image or video frame at an appropriate resolution for transmission to the client device 502. In at least one embodiment, the content application 504 on the client device 502 can also include components such as a rendering engine 510, an upsampling module 512, and a blending module 514, such that any or all of this functionality can be additionally or alternatively performed on the client device 502. In at least one embodiment, the content application 552 on the third-party content service system 550 can also include such functionality.In at least one embodiment, the location at which at least a portion of this functionality is performed may be configurable or may depend on factors such as the type of client device 502 or the availability of a network connection having an appropriate bandwidth among other such factors. In at least one embodiment, upsampling module 530 or blending module 532 may include one or more neural networks to perform or assist with this functionality, and those neural networks (or at least, the network parameters of those networks) may be provided by content server 520 or third party system 550. In at least one embodiment, a system for content generation can include any suitable combination of hardware and software at one or more locations. In at least one embodiment, the generated image or video content at one or more resolutions can also be provided to or made available for other client devices 560, such as for downloading or streaming from a media source that stores a copy of that image or video content. In at least one embodiment, this may include transmitting an image of game content for a multi-player game when different client devices can display that content at different resolutions, including one or more super-resolutions.
[0043] "Inference and Training Logic" FIG. 6A shows inference and / or training logic 615 used to perform inference and / or training operations with respect to one or more embodiments. Details regarding inference and / or training logic 615 are provided below in conjunction with FIG. 6A and / or FIG. 6B.
[0044] In at least one embodiment, the inference and / or training logic 615 may include, without limitation, code and / or data storage 601 for storing forward propagation and / or output weights, and / or input / output data, and / or other parameters for constructing neurons or layers of a neural network that are trained and / or used to infer in aspects of one or more embodiments. In at least one embodiment, the training logic 615 may include, or be coupled to, code and / or data storage 601 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 601 has weight and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 601 stores the weight parameters and / or input / output data of each layer of a neural network that is trained or used in conjunction with one or more embodiments while the input / output data and / or weight parameters are propagated forward during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 601 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.
[0045] In at least one embodiment, any portion of the code and / or data storage 601 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 601 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 601 is internal or external to, for example, a processor, or the choice of being composed of DRAM, SRAM, flash, or some other type of storage, may be determined according to the on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions being executed, the batch size of the data used in the neural network inference and / or training, or any combination of these factors.
[0046] In at least one embodiment, the inference and / or training logic 615 may include, without limitation, backward propagation and / or output weights corresponding to neurons or layers of a neural network that are trained and / or used for inference in one or more embodiments, and / or code and / or data storage 605 for storing input / output data. In at least one embodiment, the code and / or data storage 605 stores the weight parameters and / or input / output data of each layer of a neural network that is trained or used in conjunction with one or more embodiments while backpropagating the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 605 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 605 has weights and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage 605 may be included with other on-chip or off-chip data storage including the L1, L2, or L3 cache of the processor, or system memory. In at least one embodiment, any portion of the code and / or data storage 605 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 605 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the selection of whether the code and / or data storage 605 is, for example, internal or external to the processor, or the selection of whether it is composed of DRAM, SRAM, flash, or some other type of storage, may be determined according to the on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions to be executed, the batch size of the data used in neural network inference and / or training, or any combination of these factors.
[0047] In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be separate storage structures. In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be the same storage structure. In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any part of the code and / or data storage 601 and the code and / or data storage 605 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.
[0048] In at least one example, the inference and / or training logic 615 includes, without limitation, one or more arithmetic logic units (“ALUs”) 610 including integer and / or floating point units to perform logical and / or arithmetic operations based at least in part on and / or indicated by training and / or inference code (e.g., graph code), the result of which may generate activations (e.g., output values from a layer or neuron within a neural network) stored in activation storage 620, which are a function of the code and / or data stored in code and / or data storage 601 and / or the input / output and / or weight parameter data stored in code and / or data storage 605. In at least one example, the activations stored in activation storage 620 are generated in accordance with linear algebra computations and / or matrix-based computations performed by ALU 610 in response to executing instructions or other code, where the weight values stored in code and / or data storage 605 and / or code and / or data storage 601 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 605, or code and / or data storage 601, or another storage on-chip or off-chip.
[0049] In at least one embodiment, the ALU 610 is included within one or more processors, or other hardware logic devices or circuits, but in another embodiment, the ALU 610 may be external to the processor or other hardware logic device or circuit using them (e.g., a coprocessor). In at least one embodiment, the ALU 610 may be included within the execution unit of a processor, or may be otherwise included within an ALU bank accessible by an execution unit of a processor that is either within the same processor or distributed among different processors of a different type (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.) within the same or different processors. In at least one embodiment, the code and / or data storage 601, the code and / or data storage 605, and the activation storage 620 may be in the same processor or other hardware logic device or circuit, and in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in any combination of the same processor or other hardware logic device or circuit and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activation storage 620 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. Further, the inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuit, and may be fetched and / or processed using the fetch, decode, scheduling, execution, retirement, and / or other logic circuits of the processor.
[0050] In at least one embodiment, the activation storage 620 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation storage 620 may be fully or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activation storage 620 is internal or external to, for example, a processor, or the choice of being composed of DRAM, SRAM, flash, or some other type of storage, may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being executed, the batch size of the data used in neural network inference and / or training, or any combination of these factors. In at least one embodiment, the inference and / or training logic 615 shown in FIG. 6A may be used in conjunction with an application specific integrated circuit (ASIC) such as Google's TensorFlow® processing unit, Graphcore's Inference Processing Unit (IPU), or Intel Corporation's Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 615 shown in FIG. 6A may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field programmable gate array (FPGA).
[0051] FIG. 6B shows inference and / or training logic 615 according to at least one embodiment. In at least one embodiment, the inference and / or training logic 615 may include, without limiting the hardware logic, in which computing resources are dedicated to the weight values or other information corresponding to one or more layers of neurons in the neural network or are used only in conjunction with them in other ways. In at least one embodiment, the inference and / or training logic 615 shown in FIG. 6B may be used in conjunction with application-specific integrated circuits (ASICs) such as Google's TensorFlow® processing units, Graphcore's inference processing units (IPUs), or Intel Corporation's Nervana® (e.g., "Lake Crest") processors. In at least one embodiment, the inference and / or training logic 615 shown in FIG. 6B may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field programmable gate arrays (FPGAs). In at least one embodiment, the inference and / or training logic 615 includes, without limitation, code and / or data storage 601 and code and / or data storage 605, and these may be used to store code (e.g., graph code), weight values, and / or bias values, gradient information, momentum values, and / or other information including other parameters or hyperparameter information. In at least one embodiment shown in FIG. 6B, each of the code and / or data storage 601 and the code and / or data storage 605 is associated with dedicated computing resources such as computing hardware 602 and computing hardware 606, respectively. In at least one embodiment, each of the computing hardware 602 and the computing hardware 606 includes one or more ALUs that execute mathematical functions such as linear algebra functions only on the information stored in the code and / or data storage 601 and the code and / or data storage 605, respectively, and the results are stored in the activation storage 620.
[0052] In at least one embodiment, each of code and / or data storages 601 and 605, and corresponding computing hardware 602 and 606, respectively corresponds to different layers of a neural network, such that the activation resulting from one "storage / computation pair 601 / 602" of code and / or data storage 601 and computing hardware 602 is provided as an input to the next "storage / computation pair 605 / 606" of code and / or data storage 605 and computing hardware 606 in order to reflect the conceptual organization of the neural network. In at least one embodiment, the storage / computation pairs 601 / 602 and 605 / 606 may correspond to two or more layers of a neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 615 after or in parallel with the storage / computation pairs 601 / 602 and 605 / 606.
[0053] "Data Center" FIG. 7 shows an exemplary data center 700 in which at least one embodiment may be used. In at least one embodiment, the data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0054] As shown in FIG. 7, in at least one embodiment, the data center infrastructure layer 710 may include a resource orchestrator 712, grouped computing resources 714, and node computing resources (“node C.R.”) 716(1) to 716(N), where “N” represents any positive integer. In at least one embodiment, the node C.R. 716(1) to 716(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (semiconductor drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, etc., but are not limited thereto. In at least one embodiment, one or more of the node C.R. 716(1) to 716(N) may be servers having one or more of the computing resources described above.
[0055] In at least one embodiment, the grouped computing resources 714 may include separate groups of node C.R.s housed within one or more racks (not shown), or multiple racks housed in a data center at various geographic locations (also not shown). Separate groups of node C.R.s within the grouped computing resources 714 may include grouped compute resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s that include a CPU or processor may be grouped within one or more racks to provide compute resources for supporting one or more workloads. In at least one embodiment, one or more racks may also include any combination of any number of power modules, cooling modules, and network switches.
[0056] In at least one embodiment, the resource orchestrator 712 may configure or otherwise control one or more node C.R.s 716(1)-716(N) and / or the grouped computing resources 714. In at least one embodiment, the resource orchestrator 712 may include a software design infrastructure ("SDI") management entity for the data center 700. In at least one embodiment, the resource orchestrator may include hardware, software, or some combination thereof.
[0057] In at least one embodiment shown in FIG. 7, the framework layer 720 includes a job scheduler 722, a configuration manager 724, a resource manager 726, and a distributed file system 728. In at least one embodiment, the framework layer 720 may include a framework for supporting software 732 of the software layer 730 and / or one or more applications 742 of the application layer 740. In at least one embodiment, the software 732 or the application 742 may each include web-based service software or an application, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 may be a kind of free and open-source software web application framework, such as Apache Spark (trademark) (hereinafter referred to as "Spark"), which can use the distributed file system 728 for large-scale data processing (e.g., "big data"), but is not limited thereto. In at least one embodiment, the job scheduler 722 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 700. In at least one embodiment, the configuration manager 724 may be able to configure different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 728 for supporting large-scale data processing. In at least one embodiment, the resource manager 726 may be able to manage clustered or grouped computing resources mapped or allocated to support the distributed file system 728 and the job scheduler 722. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 714 in the data center infrastructure layer 710.In at least one embodiment, the resource manager 726 may manage these mapped or allocated computing resources in cooperation with the resource orchestrator 712.
[0058] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the node C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distribution file system 728 of the framework layer 720. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.
[0059] In at least one embodiment, the application 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the node C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distribution file system 728 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, recognition computing, and software for training or inference, machine learning applications including machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0060] In at least one embodiment, any one of the configuration manager 724, the resource manager 726, and the resource orchestrator 712 may implement any number and type of self-corrective measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-corrective measures may prevent the data center operator of the data center 700 from determining a configuration that may be defective and may eliminate parts of the data center that are not fully utilized and / or have low performance.
[0061] In at least one embodiment, the data center 700 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, the machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 700. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 700 by using weight parameters calculated by one or more techniques described herein.
[0062] In at least one embodiment, the data center may use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the resources described above. Further, the one or more software and / or hardware resources described above may be configured as a service to enable a user to perform training or inference of information, such as image recognition, speech recognition, or other artificial intelligence services.
[0063] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the system of FIG. 7 for inference or prediction operations, at least partially based on weight parameters calculated using the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the use cases of the neural networks.
[0064] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0065] "computer system" FIG. 8 is a block diagram showing an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or some combination 800 thereof, formed with a processor that may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, computer system 800 may include, without limitation, components such as processor 802 for using an execution unit that includes logic for executing an algorithm for processing data in accordance with the present disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 800 may include a processor such as a PENTIUM® processor family, Xeon®, Itanium®, XScale®, and / or StrongARM® available from Intel Corporation of Santa Clara, California, an Intel® Core®, or an Intel® Nervana® microprocessor, although other systems (including PCs having other microprocessors, engineering workstations, set-top boxes, etc.) may be used. In at least one embodiment, computer system 800 may execute a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and LINUX®), embedded software, and / or graphical user interfaces may be used.
[0066] Embodiments may be used in other devices such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (DSP), a system-on-chip, a network computer (NetPC), a set-top box, a network hub, a wide area network (WAN) switch, or any other system capable of executing one or more instructions according to at least one embodiment.
[0067] In at least one embodiment, computer system 800 may include a processor 802, without limitation, which may include one or more execution units 808 for training and / or inferring a machine learning model by the techniques described herein, without limitation. In at least one embodiment, computer system 800 is a single-processor desktop or server system, but in another embodiment, computer system 800 may be a multi-processor system. In at least one embodiment, processor 802 may include, without limitation, a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 802 may be coupled to a processor bus 810, which may transmit digital signals between processor 802 and other components within computer system 800.
[0068] In at least one embodiment, processor 802 may include, without limitation, a level 1 (L1) internal cache memory (cache) 804. In at least one embodiment, processor 802 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to processor 802. Other embodiments may also include a combination of both internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, register file 806 may store different types of data in various registers including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers.
[0069] In at least one embodiment, the execution unit 808, which includes without limitation logic for performing integer and floating point operations, is also in the processor 802. In at least one embodiment, the processor 802 may also include a microcode ( "u-code") read only memory ( "ROM") that stores microcode for certain macro instructions. In at least one embodiment, the execution unit 808 may include logic for handling a packed instruction set 809. In at least one embodiment, by including the packed instruction set 809 in the instruction set of the general purpose processor 802 along with the associated circuitry for executing the instructions, operations used by many multimedia applications can be executed using the packed data of the general purpose processor 802. In one or more embodiments, by performing operations on packed data using the full width of the processor's data bus, many multimedia applications can be accelerated and executed more efficiently, thereby eliminating the need to transfer smaller units of data between the processor's data buses to perform one or more operations on one data element at a time.
[0070] In at least one embodiment, the execution unit 808 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 800 may include a memory 820 without limitation. In at least one embodiment, the memory 820 may be implemented as a dynamic random access memory ( "DRAM") device, a static random access memory ( "SRAM") device, a flash memory device, or other memory device. In at least one embodiment, the memory 820 may store instructions 819 and / or data 821 represented by data signals that may be executed by the processor 802.
[0071] In at least one embodiment, a system logic chip may be coupled to a processor bus 810 and a memory 820. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (MCH) 816, and the processor 802 may communicate with the MCH 816 via the processor bus 810. In at least one embodiment, the MCH 816 may provide a high-bandwidth memory path 818 to the memory 820 for storing instructions and data and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 816 may direct data signals between the processor 802, the memory 820, and other components of the computer system 800, and may bridge data signals between the processor bus 810, the memory 820, and the system I / O 822. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 816 may be coupled to the memory 820 via the high-bandwidth memory path 818, and the graphics / video card 812 may be coupled to the MCH 816 via an accelerated graphics port (AGP) interconnect 814.
[0072] In at least one embodiment, computer system 800 may use a system I / O 822, which is a proprietary hub interface bus for coupling MCH 816 to an I / O controller hub (ICH) 830. In at least one embodiment, ICH 830 may provide direct connections to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 820, the chipset, and processor 802. By way of example, it may include, without limitation, a legacy I / O controller 823 including an audio controller 829, a firmware hub ("Flash BIOS") 828, a wireless transceiver 826, a data storage 824, a user input and keyboard interface 825, a serial expansion port 827 such as a universal serial bus ("USB"), and a network controller 834. Data storage 824 may comprise a hard disk drive, a floppy (registered trademark) disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0073] In at least one embodiment, FIG. 8 shows a system including interconnected hardware devices or "chips", while in other embodiments, FIG. 8 may show an exemplary system-on-chip ("SoC"). In at least one embodiment, the devices shown in FIG. cc may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 800 may be interconnected using a compute express link (CXL) interconnect.
[0074] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the system of FIG. 8 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations, functions, and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0075] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0076] FIG. 9 is a block diagram showing an electronic device 900 for utilizing a processor 910, according to at least one embodiment. In at least one embodiment, the electronic device 900 may be, for example, without limitation, a notebook, tower server, rack server, blade server, laptop, desktop, tablet, mobile device, phone, embedded computer, or any other suitable electronic device.
[0077] In at least one embodiment, system 900 may include, without limitation, a processor 910 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 910 is coupled using a bus or interface such as a 1°C bus, a System Management Bus (SMBus), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (SPI), a High Definition Audio (HDA) bus, a Serial Advance Technology Attachment (SATA) bus, a Universal Serial Bus (USB) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (UART) bus. In at least one embodiment, FIG. 9 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 9 may show an exemplary System-on-Chip (SoC). In at least one embodiment, the devices shown in FIG. 9 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 9 may be interconnected using a Compute Express Link (CXL) interconnect.
[0078] In at least one embodiment, FIG. 9 shows a display 924, a touch screen 925, a touch pad 930, a Near Field Communications unit (NFC) 945, a sensor hub 940, a thermal sensor 946, an Express Chipset (EC) 935, a Trusted Platform Module (TPM) 938, a BIOS / firmware / flash memory (BIOS, FW flash) 922, a DSP 960, a drive 920 such as a Solid State Disk (SSD) or a Hard Disk Drive (HDD), a wireless local area network unit (WLAN) 950, a Bluetooth unit 952, a Wireless Wide Area Network unit (WWAN) 956, a Global Positioning System (GPS) unit 955, a camera such as a USB3.0 camera (USB3.0 camera) 954, and / or a Low Power Double Data Rate (LPDDR) memory unit (LPDDR3) 915 implemented, for example, to the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0079] In at least one embodiment, other components may be communicatively coupled to the processor 910 via the components described above. In at least one embodiment, an accelerometer 941, an ambient light sensor ("ALS"), a compass 943, and a gyroscope 944 may be communicatively coupled to the sensor hub 940. In at least one embodiment, a thermal sensor 939, a fan 937, a keyboard 946, and a touch pad 930 may be communicatively coupled to the EC 935. In at least one embodiment, a speaker 963, headphones 964, and a microphone ("mic") 965 may be communicatively coupled to an audio unit ("audio codec and class D amplifier") 962, and this audio unit may be communicatively coupled to the DSP 960. In at least one embodiment, the audio unit 964 may include, without limitation, for example, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, a SIM card ("SIM") 957 may be communicatively coupled to the WWAN unit 956. In at least one embodiment, components such as the WLAN unit 950 and the Bluetooth unit 952, as well as the WWAN unit 956, may be implemented in a next generation form factor ("NGFF").
[0080] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the system of FIG. 9 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0081] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0082] FIG. 10 shows a computer system 1000 according to at least one embodiment. In at least one embodiment, the computer system 1000 is configured to implement the various processes and methods described throughout this disclosure.
[0083] In at least one embodiment, computer system 1000 includes, without limitation, at least one central processing unit (“CPU”) 1002, which is connected to a communication bus 1010 implemented using any suitable protocol, such as, without limitation, PCI: Peripheral Component Interconnect (“Peripheral Component Interconnect”), Peripheral Component Interconnect Express (“PCI-Express”: peripheral component interconnect express), AGP: Accelerated Graphics Port (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1000 includes, without limitation, main memory 1004 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in main memory 1004, which may take the form of random access memory (“RAM”: random access memory). In at least one embodiment, a network interface subsystem (“network interface”) 1022 provides an interface with other computing devices and networks for receiving data from other systems and transmitting data from computer system 1000 to other systems.
[0084] In at least one embodiment, computer system 1000 includes, without limitation in at least one embodiment, input device 1008, parallel processing system 1012, and display device 1006, and this display device can be implemented using a conventional cathode ray tube ("CRT"), liquid crystal display ("LCD"), light emitting diode ("LED"), plasma display, or other suitable display technology. In at least one embodiment, user input is received from input device 1008 such as a keyboard, mouse, touch pad, microphone, etc. In at least one embodiment, each of the above modules can be placed on a single semiconductor platform to form a processing system.
[0085] In order to perform inference and / or training operations related to one or more embodiments, inference and / or training logic 615 is used. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of FIG. 10 for inference or prediction operations, at least partially based on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0086] In order to perform inference and / or training operations related to one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used together with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0087] FIG. 11 shows a computer system 1100 according to at least one embodiment. In at least one embodiment, the computer system 1100 may include, without limitation, a computer 1110 and a USB stick 1120. In at least one embodiment, the computer 1110 may include, without limitation, any number and type of processors (not shown), as well as memory (not shown). In at least one embodiment, the computer 1110 includes, without limitation, servers, cloud instances, laptops, and desktop computers.
[0088] In at least one embodiment, the USB stick 1120 includes, without limitation, a processing unit 1130, a USB interface 1140, and USB interface logic 1150. In at least one embodiment, the processing unit 1130 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1130 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1130 comprises an application specific integrated circuit ("ASIC") optimized to perform any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing core 1130 is a tensor processing unit ("TPC") optimized to perform inference operations of machine learning. In at least one embodiment, the processing core 1130 is a vision processing unit ("VPU") optimized to perform inference operations of machine vision and machine learning.
[0089] In at least one embodiment, the USB interface 1140 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1140 is a USB3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1140 is a USB3.0 Type-A connector. In at least one embodiment, the USB interface logic 1150 may include any amount and type of logic that enables the processing unit 1130 to interface with a device (such as computer 1110) via the USB connector 1140.
[0090] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the system of FIG. 11 for inference or prediction operations, at least partially based on the training operations of the neural network, the functions and / or architectures of the neural network, or the weight parameters calculated using the use cases of the neural network described herein.
[0091] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0092] FIG. 12A shows an exemplary architecture in which a plurality of GPUs 1210-1213 are communicatively coupled to a plurality of multi-core processors 1205-1206 via high-speed links 1240-1243 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 1240-1243 support a communication throughput of 4 GB / sec, 30 GB / sec, 80 GB / sec, or more. Various interconnect protocols may be used, including but not limited to PCIe 4.0 or 5.0, and NVLink 2.0.
[0093] Further, in one embodiment, two or more of the GPUs 1210-1213 are interconnected via high-speed links 1229-1230, which may be implemented using the same or different protocols / links as those used for the high-speed links 1240-1243. Similarly, two or more of the multi-core processors 1205-1206 may be connected via high-speed link 1228, which may be a symmetric multi-processor (SMP) bus operating at 20 GB / sec, 30 GB / sec, 120 GB / sec, or more. Alternatively, all communication between the various system components shown in FIG. 12A may be realized using the same protocol / link (e.g., via a common interconnect fabric).
[0094] In one embodiment, each of the multi-core processors 1205-1206 is communicatively coupled to the processor memories 1201-1202 via the memory interconnects 1226-1227, respectively, and each of the GPUs 1210-1213 is communicatively coupled to the GPU memories 1220-1223 via the GPU memory interconnects 1250-1253, respectively. The memory interconnects 1226-1227 and 1250-1253 may utilize the same or different memory access technologies. By way of example and not limitation, the processor memories 1201-1202 and the GPU memories 1220-1223 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories such as 3D XPoint or Nano-Ram. In one embodiment, some portions of the processor memories 1201-1202 may be volatile memories and other portions may be non-volatile memories (e.g., using a two-level memory (2LM) hierarchy).
[0095] As described below, the various processors 1205-1206 and GPUs 1210-1213 may each be physically coupled to specific memories 1201-1202, 1220-1223, and an integrated memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. For example, each of the processor memories 1201-1202 may include a 64 GB system memory address space, and each of the GPU memories 1220-1223 may include a 32 GB system memory address space (resulting in a total of 256 GB of addressable memory in this example).
[0096] FIG. 12B shows further details of the interconnection between a multi-core processor 1207 and a graphics acceleration module 1246 according to one exemplary embodiment. The graphics acceleration module 1246 may include one or more GPU chips integrated on a line card coupled to the processor 1207 via a high-speed link 1240. Alternatively, the graphics acceleration module 1246 may be integrated in the same package or chip as the processor 1207.
[0097] In at least one embodiment, the illustrated processor 1207 includes a plurality of cores 1260A-1260D, each core having a translation lookaside buffer 1261A-1261D and one or more caches 1262A-1262D. In at least one embodiment, the cores 1260A-1260D may include various other components (not shown) for executing instructions and processing data. The caches 1262A-1262D may include level 1 (L1) and level 2 (L2) caches. Further, one or more shared caches 1256 may be included in the caches 1262A-1262D and shared by a set of the cores 1260A-1260D. For example, one embodiment of the processor 1207 includes 24 cores, each core having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more of the L2 and L3 caches are shared by two adjacent cores. The processor 1207 and the graphics acceleration module 1246 are connected to a system memory 1214, which may include the processor memories 1201-1202 of FIG. 12A.
[0098] For the data and instructions stored in the various caches 1262A-1262D, 1256, and the system memory 1214, coherence is maintained through inter-core communication via the coherence bus 1264. For example, each cache may have associated cache coherence logic / circuitry to communicate via the coherence bus 1264 in response to detecting a read or write to a particular cache line. In one embodiment, a cache snooping protocol is implemented via the coherence bus 1264 to monitor cache accesses.
[0099] In one embodiment, the proxy circuit 1225 communicatively couples the graphics acceleration module 1246 to the coherence bus 1264 such that the graphics acceleration module 1246 can participate in the cache coherence protocol as a peer of cores 1260A-1260D. In particular, the interface 1235 provides a connection to the proxy circuit 1225 via a high-speed link 1240 (e.g., a PCIe bus, NVLink, etc.), and the interface 1237 connects the graphics acceleration module 1246 to the link 1240.
[0100] In one embodiment, the accelerator integration circuit 1236 provides services for cache management, memory access, content management, and interrupt management instead of the plurality of graphics processing engines 1231, 1232, N of the graphics acceleration module 1246. Each of the graphics processing engines 1231, 1232, N may include a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1231, 1232, N may include different types of graphics processing engines, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines, within a GPU. In at least one example, the graphics acceleration module 1246 may be a GPU having a plurality of graphics processing engines 1231, 1232, N, or the graphics processing engines 1231, 1232, N may be individual GPUs integrated in a common package, line card, or chip.
[0101] In one embodiment, the accelerator integration circuit 1236 includes a memory management unit (MMU) 1239 for performing various memory management functions, such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation), and a memory access protocol for accessing the system memory 1214. The MMU 1239 can also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In one embodiment, the cache 1238 can store commands and data so that they can be efficiently accessed by the graphics processing engines 1231-1232, N. In one example, the data stored in the cache 1238 and the graphics memories 1233-1234, M are kept coherent with the core caches 1262A-1262D, 1256, and the system memory 1214. As described above, this can be achieved via the proxy circuit 1225 instead of the cache 1238 and the memories 1233-1234, M (for example, by sending updates regarding cache line modifications / accesses in the processor caches 1262A-1262D, 1256 to the cache 1238 and receiving updates from the cache 1238).
[0102] The set of registers 1245 stores context data for the threads executed by the graphics processing engines 1231-1232, N, and the context management circuit 1248 manages the thread contexts. For example, the context management circuit 1248 may perform save and restore operations to save and restore the contexts of various threads during a context switch (e.g., here, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1248 may store the current register values in a specified area of memory (identified, for example, by a context pointer). Then, when returning to the context, the context management circuit 1248 may restore the register values. In one embodiment, the interrupt management circuit 1247 receives and processes interrupts received from system devices.
[0103] In one embodiment, the virtual / effective addresses from the graphics processing engine 1231 are translated by the MMU 1239 into real / physical addresses of the system memory 1214. One example of the accelerator integration circuit 1236 supports a plurality (e.g., 4, 8, 16) of graphics accelerator modules 1246 and / or other accelerator devices. The graphics accelerator module 1246 may be dedicated to a single application executed on the processor 1207 or may be shared among multiple applications. In one example, there is a virtualized graphics execution environment in which the resources of the graphics processing engines 1231-1232, N are shared among multiple applications or virtual machines (VMs). In at least one example, the resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and the priorities associated with the VMs and / or applications.
[0104] In at least one embodiment, the accelerator integration circuit 1236 functions as a bridge to the system for the graphics acceleration module 1246 and provides address translation and system memory cache services. Further, the accelerator integration circuit 1236 may provide a virtualization facility for the host processor to manage the graphics processing engines 1231-1232, N virtualizations, interrupts, and memory management.
[0105] Since the hardware resources of the graphics processing engines 1231-1232, N are explicitly mapped to the physical address space seen by the host processor 1207, any host processor can directly address these resources using the effective address value. One function of the accelerator integration circuit 1236 is to physically separate the graphics processing engines 1231-1232, N so that they appear as independent units to the system.
[0106] In at least one embodiment, one or more graphics memories 1233-1234, M are each coupled to a respective one of the graphics processing engines 1231-1232, N. The graphics memories 1233-1234, M store instructions and data processed by the respective graphics processing engines 1231-1232, N. The graphics memories 1233-1234, M may be volatile memories such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memories such as 3D XPoint or Nano-Ram.
[0107] In one embodiment, in order to reduce data traffic through link 1240, data stored in graphics memories 1233-1234, M is made to be data that is most frequently used by graphics processing engines 1231-1232, N, and preferably is data that is not used (or at least not frequently used) by cores 1260A-1260D. Similarly, the bias mechanism attempts to keep data that the cores need (and thus preferably that graphics processing engines 1231-1232, N do not need) in caches 1262A-1262D, 1256 of the cores, and in system memory 1214.
[0108] FIG. 12C shows another exemplary embodiment in which the accelerator integration circuit 1236 is integrated within the processor 1207. At least in this embodiment, graphics processing engines 1231-1232, N communicate directly with the accelerator integration circuit 1236 via the high-speed link 1240 through interfaces 1237 and 1235 (again, any form of bus or interface protocol can be utilized). The accelerator integration circuit 1236 may perform the same operations as described with respect to FIG. 12B, but potentially may operate at a higher throughput considering its proximity to the coherence bus 1264 and caches 1262A-1262D, 1256. At least one embodiment supports different programming models including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1236 and a programming model controlled by the graphics acceleration module 1246.
[0109] In at least one embodiment, the graphics processing engines 1231-1232, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can concentrate other application requirements on the graphics processing engines 1231-1232, N to achieve virtualization within a VM / partition.
[0110] In at least one embodiment, the graphics processing engines 1231-1232, N may be shared by multiple VM / application partitions. In at least one embodiment, the shared model may use a system hypervisor to virtualize the graphics processing engines 1231-1232, N to enable access by each operating system. In a single partition system without a hypervisor, the graphics processing engines 1231-1232, N are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1231-1232, N to provide access to each process or application.
[0111] In at least one embodiment, the graphics acceleration module 1246 or individual graphics processing engines 1231-1232, N use a process handle to select process elements. In at least one embodiment, the process elements are stored in the system memory 1214 and can be addressed using the translation technique from the effective address to the physical address described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the context of the host process with the graphics processing engines 1231-1232, N (i.e., calling system software to add a process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element within the process element link list.
[0112] FIG. 12D shows an exemplary accelerator integration slice 1290. As used herein, a "slice" comprises a designated portion of the processing resources of accelerator integration circuit 1236. The application effective address space 1282 within system memory 1214 stores process element 1283. In at least one embodiment, process element 1283 is stored in response to a GPU call 1281 from application 1280 executing on processor 1207. Process element 1283 accommodates the process state of the corresponding application 1280. The work descriptor (WD) 1284 accommodated in process element 1283 can be a single job requested by the application, or can accommodate a pointer to a queue of jobs. In at least one embodiment, WD 1284 is a pointer to a job request queue in the application's address space 1282.
[0113] Graphics acceleration module 1246 and / or individual graphics processing engines 1231 - 1232, N can be shared by all or a subset of the processes within the system. In at least one embodiment, infrastructure for setting the process state and sending WD 1284 to graphics acceleration module 1246 to initiate a job in a virtualized environment may be included.
[0114] In at least one embodiment, a dedicated process programming model is implementation specific. In this model, a single process owns graphics acceleration module 1246 or individual graphics processing engine 1231. Since graphics acceleration module 1246 is owned by a single process, when graphics acceleration module 1246 is assigned, the hypervisor initializes accelerator integration circuit 1236 for the owning partition and the operating system initializes accelerator integration circuit 1236 for the owning process.
[0115] During operation, the WD fetch unit 1291 within the accelerator integration slice 1290 fetches the next WD 1284, including the display of work to be performed by one or more graphics processing engines of the graphics acceleration module 1246. As shown, the data from the WD 1284 is stored in the register 1245 and may be used by the MMU 1239, the interrupt management circuit 1247, and / or the context management circuit 1248. For example, one embodiment of the MMU 1239 includes a segment / page walk circuit for accessing the segment / page table 1286 within the OS virtual address space 1285. The interrupt management circuit 1247 may process the interrupt event 1292 received from the graphics acceleration module 1246. When executing a graphics operation, the effective address 1293 generated by the graphics processing engines 1231-1232, N is translated to a physical address by the MMU 1293.
[0116] In one embodiment, the same set of registers 1245 is replicated for each of the graphics processing engines 1231-1232, N, and / or the graphics acceleration module 1246 and may be initialized by the hypervisor or the operating system. Each of these replicated registers may be included in the accelerator integration slice 1290. Exemplary registers that may be initialized by the hypervisor are shown in Table 1.
Table 1
[0117] Exemplary registers that may be initialized by the operating system are shown in Table 2.
Table 2
[0118] In one embodiment, each WD1284 is specific to a particular graphics acceleration module 1246 and / or graphics processing engines 1231 - 1232 and N. WD1784 can contain all the information required for the graphics processing engines 1231 - 1232, N to perform work, or can be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0119] FIG. 12E shows further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor physical address space 1298 in which a process element list 1299 is stored. The hypervisor physical address space 1298 is accessible via a hypervisor 1296 that virtualizes the graphics acceleration module engine of the operating system 1295.
[0120] In at least one embodiment, the shared programming model enables all or a subset of processes from all or a subset of partitions within the system to use the graphics acceleration module 1246. There are two programming models in which the graphics acceleration module 1246 is shared by multiple processes and partitions: time - slice sharing and graphics - directed shared.
[0121] In this model, the system hypervisor 1296 owns the graphics acceleration module 1246 and makes its functions available to all operating systems 1295. In order for the graphics acceleration module 1246 to support the virtualization by the system hypervisor 1296, the graphics acceleration module 1246 may comply with the following: 1) The job requests of the application must be autonomous (i.e., there is no need to maintain the state between jobs), or the graphics acceleration module 1246 must provide a mechanism for saving and restoring the context. 2) The job requests of the application are guaranteed by the graphics acceleration module 1246 to be completed within a specified amount of time, including any translation errors, or the graphics acceleration module 1246 provides a function to preempt the processing of the job. 3) When the graphics acceleration module 1246 operates in a specified shared programming model, fairness must be guaranteed among processes.
[0122] In at least one embodiment, application 1280 needs to make a system call to operating system 1295, along with the type of graphics acceleration module 1246, a work descriptor (WD), a permission mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the type of graphics acceleration module 1246 describes the acceleration function targeted by the system call. In at least one embodiment, the type of graphics acceleration module 1246 may be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 1246 and can be in the form of a command for graphics acceleration module 1246, a virtual address pointer pointing to a user-defined structure, a virtual address pointer pointing to a command queue, or any other data structure for describing the work performed by graphics acceleration module 1246. In one embodiment, the AMR value is the AMR state for use by the current process. In at least one embodiment, the value passed to the operating system is the same as the application that sets the AMR. If the embodiments of accelerator integration circuit 1236 and graphics acceleration module 1246 do not support a user authority mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value and then pass the AMR to a hypervisor call. Hypervisor 1296 may optionally apply the current authority mask override register (AMOR) value and then put the AMR into process element 1283. In at least one embodiment, the CSRP is one of registers 1245 that holds the virtual address of an area within the virtual address space 1282 of the application for graphics acceleration module 1246 to save and restore the context state. This pointer is optional if there is no need to save any state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.
[0123] Upon receiving a system call, the operating system 1295 may verify that the application 1280 is registered and has been granted the right to use the graphics acceleration module 1246. Next, the operating system 1295 makes a call to the hypervisor 1296 with the information shown in Table 3.
Table 3
[0124] Upon receiving a hypervisor call, the hypervisor 1296 verifies that the operating system 1295 is registered and has been granted the right to use the graphics acceleration module 1246. Next, the hypervisor 1296 places the process element 1283 into the process element link list of the corresponding type of graphics acceleration module 1246. The process element may include the information shown in Table 4.
Table 4
[0125] In at least one embodiment, the hypervisor initializes the registers 1245 of the plurality of accelerator integration slices 1290.
[0126] As shown in FIG. 12F, in at least one embodiment, an integrated memory is used that is addressable via a common virtual memory address space used to access physical processors 1201 - 1202 and GPU memories 1220 - 1223. In this embodiment, operations executed by GPUs 1210 - 1213 utilize the same virtual / effective memory address space as accessing the processor memories 1201 - 1202, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1201, a second portion is allocated to a second processor memory 1202, and a third portion is allocated to GPU memory 1220, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thereby distributed across each of the processor memories 1201 - 1202 and GPU memories 1220 - 1223 such that any processor or GPU can access any physical memory with virtual addresses mapped to physical memory.
[0127] In one embodiment, bias / coherence management circuits 1294A - 1294E in one or more of MMUs 1239A - 1239E ensure cache coherence between the caches of one or more host processors (e.g., 1205) and the caches of GPUs 1210 - 1213, and implement a bias technique to indicate the physical memory in which a particular type of data should be stored. Multiple instances of bias / coherence management circuits 1294A - 1294E are shown in FIG. 12F, but the bias / coherence circuit may be implemented within the MMU of one or more host processors 1205 and / or within the accelerator integration circuit 1236.
[0128] One embodiment enables the memory 1220 - 1223 with GPU to be mapped as part of the system memory and to be made accessible using the Shared Virtual Memory (SVM) technique, without incurring performance degradation related to full system cache coherence. In at least one embodiment, the memory 1220 - 1223 with GPU is accessible as system memory without cumbersome cache coherence overhead, providing a beneficial operating environment for GPU offload. With this configuration, the host processor 1205 software can set operands and access computation results without the overhead of conventional I / O DMA data copying. Such conventional copying requires driver calls, interrupts, and memory - mapped I / O (MMIO) accesses, all of which are less efficient than simple memory access. In at least one embodiment, the ability to access the memory 1220 - 1223 with GPU without cache coherence overhead can be essential for the execution time of offloaded computations. For example, in the presence of significant streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 1210 - 1213. In at least one embodiment, the efficiency of operand setting, access to results, and GPU computation efficiency may help in determining the effectiveness of GPU offload.
[0129] In at least one embodiment, the selection of the GPU bias and the host processor bias is determined by a bias tracker data structure. For example, a bias table may be used, which may be a page granularity structure including one or two bits per memory page with a GPU (i.e., may be controlled at the granularity of the memory page). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPUs with memories 1220-1223 with or without a bias cache (e.g., for caching frequently used / recently used entries of the bias table) in GPUs 1210-1213. Alternatively, the entire bias table may be maintained within the GPU.
[0130] In at least one embodiment, the entry of the bias table associated with each access to the GPUs with memories 1220-1223 is accessed prior to the actual access to the GPU memory, causing the following operations. First, local requests from GPUs 1210-1213 finding their pages within the GPU bias are transferred directly to the corresponding GPUs with memories 1220-1223. Local requests from GPUs finding their pages in the host bias are transferred to the processor 1205 (e.g., via the high-speed link described above). In one embodiment, requests from the processor 1205 finding the requested page in the host processor bias complete the request in the same manner as a normal memory read. Alternatively, requests directed to GPU-biased pages may be transferred to GPUs 1210-1213. In at least one embodiment, the GPU may then migrate the page to the host processor bias if the current page is not being used. In at least one embodiment, the bias state of the page can be changed by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply a hardware-based mechanism.
[0131] One mechanism for changing the bias state utilizes an API call (e.g., OpenCL), where this API call calls the GPU's device driver, and this device driver sends a message to the GPU (or adds a command descriptor to a queue) to change the bias state, and for some transitions, guides the GPU to perform a cache flushing operation on the host. In at least one embodiment, the cache flushing operation is used for the transition from the host processor 1205's bias to the GPU bias, but not for the opposite transition.
[0132] In one embodiment, cache coherence is maintained by the host processor 1205 temporarily rendering GPU-biased pages that cannot be cached. To access these pages, the processor 1205 may request access from the GPU 1210, and the GPU 1210 may either immediately grant access or not. Thus, to reduce communication between the processor 1205 and the GPU 1210, it is beneficial to make GPU-biased pages be requested by the GPU but not by the host processor 1205, or vice versa.
[0133] Inference and / or training logic 615 is used to execute one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B.
[0134] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0135] FIG. 13 shows an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general purpose processor cores.
[0136] FIG. 13 is a block diagram showing an exemplary system-on-chip integrated circuit 1300 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1300 includes one or more application processors 1305 (e.g., CPUs), at least one graphics processor 1310, and may further include an image processor 1315 and / or a video processor 1320, any of which may be modular IP cores. In at least one embodiment, integrated circuit 1300 includes peripheral devices or bus logic including a USB controller 1325, a UART controller 1330, an SPI / SDIO controller 1335, and an I 2 S / I 2 2C controller 1340. In at least one embodiment, integrated circuit 1300 can include a display device 1345 coupled to one or more of a high-definition multimedia interface (HDMI™: high-definition multimedia interface™) controller 1350 and a mobile industry processor interface (MIPI) display interface 1355. In at least one embodiment, storage may be provided by a flash memory subsystem 1360 including a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1365 to access a SDRAM or SRAM memory device. In at least one embodiment, some integrated circuits further include an embedded security engine 1370.
[0137] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the integrated circuit 1300 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0138] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0139] FIGS. 14A - 14B show an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general - purpose processor cores.
[0140] Figures 14A - 14B are block diagrams showing exemplary graphics processors for use within a SoC according to the embodiments described herein. Figure 14A shows an exemplary graphics processor 1410 of a system - on - chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. Figure 14B shows a further exemplary graphics processor 1440 of a system - on - chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the graphics processor 1410 of Figure 14A is a low - power graphics processor core. In at least one embodiment, the graphics processor 1440 of Figure 14B is a high - performance graphics processor core. In at least one embodiment, each of the graphics processors 1410, 1440 can be a variant of the graphics processor 1310 of Figure 13.
[0141] In at least one embodiment, the graphics processor 1410 includes a vertex processor 1405 and one or more fragment processors 1415A-1415N (e.g., 1415A, 1415B, 1415C, 1415D-1415N-1, and 1415N). In at least one embodiment, the graphics processor 1410 can execute different shader programs via separate logic, such that the vertex processor 1405 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 1415A-1415N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 1405 executes the vertex processing stage of the 3D graphics pipeline to generate primitives and vertex data. In at least one embodiment, the fragment processors 1415A-1415N use the primitives and vertex data generated by the vertex processor 1405 to generate a frame buffer to be displayed on a display device. In at least one embodiment, the fragment processors 1415A-1415N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform operations similar to pixel shader programs provided in the Direct 3D API.
[0142] In at least one embodiment, the graphics processor 1410 further includes one or more memory management units (MMUs) 1420A - 1420B, caches 1425A - 1425B, and circuit interconnects 1430A - 1430B. In at least one embodiment, one or more MMUs 1420A - 1420B provide virtual to physical address mapping for the graphics processor 1410, including the vertex processor 1405 and / or the fragment processors 1415A - 1415N, and they may reference vertex or image / text data stored in memory in addition to vertex or image / text data stored in one or more caches 1425A - 1425B. In at least one embodiment, one or more MMUs 1420A - 1420B may be synchronized with one or more other MMUs in the system, including one or more MMUs associated with one or more of the application processors 1305, image processors 1315, and / or video processors 1320 of FIG. 13, such that each processor 1305 - 1320 can participate in a shared or integrated virtual memory system. In at least one embodiment, one or more circuit interconnects 1430A - 1430B enable the graphics processor 1410 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.
[0143] In at least one embodiment, the graphics processor 1440 includes one or more of the MMUs 1420A-1420B, caches 1425A-1425B, and circuit interconnects 1430A-1430B of the graphics processor 1410 of FIG. 14A. In at least one embodiment, the graphics processor 1440 includes one or more shader cores 1455A-1455N (e.g., 1455A, 1455B, 1455C, 1455D, 1455E, 1455F-1455N-1, and 1455N), which provide an integrated shader core architecture in which all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders, can be executed by a single core, or type, or core. In at least one embodiment, the number of shader cores can be varied. In at least one embodiment, the graphics processor 1440 includes an inter-core task manager 1445 that acts as a thread dispatcher for dispatching execution threads to one or more of the shader cores 1455A-1455N, and a tiling unit 1458 for accelerating tiling operations for tile-based rendering in which the rendering operation of a scene is subdivided in the image space, for example, to utilize local spatial coherence within the scene or to optimize the use of internal caches.
[0144] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in integrated circuits 14A and / or 14B for inference or prediction operations, at least in part based on weight parameters calculated using the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the use cases of the neural networks. To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0145] FIGS. 15A - 15B show further exemplary graphics processor logic according to the embodiments described herein. FIG. 15A shows a graphics core 1500, which in at least one embodiment may be included in the graphics processor 1310 of FIG. 13 and, in at least one embodiment, may be integrated shader cores 1455A - 1455N as in FIG. 14B. FIG. 15B shows a highly parallel general - purpose graphics processing unit 1530 suitable for introduction into a multi - chip module in at least one embodiment.
[0146] In at least one embodiment, the graphics core 1500 includes a shared instruction cache 1502, a texture unit 1518, and a cache / shared memory 1520, which are common to the execution resources within the graphics core 1500. In at least one embodiment, the graphics core 1500 can include a plurality of slices 1501A - 1501N, or per-core partitions, and the graphics processor can include a plurality of instances of the graphics core 1500. The slices 1501A - 1501N can include support logic that includes local instruction caches 1504A - 1504N, thread schedulers 1506A - 1506N, thread dispatchers 1508A - 1508N, and sets of registers 1510A - 1510N. In at least one embodiment, the slices 1501A - 1501N can include a set of additional functional units (AFU1512A - 1512N), floating-point units (FPU1514A - 1514N), integer arithmetic logic units (ALU1516 - 1516N), address calculation units (ACU1513A - 1513N), double-precision floating-point units (DPFPU1515A - 1515N), and matrix processing units (MPU1517A - 1517N).
[0147] In at least one embodiment, FPUs 1514A - 1514N can perform single - precision (32 - bit) and half - precision (16 - bit) floating - point operations, and DPFPU 1515A - 1515N perform double - precision (64 - bit) floating - point operations. In at least one embodiment, ALUs 1516A - 1516N can perform variable - precision integer operations with 8 - bit, 16 - bit, and 32 - bit precision and can be configured to perform mixed - precision operations. In at least one embodiment, MPU 1517A - 1517N can also be configured to perform mixed - precision matrix operations including half - precision floating - point and 8 - bit integer operations. In at least one embodiment, MPU 1517A - 1517N can perform various matrix operations for accelerating machine - learning application frameworks, including enabling support for acceleration of general matrix - matrix multiplication (GEMM). In at least one embodiment, AFU 1512A - 1512N can perform additional logical operations not supported by the floating - point unit or integer unit, including trigonometric operations (e.g., sine, cosine, etc.).
[0148] To perform inference and / or training operations related to one or more embodiments, inference and / or training logic 615 is used. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in graphics core 1500 for inference or prediction operations, at least partially based on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0149] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0150] FIG. 15B shows a general purpose processing unit (GPGPU) 1530, which can be configured to perform high parallel computing operations by an array of graphics processing units in at least one embodiment. In at least one embodiment, GPGPU 1530 can be directly linked to other instances of GPGPU 1530 to generate multiple GPU clusters to improve the training speed of a deep neural network. In at least one embodiment, GPGPU 1530 includes a host interface 1532 to enable connection with a host processor. In at least one embodiment, host interface 1532 is a PCI Express interface. In at least one embodiment, host interface 1532 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, GPGPU 1530 receives commands from a host processor and uses a global scheduler 1534 to distribute execution threads associated with these commands to a set of compute clusters 1536A - 1536H. In at least one embodiment, compute clusters 1536A - 1536H share a cache memory 1538. In at least one embodiment, cache memory 1538 can act as a high-level cache for cache memory within compute clusters 1536A - 1536H.
[0151] In at least one embodiment, the GPGPU 1530 includes memories 1544A-1544B coupled to compute clusters 1536A-1536H via a set of memory controllers 1542A-1542B. In at least one embodiment, the memories 1544A-1544B can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory.
[0152] In at least one embodiment, each of the compute clusters 1536A-1536H includes a set of graphics cores, such as the graphics core 1500 of FIG. 15A, and this set of graphics cores can include multiple types of integer and floating point logic units capable of performing computational operations with various precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating point units in each of the compute clusters 1536A-1536H can be configured to perform 16-bit or 32-bit floating point operations, while another subset of the floating point units can be configured to perform 64-bit floating point operations.
[0153] In at least one embodiment, multiple instances of GPGPU 1530 can be configured to operate as a compute cluster. In at least one embodiment, the communication used by compute clusters 1536A - 1536H for synchronization and data exchange varies across embodiments. In at least one embodiment, multiple instances of GPGPU 1530 communicate via host interface 1532. In at least one embodiment, GPGPU 1530 includes an I / O hub 1539 that couples GPGPU 1530 to a GPU link 1540 that enables direct connection to other instances of GPGPU 1530. In at least one embodiment, GPU link 1540 is coupled to a dedicated GPU - to - GPU bridge that enables communication and synchronization between multiple instances of GPGPU 1530. In at least one embodiment, GPU link 1540 is coupled to a high - speed interconnect for transmitting and receiving data to and from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 1530 are located in separate data processing systems and communicate via a network device accessible via host interface 1532. In at least one embodiment, GPU link 1540 can be configured to enable connection to a host processor in addition to, or instead of, host interface 1532.
[0154] In at least one embodiment, the GPGPU 1530 can be configured to train a neural network. In at least one embodiment, the GPGPU 1530 can be used within an inference platform. In at least one embodiment where the GPGPU 1530 is used for inference, the GPGPU may include fewer compute clusters 1536A - 1536H than when the GPGPU is used for training a neural network. In at least one embodiment, the memory technology associated with memories 1544A - 1544B may be different for the inference configuration and the training configuration, and a high - bandwidth memory technology may be applied to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 1530 can support inference - specific instructions. For example, in at least one embodiment, the inference configuration can support one or more dot - product instructions for 8 - bit integers, which may be used during the inference operation of a deployed neural network.
[0155] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the GPGPU 1530 for inference or prediction operations, based at least in part on the training operations of the neural network described herein, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network.
[0156] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel - blending weights using one or more anisotropic filters.
[0157] FIG. 16 is a block diagram showing a computing system 1600 according to at least one embodiment. In at least one embodiment, the computing system 1600 includes a processing subsystem 1601 having one or more processors 1602 and a system memory 1604 that communicate via an interconnect path that may include a memory hub 1605. In at least one embodiment, the memory hub 1605 may be a separate component within a chipset component or may be integrated within one or more processors 1602. In at least one embodiment, the memory hub 1605 is coupled to an I / O subsystem 1611 via a communication link 1606. In at least one embodiment, the I / O subsystem 1611 includes an I / O hub 1607 that enables the computing system 1600 to receive input from one or more input devices 1608. In at least one embodiment, the I / O hub 1607 can enable a display controller, which may be included in one or more processors 1602, to provide output to one or more display devices 1610A. In at least one embodiment, one or more display devices 1610A coupled to the I / O hub 1607 can include local, internal, or embedded display devices.
[0158] In at least one embodiment, the processing subsystem 1601 includes one or more parallel processors 1612 coupled to a memory hub 1605 via a bus or other communication link 1613. In at least one embodiment, the communication link 1613 may be one of a number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 1612 form a parallel or vector processing system focused on computing that can include a number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 1612 form a graphics processing subsystem that can output pixels to one of one or more display devices 1610A coupled via an I / O hub 1607. In at least one embodiment, the one or more parallel processors 1612 can also include a display controller and display interface (not shown) that enable a direct connection to one or more display devices 1610B.
[0159] In at least one embodiment, the system storage unit 1614 can be connected to the I / O hub 1607 to provide a storage mechanism for the computing system 1600. In at least one embodiment, an I / O switch 1616 can be used to provide an interface mechanism for enabling communication between the I / O hub 1607 and other components such as the network adapter 1618 and / or the wireless network adapter 1619 that may be integrated into the platform, and various other devices that can be added via one or more add-in devices 1620. In at least one embodiment, the network adapter 1618 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1619 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radios.
[0160] In at least one embodiment, the computing system 1600 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 1607. In at least one embodiment, the communication paths interconnecting the various components of FIG. 16 may be implemented using any suitable protocol such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express), or other buses or point-to-point communication interfaces such as NV-Link high-speed interconnects, or other interconnect protocols.
[0161] In at least one embodiment, one or more parallel processors 1612 incorporate circuitry optimized for graphics and video processing, including, for example, a video output circuit, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1612 incorporate circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 1600 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1612, memory hub 1605, processor 1602, and I / O hub 1607 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 1600 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 1600 can be integrated into a multi-chip module (MCM), and this module can be interconnected with other multi-chip modules to form a modular computing system.
[0162] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, inference and / or training logic 615 may be used in the system of FIG. 1600 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0163] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0164] "Processor" FIG. 17A shows a parallel processor 1700 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 1700 may be implemented using one or more integrated circuit devices such as programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 1700 is a variant of one or more parallel processors 1612 shown in FIG. 16 according to an exemplary embodiment.
[0165] In at least one embodiment, the parallel processor 1700 includes a parallel processing unit 1702. In at least one embodiment, the parallel processing unit 1702 includes an I / O unit 1704 that enables communication with other devices including other instances of the parallel processing unit 1702. In at least one embodiment, the I / O unit 1704 may be directly connected to other devices. In at least one embodiment, the I / O unit 1704 is connected to other devices via the use of a hub or switch interface such as a memory hub 1605. In at least one embodiment, the connection between the memory hub 1605 and the I / O unit 1704 forms a communication link 1613. In at least one embodiment, the I / O unit 1704 is connected to a host interface 1706 and a memory crossbar 1716, where the host interface 1706 receives commands for performing processing operations and the memory crossbar 1716 receives commands for performing memory operations.
[0166] In at least one embodiment, when host interface 1706 receives a command buffer via I / O unit 1704, host interface 1706 can direct a work operation for executing these commands towards front end 1708. In at least one embodiment, front end 1708 is coupled to scheduler 1710, and this scheduler is configured to distribute commands or other work items to processing cluster array 1712. In at least one embodiment, scheduler 1710 ensures that processing cluster array 1712 is properly configured and in an effective state before tasks are distributed to processing cluster array 1712. In at least one embodiment, scheduler 1710 is implemented via firmware logic running on a microcontroller. In at least one embodiment, microcontroller-implemented scheduler 1710 can be configured to perform complex scheduling and work distribution operations at coarse and fine granularities, enabling rapid preemption of threads running on processing array 1712 and context switching. In at least one embodiment, host software can prove the scheduling workload at processing array 1712 via one of a plurality of graphics processing doorbells. In at least one embodiment, the workload can then be automatically distributed across processing cluster array 1712 by scheduler 1710 logic within the microcontroller including scheduler 1710.
[0167] In at least one embodiment, the processing cluster array 1712 includes up to "N" processing clusters (e.g., cluster 1714A, cluster 1714B to cluster 1714N). In at least one embodiment, each of the clusters 1714A to 1714N of the processing cluster array 1712 can execute a large number of simultaneous threads. In at least one embodiment, the scheduler 1710 can use various scheduling and / or work distribution algorithms to distribute work to the clusters 1714A to 1714N of the processing cluster array 1712, and these algorithms may vary according to the workload generated for each type of program or calculation. In at least one embodiment, the scheduling may be dynamically handled by the scheduler 1710, or may be partially assisted by the compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 1712. In at least one embodiment, different clusters 1714A to 1714N of the processing cluster array 1712 can be allocated to process different types of programs or execute different types of calculations.
[0168] In at least one embodiment, the processing cluster array 1712 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 1712 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, the processing cluster array 1712 can include logic for performing processing tasks including filtering of video and / or audio data, execution of modeling operations including physical operations, and execution of data conversion.
[0169] In at least one embodiment, the processing cluster array 1712 is configured to execute parallel graphics processing operations. In at least one embodiment, the processing cluster array 1712 can include texture sampling logic for performing texture operations, as well as additional logic for supporting the execution of such graphics processing operations, including but not limited to mosiac logic and other vertex processing logic. In at least one embodiment, the processing cluster array 1712 can be configured to execute graphics processing related shader programs such as, but not limited to, vertex shaders, mosiac shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 1702 can transfer data from the system memory through the I / O unit 1704 for processing. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 1722) during processing and then written back to the system memory.
[0170] In at least one embodiment, when graphics processing is performed using the parallel processing unit 1702, the scheduler 1710 can be configured to divide the processing workload into tasks of approximately equal size so as to more effectively distribute the graphics processing operations among the plurality of clusters 1714A - 1714N of the processing cluster array 1712. In at least one embodiment, a portion of the processing cluster array 1712 can be configured to perform different types of processing. For example, in at least one embodiment, for generating and displaying a rendered image, the first portion may be configured to perform vertex shading and topology generation, the second portion may be configured to perform mosaic and geometry shading, and the third portion may be configured to perform pixel shading or other screen space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 1714A - 1714N can be stored in a buffer so that the intermediate data can be transmitted among the clusters 1714A - 1714N for further processing.
[0171] In at least one embodiment, the processing cluster array 1712 can receive the processing tasks to be executed via the scheduler 1710, and the scheduler 1710 receives commands defining the processing tasks from the front end 1708. In at least one embodiment, the processing tasks can include an index of the data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands (e.g., which program to execute) defining how the data should be processed. In at least one embodiment, the scheduler 1710 may be configured to fetch the index corresponding to the task, or may receive the index from the front end 1708. In at least one embodiment, the front end 1708 can be configured to ensure that the processing cluster array 1712 is configured in an active state before the workload specified by the incoming command buffer (e.g., batch buffer, push buffer, etc.) is started.
[0172] In at least one embodiment, each of one or more instances of the parallel processing unit 1702 can be coupled to a parallel processor - memory 1722. In at least one embodiment, the parallel processor - memory 1722 can be accessed via a memory crossbar 1716, and the memory crossbar 1716 can receive memory requests from the processing cluster array 1712 as well as the I / O unit 1704. In at least one embodiment, the memory crossbar 1716 can access the parallel processor - memory 1722 via a memory interface 1718. In at least one embodiment, the memory interface 1718 can include a plurality of partition units (e.g., partition unit 1720A, partition unit 1720B - partition unit 1720N), and each of these units can be coupled to a portion (e.g., a memory unit) of the parallel processor - memory 1722. In at least one embodiment, the number of partition units 1720A - 1720N is configured to be equal to the number of memory units, such that the first partition unit 1720A has a corresponding first memory unit 1724A, the second partition unit 1720B has a corresponding memory unit 1724B, and the Nth partition unit 1720N has a corresponding Nth memory unit 1724N. In at least one embodiment, the number of partition units 1720A - 1720N may not be equal to the number of memory devices.
[0173] In at least one embodiment, the memory units 1724A - 1724N can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory. In at least one embodiment, the memory units 1724A - 1724N may also include, but are not limited to, 3D stacked memory including high bandwidth memory (HBM). In at least one embodiment, to efficiently use the available bandwidth of the parallel processor memory 1722, a render target such as a frame buffer or texture map can be stored across the memory units 1724A - 1724N, and the partition units 1720A - 1720N may be able to write portions of each render target in parallel. In at least one embodiment, the local instance of the parallel processor memory 1722 may be excluded to be advantageous for an integrated memory design that combines system memory and local cache memory.
[0174] In at least one embodiment, any one of clusters 1714A-1714N of processing cluster array 1712 can process data to be written to any one of memory units 1724A-1724N within parallel processor memory 1722. In at least one embodiment, memory crossbar 1716 can be configured to transfer the output of each of clusters 1714A-1714N to any partition unit 1720A-1720N capable of performing further processing operations on the output, or to another one of clusters 1714A-1714N. In at least one embodiment, each of clusters 1714A-1714N can communicate with memory interface 1718 through memory crossbar 1716 to read from or write to various external memory devices. In at least one embodiment, memory crossbar 1716 has a connection to memory interface 1718 for communicating with I / O unit 1704, as well as a connection to a local instance of parallel processor memory 1722, enabling processing units within different processing clusters 1714A-1714N to communicate with system memory or other memory not local to parallel processing unit 1702. In at least one embodiment, memory crossbar 1716 can use virtual channels to separate traffic streams between clusters 1714A-1714N and partition units 1720A-1720N.
[0175] In at least one embodiment, multiple instances of the parallel processing unit 1702 may be provided on a single add-in card or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 1702 may be configured to interoperate even if they have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 1702 can include a higher precision floating point unit than other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 1702 or the parallel processor 1700 can be implemented in various configurations and form factors including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.
[0176] Figure 17B is a block diagram of a partition unit 1720 according to at least one embodiment. In at least one embodiment, the partition unit 1720 is an instance of one of the partition units 1720A - 1720N of Figure 17A. In at least one embodiment, the partition unit 1720 includes an L2 cache 1721, a frame buffer interface 1725, and a raster operation unit ("ROP") 1726. The L2 cache 1721 is a read / write cache configured to perform load and store operations received from the memory crossbar 1716 and the ROP 1726. In at least one embodiment, read misses and urgent write-back requests are output by the L2 cache 1721 to the frame buffer interface 1725 to be processed. In at least one embodiment, updates are also sent to the frame buffer via the frame buffer interface 1725 to be processed. In at least one embodiment, the frame buffer interface 1725 interfaces with one of the memory units of the parallel processor memory, such as the memory units 1724A - 1724N (e.g., within the parallel processor memory 1722 of Figure 17).
[0177] In at least one embodiment, the ROP 1726 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. In at least one embodiment, the ROP 1726 then outputs the processed graphics data stored in the graphics memory. In at least one embodiment, the ROP 1726 includes compression logic for compressing depth or color data written to the memory and decompressing depth or color data read from the memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a plurality of compression algorithms. The compression logic executed by the ROP 1726 can be changed based on the statistical characteristics of the data to be compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis for depth and color data.
[0178] In at least one embodiment, ROP1726 is included within each processing cluster (e.g., clusters 1714A - 1714N of FIG. 17A), rather than within partition unit 1720. In at least one embodiment, read and write requests for pixel data, rather than pixel fragment data, are transmitted via memory crossbar 1716. In at least one embodiment, processed graphics data may be displayed on a display device, such as one of the one or more display devices 1610 of FIG. 16, routed so as to be further processed by processor 1602, or routed so as to be further processed by one of the processing entities within parallel processor 1700 of FIG. 17A.
[0179] FIG. 17C is a block diagram of processing cluster 1714 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of the processing clusters 1714A - 1714N of FIG. 17A. In at least one embodiment, one or more of processing clusters 1714 may be configured to execute multiple threads in parallel, where a "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issue techniques are used to support parallel execution of multiple threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) techniques are used to support parallel execution of a number of threads that are overall synchronized, using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0180] In at least one embodiment, the operation of processing cluster 1714 can be controlled via a pipeline manager 1732 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 1732 receives instructions from the scheduler 1710 of FIG. 17A and manages the execution of these instructions via a graphics multiprocessor 1734 and / or a texture unit 1736. In at least one embodiment, the graphics multiprocessor 1734 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 1714. In at least one embodiment, one or more instances of the graphics multiprocessor 1734 can be included within the processing cluster 1714. In at least one embodiment, the graphics multiprocessor 1734 can process data, and a data crossbar 1740 may be used to distribute the processed data to one of a plurality of possible destinations including other shader units. In at least one embodiment, the pipeline manager 1732 can facilitate the distribution of the processed data by specifying the destination of the processed data that is to be distributed through the data crossbar 1740.
[0181] In at least one embodiment, each graphics multiprocessor 1734 within the processing cluster 1714 can include the same set of function execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner such that new instructions can be issued before the previous instruction has completed. In at least one embodiment, the function execution logic supports various operations including integer and floating point arithmetic, comparison operations, boolean operations, bit shifts, and calculations of various algebraic functions. In at least one embodiment, different operations can be executed by leveraging the hardware of the same function units, and any combination of function units may exist.
[0182] In at least one embodiment, the instructions sent to processing cluster 1714 configure threads. In at least one embodiment, a set of threads being executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within the thread group can be assigned to a different processing engine within graphics multiprocessor 1734. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within graphics multiprocessor 1734. In at least one embodiment, if the thread group includes fewer threads than the number of processing engines, one or more processing engines may be idle during the cycles in which the thread group is being processed. In at least one embodiment, the thread group may also include more threads than the number of processing engines within graphics multiprocessor 1734. In at least one embodiment, if the thread group includes more threads than the processing engines within graphics multiprocessor 1734, processing can be executed over consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 1734.
[0183] In at least one embodiment, the graphics multi-processor 1734 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multi-processor 1734 can forego the internal cache and use the cache memory (e.g., L1 cache 1748) within the processing cluster 1714. In at least one embodiment, each graphics multi-processor 1734 can also access the L2 cache within a partition unit (e.g., partition units 1720A - 1720N of FIG. 17), and these caches can be shared among all processing clusters 1714 and may be used to transfer data between threads. In at least one embodiment, the graphics multi-processor 1734 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 1702 may be used as global memory. In at least one embodiment, the processing cluster 1714 includes multiple instances of the graphics multi-processor 1734 that can share common instructions and data, which may be stored in the L1 cache 1748.
[0184] In at least one embodiment, each processing cluster 1714 may include a memory management unit (“MMU”) 1745 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 1745 may be within the memory interface 1718 of FIG. 17A. In at least one embodiment, the MMU 1745 includes a set of page table entries (PTEs) used to map virtual addresses to the physical addresses of tiles and optionally cache line indices. In at least one embodiment, the MMU 1745 may include a translation lookaside buffer (TLB) or cache, which may be within the graphics multiprocessor 1734 or L1 cache, or within the processing cluster 1714. In at least one embodiment, the physical addresses are processed to locally distribute surface data access, enabling efficient interleaving of requests among partition units. In at least one embodiment, a cache line index may be used to determine whether a cache line request is a hit or a miss.
[0185] In at least one embodiment, each graphics multi-processor 1734 is coupled to a texture unit 1736 such that the processing cluster 1714 may be configured to perform texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, the texture data is read from an internal texture L1 cache (not shown) or from the L1 cache within the graphics multi-processor 1734 and, if necessary, fetched from the L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multi-processor 1734 outputs processed tasks to the data crossbar 1740 to provide the processed tasks to another processing cluster 1714 for further processing or stores the processed tasks in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 1716. In at least one embodiment, the pre-ROP 1742 (pre-raster operation unit) is configured to receive data from the graphics multi-processor 1734 and direct the data to the ROP unit, which may be located within a partitioning unit (e.g., partitioning units 1720A - 1720N of FIG. 17A) as described herein. In at least one embodiment, the pre-ROP 1742 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.
[0186] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the graphics processing cluster 1714 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0187] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic may be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0188] FIG. 17D shows a graphics multiprocessor 1734 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 1734 is coupled to the pipeline manager 1732 of the processing cluster 1714. In at least one embodiment, the graphics multiprocessor 1734 has an execution pipeline including, but not limited to, an instruction cache 1752, an instruction unit 1754, an address mapping unit 1756, a register file 1758, one or more general-purpose graphics processing unit (GPGPU) cores 1762, and one or more load / store units 1766. The GPGPU cores 1762 and the load / store units 1766 are coupled to the cache memory 1772 and the shared memory 1770 via a memory and cache interconnect 1768.
[0189] In at least one embodiment, the instruction cache 1752 receives a stream of instructions to be executed from the pipeline manager 1732. In at least one embodiment, the instructions are cached in the instruction cache 1752 and dispatched to be executed by the instruction unit 1754. In at least one embodiment, the instruction unit 1754 can dispatch instructions as a thread group (e.g., a warp), and each thread group is assigned to different execution units within the GPGPU core 1762. In at least one embodiment, instructions can access either a local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 1756 can be used to translate an address in the unified address space to an individual memory address accessible by the load / store unit 1766.
[0190] In at least one embodiment, the register file 1758 provides a set of registers to the functional units of the graphics multiprocessor 1734. In at least one embodiment, the register file 1758 provides temporary storage for operands connected to the data paths of the functional units of the graphics multiprocessor 1734 (e.g., the GPGPU core 1762, the load / store unit 1766). In at least one embodiment, the register file 1758 is divided among each of the functional units such that each functional unit is allocated a dedicated portion of the register file 1758. In one embodiment, the register file 1758 is divided among different warps being executed by the graphics multiprocessor 1734.
[0191] In at least one embodiment, each GPGPU core 1762 can include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions of the graphics multiprocessor 1734. The GPGPU cores 1762 may have the same architecture or different architectures. In at least one embodiment, a first portion of the GPGPU core 1762 includes a single-precision FPU and an integer ALU, and a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU can perform IEEE 754-2008 standard floating-point operations or enable variable-precision floating-point operations. In at least one embodiment, the graphics multiprocessor 1734 can further include one or more fixed-function units or special-function units for executing specific functions such as rectangle copy or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores can also include fixed or special-function logic.
[0192] In at least one embodiment, the GPGPU core 1762 includes SIMD logic that can execute a single instruction on multiple data sets. In at least one embodiment, the GPGPU core 1762 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core may be generated during compilation by a shader compiler or may be automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logical unit.
[0193] In at least one embodiment, the memory and cache interconnect 1768 is an interconnect network that connects each functional unit of the graphics multiprocessor 1734 to the register file 1758 and the shared memory 1770. In at least one embodiment, the memory and cache interconnect 1768 is a crossbar interconnect that enables the load / store unit 1766 to perform load and store operations between the shared memory 1770 and the register file 1758. In at least one embodiment, the register file 1758 can operate at the same frequency as the GPGPU core 1762, and thus, the data transfer between the GPGPU core 1762 and the register file 1758 is very low latency. In at least one embodiment, the shared memory 1770 can be used to enable communication between threads executed by functional units within the graphics multiprocessor 1734. In at least one embodiment, the cache memory 1772 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 1736. In at least one embodiment, the shared memory 1770 can also be used as a program management cache. In at least one embodiment, threads executing on the GPGPU core 1762 can programmatically store data in the shared memory in addition to the automatic cache data stored in the cache memory 1772.
[0194] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated as a core on the same package or chip and communicatively coupled to the core via an internal (i.e., internal to the package or chip) processor bus / interconnect. In at least one embodiment, regardless of the method of connection of the GPU, the processor core may distribute work to such a GPU in the form of a sequence of commands / instructions included in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0195] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the graphics multiprocessor 1734 for inference or prediction operations, based at least in part on the training operations of the neural network, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network described herein.
[0196] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0197] FIG. 18 shows a multi-GPU computing system 1800 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 1800 can include a processor 1802 coupled to a plurality of general purpose graphics processing units (GPGPUs) 1806A-D via a host interface switch 1804. In at least one embodiment, the host interface switch 1804 is a PCI Express switch device that couples the processor 1802 to a PCI Express bus, via which the processor 1802 can communicate with the GPGPUs 1806A-D. The GPGPUs 1806A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 1816. In at least one embodiment, the GPU-to-GPU link 1816 is connected to each of the GPGPUs 1806A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU link 1816 enables direct communication between each of the GPGPUs 1806A-D without requiring communication via the host interface bus 1804 to which the processor 1802 is connected. In at least one embodiment, when there is GPU-to-GPU traffic directed to the P2P GPU link 1816, the host interface bus 1804 is kept available to access system memory or to communicate with other instances of the multi-GPU computing system 1800, for example, via one or more network devices. In at least one embodiment, the GPGPUs 1806A-D are connected to the processor 1802 via the host interface switch 1804, and in at least one embodiment, the processor 1802 includes direct support for the P2P GPU link 1816 and can be directly connected to the GPGPUs 1806A-D.
[0198] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the multi-GPU computing system 1800 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the use cases of the neural networks.
[0199] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0200] FIG. 19 is a block diagram of a graphics processor 1900 according to at least one embodiment. In at least one embodiment, the graphics processor 1900 includes a ring interconnect 1902, a pipeline front end 1904, a media engine 1937, and graphics cores 1980A - 1980N. In at least one embodiment, the ring interconnect 1902 couples the graphics processor 1900 to other graphics processors or other processing units including one or more general-purpose processor cores. In at least one embodiment, the graphics processor 1900 is one of a number of processors integrated within a multi-core processing system.
[0201] In at least one embodiment, the graphics processor 1900 receives a batch of commands via the ring interconnect 1902. In at least one embodiment, incoming commands are interpreted by the command streamer 1903 of the pipeline front end 1904. In at least one embodiment, the graphics processor 1900 includes scalable execution logic for performing 3D geometry processing and media processing via the graphics cores 1980A - 1980N. In at least one embodiment, for 3D geometry processing commands, the command streamer 1903 supplies the commands to the geometry pipeline 1936. In at least one embodiment, for at least some media processing commands, the command streamer 1903 supplies the commands to the video front end 1934, and the video front end 1934 is coupled to the media engine 1937. In at least one embodiment, the media engine 1937 includes a Video Quality Engine (VQE) 1930 for post - processing of video and images, and a multi - format encode / decode (MFX) 1933 engine that provides hardware - accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 1936 and the media engine 1937 each generate execution threads for the thread execution resources provided by at least one graphics core 1980A.
[0202] In at least one embodiment, the graphics processor 1900 includes a scalable thread execution resource characterized by modular cores 1980A-1980N (which may also be referred to as core slices), each modular core 1980A-1980N having a plurality of sub-cores 1950A-1950N, 1960A-1960N (which may also be referred to as core sub-slices). In at least one embodiment, the graphics processor 1900 can have any number of graphics cores 1980A-1980N. In at least one embodiment, the graphics processor 1900 includes a graphics core 1980A having at least a first sub-core 1950A and a second sub-core 1960A. In at least one embodiment, the graphics processor 1900 is a low-power processor having a single sub-core (e.g., 1950A). In at least one embodiment, the graphics processor 1900 includes a plurality of graphics cores 1980A-1980N, each including a first set of sub-cores 1950A-1950N and a second set of sub-cores 1960A-1960N. In at least one embodiment, each sub-core of the first set of sub-cores 1950A-1950N includes at least execution units 1952A-1952N and a first set of media / texture samplers 1954A-1954N. In at least one embodiment, each sub-core of the second set of sub-cores 1960A-1960N includes at least execution units 1962A-1962N and a second set of samplers 1964A-1964N. In at least one embodiment, each sub-core 1950A-1950N, 1960A-1960N shares a set of shared resources 1970A-1970N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.
[0203] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the inference and / or training logic 615 may be used in the graphics processor 1900 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functionality and / or architecture of a neural network, or use cases of a neural network, described herein.
[0204] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0205] FIG. 20 is a block diagram showing the micro-architecture of a processor 2000 that may include a logic circuit for executing instructions according to at least one embodiment. In at least one embodiment, the processor 2000 may execute instructions including x86 instructions, AMR instructions, special instructions for application specific integrated circuits (ASICs), and the like. In at least one embodiment, the processor 2000 may include registers for storing packed data, such as 64-bit wide MMX (trademark) registers in a microprocessor enabled with MMX technology by Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers available in both integer and floating point formats may operate on packed data elements with single instruction multiple data (''SIMD'') and streaming SIMD extensions (''SSE'') instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or more (collectively referred to as ''SSEx'') technologies may hold operands of such packed data. In at least one embodiment, the processor 2000 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0206] In at least one embodiment, the processor 2000 includes an in-order front end ("front end") 2001 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front end 2001 may include several units. In at least one embodiment, the instruction prefetcher 2026 fetches instructions from memory and supplies the instructions to the instruction decoder 2028, and the instruction decoder decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2028 decodes the received instruction into one or more operations called "microinstructions" or "micro-operations" that the machine can execute (also called "micro-ops" or "uops"). In at least one embodiment, the instruction decoder 2028 parses the instruction into an opcode and corresponding data, as well as a control field, such that these are used by the microarchitecture and the operations according to at least one embodiment may be executed. In at least one embodiment, the trace cache 2030 may assemble the decoded uops into a program-order sequence or trace in the uop queue 2034 for execution. In at least one embodiment, when the trace cache 2030 encounters a complex instruction, the microcode ROM 2032 provides the uops necessary for the completion of the operation.
[0207] In at least one embodiment, there are instructions that can be converted into a single micro-op, and there are also instructions that require several micro-ops to complete all operations. In at least one embodiment, if more than five micro-ops are required to complete an instruction, the instruction decoder 2028 may access the microcode ROM 2032 to execute the instruction. In at least one embodiment, the instruction may be decoded into a small number of micro-ops so that it can be processed in the instruction decoder 2028. In at least one embodiment, if a large number of micro-ops are required to complete an operation, the instruction may be stored in the microcode ROM 2032. In at least one embodiment, the trace cache 2030 determines the correct micro-instruction pointer for reading the microcode sequence by referring to an entry-point programmable logic array (PLA) to complete one or more instructions from the microcode ROM 2032 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2032 finishes sequencing the micro-ops for an instruction, the front end 2001 of the machine may resume fetching micro-ops from the trace cache 2030.
[0208] In at least one embodiment, the out-of-order execution engine ("out-of-order engine") 2003 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth the flow of instructions and change their order, optimizing performance when instructions are scheduled to flow down the pipeline and be executed. In at least one embodiment, the out-of-order execution engine 2003 includes, without limitation, an allocator / register renamer 2040, a memory uop queue 2042, an integer / floating point uop queue 2044, a memory scheduler 2046, a fast scheduler 2002, a slow / general purpose floating point scheduler ("slow / general purpose FP scheduler") 2004, and a simple floating point scheduler ("simple FP scheduler") 2006. In at least one embodiment, the fast scheduler 2002, the slow / general purpose floating point scheduler 2004, and the simple floating point scheduler 2006 are also collectively referred to herein as "uop schedulers 2002, 2004, 2006". In at least one embodiment, the allocator / register renamer 2040 allocates the machine buffers and resources required by each uop for execution. In at least one embodiment, the allocator / register renamer 2040 changes the name of the logical register upon entry into the register file. In at least one embodiment, the allocator / register renamer 2040 also distributes the entry of each uop to one of two uop queues, namely the memory uop queue 2042 for memory operations and the integer / floating point uop queue 2044 for non-memory operations, ahead of the memory scheduler 2046 and the uop schedulers 2002, 2004, 2006. In at least one embodiment, the uop schedulers 2002, 2004, 2006 determine when uops are ready for execution based on the availability of the sources of their dependent input register operands and the execution resources required by the uop to complete their operations.In at least one embodiment, the high-speed scheduler 2002 of at least one embodiment may schedule every half of the main clock cycle, and the low-speed / general-purpose floating-point scheduler 2004 and the simple floating-point scheduler 2006 may schedule once per clock cycle of the main processor. In at least one embodiment, the uop schedulers 2002, 2004, 2006 arbitrate dispatch ports to schedule uops for execution.
[0209] In at least one embodiment, the execution block 2011 includes, without limitation, the integer register file / bypass network 2008, the floating-point register file / bypass network (the "FP register file / bypass network") 2010, the address generation units ("AGU") 2012 and 2014, the high-speed arithmetic logic units (ALU) ("high-speed ALU") 2016 and 2018, the low-speed arithmetic logic units ("low-speed ALU") 2020, the floating-point ALU ("FP") 2022, and the floating-point move unit ("FP move") 2024. In at least one embodiment, the integer register file / bypass network 2008 and the floating-point register file / bypass network 2010 are also referred to herein as the "register files 2008, 2010". In at least one embodiment, the AGUs 2012 and 2014, the high-speed ALUs 2016 and 2018, the low-speed ALUs 2020, the floating-point ALU 2022, and the floating-point move unit 2024 are also referred to herein as the "execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024". In at least one embodiment, the execution block b11 may include any number and type of register files, bypass networks, address generation units, and execution units (including zero) in any combination without limitation.
[0210] In at least one embodiment, register files 2008, 2010 may be disposed between uop schedulers 2002, 2004, 2006 and execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024. In at least one embodiment, integer register file / bypass network 2008 performs integer operations. In at least one embodiment, floating-point register file / bypass network 2010 performs floating-point operations. In at least one embodiment, each of register files 2008, 2010 may include, without limitation, a bypass network that may bypass or transfer recently completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, register files 2008, 2010 may communicate data with each other. In at least one embodiment, integer register file / bypass network 2008 may include, without limitation, two separate register files, namely one register file for lower 32-bit data and a second register file for upper 32-bit data. In at least one embodiment, since floating-point instructions typically have operands with a width of 64 to 128 bits, floating-point register file / bypass network 2010 may include, without limitation, 128-bit wide entries.
[0211] In at least one embodiment, the execution units 2012, 2014, 2016, 2018, 2020, 2022, 2024 may execute instructions. In at least one embodiment, the register files 2008, 2010 store operand values of integer and floating-point data that the microinstructions need to execute. In at least one embodiment, the processor 2000 may include any number and combination of execution units 2012, 2014, 2016, 2018, 2020, 2022, 2024 without limitation. In at least one embodiment, the floating-point ALU 2022 and the floating-point shift unit 2024 may execute floating-point, MMX, SIMD, AVX, and SEE, or other operations including special machine learning instructions. In at least one embodiment, the floating-point ALU 2022 includes floating-point dividers of 64 bits each without limitation and may execute division, square root, and other micro-ops. In at least one embodiment, instructions containing floating-point values may be handled by the floating-point hardware. In at least one embodiment, the ALU operations may be passed to the fast ALUs 2016, 2018. In at least one embodiment, the fast ALUs 2016, 2018 may execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, since the slow ALU 2020 may include integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branch processing without limitation, most complex integer operations proceed to the slow ALU 2020. In at least one embodiment, the memory load / store operations may be performed by the AGUs 2012, 2014. In at least one embodiment, the fast ALU 2016, the fast ALU 2018, and the slow ALU 2020 may execute integer operations with 64-bit data operands. In at least one embodiment, the fast ALU 2016, the fast ALU 2018, and the slow ALU 2020 may be implemented to support various data bit sizes including 16, 32, 128, 256, etc. In at least one embodiment, the floating-point ALU 2022 and the floating-point shift unit 2024 may be implemented to support a wide range of operands having various bit widths.In at least one embodiment, the floating-point ALU 2022 and the floating-point shift unit 2024 may operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.
[0212] In at least one embodiment, the uop schedulers 2002, 2004, 2006 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, since uops may be scheduled and executed speculatively in the processor 2000, the processor 2000 may also include logic for handling memory misses. In at least one embodiment, when a data load misses in the data cache, there may be ongoing dependent operations in the pipeline that have passed through a scheduler with temporarily incorrect data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, dependent operations may need to be replayed, and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0213] In at least one embodiment, the term "register" may refer to a storage location of an on-board processor that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be one that can be used from outside the processor (from the perspective of a programmer). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuits within a processor using any number of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, combinations of dedicated physical registers and physically registers dynamically allocated, and the like. In at least one embodiment, an integer register stores 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packed data.
[0214] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of inference and / or training logic 615 may be incorporated into execution block 2011 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs shown in execution block 2011. Further, the weight parameters may be stored in on-chip or off-chip memories and / or registers (shown or not shown) that make up the ALU of execution block 2011 for performing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0215] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0216] FIG. 21 shows a deep learning application processor 2100 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2100 uses instructions that cause the deep learning application processor 2100 to execute some or all of the processes and techniques described throughout this disclosure when executed by the deep learning application processor 2100. In at least one embodiment, the deep learning application processor 2100 is an application specific integrated circuit (ASIC). In at least one embodiment, the application processor 2100 executes matrix multiplication operations that are "hard-wired" to hardware as a result of executing one or more instructions or both. In at least one embodiment, the deep learning application processor 2100 includes, without limitation, processing clusters 2110(1) - 2110(12), inter-chip links ("ICL") 2120(1) - 2120(12), inter-chip controllers ("ICC") 2130(1) - 2130(2), memory controllers ("Mem Ctrlr") 2142(1) - 2142(4), high bandwidth memory physical layers ("HBM PHY") 2144(1) - 2144(4), management-controller central processing unit ("management-controller CPU") 2150, peripheral component interconnect express controller and direct memory access block ("PCIe controller and DMA") 2170, and a 16-lane peripheral component interconnect express port ("PCI Expressx16") 2180.
[0217] In at least one embodiment, the processing cluster 2110 may perform deep learning operations including inference or prediction operations based on weight parameters calculated using one or more training techniques including the techniques described herein. In at least one embodiment, each processing cluster 2110 may include any number and type of processors, without limitation. In at least one embodiment, the deep learning application processor 2100 may include any number and type of processing clusters 2100. In at least one embodiment, the inter-chip link 2120 is bidirectional. In at least one embodiment, the inter-chip link 2120 and the inter-chip controller 2130 enable the plurality of deep learning application processors 2100 to exchange information including activation information obtained as a result of executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2100 may include any number and type of ICLs 2120 and ICCs 2130 (including zero).
[0218] In at least one embodiment, the HBM2 2140 provides a total of 32 gigabytes (GB) of memory. The HBM2 2140(i) is associated with both the memory controller 2142(i) and the HBM PHY 2144(i). In at least one embodiment, any number of HBM2s 2140 may provide any type and total amount of high-bandwidth memory and may be associated with any number and type of memory controllers 2142 and HBM PHYs 2144 (including zero). In at least one embodiment, the SPI, I2C, GPIO 2160, the PCIe controller and the DMA 2170, and / or the PCIe 2180 may be replaced with any number and type of blocks enabling any number and type of communication standards in any technically feasible manner.
[0219] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the deep learning application processor 2100 is used to train a machine learning model, such as a neural network, to predict or infer information provided to the deep learning application processor 2100. In at least one embodiment, the deep learning application processor 2100 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 2100 itself. In at least one embodiment, the processor 2100 may be used to execute one or more of the use cases of the one or more neural networks described herein.
[0220] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0221] FIG. 22 is a block diagram of a neuromorphic processor 2200 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2200 receives one or more inputs from a source external to the neuromorphic processor 2200. In at least one embodiment, these inputs may be sent to one or more neurons 2202 within the neuromorphic processor 2200. In at least one embodiment, the neurons 2202 and their components may be implemented using circuitry or logic that includes one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2200 may include thousands or millions of instances of neurons 2202, without limitation, although any suitable number of neurons 2202 may be used. In at least one embodiment, each instance of a neuron 2202 may include a neuron input 2204 and a neuron output 2206. In at least one embodiment, the neuron 2202 may generate an output, and this output may be sent to the inputs of other instances of the neuron 2202. For example, in at least one embodiment, the neuron input 2204 and the neuron output 2206 may be interconnected via a synapse 2208.
[0222] In at least one embodiment, neuron 2202 and synapse 2208 may be interconnected such that the neuromorphic processor 2200 operates on the information received by the neuromorphic processor 2200 to process or analyze it. In at least one embodiment, neuron 2202 may transmit an output pulse (or “fire” or “spike”) when the input received via neuron input 2204 exceeds a threshold. In at least one embodiment, neuron 2202 may sum or integrate the signals received at neuron input 2204. For example, in at least one embodiment, neuron 2202 may be implemented as a leaky integrate-and-fire neuron, where when the sum (referred to as the “membrane potential”) exceeds a threshold, neuron 2202 may use a transfer function such as a sigmoid function or a threshold function to generate an output (or “fire”). In at least one embodiment, the leaky integrate-and-fire neuron may sum the signals received at neuron input 2204 to form a membrane potential and may also apply a decay factor (or leak) to reduce the membrane potential. In at least one embodiment, the leaky integrate-and-fire neuron may fire if a plurality of input signals are received at neuron input 2204 quickly enough such that they exceed the threshold (i.e., before the decay of the membrane potential is too great to prevent firing). In at least one embodiment, neuron 2202 may be implemented using circuitry or logic that receives an input, integrates the input to form a membrane potential, and decays the membrane potential. In at least one embodiment, the input may be averaged or any other suitable transfer function may be used. Further, in at least one embodiment, neuron 2202 may include, without limitation, comparator circuitry or logic that generates an output spike at neuron 2206 when the result of applying a transfer function to neuron 2204 exceeds a threshold. In at least one embodiment, when neuron 2202 fires, it may ignore the previously received input information, for example, by resetting the membrane potential to 0 or some other suitable default value.In at least one embodiment, when the membrane potential is reset to 0, neuron 2202 may resume normal operation after a suitable period (or refractory period).
[0223] In at least one embodiment, neurons 2202 may be interconnected through synapses 2208. In at least one embodiment, synapses 2208 may be operative to transmit a signal from the output of a first neuron 2202 to the input of a second neuron 2202. In at least one embodiment, neurons 2202 may transmit information through more than one instance of synapses 2208. In at least one embodiment, one or more instances of neuron outputs 2206 may be connected to an instance of neuron inputs 2204 of the same neuron 2202 through an instance of synapses 2208. In at least one embodiment, an instance of neuron 2202 that generates an output that will be transmitted through an instance of synapses 2208 may be referred to as a “presynaptic neuron” with respect to that instance of synapses 2208. In at least one embodiment, an instance of neuron 2202 that receives an input that will be transmitted through an instance of synapses 2208 may be referred to as a “postsynaptic neuron” with respect to that instance of synapses 2208. In at least one embodiment, an instance of neuron 2202 may receive inputs from one or more instances of synapses 2208 and may also transmit outputs through one or more instances of synapses 2208, so a single instance of neuron 2202 may thus be both a “presynaptic neuron” and a “postsynaptic neuron” with respect to various instances of synapses 2208.
[0224] In at least one embodiment, neurons 2202 may be organized into one or more layers. Each instance of neuron 2202 may have one neuron output 2206 that can fan out to one or more neuron inputs 2204 through one or more synapses 2208. In at least one embodiment, the neuron output 2206 of neurons 2202 in the first layer 2210 may be connected to the neuron inputs 2204 of neurons 2202 in the second layer 2212. In at least one embodiment, layer 2210 may be referred to as a "feed-forward layer". In at least one embodiment, each instance of neuron 2202 in an instance of the first layer 2210 may fan out to each instance of neuron 2202 in the second layer 2212. In at least one embodiment, the first layer 2210 may be referred to as a "fully-connected feed-forward layer". In at least one embodiment, each instance of neuron 2202 in an instance of the second layer 2212 may fan out to fewer instances of neuron 2202 in the third layer 2214 than all instances of neuron 2202 in the third layer 2214. In at least one embodiment, the second layer 2212 may be referred to as a "sparsely-connected feed-forward layer". In at least one embodiment, neurons 2202 in the second layer 2212 may fan out to neurons 2202 in a plurality of other layers, including neurons 2202 in the (same) second layer 2212. In at least one embodiment, the second layer 2212 may be referred to as a "recurrent layer". In at least one embodiment, neuromorphic processor 2200 may include any suitable combination of recurrent layers and feed-forward layers, including both sparsely-connected feed-forward layers and fully-connected feed-forward layers, without limitation.
[0225] In at least one embodiment, the neuromorphic processor 2200 may include, without limitation, a reconfigurable interconnect architecture for connecting synapses 2208 to neurons 2202, or dedicated hard-wired interconnects. In at least one embodiment, the neuromorphic processor 2200 may include, without limitation, circuitry or logic that can distribute synapses to different neurons 2202 as needed, based on a neural network topology and the fan-in / fan-out of the neurons. For example, in at least one embodiment, the synapses 2208 may be connected to the neurons 2202 using an interconnect fabric such as a network-on-chip or using dedicated connections. In at least one embodiment, the synapse interconnects and their components may be implemented using circuitry or logic.
[0226] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, this logic may be used with the components of these figures to determine one or more pixel blending weights using one or more anisotropic filters.
[0227] FIG. 23 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, the system 2300 includes one or more processors 2302 and one or more graphics processors 2308 and may be a single-processor desktop system, a multiprocessor workstation system, or a server system having a large number of processors 2302 or processor cores 2307. In at least one embodiment, the system 2300 is a processing platform integrated within a system-on-chip (SoC) integrated circuit for use in a mobile device, a portable device, or an embedded device.
[0228] In at least one embodiment, system 2300 may include, or be incorporated in, a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable gaming console, or an online gaming console. In at least one embodiment, system 2300 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, processing system 2300 may also include, be coupled to, or be integrated within wearable devices such as a smartwatch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 2300 is a television or set-top box device having one or more processors 2302 and a graphical interface generated by one or more graphics processors 2308.
[0229] In at least one embodiment, each of the one or more processors 2302 includes one or more processor cores 2307 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 2307 is configured to process a particular instruction set 2309. In at least one embodiment, instruction set 2309 may facilitate computing via a complex instruction set computing (CISC), a reduced instruction set computing (RISC), or a very long instruction word (VLIW). In at least one embodiment, the processor cores 2307 may each process different instruction sets 2309, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, processor core 2307 may also include other processing devices such as a digital signal processor (DSP).
[0230] In at least one embodiment, the processor 2302 includes a cache memory 2304. In at least one embodiment, the processor 2302 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of the processor 2302. In at least one embodiment, the processor 2302 may also use an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), and this cache may be shared among the processor cores 2307 using known cache coherence techniques. In at least one embodiment, a register file 2306 is further included in the processor 2302, and this register file may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, the register file 2306 may include general purpose registers or other registers.
[0231] In at least one embodiment, one or more processors 2302 are coupled to one or more interface buses 2310 to transmit communication signals, such as address, data, or control signals, between the processor 2302 and other components within the system 2300. In at least one embodiment, the interface bus 2310 can be a processor bus, such as a version of a Direct Media Interface (DMI) bus, in one embodiment. In at least one embodiment, the interface 2310 is not limited to a DMI bus and may include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 2302 includes an integrated memory controller 2316 and a platform controller hub 2330. In at least one embodiment, the memory controller 2316 facilitates communication between the memory device and other components of the system 2300, while the platform controller hub (PCH) 2330 provides connections to I / O devices via a local I / O bus.
[0232] In at least one embodiment, the memory device 2320 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or any other memory device having suitable performance to serve as a process memory. In at least one embodiment, the memory device 2320 operates as a system memory for the system 2300 and can store data 2322 and instructions 2321 for use when one or more processors 2302 execute an application or process. In at least one embodiment, the memory controller 2316 is also coupled to an optional external graphics processor 2312, which may communicate with one or more graphics processors 2308 within the processor 2302 to perform graphics and media operations. In at least one embodiment, the display device 2311 can be connected to the processor 2302. In at least one embodiment, the display device 2311 can include one or more of an internal display device such as a mobile electronic device or a laptop device, or an external display device attached via a display interface (e.g., a display port, etc.). In at least one embodiment, the display device 2311 can include a head-mounted display (HMD) such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.
[0233] In at least one embodiment, the platform controller hub 2330 enables peripheral devices to be connected to the memory device 2320 and the processor 2302 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 2346, a network controller 2334, a firmware interface 2328, a wireless transceiver 2326, a touch sensor 2325, a data storage device 2324 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 2324 can be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 2325 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 2326 can be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2328 enables communication with system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 2334 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 2310. In at least one embodiment, the audio controller 2346 is a multi-channel high-definition audio controller. In at least one embodiment, the system 2300 includes an optional legacy I / O controller 2340 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.In at least one embodiment, the platform controller hub 2330 can also be connected to connection input devices of one or more universal serial bus (USB) controllers 2342, such as a combination of a keyboard and a mouse 2343, a camera 2344, or other USB input devices.
[0234] In at least one embodiment, instances of the memory controller 2316 and the platform controller hub 2330 may be integrated into a separate external graphics processor, such as the external graphics processor 2312. In at least one embodiment, the platform controller hub 2330 and / or the memory controller 2316 may be external to one or more processors 2302. For example, in at least one embodiment, the system 2300 can include an external memory controller 2316 and a platform controller hub 2330, which may be configured as a memory controller hub and a peripheral device controller hub in a system chipset that communicates with the processor 2302.
[0235] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated within the graphics processor 2300. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more of the ALUs embodied within the graphics processor 2312. Additionally, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that shown in FIGS. 6A or 6B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the graphics processor 2300 to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0236] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0237] FIG. 24 is a block diagram of a processor 2400 having at least one embodiment with one or more processor cores 2402A - 2402N, an integrated memory controller 2414, and an integrated graphics processor 2408. In at least one embodiment, processor 2400 can include a lesser number of additional cores including additional core 2402N represented by the dashed rectangle. In at least one embodiment, each of processor cores 2402A - 2402N includes one or more internal cache units 2404A - 2404N. In at least one embodiment, each processor core can also access one or more shared cache units 2406.
[0238] In at least one embodiment, internal cache units 2404A - 2404N and shared cache unit 2406 represent a cache memory hierarchy within processor 2400. In at least one embodiment, cache memory units 2404A - 2404N can include at least one level of instruction and data cache within each processor core, and one or more levels of shared intermediate level caches such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence among the various cache units 2406 and 2404A - 2404N.
[0239] In at least one embodiment, the processor 2400 may also include a set of one or more bus controller units 2416 and a system agent core 2410. In at least one embodiment, the one or more bus controller units 2416 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 2410 provides management functions for various processor components. In at least one embodiment, the system agent core 2410 includes one or more integrated memory controllers 2414 for managing access to various external memory devices (not shown).
[0240] In at least one embodiment, one or more of the processor cores 2402A - 2402N include support for simultaneous multithreading. In at least one embodiment, the system agent core 2410 includes components for coordinating and operating the cores 2402A - 2402N during multithreaded processing. In at least one embodiment, the system agent core 2410 may further include a power control unit (PCU) that includes logic and components for adjusting the power state of one or more of the processor cores 2402A - 2402N and the graphics processor 2408.
[0241] In at least one embodiment, the processor 2400 further includes a graphics processor 2408 for performing graphics processing operations. In at least one embodiment, the graphics processor 2408 is coupled to a shared cache unit 2406 and a system agent core 2410 including one or more integrated memory controllers 2414. In at least one embodiment, the system agent core 2410 also includes a display controller 2411 for outputting the output of the graphics processor towards one or more coupled displays. In at least one embodiment, the display controller 2411 may also be a separate module coupled to the graphics processor 2408 via at least one interconnect, or may be integrated within the graphics processor 2408.
[0242] In at least one embodiment, a ring-based interconnect unit 2412 is used to couple the internal components of the processor 2400. In at least one embodiment, alternative interconnect units such as point-to-point interconnects, switch interconnects, or other techniques may be used. In at least one embodiment, the graphics processor 2408 is coupled to the ring interconnect 2412 via an I / O link 2413.
[0243] In at least one embodiment, the I / O link 2413 represents at least one of a variety of I / O interconnects including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 2418 such as an eDRAM module. In at least one embodiment, each of the processor cores 2402A - 2402N and the graphics processor 2408 use the embedded memory module 2418 as a shared last-level cache.
[0244] In at least one embodiment, the processor cores 2402A - 2402N are of the same type executing a common instruction set architecture. In at least one embodiment, the processor cores 2402A - 2402N are heterogeneous from the perspective of an instruction set architecture (ISA), where one or more of the processor cores 2402A - 2402N execute a common instruction set, but one or more other cores of the processor cores 2402A - 2402N execute a subset of the common instruction set, or a different instruction set. In at least one embodiment, the processor cores 2402A - 2402N are heterogeneous from the perspective of a microarchitecture, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, the processor 2400 can be implemented on one or more chips, or implemented as a SoC integrated circuit.
[0245] In order to perform inference and / or training operations related to one or more embodiments, the inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the processor 2400. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in the graphics processor 2312, the graphics cores 2402A - 2402N, or other components of FIG. 24. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIGS. 6A or 6B. In at least one embodiment, the weight parameters may be stored in on - chip or off - chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 2400 for executing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0246] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0247] FIG. 25 is a block diagram of the hardware logic of a graphics processor core 2500 according to at least one embodiment described herein. In at least one embodiment, the graphics processor core 2500 is included within a graphics core array. In at least one embodiment, the graphics processor core 2500, which may also be referred to as a core slice, can be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 2500 is an example of one graphics core slice, and the graphics processor described herein may include multiple graphics core slices based on the target power and performance envelope. In at least one embodiment, each graphics core 2500 can include a fixed function block 2530 coupled to a plurality of sub-cores 2501A - 2501F, also referred to as sub-slices, which include modular blocks of general purpose and fixed function logic.
[0248] In at least one embodiment, the fixed function block 2530 includes a geometry / fixed function pipeline 2536 that can be shared by all sub-cores within the graphics processor 2500, for example, in a low performance and / or low power graphics processor embodiment. In at least one embodiment, the geometry / fixed function pipeline 2536 includes a 3D fixed function pipeline, a video front end unit, a thread spawner and thread dispatcher, and an integrated return buffer manager that manages the integrated return buffer.
[0249] In at least one embodiment, the fixed function block 2530 also includes a graphics SoC interface 2537, a graphics microcontroller 2538, and a media pipeline 2539. In at least one embodiment, the fixed graphics SoC interface 2537 provides an interface between the graphics core 2500 and other processor cores within the system-on-chip integrated circuit. In at least one embodiment, the graphics microcontroller 2538 is a programmable sub-processor configurable to manage various functions of the graphics processor 2500, including thread dispatch, scheduling, and preemption. In at least one embodiment, the media pipeline 2539 includes logic to facilitate decoding, encoding, preprocessing, and / or postprocessing of multimedia data, including image and video data. In at least one embodiment, the media pipeline 2539 performs media operations via requests to compute logic or sampling logic within sub-cores 2501-2501F.
[0250] In at least one embodiment, the SoC interface 2537 enables the graphics core 2500 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, and other components within the SoC include memory hierarchy elements such as a shared last-level cache memory, system RAM, and / or embedded on-chip or on-package DRAM. In at least one embodiment, the SoC interface 2537 also enables communication with fixed-function devices within the SoC, such as a camera imaging pipeline, enables and / or implements the use of global memory atomics that can be shared between the graphics core 2500 and the CPU within the SoC. In at least one embodiment, the SoC interface 2537 can also implement power management control of the graphics core 2500 and interface between the clock domain of the graphics core 2500 and other clock domains within the SoC. In at least one embodiment, the SoC interface 2537 is configured to receive a command buffer from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, the commands and instructions can be dispatched to the media pipeline 2539 when a media operation is executed, or to a geometry and fixed-function pipeline (e.g., geometry and fixed-function pipeline 2536, geometry and fixed-function pipeline 2514) when a graphics processing operation is executed.
[0251] In at least one embodiment, the graphics microcontroller 2538 can be configured to perform various scheduling and management tasks for the graphics core 2500. In at least one embodiment, the graphics microcontroller 2538 can execute graphics and / or compute workload scheduling in the execution unit (EU) arrays 2502A - 2502F, 2504A - 2504F within the sub - cores 2501A - 2501F. In at least one embodiment, the host software running on the CPU core of the SoC including the graphics core 2500 can send a workload to one of the multiple graphics processor doorbells, and this path invokes the scheduling operation for the appropriate graphics engine. In at least one embodiment, the scheduling operation includes determining which workload to execute next, sending the workload to the command streamer, pre - empting existing workloads running on the engine, managing the progress of the workload, and notifying the host software when the workload is completed. In at least one embodiment, the graphics microcontroller 2538 can also facilitate the low - power or idle state of the graphics core 2500 and provide the graphics core 2500 with the function of saving and restoring registers within the graphics core 2500 throughout the transition to the low - power state, independent of the operating system and / or the graphics driver software on the system.
[0252] In at least one embodiment, the graphics core 2500 may have more than, or less than, the illustrated sub-cores 2501A - 2501F, up to N modular sub-cores. For each set of N sub-cores, in at least one embodiment, the graphics core 2500 may also include shared function logic 2510, shared and / or cache memory 2512, geometry / fixed function pipeline 2514, and additional fixed function logic 2516 for accelerating various graphics and computing processing operations. In at least one embodiment, the shared function logic 2510 may include logical units (e.g., sampler, math, and / or inter-thread communication logic) that can be shared by each of the N sub-cores within the graphics core 2500. In at least one embodiment, the fixed shared and / or cache memory 2512 can serve as the last-level cache for the N sub-cores 2501A - 2501F within the graphics core 2500, and can also act as shared memory accessible by multiple sub-cores. In at least one embodiment, the geometry / fixed function pipeline 2514 may be included instead of the geometry / fixed function pipeline 2536 within the fixed function block 2530, and can include the same or similar logical units.
[0253] In at least one embodiment, the graphics core 2500 includes additional fixed function logic 2516 that can include various fixed function acceleration logic for use by the graphics core 2500. In at least one embodiment, the additional fixed function logic 2516 includes an additional geometry pipeline for use in position only shading. In position only shading, there are at least two geometry pipelines, but in the full geometry pipeline and the cull pipeline within the geometry / fixed function pipelines 2516, 2536, and this cull pipeline is an additional geometry pipeline that may be included within the additional fixed function logic 2516. In at least one embodiment, the cull pipeline is a scaled down version of the full geometry pipeline. In at least one embodiment, the full pipeline and the cull pipeline can execute different instances of an application, and each instance has a separate context. In at least one embodiment, position only shading can hide long cull runs of discarded triangles and can complete shading faster in some instances. For example, in at least one embodiment, the cull pipeline fetches and shades the vertex position attributes without rasterizing and rendering the pixels to the frame buffer, so the cull pipeline logic within the additional fixed function logic 2516 can execute the position shader in parallel with the main application and generate critical results overall faster than the full pipeline. In at least one embodiment, the cull pipeline can use the generated critical results to compute visibility information for all triangles, regardless of whether these triangles are culled. In at least one embodiment, the full pipeline (which may be called the replay pipeline in this instance) can consume the visibility information, skip the culled triangles, and shade only the visible triangles, and these visible triangles are ultimately passed to the rasterization phase.
[0254] In at least one embodiment, the additional fixed function logic 2516 can also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for embodiments that include optimization of machine learning training or inference.
[0255] In at least one embodiment, within each of the graphics sub-cores 2501A - 2501F, a set of execution resources is included, and this set may be used to execute graphics operations, media operations, and compute operations in response to requests from a graphics pipeline, a media pipeline, or a shader program. In at least one embodiment, the graphics sub-cores 2501A - 2501F include a plurality of EU arrays 2502A - 2502F, 2504A - 2504F, thread dispatch and inter-thread communication (TD / IC) logic 2503A - 2503F, 3D (e.g., texture) samplers 2505A - 2505F, media samplers 2506A - 2506F, shader processors 2507A - 2507F, and shared local memory (SLM) 2508A - 2508F. Each of the EU arrays 2502A - 2502F, 2504A - 2504F includes a plurality of execution units, and these are general-purpose graphics processing units that can perform floating-point and integer / fixed-point logical operations in the service of graphics operations, media operations, or compute operations including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 2503A - 2503F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between threads executing on the execution units of the sub-core. In at least one embodiment, the 3D samplers 2505A - 2505F can read texture or other 3D graphics-related data into memory. In at least one embodiment, the 3D sampler can read texture data in different ways based on the configured sample state and texture format associated with a given texture. In at least one embodiment, the media samplers 2506A - 2506F can perform similar read operations based on the type and format associated with media data.In at least one embodiment, each of the graphics sub-cores 2501A-2501F can alternatively include an integrated sampler for 3D and media. In at least one embodiment, threads executing on execution units within each sub-core 2501A-2501F can utilize the shared local memories 2508A-2508F within each sub-core so that threads executing within a thread group can execute using a common pool of on-chip memory.
[0256] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the graphics processor 2510. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in the graphics processor 2312, the graphics microcontroller 2538, the geometry and fixed function pipelines 2514 and 2536, or other logic of FIG. 24. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIGS. 6A or 6B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 2500 for performing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0257] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0258] Figures 26A-26B illustrate thread execution logic 2600 including an array of processing elements of a graphics processor core, according to at least one embodiment. FIG. 26A illustrates at least one embodiment in which the thread execution logic 2600 is used. FIG. 26B is a diagram illustrating exemplary inner details of execution units, according to at least one embodiment.
[0259] As shown in FIG. 26A, in at least one embodiment, the thread execution logic 2600 includes a shader processor 2602, a thread dispatcher 2604, an instruction cache 2606, a scalable execution unit array including a plurality of execution units 2608A-2608N, a sampler 2610, a data cache 2612, and a data port 2614. In at least one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any of execution units 2608A, 2608B, 2608C, 2608D-2608N-1, and 2608N), for example, based on the computational requirements of a workload. In at least one embodiment, the scalable execution units are interconnected via an interconnect fabric linked to each of the execution units. In at least one embodiment, the thread execution logic 2600 includes one or more connections to memory, such as system memory or cache memory, via the instruction cache 2606, the data port 2614, the sampler 2610, and one or more of the execution units 2608A-2608N. In at least one embodiment, each execution unit (e.g., 2608A) is a stand-alone programmable general-purpose computing unit capable of executing a plurality of simultaneous hardware threads while processing a plurality of data elements in parallel for each thread. In at least one embodiment, the array of execution units 2608A-2608N is scalable to include any number of individual execution units.
[0260] In at least one embodiment, execution units 2608A - 2608N are primarily used to execute shader programs. In at least one embodiment, shader processor 2602 can process various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 2604. In at least one embodiment, thread dispatcher 2604 arbitrates thread start requests from the graphics and media pipeline and includes logic for instantiating the requested threads on one or more of execution units 2608A - 2608N. For example, in at least one embodiment, the geometry pipeline can dispatch a vertex shader, a mosaic shader, or a geometry shader to the thread execution logic for processing. In at least one embodiment, thread dispatcher 2604 can also process runtime thread spawning requests from the executing shader program.
[0261] In at least one embodiment, execution units 2608A - 2608N support an instruction set that includes native support for many standard 3D graphics shader instructions, whereby shader programs from graphics libraries (e.g., Direct3D and OpenGL) are executed with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders). In at least one embodiment, each of execution units 2608A - 2608N, which includes one or more arithmetic - logic units (ALUs), can issue multiple single - instruction multiple - data (SIMD) executions, enabling an efficient execution environment despite high memory - access latency through multithreaded operation. In at least one embodiment, each hardware thread within each execution unit has a dedicated high - bandwidth register file and associated independent thread state. In at least one embodiment, execution is issued multiple times per clock to a pipeline that can perform integer arithmetic, single - precision and double - precision floating - point arithmetic, SIMD branch performance, logical operations, transcendental operations, and various other operations. In at least one embodiment, while waiting for data from memory or one of the shared functions, the dependent logic within execution units 2608A - 2608N puts the waiting threads into a sleep state until the requested data is returned. In at least one embodiment, the hardware resources may be dedicated to the processing of other threads while the waiting threads are in the sleep state. For example, in at least one embodiment, during the latency associated with vertex - shader operation, the execution unit can execute a different type of shader program, including a pixel shader, a fragment shader, or a different vertex shader.
[0262] In at least one embodiment, each of execution units 2608A - 2608N operates on an array of data elements. In at least one embodiment, the number of data elements is the "execution size" or the number of channels for an instruction. In at least one embodiment, an execution channel is a logical unit of execution related to access, masking, and flow control within an instruction for data elements. In at least one embodiment, the number of channels may be independent of the number of physical arithmetic logic units (ALUs) or floating - point units (FPUs) for a particular graphics processor. In at least one embodiment, execution units 2608A - 2608N may support integer and floating - point data types.
[0263] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements may be stored in registers as a packed data type, and the execution unit processes various elements based on the data size of the elements. For example, in at least one embodiment, when operating on a 256 - bit - wide vector, the 256 bits of the vector are stored in a register, and the execution unit operates on the vector as 4 separate 64 - bit packed data elements (data elements of quad - word (QW) size), 8 separate 32 - bit packed data elements (data elements of double - word (DW) size), 16 separate 16 - bit packed data elements (data elements of word (W) size), or 32 separate 8 - bit data elements (data elements of byte (B) size). However, in at least one embodiment, different vector widths and register sizes are possible.
[0264] In at least one embodiment, one or more execution units can be combined to form fused execution units 2609A - 2609N having thread control logic (2607A - 2607N) common to the fused EUs. In at least one embodiment, multiple EUs can be fused to form an EU group. In at least one embodiment, each EU in the fused EU group can be configured to execute separate SIMD hardware threads. The number of EUs in the fused EU group may vary according to various embodiments. In at least one embodiment, various SIMD widths including, but not limited to, SIMD8, SIMD16, and SIMD32 can be executed per EU. In at least one embodiment, each fused graphics execution unit 2609A - 2609N includes at least two execution units. For example, in at least one embodiment, the fused execution unit 2609A includes a first EU 2608A, a second EU 2608B, and thread control logic 2607A common to the first EU 2608A and the second EU 2608B. In at least one embodiment, the thread control logic 2607A controls the threads executed in the fused graphics execution unit 2609A to enable each EU within the fused execution units 2609A - 2609N to be executed using a common instruction pointer register.
[0265] In at least one embodiment, one or more internal instruction caches (e.g., 2606) are included in the thread execution logic 2600 to cache thread instructions for the execution units. In at least one embodiment, one or more data caches (e.g., 2612) are included to cache thread data during thread execution. In at least one embodiment, the sampler 2610 is included to perform texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, the sampler 2610 includes special texture or media sampling functions and processes texture or media data during sampling before providing the sampled data to the execution units.
[0266] During execution, in at least one embodiment, the graphics and media pipeline sends thread start requests to thread execution logic 2600 via thread spawning and dispatch logic. In at least one embodiment, when a group of geometric objects is processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 2602 is called to further compute output information and write the results to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader computes the values of various vertex attributes that will be interpolated between rasterized objects. In at least one embodiment, the pixel processor logic within shader processor 2602 then executes a pixel shader program or fragment shader program with an application programming interface (API). In at least one embodiment, to execute the shader program, shader processor 2602 dispatches threads to execution units (e.g., 2608A) via thread dispatcher 2604. In at least one embodiment, shader processor 2602 uses the texture sampling logic of sampler 2610 to access the texture data of a texture map stored in memory. In at least one embodiment, through arithmetic operations on the texture data and input geometry data, the pixel color data for each geometry fragment is computed, or one or more pixels are discarded so that they are not further processed.
[0267] In at least one embodiment, data port 2614 provides a memory access mechanism for thread execution logic 2600 to output processed data to memory so that it can be further processed in the graphics processor output pipeline. In at least one embodiment, data port 2614 includes or is coupled to one or more cache memories (e.g., data cache 2612) to cache data for memory access through the data port.
[0268] As shown in FIG. 26B, in at least one embodiment, graphics execution unit 2608 can include an instruction fetch unit 2637, a general register file array (GRF) 2624, an architecture register file array (ARF) 2626, a thread arbiter 2622, a sending unit 2630, a branch unit 2632, a set of SIMD floating point units (FPUs) 2634, and in at least one embodiment, a set of dedicated integer SIMD ALUs 2635. In at least one embodiment, GRF 2624 and ARF 2626 include a set of general register files and architecture register files associated with each simultaneous hardware thread, and this hardware thread may be active in graphics execution unit 2608. In at least one embodiment, the architecture state per thread is maintained in ARF 2626, and the data used during thread execution is stored in GRF 2624. In at least one embodiment, the execution state of each thread, including the instruction pointer for each thread, can be held in the thread-specific registers of ARF 2626.
[0269] In at least one embodiment, the graphics execution unit 2608 has an architecture that is a combination of simultaneous multi-threading (SMT) and interleaved multi-threading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, where the resources of the execution unit are divided across the logic used to execute multiple simultaneous threads.
[0270] In at least one embodiment, the graphics execution unit 2608 can co-issue multiple instructions, which may each be different instructions. In at least one embodiment, the thread arbiter 2622 of the graphics execution unit thread 2608 can enable instructions to be dispatched for execution to one of the transmission unit 2630, the branch unit 2642, or the SIMD FPU 2634. In at least one embodiment, each execution thread can access 128 general-purpose registers in the GRF 2624, where each register can store 32 bytes that are accessible as a vector of SIMD8 elements of 32-bit data elements. In at least one embodiment, each execution unit thread can access 4K bytes in the GRF 2624, but the embodiments are not so limited, and more or fewer resources may be provided in other embodiments. In at least one embodiment, up to seven threads can be executed simultaneously, but the number of threads per execution unit can also vary according to the embodiment. In at least one embodiment where seven threads can access 4K bytes, the GRF 2624 can store a total of 28K bytes. In at least one embodiment, a flexible addressing mode can enable multiple registers to be addressed together to construct wider registers or represent strided rectangular block data structures.
[0271] In at least one embodiment, memory operations, sampler operations, and other long-latency system communications are dispatched via "send" instructions executed by the message delivery sending unit 2630. In at least one embodiment, branch instructions are dispatched to a dedicated branch unit 2632 to facilitate SIMD divergence and eventual convergence.
[0272] In at least one embodiment, the graphics execution unit 2608 includes one or more SIMD floating point units (FPUs) 2634 for performing floating point operations. In at least one embodiment, the FPU 2634 also supports integer calculations. In at least one embodiment, the FPU 2634 can perform up to M 32-bit floating point (or integer) operations in SIMD, or up to 2M 16-bit integer operations, or 16-bit floating point operations in SIMD. In at least one embodiment, at least one of the FPUs provides extended mathematical functionality to support high-throughput transcendental mathematical functions and double-precision 64-bit floating point. In at least one embodiment, a set of 8-bit integer SIMD ALUs 2635 may also be present and may be specifically optimized to perform operations related to machine learning calculations.
[0273] In at least one embodiment, an array of multiple instances of the graphics execution unit 2608 may be instantiated in a graphics sub-core group (e.g., a sub-slice). In at least one embodiment, the execution unit 2608 can execute instructions across multiple execution channels. In at least one embodiment, each thread executed by the graphics execution unit 2608 is executed on a different channel.
[0274] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into execution logic 2600. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that shown in FIGS. 6A or 6B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the execution logic 2600 to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0275] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0276] FIG. 27 shows a parallel processing unit (PPU) 2700 according to at least one embodiment. In at least one embodiment, the PPU 2700 is composed of machine-readable code that causes the PPU 2700 to execute some or all of the processes and techniques described throughout this disclosure when executed by the PPU 2700. In at least one embodiment, the PPU 2700 is a multi-threaded processor, and this processor is implemented on one or more integrated circuit devices and utilizes multi-threading as a latency hiding technique designed to process computer-readable instructions (also called machine-readable instructions or simply instructions) in parallel across multiple threads. In at least one embodiment, a thread refers to a thread of execution and is an instantiation of a set of instructions configured to be executed by the PPU 2700. In at least one embodiment, the PPU 2700 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data in order to generate two-dimensional (2D) image data that can be displayed on a display device such as a liquid crystal display (LCD) device. In at least one embodiment, the PPU 2700 is utilized to perform calculations such as linear algebra operations and machine learning operations. FIG. 27 shows an exemplary parallel processor for illustrative purposes only and should be interpreted as a non-limiting example of a processor architecture contemplated within the scope of this disclosure, and it should be interpreted that any suitable processor may be utilized to add to and / or replace the same processor.
[0277] In at least one embodiment, one or more PPU2700s are configured to accelerate applications in high performance computing ("HPC"), data centers, and machine learning. In at least one embodiment, PPU2700 is configured to accelerate deep learning systems and applications including the following non-limiting examples: autonomous vehicle platforms, deep learning, high-precision audio, images, text recognition systems, intelligent video analysis, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, etc.
[0278] In at least one embodiment, PPU2700 includes, without limitation, an input / output ("I / O") unit 2706, a front-end unit 2710, a scheduler unit 2712, a work distribution unit 2714, a hub 2716, a crossbar ("Xbar") 2720, one or more general processing clusters ("GPC") 2718, and one or more partition units ("memory partition units") 2722. In at least one embodiment, PPU2700 is connected to a host processor or another PPU2700 via one or more high-speed GPU interconnects ("GPU interconnects") 2708. In at least one embodiment, PPU2700 is connected to a host processor or other peripheral devices via interconnect 2702. In at least one embodiment, PPU2700 is connected to a local memory with one or more memory devices ("memory") 2704. In at least one embodiment, memory device 2704 includes, without limitation, one or more dynamic random access memory ("DRAM") devices. In at least one embodiment, one or more DRAM devices may be configured as, and / or be configurable as, a high bandwidth memory ("HBM") subsystem in which multiple DRAM dies are stacked within each device.
[0279] In at least one embodiment, the high-speed GPU interconnect 2708 may refer to a wired-based multi-lane communication link that is used by the system to scale and that includes one or more PPUs 2700 combined with one or more central processing units (“CPUs”), and supports cache coherence and CPU mastering between the PPU 2700 and the CPU. In at least one embodiment, data and / or commands are transmitted to / from another unit of the PPU 2700, such as one or more copy engines, video encoders, video decoders, power management units, and other components that may not be explicitly shown in FIG. 27, via the high-speed GPU interconnect 2708 and via the hub 2716.
[0280] In at least one embodiment, the I / O unit 2706 is configured to communicate (e.g., commands, data) to / from a host processor (not shown in FIG. 27) via the system bus 2702. In at least one embodiment, the I / O unit 2706 communicates with the host processor directly via the system bus 2702 or via one or more intermediate devices such as one or more memory bridges. In at least one embodiment, the I / O unit 2706 may communicate with one or more other processors such as one or more of the PPUs 2700 via the system bus 2702. In at least one embodiment, the I / O unit 2706 implements a peripheral component interconnect express (“PCIe”) interface to enable communication via a PCIe bus. In at least one embodiment, the I / O unit 2706 implements an interface for communicating with external devices.
[0281] In at least one embodiment, the I / O unit 2706 decodes packets received via the system bus 2702. In at least one embodiment, at least some of the packets represent commands configured to cause the PPU 2700 to perform various operations. In at least one embodiment, the I / O unit 2706 transmits the decoded commands to various other units of the PPU 2700 specified by the commands. In at least one embodiment, the commands are transmitted to the front-end unit 2710 and / or to the hub 2716 or to other units of the PPU 2700 such as one or more copy engines, video encoders, video decoders, power management units (not explicitly shown in FIG. 27). In at least one embodiment, the I / O unit 2706 is configured to route communications between various logical units of the PPU 2700.
[0282] In at least one embodiment, a program executed by the host processor encodes a command stream in a buffer that enables the PPU 2700 to be provided with a workload for processing. In at least one embodiment, the workload includes instructions and data to be processed by these instructions. In at least one embodiment, the buffer is an area within a memory that is accessible (e.g., writable / readable) by both the host processor and the PPU 2700, and the host interface unit may be configured to access the buffer in the system memory connected to the system bus 2702 via memory requests transmitted via the system bus 2702 by the I / O unit 2706. In at least one embodiment, the host processor writes the command stream to the buffer and then transmits a pointer indicating the starting point of the command stream to the PPU 2700, whereby the front-end unit 2710 receives a pointer indicating one or more command streams, manages the one or more command streams, reads commands from the command streams, and transfers the commands to various units of the PPU 2700.
[0283] In at least one embodiment, the front-end unit 2710 is coupled to a scheduler unit 2712 that configures various GPCs 2718 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 2712 is configured to track state information associated with various tasks managed by the scheduler unit 2712, where the state information may indicate, for example, which GPC 2718 a task is assigned to, whether the task is active or inactive, the priority level associated with the task, and so on. In at least one embodiment, the scheduler unit 2712 manages the execution of multiple tasks in one or more of the GPCs 2718.
[0284] In at least one embodiment, the scheduler unit 2712 is coupled to a work distribution unit 2714 configured to dispatch tasks for execution on the GPCs 2718. In at least one embodiment, the work distribution unit 2714 tracks the number of scheduled tasks received from the scheduler unit 2712, and the work distribution unit 2714 manages a pending task pool and an active task pool for each of the GPCs 2718. In at least one embodiment, the pending task pool comprises a number of slots (e.g., 32 slots) for tasks assigned to be processed by a particular GPC 2718, and the active task pool comprises a number of slots (e.g., 4 slots) for tasks being actively processed by the GPC 2718, such that when one of the GPCs 2718 completes execution of a task, that task is removed from the active task pool of the GPC 2718 and one of the other tasks from the pending task pool is selected and scheduled for execution on the GPC 2718. In at least one embodiment, if an active task is idle on the GPC 2718, such as while waiting for data dependencies to be resolved, the active task is removed from the GPC 2718 and returned to the pending task pool, during which another task from the pending task pool is selected and scheduled for execution on the GPC 2718.
[0285] In at least one embodiment, the work distribution unit 2714 communicates with one or more GPCs 2718 via an X-bar 2720. In at least one embodiment, the X-bar 2720 is an interconnect network that couples many of the units of the PPU 2700 to other units of the PPU 2700 and can be configured to couple the work distribution unit 2714 to a particular GPC 2718. In at least one embodiment, one or more other units of the PPU 2700 may also be connected to the X-bar 2720 via a hub 2716.
[0286] In at least one embodiment, the task is managed by a scheduler unit 2712 and dispatched to one of the GPCs 2718 by a work distribution unit 2714. The GPC 2718 is configured to process the task and generate a result. In at least one embodiment, the result may be consumed by other tasks within the GPC 2718, routed to a different GPC 2718 via the crossbar 2720, or stored in the memory 2704. In at least one embodiment, the result can be written to the memory 2704 via a partition unit 2722, which implements a memory interface for reading and writing data to / from the memory 2704. In at least one embodiment, the result can be sent to another PPU 2704 or CPU via the high-speed GPU interconnect 2708. In at least one embodiment, the PPU 2700 includes, without limitation, U partition units 2722 equal in number to the number of separate individual memory devices 2704 coupled to the PPU 2700. In at least one embodiment, the partition unit 2722 is further described in more detail below in conjunction with FIG. 29.
[0287] In at least one embodiment, the host processor runs a driver kernel that implements an application programming interface (API) that enables scheduling of operations for one or more applications running on the host processor to be executed on the PPU 2700. In at least one embodiment, multiple compute applications are executed simultaneously by the PPU 2700, and the PPU 2700 provides isolation, quality of service (“QoS”), and independent address spaces for the multiple compute applications. In at least one embodiment, an application generates instructions (e.g., in the form of API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 2700, and the driver kernel outputs the tasks to one or more streams being processed by the PPU 2700. In at least one embodiment, each task comprises one or more groups of related threads, which may be referred to as warps. In at least one embodiment, a warp comprises multiple related threads (e.g., 32 threads) that can be executed in parallel. In at least one embodiment, cooperating threads may refer to multiple threads that contain instructions for executing a task and exchange data via shared memory. In at least one embodiment, threads and cooperating threads are described in further detail by at least one embodiment in conjunction with FIG. 29.
[0288] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to PPU 2700. In at least one embodiment, PPU 2700 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by PPU 2700. In at least one embodiment, PPU 2700 may be used to execute one or more use cases of the one or more neural networks described herein.
[0289] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0290] FIG. 28 shows a general-purpose processing cluster (“GPC”) 2800 according to at least one embodiment. In at least one embodiment, GPC 2800 is GPC 2718 of FIG. 27. In at least one embodiment, each GPC 2800 includes, without limitation, several hardware units for processing tasks, and each GPC 2800 includes, without limitation, a pipeline manager 2802, a pre-raster operations unit (“PROP”) 2804, a raster engine 2808, a work distribution crossbar (“WDX”) 2816, a memory management unit (“MMU”) 2818, one or more data processing clusters (“DPC”) 2806, and any suitable combination of parts.
[0291] In at least one embodiment, the operation of the GPC2800 is controlled by a pipeline manager 2802. In at least one embodiment, the pipeline manager 2802 manages the configuration of one or more DPCs 2806 to process tasks allocated to the GPC2800. In at least one embodiment, the pipeline manager 2802 configures at least one of the one or more DPCs 2806 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, the DPC 2806 is configured to execute a vertex shader program on a programmable streaming multi-processor ("SM") 2814. In at least one embodiment, the pipeline manager 2802 is configured to route packets received from a work distribution unit to appropriate logical units within the GPC2800, and some packets may be routed to the fixed function hardware unit of the PROP2804 and / or the raster engine 2808, and other packets may be routed to the DPC 2806 to be processed by the primitive engine 2812 or the SM 2814. In at least one embodiment, the pipeline manager 2802 configures at least one of the DPCs 2806 to implement a neural network model and / or a computing pipeline.
[0292] In at least one embodiment, the PROP unit 2804 is configured to route data generated by the raster engine 2808 and the DPC 2806 to the raster operation (“ROP”) unit of the partition unit 2722, which was described in more detail above in conjunction with FIG. 27. In at least one embodiment, the PROP unit 2804 is configured to perform optimizations for color blending, organize pixel data, perform address translation, and perform other operations. In at least one embodiment, the raster engine 2808 includes, without limitation, several fixed-function hardware units configured to perform various raster operations in at least one embodiment, and the raster engine 2808 includes, without limitation, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile coalescing engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives the transformed vertices, generates plane equations associated with the geometric primitives defined by the vertices, and the plane equations are sent to the coarse raster engine to generate coverage information for the primitives (e.g., the x, y coverage masks of the tiles), and the output of the coarse raster engine is sent to the culling engine, where fragments associated with primitives that failed the z-test are culled and sent to the clipping engine, where fragments outside the frustum are clipped. In at least one embodiment, the fragments that pass through clipping and culling are passed to the fine raster engine to generate attributes for the pixel fragments based on the plane equations generated by the setup engine. In at least one embodiment, the output of the raster engine 2808 includes fragments that will be processed by any suitable entity, such as by a fragment shader implemented within the DPC 2806.
[0293] In at least one embodiment, each DPC2806 included in GPC2800 includes, without limitation, an M pipe controller ("MPC": M-Pipe Controller) 2810, a primitive engine 2812, one or more SMs 2814, and any suitable combination thereof. In at least one embodiment, MPC2810 controls the operation of DPC2806 to route the packets received from pipeline manager 2802 to the appropriate units within DPC2806. In at least one embodiment, packets associated with vertices are routed to primitive engine 2812 configured to fetch vertex attributes associated with the vertices from memory, whereas, in contrast, packets associated with shader programs may be sent to SM2814.
[0294] In at least one embodiment, SM2814 includes, without limitation, a programmable streaming processor configured to process tasks represented by a number of threads. In at least one embodiment, SM2814 is multi-threaded and configured to simultaneously execute a plurality of threads (e.g., 32 threads) from a particular group of threads, implementing a single instruction multiple data (SIMD) architecture, where each thread within a group of threads (warp) is configured to process different data sets based on the same instruction set. In at least one embodiment, all threads within a thread group execute the same instruction. In at least one embodiment, SM2814 implements a single instruction multiple thread (SIMT) architecture, where each thread of a thread group is configured to process different data sets based on the same set of instructions, but individual threads within a thread group are allowed to diverge during execution. In at least one embodiment, the program counter, call stack, and execution state are maintained on a per-warp basis, enabling simultaneous processing between warps and serial execution within a warp when threads within a warp diverge. In another embodiment, the program counter, call stack, and execution state are maintained on a per-individual thread basis, enabling equal simultaneous processing between all threads, within a warp, and between warps. In at least one embodiment, the execution state is maintained on a per-individual thread basis, and threads executing the same instruction may be converged and executed in parallel for greater efficiency. At least one embodiment of SM2814 is described in further detail below.
[0295] In at least one embodiment, the MMU 2818 provides an interface between the GPC 2800 and a memory partition unit (e.g., partition unit 2722 of FIG. 27), and the MMU 2818 provides virtual address to physical address translation, memory protection, and mediation of memory requests. In at least one embodiment, the MMU 2818 provides one or more translation lookaside buffers ("TLBs") for performing translation of virtual addresses to physical addresses of memory.
[0296] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, a deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the GPC 2800. In at least one embodiment, the GPC 2800 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system, or by the GPC 2800 itself. In at least one embodiment, the GPC 2800 may be used to execute one or more use cases of the neural networks described herein.
[0297] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0298] FIG. 29 shows a memory partition unit 2900 of a parallel processing unit (``PPU'') according to at least one embodiment. In at least one embodiment, the memory partition unit 2900 includes, without limitation, a raster operation (``ROP'') unit 2902, a level 2 (``L2'') cache 2904, a memory interface 2906, and any suitable combination thereof. In at least one embodiment, the memory interface 2906 is coupled to the memory. In at least one embodiment, the memory interface 2906 may implement a 32, 64, 128, 1024-bit data bus, or a similar embodiment, for high-speed data transfer. In at least one embodiment, the PPU incorporates U memory interfaces 2906 into the partition unit 2900, one memory interface 2906 per pair of partition units 2900, where each pair of partition units 2900 is connected to a corresponding memory device. For example, in at least one embodiment, the PPU may be connected to up to Y memory devices, such as a high-bandwidth memory stack, or graphics double data rate, version 5, synchronous dynamic random access memory (``GDDR5 SDRAM'').
[0299] In at least one embodiment, the memory interface 2906 implements a second generation high bandwidth memory (“HBM2”) memory interface, and Y is equal to half of U. In at least one embodiment, the HBM2 memory stack is positioned in the same physical package as the PPU, realizing substantial power and area savings compared to a conventional GDDR5 SDRAM system. In at least one embodiment, each HBM2 stack includes, without limitation, four memory dies, Y is equal to 4, and each HBM2 stack includes a total of eight channels of two 128-bit channels per die and a data bus width of 1024 bits. In at least one embodiment, the memory supports a single-bit error correction two-bit error detection (“SECDED”) error correction code (“ECC”) to protect data. In at least one embodiment, the ECC provides higher reliability to compute applications that are vulnerable to data corruption.
[0300] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partition unit 2900 supports integrated memory to provide a single integrated virtual address space to the central processing unit (“CPU”) and PPU memory, enabling sharing of data between virtual memory systems. In at least one embodiment, the frequency of access by the PPU to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the pages more frequently. In at least one embodiment, the high-speed GPU interconnect 2708 supports an address translation service to enable the PPU to directly access the CPU's page table, realizing full access by the PPU to CPU memory.
[0301] In at least one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In at least one embodiment, the copy engine can generate a page error for an address that is not mapped to a page table, and then the memory partition unit 2900 maps the address to the page table in response to the page error, and then the copy engine executes the transfer. In at least one embodiment, the memory is pinned (e.g., made non-pageable) for multiple operations of the copy engine between multiple processors, reducing the substantially available memory. In at least one embodiment, in the case of a hardware page error, the address can be passed to the copy engine regardless of whether the memory page is resident, and the copy process is transparent.
[0302] According to at least one embodiment, data from the memory 2704 of FIG. 27 or other system memory is fetched by the memory partition unit 2900 and stored in the L2 cache 2904, which is located on-chip and shared among various GPCs. In at least one embodiment, each memory partition unit 2900 includes, without limitation, at least a portion of the L2 cache associated with the corresponding memory device. In at least one embodiment, a lower level of cache is implemented in various units within the GPC. In at least one embodiment, each of the SM2814s may implement a level 1 ("L1") cache, where the L1 cache is private memory dedicated to a particular SM2814, and data from the L2 cache 2904 is fetched and stored in each of the L1 caches for processing by the functional units of the SM2814. In at least one embodiment, the L2 cache 2904 is coupled to the memory interface 2906 and the X bar 2720.
[0303] In at least one embodiment, the ROP unit 2902 performs graphics raster operations related to pixel colors, such as color compression and pixel blending. The ROP unit 2902, in at least one embodiment, performs a depth test in conjunction with the raster engine 2808 to receive the depth of the sample location associated with the pixel fragment from the culling engine of the raster engine 2808. In at least one embodiment, the depth is tested against the corresponding depth in the depth buffer of the sample location associated with the fragment. In at least one embodiment, when the fragment passes the depth test of the sample location, the ROP unit 2902 updates the depth buffer and transmits the result of the depth test to the raster engine 2808. It will be appreciated that the number of partition units 2900 may be different from the number of GPCs, and thus each ROP unit 2902 may be coupled to each of the GPCs in at least one embodiment. In at least one embodiment, the ROP unit 2902 tracks packets received from different GPCs and determines to which of them to route the results generated by the ROP unit 2902 through the X bar 2720.
[0304] FIG. 30 shows a streaming multiprocessor (“SM”) 3000 according to at least one embodiment. In at least one embodiment, the SM 3000 is the SM 2814 of FIG. 28. In at least one embodiment, the SM 3000 includes, without limitation, an instruction cache 3002, one or more scheduler units 3004, a register file 3008, one or more processing cores (“cores”) 3010, one or more special function units (“SFU”: special function unit) 3012, one or more load / store units (“LSU” load / store unit) 3014, an interconnect network 3016, a shared memory / level 1 (“L1”) cache 3018, and / or any suitable combination thereof. In at least one embodiment, the work distribution unit dispatches tasks for execution in a general-purpose processing cluster (“GPC”) of a parallel processing unit (“PPU”), each task is distributed to a specific data processing cluster (“DPC”) within the GPC, and when the task is related to a shader program, the task is distributed to one of the SMs 3000. In at least one embodiment, the scheduler unit 3004 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks assigned to the SM 3000. In at least one embodiment, the scheduler unit 3004 schedules thread blocks so that they can be executed as warps of parallel threads, where each thread block is distributed to at least one warp. In at least one embodiment, each warp executes threads. In at least one embodiment, the scheduler unit 3004 manages a plurality of different thread blocks, distributes warps to different thread blocks, and then dispatches instructions from a plurality of different cooperating groups to various functional units (e.g., processing cores 3010, SFUs 3012, and LSUs 3014) during each clock cycle.
[0305] In at least one embodiment, a cooperative group refers to a programming model for organizing a group of communicating threads, which enables a developer to express the granularity at which threads communicate and allows for a richer and more efficient representation of parallel decomposition. In at least one embodiment, the cooperative launch API supports synchronization between thread blocks so as to execute parallel algorithms. In at least one embodiment, applications of conventional programming models provide a single simple construct for synchronizing cooperating threads, namely a barrier (e.g., the syncthreads() function) across all threads of a thread block. However, in at least one embodiment, a programmer may define thread groups smaller than the granularity of a thread block, synchronize within the defined groups, and enable higher performance, design flexibility, and software reuse in the form of a collective functional interface across the entire collective group. In at least one embodiment, the cooperative group enables a programmer to explicitly define groups of threads at the granularity of sub-blocks (i.e., the same size as a single thread) and at the granularity of multi-blocks, and to perform collective operations such as synchronization on the threads within the cooperative group. In at least one embodiment, the programming model supports clean composition across software boundaries, thereby enabling libraries and utility functions to synchronize safely within their local contexts without the need to make assumptions about convergence. In at least one embodiment, the primitives of the cooperative group enable a new pattern of cooperative parallelism that includes producer-consumer parallelism, opportunistic parallelism, and global synchronization across the grid of thread blocks without limiting them.
[0306] In at least one embodiment, the dispatch unit 3006 is configured to send instructions to one or more of the functional units, and the scheduler unit 3004 includes two dispatch units 3006 without limitation that enable two different instructions from the same warp to be dispatched during each clock cycle. In at least one embodiment, each scheduler unit 3004 includes a single dispatch unit 3006 or an additional dispatch unit 3006.
[0307] In at least one embodiment, each SM3000 includes, without limitation, a register file 3008 that provides a set of registers to the functional units of the SM3000 in at least one embodiment. In at least one embodiment, the register file 3008 is divided among each of the functional units such that each functional unit is allocated a dedicated portion of the register file 3008. In at least one embodiment, the register file 3008 is divided among different warps being executed by the SM3000, and the register file 3008 provides temporary storage for operands connected to the data paths of the functional units. In at least one embodiment, each SM3000 includes, without limitation, a plurality of L processing cores 3010. In at least one embodiment, each SM3000 includes, without limitation, a large number (e.g., 128 or more) of individual processing cores 3010. In at least one embodiment, each processing core 3010 includes, without limitation, a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit that includes, without limitation, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point operations. In at least one embodiment, the processing core 3010 includes, without limitation, 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.
[0308] The tensor core is configured to perform matrix operations according to at least one embodiment. In at least one embodiment, one or more tensor cores are included in the processing core 3010. In at least one embodiment, the tensor core is configured to perform deep learning matrix operations such as convolution operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4×4 matrix and performs a matrix multiply and accumulate operation D = A×B + C, where A, B, C, and D are 4×4 matrices.
[0309] In at least one embodiment, the input A and B of the matrix multiplication are 16-bit floating-point matrices, and the sum matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor core operates on 16-bit floating-point input data with a 32-bit floating-point sum. In at least one embodiment, the 16-bit floating-point multiplication uses 64 operations, resulting in a full-precision product, which is then added using 32-bit floating-point addition with other intermediate products of a 4×4×4 matrix multiplication. Using the tensor core, in at least one embodiment, much larger two-dimensional or even higher-dimensional matrix operations constructed from these small elements are performed. In at least one embodiment, an API such as the CUDA9 C++ API exposes special matrix load operations, matrix multiply and accumulate operations, and matrix store operations to efficiently use the tensor core from a CUDA-C++ program. In at least one embodiment, at the CUDA level, the warp-level interface assumes a matrix of size 16×16 across all 32 threads of a warp.
[0310] In at least one embodiment, each SM3000 includes, without limitation, M SFU3012s that execute special functions (such as attribute evaluation, reciprocal square root, etc.). In at least one embodiment, the SFU3012s include, without limitation, a tree traversal unit configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFU3012s include, without limitation, a texture unit configured to perform a filtering operation on a texture map. In at least one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texels) from a memory and a sample texture map and generate sampled texture values for use in a shader program executed by the SM3000. In at least one embodiment, the texture map is stored in the shared memory / L1 cache 3018. In at least one embodiment, the texture unit performs texture operations such as filtering operations using mip maps (e.g., texture maps with different levels of detail) according to at least one embodiment. In at least one embodiment, each SM3000 includes, without limitation, two texture units.
[0311] In at least one embodiment, each SM3000 includes, without limitation, N LSU3014s that perform load and store operations between the shared memory / L1 cache 3018 and the register file 3008. In at least one embodiment, each SM3000 includes, without limitation, an interconnect network 3016 that connects each of the functional units to the register file 3008 and connects the LSU3014 to the register file 3008, and the shared memory / L1 cache 3018. In at least one embodiment, the interconnect network 3016 is a crossbar, and this crossbar may be configured to connect any of the functional units to any of the registers of the register file 3008 and connect the LSU3014 to the register file 3008 and the memory locations of the shared memory / L1 cache 3018.
[0312] In at least one embodiment, the shared memory / L1 cache 3018 is, in at least one embodiment, an array of on-chip memory that enables data storage and communication between the SM3000 and the primitive engine, and between the threads of the SM3000. In at least one embodiment, the shared memory / L1 cache 3018 has, without limitation, a storage capacity of 128 KB and is on the path from the SM3000 to the partition unit. In at least one embodiment, the shared memory / L1 cache 3018 is, in at least one embodiment, used to cache reads and writes. In at least one embodiment, one or more of the shared memory / L1 cache 3018, the L2 cache, and the memory are auxiliary storage.
[0313] In at least one embodiment, by combining a data cache and a shared memory function into a single memory block, the performance for both types of memory access is improved. In at least one embodiment, the capacity can be used as, or is available as, a cache by programs that do not use the shared memory. Thus, when the shared memory is configured to use half of the capacity, texture and load / store operations can use the remaining capacity. According to at least one embodiment, by integrating into the shared memory / L1 cache 3018, the shared memory / L1 cache 3018 can function as a high-throughput pipe for streaming data while at the same time providing high-bandwidth and low-latency access to frequently reused data. In at least one embodiment, when configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function graphics processing unit is bypassed to create a much simpler programming model. In the general-purpose parallel computing configuration, the work distribution unit directly assigns and distributes thread blocks to the DPC in at least one embodiment. In at least one embodiment, the threads within a block execute the same program using unique thread IDs in the calculation so that each thread is guaranteed to generate a unique result, use the SM3000 to execute the program and perform the calculation, use the shared memory / L1 cache 3018 to communicate between threads, and use the LSU3014 to read from and write to the global memory via the shared memory / L1 cache 3018 and the memory partition unit. In at least one embodiment, when configured for general-purpose parallel computing, the SM3000 writes commands that can be used by the scheduler unit 3004 to launch new work on the DPC.
[0314] In at least one embodiment, the PPU is included in or coupled to a desktop computer, laptop computer, tablet computer, server, supercomputer, smartphone (e.g., a wireless portable device), personal digital assistant ("PDA"), digital camera, vehicle, head-mounted display, portable electronic device, etc. In at least one embodiment, the PPU is embodied on a single semiconductor substrate. In at least one embodiment, the PPU is included in a system-on-chip ("SoC") with one or more other devices such as an additional PPU, memory, reduced instruction set computer ("RISC") CPU, memory management unit ("MMU"), digital-to-analog converter ("DAC").
[0315] In at least one embodiment, the PPU may be included in a graphics card that includes one or more memory devices. The graphics card may be configured to interface with a PCIe slot on the motherboard of a desktop computer. In at least one embodiment, the PPU may be an integrated graphics processing unit ("iGPU") included in the motherboard chipset.
[0316] To perform inference and / or training operations related to one or more embodiments, inference and / or training logic 615 is used. Details regarding the inference and / or training logic 615 are provided below in conjunction with FIGS. 6A and / or 6B. In at least one embodiment, a deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the SM3000. In at least one embodiment, the SM3000 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the SM3000 itself. In at least one embodiment, the SM3000 may be used to execute one or more use cases of the one or more neural networks described herein.
[0317] To perform inference and / or training operations related to one or more embodiments, inference and / or training logic 615 is used. In at least one embodiment, this logic can be used with the components of these drawings to determine one or more pixel blending weights using one or more anisotropic filters.
[0318] At least one embodiment can be described in light of the following sections.
[0319] 1. A processor comprising: one or more circuits for determining one or more pixel blending weights using one or more anisotropic filters A processor comprising the same.
[0320] 2. The processor of item 1, wherein the one or more circuits are further for inferring one or more anisotropic filters based at least in part on one or more image features identified within input image data using one or more neural networks.
[0321] 3. The processor according to item 2, wherein one or more circuits are further for applying one or more anisotropic filters to one or more initial blending weights inferred by one or more neural networks to generate one or more pixel blending weights.
[0322] 4. The processor according to item 3, wherein one or more circuits are further for generating one or more output images by blending a current color value and a warping history color value at least partially according to one or more pixel blending weights.
[0323] 5. The processor according to item 4, wherein one or more circuits are further for upsampling a current image including a current color value using one or more anisotropic filters.
[0324] 6. The processor according to item 4, wherein a current image is to be generated by a rendering engine at a first resolution, and the rendering engine is for generating a set of motion vectors to be applied to history image data to generate warping history color values.
[0325] 7. A system, one or more processors for determining one or more pixel blending weights using one or mo...
Claims
Claim 1 A processor comprising: one or more circuits for using one or more neural networks to generate one or more anisotropic filters, thereby generating one or more pixel blending weights for blending two or more pixels of a previous video frame and a current upsampled video frame A processor having the same. Claim 2 The processor according to claim 1, wherein the one or more circuits are further configured to use the one or more neural networks to infer the one or more anisotropic filters based at least in part on one or more image features identified in input image data. Claim 3 The processor according to claim 2, wherein the one or more circuits are further configured to apply the one or more anisotropic filters to one or more initial blending weights inferred by the one or more neural networks to generate the one or more pixel blending weights. Claim 4 The processor according to claim 3, wherein the one or more circuits are further configured to generate one or more output images by blending a current color value and a warping history color value at least in part according to the one or more pixel blending weights. Claim 5 The processor according to claim 4, wherein the one or more circuits are further configured to upsample a current image including the current color value using the one or more anisotropic filters. Claim 6 The processor according to claim 4, wherein a current image is generated by a rendering engine at a first resolution, and the rendering engine is configured to generate a set of motion vectors applied to history image data to generate the warping history color value. Claim 7 A system comprising: One or more processors for using one or more neural networks to generate one or more anisotropic filters, whereby one or more pixel blending weights are generated to blend two or more pixels of a previous video frame and a current upsampled video frame A system having the same. **Claim 8** The system according to claim 7, wherein the one or more processors are further configured to use the one or more neural networks to infer the one or more anisotropic filters based at least in part on one or more image features identified in the input image data. **Claim 9** The system according to claim 8, wherein the one or more processors are further configured to apply the one or more anisotropic filters to one or more initial blending weights inferred by the one or more neural networks to generate the one or more pixel blending weights. **Claim 10** The system according to claim 9, wherein the one or more circuits are further configured to generate one or more output images by blending a current color value and a warping history color value at least in part according to the one or more pixel blending weights. **Claim 11** The system according to claim 10, wherein the one or more circuits are further configured to upsample a current image including the current color value using the one or more anisotropic filters. **Claim 12** The system according to claim 10, wherein a current image is generated by a rendering engine at a first resolution, and the rendering engine is configured to generate a set of motion vectors applied to history image data to generate the warping history color value. **Claim 13** A non-transitory machine-readable medium storing a set of instructions that, when executed by one or more processors, cause the one or more processors to at least A non - transitory machine - readable medium that causes one or more neural networks to generate one or more anisotropic filters, thereby generating one or more pixel - blending weights to blend two or more pixels of a previous video frame and a current upsampled video frame. **Claim 14** When executed, the instructions further cause the one or more processors to use the one or more neural networks to infer the one or more anisotropic filters, at least in part based on one or more image features identified within input image data. The non - transitory machine - readable medium according to claim 13. **Claim 15** When executed, the instructions further cause the one or more processors to apply the one or more anisotropic filters to one or more initial blending weights inferred by the one or more neural networks to generate the one or more pixel - blending weights. The non - transitory machine - readable medium according to claim 14. **Claim 16** When executed, the instructions further cause the one or more processors to generate one or more output images by blending a current color value and a warping - history color value, at least in part according to the one or more pixel - blending weights. The non - transitory machine - readable medium according to claim 15. **Claim 17** When executed, the instructions further cause the one or more processors to upsample a current image including the current color value using the one or more anisotropic filters. The non - transitory machine - readable medium according to claim 16. **Claim 18** The current image is adapted to be generated by a rendering engine at a first resolution, and the rendering engine is adapted to generate a set of motion vectors applied to history image data to generate the warping - history color values. The non - transitory machine - readable medium according to claim 16. **Claim 19** An image reconstruction system, One or more processors for using one or more neural networks to generate one or more anisotropic filters, whereby one or more pixel blending weights are generated to blend two or more pixels of a previous video frame and a current upsampled video frame A memory for storing the one or more anisotropic filters An image reconstruction system having the same. **Claim 20** The image reconstruction system according to claim 19, wherein the one or more processors are further configured to use the one or more neural networks to infer the one or more anisotropic filters based at least in part on one or more image features identified in input image data. **Claim 21** The image reconstruction system according to claim 20, wherein the one or more processors are further configured to apply the one or more anisotropic filters to one or more initial blending weights inferred by the one or more neural networks to generate the one or more pixel blending weights. **Claim 22** The image reconstruction system according to claim 21, wherein the one or more processors are further configured to generate one or more output images by blending a current color value and a warping history color value according at least in part to the one or more pixel blending weights. **Claim 23** The image reconstruction system according to claim 22, wherein the one or more processors are further configured to upsample a current image including the current color value using the one or more anisotropic filters. **Claim 24** The image reconstruction system according to claim 22, wherein a current image is generated by a rendering engine at a first resolution, and the rendering engine is configured to generate a set of motion vectors applied to history image data to generate the warping history color value. **Claim 25** The processor according to claim 1, wherein the previous video frame includes a previous upsampled video frame. **Claim 26** The processor according to claim 1, wherein different anisotropic filters among the one or more anisotropic filters are used for different pixels.
27. The processor according to claim 1, wherein the previous video frame is consecutive to the current upsampled video frame.
Citation Information
Patent Citations
Method and system for filtering medical image
CN101877124A
Method and system for transforming graphical object to image chunk and combining image layers into display image
JP2004046886A
Image processor, image processing method and program
JP2012213116A
Image processing apparatus, method, and program
JP2016051188A
Image processing apparatus, image processing method and image processing program
JP2020017229A