Pixel generation techniques

The processor uses prefixes and technology to identify pixels in the polygon, which solves the problem of insufficient computing resources during the rasterization process and improves processing efficiency and performance.

CN120339033APending Publication Date: 2025-07-18NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510079081.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-17
Filing Date
2025-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art requires a large amount of computing resources during the rasterization process, which leads to difficulty in improving performance.

Method used

By using a processor to identify pixels within a polygon based on prefixes and techniques, reducing operands and improving efficiency, including calculating windings and gradient information to identify whether the pixels are within a polygon.

Benefits of technology

It improves the performance of the rasterization process, reduces the operand and time of the processor, and improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339033A_ABST
    Figure CN120339033A_ABST
Patent Text Reader

Abstract

The invention relates to pixel generation techniques. Apparatus, systems, and techniques for converting polygon data to pixels. In at least one embodiment, the processor converts polygonal input data into pixels by using one or more prefix sums to generate one or more revolutions to indicate whether the pixels are within the polygon.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment relates to processing resources for converting data to pixels during rasterization. For example, one or more processors including one or more circuits convert vector data to pixels at least in part based on performing a prefix sum. Background Art

[0002] Rasterization uses a large amount of computing resources. For example, a processor may use thousands of processing cores to convert vertex data to pixels. Thus, rasterization performance can be improved. Brief Description of the Drawings

[0003] Figure 1 A block diagram showing a system including one or more processors according to at least one embodiment, the processor converting vertex data to pixels at least in part based on a prefix sum;

[0004] Figure 2 A block diagram showing a system including one or more processors according to at least one embodiment, the processor converting vertex data to pixels at least in part based on assigning gradients to corners in the pixels;

[0005] Figure 3 A block diagram showing a system including one or more processors according to at least one embodiment, the processor converting vertex data to pixels at least in part based on calculating a prefix sum;

[0006] Figure 4 A block diagram showing a process for converting vertex data to pixels at least in part based on a prefix sum according to at least one embodiment;

[0007] Figure 5 A block diagram showing a process for causing an application programming interface (API) to convert vertex data to pixels at least in part based on a prefix sum according to at least one embodiment;

[0008] Figure 6 A block diagram showing a driver and / or runtime for converting vertex data to pixels at least in part based on a prefix sum according to at least one embodiment;

[0009] Figure 7 A block diagram showing an exemplary data center according to at least one embodiment;

[0010] Figure 8 A block diagram showing a processing system according to at least one embodiment;

[0011] Figure 9 A block diagram showing a computer system according to at least one embodiment;

[0012] Figure 10 Shows a system according to at least one embodiment;

[0013] Figure 11 Shows an exemplary integrated circuit according to at least one embodiment;

[0014] Figure 12 Shows a computing system according to at least one embodiment;

[0015] Figure 13 Shows an APU according to at least one embodiment;

[0016] Figure 14 Shows a CPU according to at least one embodiment;

[0017] Figure 15 Shows an exemplary accelerator integration slice according to at least one embodiment;

[0018] Figures 16A to 16B Shows an exemplary graphics processor according to at least one embodiment;

[0019] Figure 17A Shows a graphics core according to at least one embodiment;

[0020] Figure 17B Shows a GPGPU according to at least one embodiment;

[0021] Figure 18A Shows a parallel processor according to at least one embodiment;

[0022] Figure 18B Shows a processing cluster according to at least one embodiment;

[0023] Figure 18C Shows a graphics multiprocessor according to at least one embodiment;

[0024] Figure 19 Shows a graphics processor according to at least one embodiment;

[0025] Figure 20 Shows a processor according to at least one embodiment;

[0026] Figure 21 Shows a processor according to at least one embodiment;

[0027] Figure 22 Shows a graphics processor core according to at least one embodiment;

[0028] Figure 23 Shows a PPU according to at least one embodiment;

[0029] Figure 24Shows a GPC according to at least one embodiment;

[0030] Figure 25 Shows a streaming multiprocessor according to at least one embodiment;

[0031] Figure 26 Shows the software stack of a programming platform according to at least one embodiment;

[0032] Figure 27 Shows according to at least one embodiment Figure 26 of the CUDA implementation of the software stack;

[0033] Figure 28 Shows according to at least one embodiment Figure 26 of the ROCm implementation of the software stack;

[0034] Figure 29 Shows according to at least one embodiment Figure 26 of the OpenCL implementation of the software stack;

[0035] Figure 30 Shows software supported by a programming platform according to at least one embodiment;

[0036] Figure 31 Shows according to at least one embodiment in Figures 26 to 29 compiled code executed on the programming platform;

[0037] Figure 32 Shows according to at least one embodiment in Figures 26 to 29 more detailed compiled code executed on the programming platform;

[0038] Figure 33 Shows converting source code before compiling the source code according to at least one embodiment;

[0039] Figure 34A Shows a system configured to compile and execute CUDA source code using different types of processing units according to at least one embodiment;

[0040] Figure 34B Shows according to at least one embodiment configured to use a CPU and a CUDA-enabled GPU to compile and execute Figure 34A the CUDA source code of;

[0041] Figure 34C Shows according to at least one embodiment configured to use a CPU and a non-CUDA-enabled GPU to compile and execute Figure 34A the CUDA source code of;

[0042] Figure 35Shows an exemplary kernel converted by the CUDA to HIP conversion tool of Figure 34C ;

[0043] Figure 36 Shows in more detail a CUDA - disabled GPU of Figure 34C according to at least one embodiment;

[0044] Figure 37 Shows how threads of an exemplary CUDA grid are mapped to different compute units of Figure 36 according to at least one embodiment;

[0045] Figure 38 Shows how to migrate existing CUDA code to data - parallel C++ code according to at least one embodiment; and

[0046] Figure 39 Shows components of a system for accessing large language models according to at least one embodiment. Detailed Description

[0047] In the following description, numerous specific details are set forth to provide a more thorough understanding of at least one embodiment. However, those skilled in the art will appreciate that the inventive concepts described herein may be practiced without one or more of these specific details, and that two or more aspects of any one or more of the embodiments described herein may be combined.

[0048] In at least one embodiment, one or more processors including one or more circuits are used to identify one or more pixels within one or more polygons based at least in part on whether one or more pixels adjacent to the one or more pixels within the one or more polygons are partially outside the one or more polygons. In at least one embodiment, a pixel is a two - dimensional (2D) element that represents a corresponding display element in a display device, such as the screen of a mobile device, a computer monitor, or a television. In at least one embodiment, one or more processors including one or more circuits are used to identify one or more pixels within one or more polygons to generate, display, or otherwise activate pixels to be displayed on a display device.

[0049] In at least one embodiment, one or more processors including one or more circuits are used to calculate one or more winding numbers on a pixel grid, pixel dot grid, or gradient field starting with all zeros, as described herein at least in conjunction with Figure 4Further description. In at least one embodiment, one or more processors including one or more circuits add +1 to each grid point immediately to the left of an upward-facing (pointing upward) edge of a polygon, where the grid points must be the first grid points encountered by a horizontal ray starting from the edge, and as described elsewhere herein. In at least one embodiment, one or more processors including one or more circuits subtract -1 from all grid points to the left of each downward-facing edge, and as described elsewhere herein. In at least one embodiment, one or more processors including one or more circuits perform an inclusive prefix sum on the rows of an array (or other tensor) that includes the +1 and -1 values assigned to grid points. In at least one embodiment, one or more processors including one or more circuits assign values to one or more grid points for each edge, regardless of which polygon the one or more grid points belong to. In at least one embodiment, one or more processors including one or more circuits receive and / or otherwise obtain input data representing the edges, without including information about the polygon and / or the polygon bounding box, and as described elsewhere herein. In at least one embodiment, one or more processors including one or more circuits identify one or more pixels within one or more polygons by using information about the edges rather than using information about the polygon and / or the polygon bounding box, and as described elsewhere herein.

[0050] In at least one embodiment, one or more processors including one or more circuits are used to determine whether a portion (e.g., a corner) of a pixel is within a shape (e.g., a polygon) based at least in part on information indicating whether another corner of the pixel is within the shape. In at least one embodiment, one or more processors including one or more circuits are used to identify a corner (or other portion) of a pixel adjacent to an edge of the shape and then propagate that information to adjacent corners. In at least one embodiment, one or more processors including one or more circuits will determine that a corner of a pixel is within the shape and then move along a column or row, identifying that a corner of a second pixel adjacent to the first pixel is within the shape, identifying that a corner of a third pixel adjacent to the second pixel is within the shape, and so on until an edge is identified that lies between the corner of the pixel identified as being within the shape and the next corner. In at least one embodiment, one or more processors including one or more circuits perform these identifications based at least in part on one or more gradients, one or more prefix sums, and other things described herein. In at least one embodiment, one or more processors including one or more circuits perform these identifications in parallel, for example, by dividing rows among multiple parallel threads.

[0051] In at least one embodiment, advantages of the techniques described herein include writing (or loading / storing) data values into an array corresponding to a polygon perimeter. In at least one embodiment, advantages of the techniques described herein include reducing conflicts where a processor attempts to simultaneously write multiple data values generated by atomic operations into memory locations representing elements (e.g., pixel points (grid points)) of a grid.

[0052] In at least one embodiment, one or more processors including one or more circuits identify one or more pixels within the one or more polygons by using information indicating whether a portion of another pixel is within the one or more polygons. In at least one embodiment, such information about these other pixels is used to compute prefix sum values that are used to identify whether any pixel is inside or outside the polygon. In at least one embodiment, the techniques described herein (including the computation of prefix sums) allow one or more processors to convert vertex data to pixels using fewer operations and / or faster operations by writing information only to grid points on the polygon perimeter and using prefix sums, which are relatively simple mathematical operations that processors are optimized to perform. In at least one embodiment, the grid points are referred to as pixel points.

[0053] In at least one embodiment, one or more processors including one or more circuits are used to identify the corners of pixels located at a specific position relative to the edges of a polygon. In at least one embodiment, the edge is a vector. In at least one embodiment, one or more processors including one or more circuits are used to identify the corners of the pixels closest to and to the left of the edge of the polygon (from the perspective of a person viewing the pixel grid). In at least one embodiment, one or more processors including one or more circuits are used to assign a value that indicates whether the identified corner is inside or outside the polygon. In at least one embodiment, one or more processors including one or more circuits are used to propagate and assign this value to other identified corners along the edge. In at least one embodiment, one or more processors including one or more circuits are used to perform a prefix sum row-by-row in the pixel grid using the value assigned to the identified corners as some of its inputs. In at least one embodiment, one or more processors including one or more circuits are used to generate a value field for indicating whether a pixel is inside (interior) the polygon using the result of the prefix sum. In at least one embodiment, one or more processors including one or more circuits are used to determine whether a pixel is inside the polygon at least in part based on the number of corners inside the polygon. In at least one embodiment, one or more processors including one or more circuits are used to generate the pixels determined to be inside the polygon. In at least one embodiment, the one or more processors that perform the operations of converting vertex data to the pixels described herein (including the operations for calculating the prefix sum) cause these one or more processors to perform fewer operations and / or perform the operations faster. In at least one embodiment, at least in part based on the prefix sum and the techniques for converting vertex data to pixels described herein are applied to fields such as graphics rendering, artificial intelligence (AI)-assisted computer vision, lithography for manufacturing computing components, or some combination thereof.

[0054] In at least one embodiment, one or more processors including one or more circuits perform one or more of the techniques described herein to identify whether one or more pixels are within a polygon based at least in part on information determined about other adjacent pixels. In at least one embodiment, the information about adjacent pixels includes information about pixels next to other pixels, pixels in the same row as other pixels, pixels in the same column as other pixels, or some combination of such pixels, as well as information described elsewhere herein. In at least one embodiment, the information about adjacent pixels refers to information about the corners and / or portions of these adjacent pixels, such as gradients further described herein. In at least one embodiment, the information about adjacent pixels includes information represented as values assigned to one or more portions of these pixels, such as gradients further described herein, as well as information described elsewhere herein. In at least one embodiment, one or more portions (such as corners) of adjacent pixels are represented as points on a grid. In at least one embodiment, the information about adjacent pixels includes values indicating whether a portion of these pixels is inside (within) or outside the polygon, as well as information described elsewhere herein. In at least one embodiment, one or more processors including one or more circuits use the information about adjacent pixels to determine whether adjacent pixels are within a polygon based at least in part on prefix sums, as described elsewhere herein.

[0055] Figure 1 FIG. 4 shows a block diagram of a system 100 including one or more processors, the processors including one or more circuits for identifying one or more pixels within one or more polygons based at least in part on whether one or more pixels adjacent to one or more pixels within the one or more polygons are partially outside the one or more polygons. In at least one embodiment, one or more aspects of one or more embodiments described herein in connection with Figure 1 are combined with one or more aspects of one or more embodiments described herein, including at least in connection with Figures 2 to 6 those aspects described. In at least one embodiment, one or more processors perform one or more operations of system 100. In at least one embodiment, any one or more of the processors described herein includes one or more circuits. In at least one embodiment, the one or more processors performing one or more operations of system 100 are any one processor or combination of processors described herein, including Figure 1 processor 104 of Figure 14 CPU 1400 of Figure 16B graphics processor 1640 of Figure 17A graphics processor 1710 described in connection withFigure 23 parallel processing unit (PPU) 2300. In at least one embodiment, processor 104 performs operations used by system 100, such as loading / storing winding values in an arithmetic logic unit (ALU) (e.g., Figure 7 ALU 710). In at least one embodiment, processor 104 performs one or more operations associated with Figure 2 as described, such as assigning gradients to corners in a pixel. In at least one embodiment, processor 104 performs one or more operations associated with Figure 3 as described, such as calculating a prefix sum. In at least one embodiment, processor 104 performs one or more operations associated with Figure 4 as described, such as assigning a value to a pixel point using operation 406. In at least one embodiment, processor 104 performs one or more operations associated with Figure 5 as described, such as performing an application programming interface (API) function for calculating winding numbers using prefix sums using operation 504. In at least one embodiment, processor 104 performs one or more operations associated with Figure 6 as described, such as operations of API 610.

[0056] In at least one embodiment, system 100 includes and / or otherwise obtains vertex data as input data, depicted as vertex input data 102. In at least one embodiment, processor 104 is used to perform operations to convert any representation of an environment, model, object, or some combination thereof into a pixel representation. In at least one embodiment, any operation performed by processor 104 that at least partially converts any data representation of an environment, model, object, or some combination thereof into a pixel representation is referred to as rasterization. In at least one embodiment, one or more aspects of the rasterization process are referred to as shading. In at least one embodiment, one or more of the operations described herein are performed by one or more shader cores, such as shader cores 1655A - 1655N of graphics processor 1640.

[0057] In at least one embodiment, vertex input data 102 includes information about vertices (or points) of polygons (e.g., triangles) used to represent a 3D model and / or object. In at least one embodiment, vertex input data 102 is polygon input data that includes data representing the edges and vertices of the polygons. In at least one embodiment, vertex input data 102 defines the edges of all polygons to be represented by pixels of a pixel grid. In at least one embodiment, vertex input data 102 includes data that defines the edges of the polygons based at least in part on the nature (characteristics) of those edges (e.g., position, start point, end point, direction, attributes, or some combination thereof) and other things described herein. In at least one embodiment, vertex input data 102 includes information about vertices, such as their position, color, reflectivity, texture, or some combination thereof. In at least one embodiment, vertex input data 102 uses vertices to represent the edges (or vectors) of the polygons that are to be represented as pixels. In at least one embodiment, vertex input data 102 includes information about the edges, such as their direction, which by convention is used to indicate whether the space and / or pixels located on one side of the edge are inside or outside the polygon containing the edge. In at least one embodiment, vertex input data 102 is stored as one or more tensors.

[0058] In at least one embodiment, processor 104 is any one of the processors described herein or a combination of multiple processors, including Figure 1 processor 104 of Figure 14 CPU 1400 of Figure 16B graphics processor 1640 of Figure 17A graphics processor 1710 described in conjunction with Figure 23 parallel processing unit (PPU) 2300 of. In at least one embodiment, any module described as being implemented on processor 104 is implemented on any one of the processors or a combination of multiple processors. In at least one embodiment, processor 104 performs at least part of any operation for rasterization.

[0059] In at least one embodiment, as used in any of the embodiments described herein, unless the context clearly dictates otherwise or is clearly contrary, terms such as "system", "device", "component", or "module" and nominalized verbs (e.g., compiler, shader, and / or other terms) each refer to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functions described herein. In at least one embodiment, any system, device, component, and module described herein is combined and / or communicatively coupled with at least one other component, system, device, component, and module, regardless of how these components are combined and / or communicatively coupled in other embodiments. In at least one embodiment, software may be embodied as a software package, code, and / or instruction set or instructions. In at least one embodiment, hardware includes hardwired circuitry, programmable circuitry, state machine circuitry, fixed function circuitry, execution unit circuitry, and / or firmware that stores instructions executed by the programmable circuitry, either alone or in any combination. In at least one embodiment, a module may be embodied, either collectively or individually, as circuitry that forms part of a larger system, such as an integrated circuit (IC), system on a chip (SoC), etc. In at least one embodiment, any one or more architectures of any circuitry of one or more modules are represented as a register transfer level (RTL) representation and / or another fabless representation, which may be licensed and / or used for tape-out, which is the final stage of IC design before it is used to manufacture the IC.

[0060] In at least one embodiment, the processor 104 performs the operations of the vertex transformation module 106 to transform (or modify) the vertex input data 102 into an array (or tensor) of data that can later be converted into pixels. In at least one embodiment, the processor 104 performs the operations of the vertex transformation module 106 to convert the three-dimensional (3D) position coordinates of the vertices in the vertex input data 102 into two-dimensional (2D) position coordinates corresponding to the pixel grid and / or display device. In at least one embodiment, the processor 104 performs the operations of the vertex transformation module 106 to transform the vertex input data 102 into a format suitable for processing with other tasks of the rasterization process, such as lighting, projection, and clipping. In at least one embodiment, the processor 104 performs the operations of the vertex transformation module 106 and outputs the transformed vertex information to be received and / or otherwise obtained by the primitive assembly module 108.

[0061] In at least one embodiment, the processor performs the operations of the primitive assembly module 108 to modify the transformed vertex input data into primitives (geometric primitives), where the primitives are edges, points, lines, triangles, or some combination thereof. In at least one embodiment, the primitives represent the edges that make up a polygon. In at least one embodiment, the edges of the polygon represent an object. In at least one embodiment, the processor outputs data representing the edges, but does not include any information about which polygons these edges form and / or help represent. In at least one embodiment, if the primitive assembly module 108 is configured to output data including the edges and information about the polygons formed by these edges and / or the polygons they help represent, one or more of the constituent primitives (e.g., polygons) that together represent an object will be ignored and / or removed to create one or more arrays (or tensors) of data values that represent the polygon bounding boxes, which include the outer edges (bounds) of the represented object. In at least one embodiment, the polygon represents the outer edge of the object. In at least one embodiment, the processor performs the operations of the primitive assembly module 108 to output the polygon bounding boxes and further processes these polygon bounding boxes to output one or more arrays of edge data, without including data representing the polygons and / or the polygon bounding boxes. In at least one embodiment, the processor performs the operations of the primitive assembly module 108 to group vertices into primitives and perform operations on each primitive, such as setting edge and plane equations, removing primitives that would be invisible if transformed into pixels, clipping primitives that intersect the view frustum, or some combination of these operations. In at least one embodiment, the processor performs the operations of the primitive assembly module 108 to indicate the edges of the polygon and their directions.

[0062] In at least one embodiment, the processor 104 performs the operations of the prefix sum rasterization module 120 to receive and / or otherwise obtain the primitives and other associated data output by the primitive assembly module 108. In at least one embodiment, the processor 103 performs the operations of the prefix sum rasterization module 108 to receive and / or otherwise obtain primitives that include edges. In at least one embodiment, the processor 104 performs the algorithm of the prefix sum rasterization module 110 to rasterize the primitives. In at least one embodiment, the processor 104 performs the operations of the prefix sum rasterization module 110 to rasterize the primitives, which includes assigning one or more values to pixel points (grid points) and calculating the prefix sum, and as described in more detail herein at least in conjunction with Figures 2 to 6 is described in more detail. In at least one embodiment, the prefix sum rasterization module 110 is further described and depicted as Figure 2 the prefix sum rasterization module 210 of Figure 3 and the prefix sum rasterization module 310 of

[0063] In at least one embodiment, the processor 104 performs the operations of the prefix sum rasterization module 110 to receive and / or otherwise obtain data regarding the edges that make up one or more polygons, without information identifying the individual polygons and / or polygon bounding boxes. In at least one embodiment, the processor 104 performs the operations of the prefix sum rasterization module 110 to receive and / or otherwise obtain data regarding the edges that make up one or more polygons, where the data indicates whether the region to the left or right of the edge is inside or outside the polygon, but does not indicate overall information about the polygon, such as an identifier of the polygon and / or polygon bounding box. In at least one embodiment, the processor performs the operations of the prefix sum rasterization module 110 while being agnostic to any data that represents the polygon as a complete whole. In at least one embodiment, the processor 104 performs the operations of the prefix sum rasterization module 110 to identify the point (pixel point) of the pixel closest to the edge on one side of the edge. In at least one embodiment, the point of the pixel is a corner of the pixel. In at least one embodiment, any number of points are used to identify a portion of the pixel. In at least one embodiment, for example, a pixel is represented by 144 equally spaced points, where 16 points are in 9 equally sized segments of the pixel. In at least one embodiment, the representation of a pixel by points is called sub-pixel mapping, and the pixel grid depicting the pixel points is called a pixel point grid or sub-pixel map. In at least one embodiment, the pixel points are called grid points.

[0064] In at least one embodiment, one or more processors 104 perform the operations of the prefix sum rasterization module 110 to assign a value to each point (pixel point) of a pixel in a pixel grid and / or sub-pixel map, where each point of the pixel in the pixel grid and / or sub-pixel map is on one side of and closest to each edge in one or more edge datasets, and the value indicates whether the pixel point is inside or outside the polygon. In at least one embodiment, the pixel point is closest to the edge if there is no intermediate pixel point in the row and / or column of the pixel point. In at least one embodiment, one side of the edge is based on an assumed human viewing point facing the pixel grid. In at least one embodiment, one side of the edge is based on the direction of the edge.

[0065] In at least one embodiment, one or more processors 104 perform the operations of the prefix sum rasterization module 110 to identify the point (pixel point) of the pixel closest to the left side of the edge, where the sides are based on an assumed human viewing point facing the pixel grid, and in combination with Figure 2Further detailed description. In at least one embodiment, any side (e.g., the right side of an edge) is used to identify a point of a pixel. In at least one embodiment, one or more processors 104 perform the operations of the prefix sum rasterization module 110 to determine whether the identified pixel point is inside or outside a polygon by using the direction information about the edge. In at least one embodiment, during vertex transformation or primitive assembly, direction information is assigned to an edge by convention, for example, the convention that the left side of an upward-pointing edge is defined as being inside the polygon and the left side of a downward-pointing edge is defined as being outside the polygon. In at least one embodiment, determining whether a pixel point is inside a polygon is at least partially based on comparing the position of the pixel point in the pixel grid with the polygon as if the polygon were overlaid on the pixel grid. In at least one embodiment, the direction of one or more edges of one or more polygons is used to determine whether a pixel is partially outside these polygons. In at least one embodiment, determining whether these pixels are partially outside these polygons is used to identify other pixels inside these polygons by calculating the winding number of points in these pixels using prefix sum, which is at least combined with Figure 3 is described in more detail.

[0066] In at least one embodiment, one or more processors 104 perform the pixel point identification operation of the prefix sum rasterization module 110 to assign a value to the identified pixel point, the value indicating whether the pixel point is inside or outside the polygon. In at least one embodiment, the prefix sum rasterization module 110 uses any suitable algorithm to identify pixel points and assign values to these pixel points. In at least one embodiment, assigning a value to the identified pixel points starts from the origin of each edge and then, by moving along the length of the edge and propagating the value to other pixel points, assigns the value to these other pixel points on the same side of the edge. In at least one embodiment, the assigned value of the identified pixel point is partially used to identify whether adjacent pixels are inside the polygon, as described at least in conjunction with Figure 3 is further described. In at least one embodiment, one or more processors 104 perform the pixel point identification operation of the prefix sum rasterization module 110 by performing this identification in parallel for each edge, identifying for each edge all pixel points on one side and closest to the edge.

[0067] In at least one embodiment, the processor 104 performs a pixel identification operation of the prefix sum and rasterization module 110 to perform a prefix sum including values assigned row by row to the identified pixels. In at least one embodiment, performing the prefix sum uses the information (gradient) assigned to the pixels close to the edge to identify whether adjacent pixels along the same row of pixels are inside the polygon. Thus, performing the prefix sum helps to identify whether pixels adjacent to pixels close to the edge are inside the polygon. In at least one embodiment, the value assigned to the identified pixel may alternatively be referred to as a gradient, as further described herein at least in connection with Figure 2 the values 214a - 214c. In at least one embodiment, the prefix sum is at least partially used to identify the one or more pixels within the one or more polygons by using information indicating whether one or more pixels adjacent to one or more pixels within one or more polygons are partially outside the one or more polygons. In at least one embodiment, the information indicating whether one or more pixels are partially outside the one or more polygons includes the value assigned to the identified pixel. In at least one embodiment, the prefix sum is further described herein at least in connection with Figure 3 In this article.

[0068] In at least one embodiment, the processor 104 performs a pixel identification operation of the prefix sum and rasterization module 110 to assign a default value of 0 to pixels that have not been assigned a value previously. In at least one embodiment, performing the prefix sum using the value assigned to the identified pixel results in calculating the winding number for each pixel in the pixel grid and / or sub - pixel map, which will be shown and described in further detail in connection with Figure 3 In further detail. In at least one embodiment, calculating the winding number for each pixel results in a winding number field, and the prefix sum and rasterization module 110 uses the winding number field to identify the pixels that are considered to be inside the polygon and thus should be generated, displayed, or otherwise activated on the display device.

[0069] Figure 2 A block diagram of a system 200 according to at least one embodiment is shown, the system including one or more processors, the processors including one or more circuits for assigning values to the pixels on one side and along one side of each edge that are closest to each edge. In at least one embodiment, one or more aspects of one or more embodiments described herein in connection with Figure 2 are combined with one or more aspects of one or more embodiments described herein, including at least in connection with Figure 1 and Figures 3 to 6Described aspects. In at least one embodiment, one or more processors perform one or more operations of system 200. In at least one embodiment, any one or more of the processors described herein include one or more circuits. In at least one embodiment, the one or more processors that perform one or more operations of system 200 are any one processor or combination of processors described herein, including Figure 1 processor 104 of Figure 14 CPU 1400 of Figure 16B graphics processor 1640 of Figure 17A graphics processor 1710 described in conjunction with Figure 23 parallel processing unit (PPU) 2300 of Figure 7 ALU 710 of Figure 3 In at least one embodiment, processor 104 performs operations of system 200, such as loading / storing winding values in an arithmetic logic unit (ALU), such as Figure 4 In at least one embodiment, one or more processors of system 200 perform one or more operations described in conjunction with Figure 5 such as calculating a prefix sum. In at least one embodiment, one or more processors of system 200 perform one or more operations described in conjunction with Figure 6 such as assigning a value to a pixel point using operation 406. In at least one embodiment, one or more processors of system 200 perform one or more operations described in conjunction with

[0070] In at least one embodiment, system 200 includes a prefix sum rasterization module 210 that shares Figure 1 one or more components of prefix sum rasterization module 110 of Figure 3 and / or prefix sum rasterization module 310 of Figure 2 In at least one embodiment, one or more processors generate a pixel point grid 212, which is depicted as a 2D point grid in

[0071] In at least one embodiment, the pixel grid 212 includes points representing the corners of each pixel in the grid, which are depicted as black dots. In at least one embodiment, the pixel grid 212 matches the dimensions of a display device (such as a mobile phone screen, a computer monitor, or a television). In at least one embodiment, the edges are depicted as arrows. In at least one embodiment, although the edges are depicted as forming polygons, one or more processors do not know and / or do not have access to data indicating how these polygons are constructed as a whole. In at least one embodiment, one or more processors can only access each edge of the polygon and not information about the polygon as a whole. In at least one embodiment, any non-horizontal edge (defined by the user) is considered an edge pointing up or down. In at least one embodiment, whether an edge is horizontal is defined by the user as an edge that forms an angle within a degree range (such as 0.2 degrees to -0.2 degrees) with the horizontal axis.

[0072] In at least one embodiment, one or more processors have performed the operations of the prefix sum rasterization module 210 to identify the pixel points closest to each edge of two polygons (including polygon 216a) and to the left of each edge of the two polygons (including polygon 216a). In at least one embodiment, one or more processors have performed the operations of the prefix sum rasterization module 210 to assign the value +1 or -1 to those identified pixel points. In at least one embodiment, the left side of the edge is based on the assumption that a person is observing the pixel grid 212. In at least one embodiment, as Figure 2 depicted, the value 214a is assigned to the pixel point closest to and to the left of the edge 215a. In at least one embodiment, as Figure 2 depicted, the value 214c is assigned to the pixel point closest to and to the left of the downward-pointing edge 215b.

[0073] In at least one embodiment, for example, the value 214a (+1) is assigned to the pixel point closest to and to the left of the upward-pointing edge 215a. In at least one embodiment, +1 indicates that the pixel point is inside the polygon. In at least one embodiment, the value 214b (-1) is assigned to the pixel point closest to and to the left of the downward-pointing edge 215c. In at least one embodiment, -1 indicates that the pixel point is outside the polygon. In at least one embodiment, due to the arrangement of its edges, polygon 216b contains a hole within polygon 216a. In at least one embodiment, as Figure 2 shown, any horizontal or near-horizontal edge is not assigned any value.

[0074] In at least one embodiment, the value assigned to an identified pixel point is referred to as a gradient or gradient value. In at least one embodiment, a pixel grid including information about the gradient assigned to pixel points is referred to as a gradient field. In at least one embodiment, one or more processors perform operations to assign gradients to pixel points in parallel, for example, by assigning different edges to multiple parallel threads. In at least one embodiment, the gradient is referred to as a winding number gradient, in part because the gradient is used to calculate the winding number as a prefix sum value, which is described further herein at least in conjunction with Figure 3 Further described. In at least one embodiment, one or more processors perform operations to identify the position of a pixel corner determined relative to an edge, as further described herein at least in conjunction with the description of the left / right side of the edge and the edge direction. In at least one embodiment, one or more processors assign a gradient to the corners of a pixel to at least partially identify one or more pixels within one or more polygons, where the assigned gradients are those belonging to the corners of one or more adjacent pixels. In at least one embodiment, one or more processors have performed the operations of the prefix sum rasterization module 210 to identify the corners (pixel points) in a pixel and assign gradients to the corners. In at least one embodiment, the values assigned to the identified pixel points are referred to as gradients because these values represent the change in the edge direction at the pixel points. In at least one embodiment, there is no edge to the right of the upward edge 215a along the pixel point row 217, and thus there is no gradient. In at least one embodiment, moving left along the pixel point row 217, the upward edge 215a intersects the row, and thus the gradient (value 214a)+1 is applied. In at least one embodiment, moving left from the value 214a, the downward edge 215c intersects the pixel point row 217, and thus the gradient (value 214b)-1 is applied. In at least one embodiment, one or more processors perform the operations of the prefix sum rasterization module 210 to store the gradients as an array. In at least one embodiment, the gradient is represented as grad(winding number) in pseudocode. In at least one embodiment, one or more processors perform the operations of the prefix sum rasterization module 210 to calculate the dot product of grad(winding number) and the (-1,0) vector.

[0075] In at least one embodiment, rather than assigning values to pixel points on the left side of each edge, one or more processors identify the interior of a polygon based on edge data and only assign values to pixel points within the interior of the polygon, and do not assign values to pixel points outside the polygon as long as the pixel points outside the polygon are not within another polygon. In at least one embodiment, only assigning values to pixel points within a polygon requires additional operations for calculating and / or determining the characteristics of the polygon and traversing the line segments and / or edges of the polygon in a polygon line segment loop.

[0076] Figure 3 FIG. 300 is a block diagram of a system including one or more processors according to at least one embodiment, the processors including one or more circuits for calculating a prefix sum and generating a winding number. In at least one embodiment, one or more aspects of one or more embodiments described herein in connection with Figure 3 are combined with one or more aspects of one or more embodiments described herein, including at least in combination with Figures 1 to 2 and Figures 4 to 6 the aspects described. In at least one embodiment, one or more processors perform one or more operations of system 300. In at least one embodiment, any one or more of the processors described herein include one or more circuits. In at least one embodiment, the one or more processors performing one or more operations of system 300 are any one processor or combination of processors described herein, including Figure 1 processor 104 of Figure 14 CPU 1400 of Figure 16B graphics processor 1640 of Figure 17A graphics processor 1710 described in connection with Figure 23 and parallel processing unit (PPU) 2300 of Figure 7 ALU 710 of Figure 2 In at least one embodiment, one or more processors of system 300 perform one or more operations described in connection with Figure 4 such as assigning a gradient to an identified pixel point. In at least one embodiment, one or more processors of system 300 perform one or more operations described in connection with Figure 5 such as assigning a value to a pixel point using operation 406. In at least one embodiment, one or more processors of system 300 perform one or more operations described in connection with Figure 6 such as performing an operation of an application programming interface (API) function for calculating a winding number using a prefix sum using operation 504. In at least one embodiment, one or more processors of system 300 perform one or more operations described in connection with

[0077] In at least one embodiment, system 300 includes a prefix sum rasterization module 310 that shares Figure 1 prefix sum rasterization module 110 of Figure 2and one or more components of the prefix sum rasterization module 210. In at least one embodiment, one or more processors generate a pixel grid 312. In at least one embodiment, one or more processors perform the operations of the prefix sum rasterization module 310 to calculate the value of the inclusive prefix sum 311 and generate the pixel grid 312.

[0078] In at least one embodiment, one or more processors perform the operations of the prefix sum rasterization module 310 to calculate the inclusive prefix sum 311. In at least one embodiment, the prefix sum is referred to as a cumulative sum, an inclusive scan, or a scan. In at least one embodiment, the prefix sum (or prefix sums) is a sequence of numbers (values) where at least in part, each number is the sum of the numbers before it. In at least one embodiment, one or more processors perform the operations of the prefix sum rasterization module 310 to use any prefix sum type or technique, such as an inclusive prefix sum. In at least one embodiment, one or more processors execute one or more prefix sum functions to output a sequence of numbers as the prefix sum.

[0079] In at least one embodiment, the inclusive prefix sum 311 is used as a winding number. In at least one embodiment, the winding number indicates the number of times a polygon wraps around a pixel. In at least one embodiment, by performing the operations of the inclusive prefix sum 311 as further described herein, one or more processors can identify the winding number of one or more pixels at least in part based on edges and gradients, as further described herein. In at least one embodiment, one or more processors perform the operations of the inclusive prefix sum 311 to identify the winding number of one or more pixels at least in part based on data representing edges, but not including data representing the entirety of one or more polygons and / or polygon bounding boxes.

[0080] In at least one embodiment, the winding number indicates whether one or more pixels are located within one or more polygons, as further described herein. In at least one embodiment, the winding number is depicted in boxes in the pixel grid 312. In at least one embodiment, the winding number is used to identify which pixels should be generated, displayed, or otherwise activated. In at least one embodiment, the inclusive prefix sum 311 depicts the values used to calculate the prefix sum for row 315a. In at least one embodiment, rows 315a and 315b represent the order in which one or more processors calculate one or more prefix sums. In at least one embodiment, one or more processors execute the prefix sum rasterization module 310 to calculate the prefix sum for each row of the pixel grid 312. In at least one embodiment, other methods or algorithms may be used to calculate the prefix sum, such as row by row, left to right, or column by column from top to bottom, depending on the convention and / or algorithm used to assign gradients to the pixels. In at least one embodiment, row 315a depicts how one or more processors calculate the inclusive prefix sum 311. In at least one embodiment, the gradient values (e.g., Figure 2 the gradient values calculated and described therein) are used as input values, as shown in the rows of the table depicted in the inclusive prefix sum 311. In at least one embodiment, one or more processors perform the operations of the prefix sum rasterization module 310 to output prefix sum values, such as outputting the winding numbers 316a - 316c. In at least one embodiment, the output winding number 316b is 1 because the input gradient 314a is +1 and the previous output winding number was 0. In at least one embodiment, the output winding number 316c is 0 because the input gradient 314b is -1 and the previous output winding number was 1. In at least one embodiment, the input gradients 314a - 314b and the output winding numbers 316a - 316c are both depicted in the Figure 3 inclusive prefix sum 311 and the pixel grid 312 for reference.

[0081] In at least one embodiment, one or more processors perform an inclusive prefix sum on each row of the pixel grid 312, resulting in a pixel grid in which each point is assigned a winding number. In at least one embodiment, the pixel grid in which each point is assigned a winding number is referred to as a winding number field. In at least one embodiment, one or more processors perform operations to identify whether one or more pixels are within one or more polygons based at least in part on one or more winding domains generated using one or more prefix sums and gradients. In at least one embodiment, performing an inclusive prefix sum on each row of the pixel grid 312 is referred to as integrating the gradient field to generate a winding number field. In at least one embodiment, one or more processors perform the operations of the prefix sum rasterization module 310 to use the winding number field to identify pixels as being within a polygon and thus should be generated, displayed, or otherwise activated on a display device. In at least one embodiment, one or more processors determine that a pixel is within a polygon by adding / summing the winding numbers of the pixel and further determining whether the resulting sum reaches or exceeds a user-set threshold. In at least one embodiment, if the sum of the winding numbers of a pixel is 2 or greater and the threshold is 2, then it is determined that the pixel is within the polygon. In at least one embodiment, if the sum of the winding numbers is 1 and the threshold is 2, then the pixel is considered to be outside the polygon even if a portion of the pixel is within the polygon.

[0082] Figure 4 FIG. 400 is a block diagram illustrating a process 400 for rasterizing vertex data based at least in part on prefix sums according to at least one embodiment. In at least one embodiment, in combination with Figure 4 one or more aspects of one or more embodiments described herein are combined with one or more aspects of one or more embodiments described herein, including at least in combination with Figures 1 to 3 and Figures 5 to 6 the aspects described. In at least one embodiment, one or more processors perform one or more operations of the system 400. In at least one embodiment, any one or more processors described herein include one or more circuits. In at least one embodiment, the one or more processors performing one or more operations of the process 400 are any one processor or combination of processors described herein, including Figure 1 processor 104 of Figure 14 CPU 1400 of Figure 16B graphics processor 1640 of Figure 17A graphics processor 1710 described in combination with Figure 23parallel processing unit (PPU) 2300. In at least one embodiment, the processor 104 performs the operations of process 400, such as assigning gradients to the pixel points of operation 406. In at least one embodiment, one or more processors of system 300 perform one or more operations described in conjunction with Figure 2 , such as assigning gradients to the identified pixel points. In at least one embodiment, one or more processors of processor 400 perform one or more operations described in conjunction with Figure 3 , such as calculating a prefix sum. In at least one embodiment, one or more processors of process 400 perform one or more operations described in conjunction with Figure 5 , such as performing an application programming interface (API) function for calculating the winding number using the prefix sum with operation 504. In at least one embodiment, one or more processors of process 400 perform one or more operations described in conjunction with Figure 6 , such as the operations of API 610.

[0083] In at least one embodiment, one or more processors begin process 400 at operation 402 by receiving and / or otherwise obtaining polygon input data. In at least one embodiment, the polygon input data is any data representing a polygon for representing 2D and / or 3D objects. In at least one embodiment, the polygon input data includes data about a closed polygon. In at least one embodiment, the polygon input data is Figure 1 vertex input data 102, 3D mesh data, point cloud data, or some combination thereof. In at least one embodiment, the vertex data is further described herein at least in conjunction with Figure 1 . In at least one embodiment, the vertex data is generated by one or more processors that modify vector-based data representing one or more objects (such as a 3D mesh). In at least one embodiment, the vertex data includes the vertices of triangles representing one or more objects and their attributes.

[0084] In at least one embodiment, one or more processors continue operation 404 of process 400, which processes polygon input data to output data representing edges (edge data) without outputting data representing each polygon and / or the overall polygon bounding box. In at least one embodiment, one or more processors perform operation 404 by loop scanning the polygon to generate one or more arrays of one or more vertices and / or edges of one or more polygons, where loop scanning means following the edges of the polygon. In at least one embodiment, loop scanning the polygon means scanning the intersections between the polygon and pixels row by row or column by column. In at least one embodiment, one or more processors skip operation 404 because the polygon input data of operation 402 includes a linked list representing edges and / or vertex pairs. In at least one embodiment, one or more processors perform operation 404 to generate an array sized equal to the number of edges in the polygon. In at least one embodiment, the array representing the edges of the polygon includes one or more zeros.

[0085] In at least one embodiment, one or more processors perform operation 404 to output data representing only those edges that form the polygon bounding box. In at least one embodiment, Figure 1 the primitive assembly module 108 performs operation 404 to output data representing only those edges that form the polygon bounding box. In at least one embodiment, one or more processors transform, assemble, or otherwise modify the polygon input data to output edge data, which will be described in further detail herein at least in conjunction with Figure 2 In at least one embodiment, techniques for rasterizing vertex data at least in part based on prefix sums (as described herein) are capable of generating and assigning winding numbers to pixel points, even though the polygon bounding box is large, because these techniques generate winding numbers based on edge data and are independent of polygon size. In at least one embodiment, as further described herein, techniques for rasterizing vertex data are at least in part based on each edge of each polygon, regardless of which polygon the edges belong to.

[0086] In at least one embodiment, one or more processors continue operation 406 of process 400 by performing an operation to assign a gradient of -1 to the pixel point immediately to the left of each downward edge in the gradient field. In at least one embodiment, the pixel point immediately to the left of each downward edge is the pixel point closest to the downward edge, as described herein at least in conjunction with Figure 2 In at least one embodiment, the gradient of -1 assigned to the pixel point represents a pixel point located outside the polygon, as described herein at least in conjunction with Figure 2 In at least one embodiment, the gradient of -1 assigned to the pixel point represents a pixel point located outside the polygon, as described herein at least in conjunction with

[0087] In at least one embodiment, one or more processors continue operation 408 of process 400 by performing an operation to assign a gradient of +1 to the pixel points immediately to the left of each upward edge in the gradient field. In at least one embodiment, the pixel points immediately to the right of each downward edge are the pixel points closest to the upward edge, as described at least in conjunction with Figure 2 as further described herein. In at least one embodiment, the gradient of +1 assigned to a pixel point indicates that the pixel point is inside the polygon, as described at least in conjunction with Figure 2 as further described herein.

[0088] In at least one embodiment, one or more processors continue operation 410 of process 400 by performing an operation to perform an inclusive prefix sum on each row of the gradient field from right to left to integrate the winding number gradient, as described at least in conjunction with Figure 3 as further described herein. In at least one embodiment, integrating the winding number gradient refers to a running sum based on prefix sum using the gradient as an input, as described at least in conjunction with Figure 3 and further described herein.

[0089] In at least one embodiment, one or more processors continue operation 412 of process 400 by performing an operation to identify the winding number of each pixel in a pixel grid to determine whether to turn on the pixel. In at least one embodiment, turning on a pixel means that one or more processors generate, display, or otherwise activate the pixel on a display device. In at least one embodiment, one or more processors use the winding number of the pixel to determine / identify how many portions of the pixel are inside or outside the polygon based on edge information about the polygon. In at least one embodiment, if the number of portions of a pixel inside the polygon exceeds a threshold number, the pixel is considered to be inside the polygon. In at least one embodiment, one or more processors output pixel data that indicates which pixels are to be turned on and what attributes they have.

[0090] Figure 5 FIG. shows a block diagram of process 500 of using a processor to execute an API function according to at least one embodiment, the API function causing one or more processors to perform rasterization using prefix sum. In at least one embodiment, one or more aspects of one or more embodiments described herein in conjunction with Figure 5 are combined with one or more aspects of one or more embodiments described herein, including at least in conjunction with Figures 1 to 4Aspects described in 6. In at least one embodiment, one or more processors perform one or more operations of process 500. In at least one embodiment, any one or more of the processors described herein include one or more circuits. In at least one embodiment, one or more processors that perform one or more operations of process 500 are any one processor or combination of processors described herein, including Figure 1 processor 104 of Figure 14 CPU 1400 of Figure 16B graphics processor 1640 of Figure 17A graphics processor 1710 described in conjunction with Figure 23 and parallel processing unit (PPU) 2300 of Figure 2 In at least one embodiment, processor 102 performs one or more operations of process 500, such as performing the function of rasterizing data using prefix sum with operation 504. In at least one embodiment, one or more processors of process 500 perform Figure 3 one or more operations of system 200 of

[0091] such as calculating the gradient of a pixel point. In at least one embodiment, one or more operations of process 500 are combined with Figure 6 one or more operations of system 300 of and In at least one embodiment, the user interface is a mobile application, a website, a customer portal, an AI-assisted conversational user interface, or some combination thereof. In at least one embodiment, the user causes the processor to automatically input vertex data into the API function, such as based on sensor streaming vertex data installed on autonomous vehicle 1000 described in conjunction with Figure 10 A. In at least one embodiment, inputting vertex data or an indication thereof into the API causes one or more processors to input the data into Figure 1 one or more modules of

[0092] In at least one embodiment, one or more processors use prefix sums to rasterize data (e.g., vertex data) by executing one or more API functions, thereby continuing operation 504 of process 500. In at least one embodiment, the one or more API functions for rasterizing data (e.g., vertex data) are one or more functions for identifying pixel points relative to the edges of a polygon, assigning gradients to the identified pixel points, calculating winding numbers using prefix sums, integrating the winding number gradient field using prefix sums, or some combination thereof. In at least one embodiment, the one or more API functions for rasterizing data cause one or more processors to assign memory locations in one or more data storage locations to store gradients, winding numbers, pixel point locations, pixel point attributes, or some combination thereof.

[0093] In at least one embodiment, one or more API functions of process 500 cause one or more processors to identify one or more pixels within one or more polygons based at least in part on whether one or more pixels adjacent to the one or more pixels within the one or more polygons are partially outside the one or more polygons. In at least one embodiment, identifying one or more pixels is based on one or more API functions of process 500 that determine the gradients of the corners in these one or more adjacent pixels, which is at least combined with Figure 2 described in more detail. In at least one embodiment, identifying the one or more pixels is based on one or more API functions of process 500 that calculate prefix sums to indicate whether one or more adjacent pixels are partially outside one or more polygons, where a pixel is partially outside a polygon based on the pixel point being outside the polygon, as described herein at least in combination with Figure 2 described. In at least one embodiment, identifying one or more pixels is based on one or more API functions of process 500 that use one or more directions of one or more edges of one or more polygons, where the directions are used to determine whether a pixel point is within the polygon, as described herein at least in combination with Figure 2 further described. In at least one embodiment, identifying one or more pixels is based on one or more API functions of process 500 that generate one or more winding numbers that indicate whether the one or more pixels are within one or more polygons, where the value of the winding number is at least combined with Figure 3Further described herein. In at least one embodiment, identifying these one or more pixels is based on one or more API functions of process 500 that identify portions of one or more pixels within one or more polygons, where if a pixel point of a pixel is within a polygon, then the portion of the pixel is within the polygon, as further described herein at least in conjunction with Figure 2 the values 214a - 214c. In at least one embodiment, identifying these one or more pixels is based on one or more API functions of process 500 that use different combinations of hardware resources for different pixels as part of a variable rate shading process. In at least one embodiment, the variable rate shading process uses different combinations of hardware resources (e.g., processors) to rasterize data into pixels to perform one or more operations described herein in parallel and / or to perform one or more operations for different groups of pixels in a pixel grid.

[0094] In at least one embodiment, one or more processors continue operation 506 of process 500 by outputting pixel data that includes an indication of whether a pixel is covered by a polygon. In at least one embodiment, a pixel being covered by a polygon is alternatively referred to as a pixel being within a polygon, as further described herein. In at least one embodiment, one or more processors perform an operation that causes one or more pixels to be generated, displayed, or activated on a display device using the winding numbers assigned to those pixels, and as described elsewhere herein at least in conjunction with Figure 3 is described.

[0095] Figure 6 A block diagram of a driver and / or runtime according to at least one embodiment is shown that includes one or more libraries to provide one or more application programming interfaces (APIs). In at least one embodiment, any one processor or combination of processors executes API 610, including Figure 1 processor 104 of Figure 14 CPU 1400 of Figure 16B graphics processor 1640 of Figure 17A graphics processor 1710 described in conjunction with Figure 23 and parallel processing unit (PPU) 2300 of Figures 1 to 3 In at least one embodiment, API 610 is further described herein. In at least one embodiment, a call to API 610 causes Figures 1 to 5Any one or more of the described operations. In at least one embodiment, the API 610 receives vertex data or an indication thereof as input and causes Figure 1 the prefix sum and rasterization module 110 to perform an operation of rasterizing vertex data based at least in part on the prefix sum. In at least one embodiment, the API 610 receives vertex data or an indication thereof and causes the prefix sum rasterization module 210 to assign gradients to pixel points, as further described herein at least in conjunction with Figure 2 as further described. In at least one embodiment, the invocation of the API 610 causes a processor to perform one or more operations of the prefix sum rasterization module 310, such as calculating a prefix sum, as further described herein at least in conjunction with Figure 3 as further described. In at least one embodiment, one or more operations to be performed by a processor when the API 610 is invoked are described in the Figure 4 process 400 and Figure 5 process 500.

[0096] In at least one embodiment, the software program 602 is a software module. In at least one embodiment, the software program 602 includes one or more software modules. In at least one embodiment, one or more of the APIs 610 are software instruction sets that, if executed, cause one or more processors to perform one or more computing operations. In at least one embodiment, one or more of the APIs 610 are distributed or otherwise provided as part of one or more libraries 606, runtimes 604, drivers 604, and / or any other software and / or executable code groups further described herein. In at least one embodiment, one or more of the APIs 610 perform one or more computing operations in response to an invocation of the software program 602. In at least one embodiment, the software program 602 is a collection of software code, commands, instructions, or other text sequences for instructing a computing device to perform one or more computing operations and / or to invoke one or more other instruction sets, such as the APIs 610 or functions 612 to be executed. In at least one embodiment, the functions provided by one or more of the APIs 610 include software functions, such as software functions that can be used to accelerate one or more portions of the software program 602 using one or more parallel processing units (PPUs) (e.g., graphics processing units (GPUs)). In at least one embodiment, the software program is a compiler.

[0097] In at least one embodiment, API 610 is a hardware interface of one or more circuits for performing one or more computing operations. In at least one embodiment, one or more of the software APIs 610 described herein are implemented as one or more circuits to perform one or more of the techniques described herein. In at least one embodiment, one or more software programs 602 include instructions that, if executed, cause one or more hardware devices and / or circuits to perform one or more of the techniques further described herein.

[0098] In at least one embodiment, the software program 602 (e.g., a user-implemented software program) utilizes one or more application programming interfaces (APIs) 610 to perform various computing operations, such as memory reservation, matrix multiplication, arithmetic operations, or any computing operations performed by a parallel processing unit (PPU) (e.g., a graphics processing unit (GPU)), as further described herein. In at least one embodiment, one or more of the APIs 610 provide a set of callable functions 612 (referred to herein as APIs, API functions, and / or functions) that individually perform one or more computing operations, such as computing operations related to parallel computing. For example, in one embodiment, one or more of the APIs 610 provide functions 612 to cause a scheduler to schedule instructions to be executed by processors based on the latency of the interconnects coupled to those processors. In at least one embodiment, the API 610 provides one or more functions 612 that are one or more neural networks, e.g., neural networks trained to improve the efficiency of processor use during rasterization.

[0099] In at least one embodiment, one or more software programs 602 interact or otherwise communicate with one or more APIs 610 to perform one or more computing operations using one or more PPUs (e.g., GPUs). In at least one embodiment, one or more of the computing operations using one or more PPUs include at least one set or more of computing operations that are at least partially executed by the one or more PPUs to be accelerated. In at least one embodiment, one or more software programs 602 interact with one or more APIs 610 to facilitate parallel computing using a remote or local interface.

[0100] In at least one embodiment, the interface is software instructions that, if executed, provide access to one or more functions 612 provided by one or more APIs 610. In at least one embodiment, when a software developer compiles one or more software programs 602 in conjunction with one or more libraries 606 that include or otherwise provide access to one or more APIs 610, the software programs 602 use a native interface. In at least one embodiment, one or more software programs 602 are statically compiled in conjunction with a precompiled library 606 that includes instructions to execute one or more APIs 610 or uncompiled source code. In at least one embodiment, one or more software programs 602 are dynamically compiled and the one or more software programs are linked by a linker to one or more precompiled libraries 606 that include one or more APIs 610.

[0101] In at least one embodiment, when a software developer executes a software program that utilizes or includes a library 606 that includes one or more APIs 610 and otherwise communicates with a library 606 that includes one or more APIs 610 via a network or other remote communication medium, the software program 602 uses a remote interface. In at least one embodiment, one or more libraries 606 that include one or more APIs 610 are executed by a remote computing service (e.g., a computing resources service provider). In another embodiment, one or more libraries 606 that include one or more APIs 610 are executed by any other computing host that provides the one or more APIs 610 to one or more software programs 602.

[0102] In at least one embodiment, a processor that executes or uses one or more software programs 602 invokes, uses, executes, or otherwise implements one or more APIs 610 to allocate and otherwise manage the memory to be used by the software program 602. In at least one embodiment, one or more software programs 602 utilize one or more APIs 610 to allocate and otherwise manage the memory to be used by one or more portions of the software program 602 for acceleration using one or more PPUs (e.g., GPUs or any other accelerator or processor further described herein). These software programs 602 may be executed by one or more processors using functions 612 provided by one or more APIs 610 in at least part based on the latency of the interconnect coupled to the one or more processors.

[0103] In at least one embodiment, API 610 is an API for facilitating parallel computing. In at least one embodiment, API 610 is any other API further described herein. In at least one embodiment, API 610 is provided by driver and / or runtime 604. In at least one embodiment, API 610 is provided by the CUDA user mode driver. In at least one embodiment, API 610 is provided by the CUDA runtime. In at least one embodiment, driver 604 is data values and software instructions that, if executed, perform or otherwise facilitate the operation of one or more functions 612 of API 610 during the loading and execution of one or more portions of software program 602. In at least one embodiment, runtime 604 is data values and software instructions that, if executed, perform or otherwise facilitate the operation of one or more functions 612 of API 610 during the execution of software program 602. In at least one embodiment, one or more software programs 602 utilize one or more APIs 610 implemented or otherwise provided by driver and / or runtime 604 to perform combined arithmetic operations during execution by one or more PPUs (e.g., GPUs).

[0104] In at least one embodiment, one or more software programs 602 utilize one or more APIs 610 provided by driver and / or runtime 604 to perform combined arithmetic operations on one or more PPUs (e.g., GPUs). In at least one embodiment, one or more APIs 610 provide combined arithmetic operations through driver and / or runtime 604, as described above. In at least one embodiment, one or more software programs 602 utilize one or more APIs 610 provided by driver and / or runtime 604 to allocate or otherwise reserve one or more memory blocks 614 of one or more PPUs (e.g., GPUs). In at least one embodiment, one or more software programs 602 utilize one or more APIs 610 provided by driver and / or runtime 604 to allocate or otherwise reserve memory blocks. In at least one embodiment, one or more APIs 610 are used to perform combined mathematical functions as described herein.

[0105] In at least one embodiment, to improve the usability of software program 602 and / or optimize one or more portions of the software program 602 for acceleration by one or more PPUs (e.g., GPUs), one or more APIs 610 provide one or more API functions 612 to perform a scheduling system that can be used by or is used by one or more of the computing devices described herein. In at least one embodiment, a processor executes one or more software programs to combine two or more application programming interfaces (APIs) into a single API. In at least one embodiment, the processor uses an API to cause a scheduler to select a thread selection mechanism and / or otherwise perform the operations described herein. In at least one embodiment, an API calls a scheduler to cause resource allocation. In at least one embodiment, the processor uses an exemplary API to schedule one or more instructions to be executed by one or more processors at least partially based on the latency of one or more interconnects coupled to these one or more processors.

[0106] In at least one embodiment, memory 614 is system memory 1206 of computing system 1200. In at least one embodiment, memory 614 is any form of hardware that stores data and is referred to as storage or data storage. In at least one embodiment, memory 614 stores data associated with Figure 4 any one or more of the arrays described. In at least one embodiment, memory 614 stores data of a voxelized hash table, as described in conjunction with Figure 2 and otherwise described herein. In at least one embodiment, memory 614 stores data used in the various operations described herein, including Figure 1 vertex input data 102 of Figure 2 gradients of Figure 3 prefix sums of Figure 4 winding numbers of operation 410 of Figure 5 and vertex data input to one or more APIs using operation 502 of

[0107] In at least one embodiment, memory 614 is a computer-readable storage medium and / or code stored on the computer-readable storage medium in the form of a computer program that includes a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the instructions are available for execution in connection with Figure 1The computer-readable instructions for the associated operations are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). In at least one embodiment, the non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, the memory 614 is implemented as a non-transitory computer-readable storage medium storing executable instructions that, if executed by one or more processors of a computer system, cause the computer system to infer a computer system architecture design as at least in conjunction with Figures 1 to 6 Further described.

[0108] Data center

[0109] Figure 7 FIG. 9 illustrates an example data center 700 according to at least one embodiment. In at least one embodiment, the data center 700 includes, but is not limited to, a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0110] In at least one embodiment, as Figure 7 shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources (“node C.R.”) 716(1)-716(N), where “N” represents any whole positive integer. In at least one embodiment, the node C.R. 716(1)-716(N) may include, but is not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (“FPGAs”), data processing units (DPUs) in network devices, graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node C.R. 716(1)-716(N) may be a server having one or more of the above computing resources.

[0111] In at least one embodiment, Figure 7 at least one component shown or described in is used to implement the techniques and / or functions described in conjunction with Figures 1 to 6 described. In at least one embodiment, the node computing resources (“node C.R.”) 716(1)-716(N) perform one or more operations of converting vertex data to pixels based at least in part on the prefix sums described in conjunction with Figure 3 described, as well as the operations described herein.

[0112] In at least one embodiment, the grouped computing resources 714 can include separate groupings (not shown) of node C.R.s housed within one or more racks, or numerous racks (also not shown) within data centers located in various geographical locations. Separate groupings of node C.R.s within the grouped computing resources 714 can include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s including CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any number of power modules, cooling modules, and network switches, in any combination.

[0113] In at least one embodiment, the resource coordinator 712 can configure or otherwise control one or more node C.R.s 716(1)-716(N) and / or the grouped computing resources 714. In at least one embodiment, the resource coordinator 712 can include a software design infrastructure (“SDI”) management entity for the data center 700. In at least one embodiment, the resource coordinator 712 can include hardware, software, or some combination thereof.

[0114] In at least one embodiment, as Figure 7 shown, the framework layer 720 includes, but is not limited to, a job scheduler 732, a configuration manager 734, a resource manager 736, and a distributed file system 738. In at least one embodiment, the framework layer 720 can include a framework for the software 752 of the support software layer 730 and / or one or more applications 742 of the application layer 740. In at least one embodiment, the software 752 or the application 742 can respectively include web-based service software or applications, such as the services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 can be, but is not limited to, a free and open-source software web application framework, such as Apache Spark which can utilize the distributed file system 738 for large-scale data processing (e.g., “big data”). TM(hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 732 may include a Spark driver to facilitate scheduling of workloads supported by the various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be able to configure the different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 738 for supporting large-scale data processing. In at least one embodiment, the resource manager 736 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 738 and the job scheduler 732. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 714 on the data center infrastructure layer 710. In at least one embodiment, the resource manager 736 may coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.

[0115] In at least one embodiment, the software 752 included in the software layer 730 may include software used by at least a portion of the nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.

[0116] In at least one embodiment, one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of applications may include, but are not limited to, CUDA applications.

[0117] In at least one embodiment, any one of the configuration manager 734, the resource manager 736, and the resource coordinator 712 may perform any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modifying actions may relieve the data center operator of the data center 700 from making potentially bad configuration decisions and may avoid underutilization and / or poorly performing parts of the data center.

[0118] Computer-based system

[0119] The following figures present, but are not limited to, exemplary computer-based systems that may be used to implement at least one embodiment.

[0120] Figure 8FIG. 800 shows a processing system according to at least one embodiment. In at least one embodiment, system 800 includes one or more processors 802 and one or more graphics processors 808, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 802 or processor cores 807. In at least one embodiment, processing system 800 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices. In at least one embodiment, processor core 807 is referred to as a computing unit or arithmetic unit.

[0121] In at least one embodiment, Figure 8 at least one component shown or described herein is used to implement techniques and / or functions described in connection with Figures 1 to 6 the description. In at least one embodiment, processing system 800 performs one or more operations that convert vertex data to pixels based at least in part on prefix sums described in connection with Figure 3 the description, as well as operations described elsewhere herein.

[0122] In at least one embodiment, processing system 800 may be included in or incorporated within a server-based gaming platform, a gaming console including a game and media console, a mobile gaming console, a handheld gaming console, or an online gaming console. In at least one embodiment, processing system 800 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, processing system 800 may also be coupled to or integrated within a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 800 is a television or set-top box device having one or more processors 802 and a graphical interface generated by one or more graphics processors 808.

[0123] In at least one embodiment, each of the one or more processors 802 includes one or more processor cores 807 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 807 is configured to process a particular instruction set 809. In at least one embodiment, instruction set 809 may facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, multiple processor cores 807 may each process a different instruction set 809, which may include instructions that help to emulate other instruction sets. In at least one embodiment, processor core 807 may also include other processing devices, such as a digital signal processor (“DSP”).

[0124] In at least one embodiment, the processor 802 includes a cache memory 804. In at least one embodiment, the processor 802 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among the various components of the processor 802. In at least one embodiment, the processor 802 also uses an external cache (e.g., a level three (L3) cache or a last level cache (LLC)) (not shown), which may share this logic among the processor cores 807 using known cache coherence techniques. In at least one embodiment, the processor 802 further includes a register file 806, and the processor 802 may include different types of registers for storing different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, the register file 806 may include general purpose registers or other registers.

[0125] In at least one embodiment, one or more processors 802 are coupled to one or more interface buses 810 to transfer communication signals, such as address, data, or control signals, between the processor 802 and other components in the system 800. In at least one embodiment, the interface bus 810 may be a processor bus, such as a version of the direct media interface (DMI) bus, in one embodiment. In at least one embodiment, the interface bus 810 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 802 includes an integrated memory controller 816 and a platform controller hub 830. In at least one embodiment, the memory controller 816 facilitates communication between the storage device and other components of the processing system 800, while the platform controller hub (PCH) 830 provides connections to input / output (I / O) devices via a local I / O bus. In at least one embodiment, one or more peripheral component interconnect buses include PCIe Gen 5, which provides an interface for the processor.

[0126] In at least one embodiment, the memory device 820 can be a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, a phase change memory device, or have suitable performance to be used as a processor memory. In at least one embodiment, the memory device 820 can be used as the system memory of the processing system 800 to store data 822 and instructions 821 for use when one or more processors 802 execute an application or process. In at least one embodiment, the memory controller 816 is also coupled to an optional external graphics processor 812, which can communicate with one or more graphics processors 808 in the processor 802 to perform graphics and media operations. In at least one embodiment, the display device 811 can be connected to the processor 802. In at least one embodiment, the display device 811 can include one or more of an internal display device, such as in a mobile electronic device or a portable computer device, or an external display device connected through a display interface (such as DisplayPort, etc.). In at least one embodiment, the display device 811 can include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) applications or augmented reality (AR) applications.

[0127] In at least one embodiment, the platform controller hub 830 enables peripheral devices to be connected to the memory device 820 and the processor 802 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 846, a network controller 834, a firmware interface 828, a wireless transceiver 826, a touch sensor 825, a data storage device 824 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 824 can be connected via a memory interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 825 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 826 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 828 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 834 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 810. In at least one embodiment, the audio controller 846 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 800 includes an optional legacy I / O controller 840 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the processing system 800. In at least one embodiment, the platform controller hub 830 can also be connected to one or more Universal Serial Bus (USB) controllers 842, which connect input devices, such as a keyboard and mouse 843 combination, a camera 844, or other USB input devices.

[0128] In at least one embodiment, instances of the memory controller 816 and the platform controller hub 830 can be integrated into a discrete external graphics processor, such as the external graphics processor 812. In at least one embodiment, the platform controller hub 830 and / or the storage controller 816 can be external to one or more processors 802. For example, in at least one embodiment, the processing system 800 can include an external storage controller 816 and a platform controller hub 830, which can be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 802.

[0129] Figure 9FIG. 900 shows a computer system according to at least one embodiment. In at least one embodiment, computer system 900 may be a system having interconnected devices and components, a SOC, or some combination thereof. In at least one embodiment, computer system 900 is formed by a processor 902, which may include execution units for executing instructions. In at least one embodiment, computer system 900 may include, but is not limited to, components such as processor 902, which employs execution units including logic for executing algorithms for processing data. In at least one embodiment, computer system 900 may include a processor, such as a processor family, XeonTM, XScaleTM, and / or StrongARMTM Core TM or Nervana TM microprocessor, although other systems (including PCs, engineering workstations, set-top boxes, etc. having other microprocessors) may also be used. In at least one embodiment, computer system 900 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0130] In at least one embodiment, Figure 9 at least one of the components shown or described is used to implement the techniques and / or functions described in connection with Figures 1 to 6 this description. In at least one embodiment, computer system 900 performs one or more operations for converting vertex data to pixels based at least in part on prefix sums described in connection with Figure 3 this description, as well as the operations described herein.

[0131] In at least one embodiment, computer system 900 can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include microcontrollers, digital signal processors (“DSPs”), system-on-chips (“SoCs”), network computers (“NetPCs”), set-top boxes, network hubs, wide area network (“WAN”) switches, or any other system that can execute one or more instructions according to at least one embodiment.

[0132] In at least one embodiment, computer system 900 can include, but is not limited to, a processor 902, which can include, but is not limited to, one or more execution units 908 that can be configured to execute Compute Unified Device Architecture (“CUDA”) (developed by NVIDIA Corporation of Santa Clara, California) programs. In at least one embodiment, a CUDA program is at least a portion of a software application written in the CUDA programming language. In at least one embodiment, computer system 900 is a single-processor desktop or server system. In at least one embodiment, computer system 900 can be a multi-processor system. In at least one embodiment, processor 902 can include, but is not limited to, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 can be coupled to a processor bus 910 that can transfer data signals between processor 902 and other components in computer system 900.

[0133] In at least one embodiment, processor 902 can include, but is not limited to, a level 1 (“L1”) internal cache memory (“cache”) 904. In at least one embodiment, processor 902 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory can reside external to processor 902. In at least one embodiment, processor 902 can include a combination of internal and external caches. In at least one embodiment, register file 906 can store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.

[0134] In at least one embodiment, an execution unit 908, including but not limited to logic for performing integer and floating point operations, is also located within the processor 902. The processor 902 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 908 may include logic for processing a packet instruction set 909. In at least one embodiment, by including the packet instruction set 909 in the instruction set of the general-purpose processor 902 and the associated circuitry for the instructions to be executed, operations used by many multimedia applications can be performed using packet data in the general-purpose processor 902. In at least one embodiment, operations can be performed on packet data by using the full width of the processor's data bus to accelerate and more efficiently execute many multimedia applications, which may not require transferring smaller data units on the processor's data bus to perform one or more operations on one data element at a time.

[0135] In at least one embodiment, the execution unit 908 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 900 may include but not be limited to a memory 920. In at least one embodiment, the memory 920 can be implemented as a DRAM device, an SRAM device, a flash memory device, or other storage devices. The memory 920 may store instructions 919 and / or data 921 represented by data signals that can be executed by the processor 902.

[0136] In at least one embodiment, the system logic chip can be coupled to the processor bus 910 and the memory 920. In at least one embodiment, the system logic chip may include but not be limited to a memory controller hub ("MCH") 916, and the processor 902 can communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 can provide a high-bandwidth memory path 918 to the memory 920 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 916 can initiate data signals among the processor 902, the memory 920, and other components in the computer system 900, and bridge data signals among the processor bus 910, the memory 920, and the system I / O 922. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 can be coupled to the memory 920 via the high-bandwidth memory path 918, and the graphics / video card 912 can be coupled to the MCH 916 via an Accelerated Graphics Port ("AGP") interconnect 914.

[0137] In at least one embodiment, the computer system 900 may use the system I / O 922 as a proprietary hub interface bus to couple the MCH 916 to an I / O controller hub (“ICH”) 930. In at least one embodiment, the ICH 930 may provide direct connections to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 920, the chipset, and the processor 902. Examples may include, but are not limited to, an audio controller 929, a firmware hub (“Flash BIOS”) 928, a wireless transceiver 926, a data storage 924, a legacy I / O controller 923 that includes user input 925 and a keyboard interface, a serial expansion port 927 (such as USB), and a network controller 934. The data storage 924 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage devices.

[0138] In at least one embodiment, Figure 9 A system including interconnected hardware devices or “chips” is shown. In at least one embodiment, Figure 9 An exemplary SoC may be shown. In at least one embodiment, Figure 9 The devices shown therein may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the system 900 use a Compute Express Link (“CXL”) interconnect to interconnect.

[0139] Figure 10 System 1000 according to at least one embodiment is shown. In at least one embodiment, system 1000 is an electronic device utilizing a processor 1010. In at least one embodiment, system 1000 may be, for example but not limited to, a laptop computer, a tower server, a rack server, a blade server, an edge device communicatively coupled to one or more local or cloud service providers, a notebook computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0140] In at least one embodiment, Figure 10 At least one component shown or described therein is used to implement the techniques and / or functions described in connection with Figures 1 to 6 described. In at least one embodiment, system 1000 performs one or more operations for converting vertex data to pixels based at least in part on the prefix sum described in connection with Figure 3 described, as well as operations described elsewhere herein.

[0141] In at least one embodiment, system 1000 may include, but is not limited to, a processor 1010 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1010 is coupled using a bus or interface, such as an I 2 C bus, a system management bus (“SMBus”), a low pin count (LPC) bus, a serial peripheral interface (“SPI”), a high definition audio (“HDA”) bus, a serial advanced technology attachment (“SATA”) bus, a USB (versions 1, 2, 3) or a universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, Figure 10 A system is shown that includes interconnected hardware devices or “chips”. In at least one embodiment, Figure 10 An exemplary SoC may be shown. In at least one embodiment, Figure 10 The devices shown in may be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof.

[0142] In at least one embodiment, Figure 10 One or more components of are interconnected using Compute Express Link (CXL) interconnects.

[0143] In at least one embodiment, Figure 10 may include a display 1024, a touch screen 1025, a touchpad 1030, a near field communication unit (“NFC”) 1045, a sensor hub 1040, a thermal sensor 1046, an embedded controller (“EC”) 1035, a trusted platform module (“TPM”) 1038, a BIOS / firmware / flash (“BIOS, FW Flash”) 1022, a DSP 1060, a solid state drive (“SSD”) or a hard disk drive (“HDD”) 1020, a wireless local area network unit (“WLAN”) 1050, a Bluetooth unit 1052, a wireless wide area network unit (“WWAN”) 1056, a global positioning system (GPS) 1055, a camera (“USB 3.0 camera”) 1054 (e.g., a USB3.0 camera), or a low power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1015 implemented, for example, to the LPDDR3 standard. These components may each be implemented in any suitable manner.

[0144] In at least one embodiment, other components may be communicatively coupled to the processor 1010 through the components discussed above. In at least one embodiment, an accelerometer 1041, an ambient light sensor (“ALS”) 1042, a compass 1043, and a gyroscope 1044 may be communicatively coupled to the sensor hub 1040. In at least one embodiment, a thermal sensor 1039, a fan 1037, a keyboard 1036, and a touchpad 1030 may be communicatively coupled to the EC 1035. In at least one embodiment, a speaker 1063, headphones 1064, and a microphone (“mic”) 1065 may be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 1062, which may in turn be communicatively coupled to the DSP 1060. In at least one embodiment, the audio unit 1062 may include, for example but not limited to, an audio encoder / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a subscriber identity module (“SIM”) 1057 may be communicatively coupled to the WWAN unit 1056. In at least one embodiment, components such as the WLAN unit 1050, the Bluetooth unit 1052, and the WWAN unit 1056 may be implemented in a next-generation form factor (NGFF).

[0145] Figure 11 An exemplary integrated circuit 1100 is shown in accordance with at least one embodiment. In at least one embodiment, the exemplary integrated circuit 1100 is a system-on-a-chip (SoC) that may be fabricated using one or more IP cores. In at least one embodiment, the integrated circuit 1100 includes one or more application processors 1105 (e.g., a CPU, a DPU), at least one graphics processor 1110, and may additionally include an image processor 1115 and / or a video processor 1120, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 1100 includes peripheral or bus logic that includes a USB controller 1125, a UART controller 1130, an SPI / SDIO controller 1135, and an 2 S / I 2 2C controller 1140. In at least one embodiment, the integrated circuit 1100 may include a display device 1145 coupled to one or more of a high-definition multimedia interface (“HDMI”) controller 1150 and a mobile industry processor interface (“MIPI”) display interface 1155. In at least one embodiment, storage may be provided by a flash memory subsystem 1160 that includes flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1165 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1170.

[0146] In at least one embodiment,Figure 11 At least one component shown or described herein is used to implement the techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, IC 1100 performs one or more operations to convert vertex data into pixels that are at least partially based on the prefix sums associated with Figure 3 described, as well as the operations described herein.

[0147] Figure 12 FIG. 1200 shows a computing system 1200 in accordance with at least one embodiment. In at least one embodiment, computing system 1200 includes a processing subsystem 1201 having one or more processors 1202 and system memory 1204 that communicate via an interconnect path that can include a memory hub 1205. In at least one embodiment, memory hub 1205 can be a separate component within a chipset component or integrated within one or more of processors 1202. In at least one embodiment, memory hub 1205 is coupled to an I / O subsystem 1211 via a communication link 1206. In at least one embodiment, I / O subsystem 1211 includes an I / O hub 1207 that can enable computing system 1200 to receive input from one or more input devices 1208. In at least one embodiment, I / O hub 1207 can enable a display controller that is included within one or more of processors 1202 to provide output to one or more display devices 1210A. In at least one embodiment, one or more display devices 1210A coupled to I / O hub 1207 can include a local, internal, or embedded display device.

[0148] In at least one embodiment, Figure 12 At least one component shown or described herein is used to implement the techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, computing system 1200 performs one or more operations to convert vertex data into pixels that are at least partially based on the prefix sums associated with Figure 3 described, as well as the operations described herein.

[0149] In at least one embodiment, the processing subsystem 1201 includes one or more parallel processors 1212 coupled to the memory hub 1205 via a bus or other communication link 1213. In at least one embodiment, the communication link 1213 can be one of many standard-based communication link technologies or protocols, such as, but not limited to, PCIe, or can be a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 1212 form a parallel or vector processing system in a computing focus, which can include a large number of processing cores and / or processing clusters, such as a Many Integrated Core (MIC) processor or a computing unit. In at least one embodiment, the one or more parallel processors 1212 form a graphics processing subsystem that can output pixels to one of the one or more display devices 1210A coupled via the I / O hub 1207. In at least one embodiment, the one or more parallel processors 1212 can also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 1210B.

[0150] In at least one embodiment, the system storage unit 1214 can be connected to the I / O hub 1207 to provide a storage mechanism for the computing system 1200. In at least one embodiment, the I / O switch 1216 can be used to provide an interface mechanism to enable connections between the I / O hub 1207 and other components, such as a network adapter 1218 and / or a wireless network adapter 1219 that can be integrated into the platform, and various other devices that can be added via one or more additional devices 1220. In at least one embodiment, the network adapter 1218 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1219 can include one or more of Wi-Fi, Bluetooth, NFC, or other network devices including one or more radios.

[0151] In at least one embodiment, the computing system 1200 can include other components not explicitly shown, including USB or other port connections, an optical storage drive, a video capture device, etc., which can also be connected to the I / O hub 1207. In at least one embodiment, Figure 12 the communication paths interconnecting the various components can be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCIe), or other bus or point-to-point communication interfaces and / or protocols (e.g., NVLink high-speed interconnect or interconnect protocol).

[0152] In at least one embodiment, one or more parallel processors 1212 include circuitry optimized for graphics and video processing (including, for example, video output circuitry) and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1212 include circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 1200 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1212, memory hub 1205, processor 1202, and I / O hub 1207 may be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 1200 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 1200 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules into a modular computing system. In at least one embodiment, I / O subsystem 1211 and display device 1210B are omitted from computing system 1200. In at least one embodiment, one or more parallel processors 1212 include one or more tensor memory accelerators (TMA) units that can transfer data blocks between global memory and shared memory. In at least one embodiment, one or more processors use or access one or more TMAs to perform bidirectional copy operations, such as from global memory to shared memory and vice versa.

[0153] Processing system

[0154] The following figures illustrate, but are not limited to, exemplary processing systems that may be used to implement at least one embodiment.

[0155] Figure 13Shows an accelerated processing unit (“APU”) 1300 according to at least one embodiment. In at least one embodiment, the APU 1300 is developed by AMD Corporation of Santa Clara, California. In at least one embodiment, the APU 1300 can be configured to execute applications, such as CUDA programs. In at least one embodiment, the APU 1300 includes, but is not limited to, a core complex 1310, a graphics complex 1340, a fabric 1360, an I / O interface 1370, a memory controller 1380, a display controller 1392, and a multimedia engine 1394. In at least one embodiment, the APU 1300 can include any combination of any number of core complexes 1310, any number of graphics complexes 1350, any number of display controllers 1392, and any number of multimedia engines 1394. For illustrative purposes, multiple instances of similar objects are denoted by reference numerals herein, where the reference numeral identifies the object and the number in parentheses identifies the instance required.

[0156] In at least one embodiment, Figure 13 at least one component shown or described in is used to implement the techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, the APU 1300 performs one or more operations that convert vertex data to pixels based at least in part on the prefix sum associated with Figure 3 described, as well as the operations described herein.

[0157] In at least one embodiment, the core complex 1310 is a CPU, the graphics complex 1340 is a GPU, and the APU 1300 is a processing unit that integrates, among other things, the core complex 1310 and the graphics complex 1340 onto a single chip. In at least one embodiment, some tasks can be assigned to the core complex 1310 while other tasks can be assigned to the graphics complex 1340. In at least one embodiment, the core complex 1310 is configured to execute the main control software associated with the APU 1300, such as an operating system. In at least one embodiment, the core complex 1310 is the main processor of the APU 1300, which controls and coordinates the operation of other processors. In at least one embodiment, the core complex 1310 issues commands that control the operation of the graphics complex 1340. In at least one embodiment, the core complex 1310 can be configured to execute host-executable code derived from CUDA source code, and the graphics complex 1340 can be configured to execute device-executable code derived from CUDA source code.

[0158] In at least one embodiment, the core complex 1310 includes, but is not limited to, cores 1320(1)-1320(4) and the L3 cache 1330. In at least one embodiment, the core complex 1310 may include, but is not limited to, any combination of any number of cores 1320 and any number and type of caches. In at least one embodiment, the cores 1320 are configured to execute instructions of a specific instruction set architecture (“ISA”). In at least one embodiment, each core 1320 is a CPU core. In at least one embodiment, the cores 1320 are referred to as computing units or arithmetic units.

[0159] In at least one embodiment, each core 1320 includes, but is not limited to, an instruction fetch / decode unit 1322, an integer execution engine 1324, a floating-point execution engine 1326, and an L2 cache 1328. In at least one embodiment, the instruction fetch / decode unit 1322 fetches instructions, decodes these instructions, generates micro-operations, and dispatches individual micro-instructions to the integer execution engine 1324 and the floating-point execution engine 1326. In at least one embodiment, the instruction fetch / decode unit 1322 may dispatch one micro-instruction to the integer execution engine 1324 and another micro-instruction to the floating-point execution engine 1326 simultaneously. In at least one embodiment, the integer execution engine 1324 executes operations not limited to integers and memory. In at least one embodiment, the floating-point engine 1326 executes operations not limited to floating-point and vector operations. In at least one embodiment, the instruction fetch-decode unit 1322 dispatches micro-instructions to a single execution engine that replaces both the integer execution engine 1324 and the floating-point execution engine 1326.

[0160] In at least one embodiment, each core 1320(i) may access the L2 cache 1328(i) included in the core 1320(i), where i is an integer representing a specific instance of the core 1320. In at least one embodiment, each core 1320 included in the core complex 1310(j) is connected to other cores 1320 included in the core complex 1310(j) via the L3 cache 1330(j) included in the core complex 1310(j), where j is an integer representing a specific instance of the core complex 1310. In at least one embodiment, the cores 1320 included in the core complex 1310(j) may access all the L3 caches 1330(j) included in the core complex 1310(j), where j is an integer representing a specific instance of the core complex 1310. In at least one embodiment, the L3 cache 1330 may include, but is not limited to, any number of slices.

[0161] In at least one embodiment, the graphics complex 1340 may be configured to perform computational operations in a highly parallel manner. In at least one embodiment, the graphics complex 1340 is configured to perform graphics pipeline operations, such as draw commands, pixel operations, geometric calculations, and other operations associated with rendering an image to a display. In at least one embodiment, the graphics complex 1340 is configured to perform operations that are not related to graphics. In at least one embodiment, the graphics complex 1340 is configured to perform operations related to graphics and operations not related to graphics.

[0162] In at least one embodiment, the graphics complex 1340 includes, but is not limited to, any number of compute units 1350 and an L2 cache 1342. In at least one embodiment, the compute units 1350 share the L2 cache 1342. In at least one embodiment, the L2 cache 1342 is partitioned. In at least one embodiment, the graphics complex 1340 includes, but is not limited to, any number of compute units 1350 and any number (including zero) and type of cache. In at least one embodiment, the graphics complex 1340 includes, but is not limited to, any number of dedicated graphics hardware.

[0163] In at least one embodiment, each computing unit 1350 includes, but is not limited to, any number of SIMD units 1352 and a shared memory 1354. In at least one embodiment, each SIMD unit 1352 implements a SIMD architecture and is configured to perform operations in parallel. In at least one embodiment, each computing unit 1350 may execute any number of thread blocks, but each thread block is executed on a single computing unit 1350. In at least one embodiment, a thread block includes, but is not limited to, any number of execution threads. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 1352 executes a different warp. In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in the warp belongs to a single thread block and is configured to process different data sets based on a single instruction set. In at least one embodiment, predication may be used to disable one or more threads in a warp. In at least one embodiment, a lane is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a warp. In at least one embodiment, different wavefronts in a thread block may synchronize together and communicate via the shared memory 1354. In at least one embodiment, each computing unit 1350 includes one or more thread block clusters, where the thread block clusters may implement locality programming control at a coarser granularity than a single thread block of a single streaming multi-processor (SM). In at least one embodiment, the thread block clusters (also referred to as "clusters") enable multiple thread blocks running concurrently across streaming multi-processors to synchronize and collaboratively acquire, exchange, or otherwise use data.

[0164] In at least one embodiment, the structure 1360 is a system interconnect that facilitates data and control transfer across the core complex 1310, the graphics complex 1340, the I / O interface 1370, the memory controller 1380, the display controller 1392, and the multimedia engine 1394. In at least one embodiment, in addition to or instead of the structure 1360, the APU 1300 may also include, but is not limited to, any number and type of system interconnects that facilitate data and control transfer across any number and type of directly or indirectly linked components that may be internal or external to the APU 1300. In at least one embodiment, the I / O interface 1370 represents any number and type of I / O interfaces (e.g., PCI, PCI-Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to the I / O interface 1370. In at least one embodiment, the peripheral devices coupled to the I / O interface 1370 may include, but are not limited to, keyboards, mice, printers, scanners, joysticks or other types of game controllers, media recording devices, external storage devices, network interface cards, etc.

[0165] In at least one embodiment, the display controller AMD92 displays images on one or more display devices (e.g., liquid crystal display (LCD) devices). In at least one embodiment, the multimedia engine 1394 includes, but is not limited to, any number and type of multimedia-related circuits, such as video decoders, video encoders, image signal processors, etc. In at least one embodiment, the memory controller 1380 facilitates data transfer between the APU 1300 and the unified system memory 1390. In at least one embodiment, the core complex 1310 and the graphics complex 1340 share the unified system memory 1390.

[0166] In at least one embodiment, the APU 1300 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 1380 and memory devices (e.g., shared memory 1354) that may be dedicated to one component or shared among multiple components. In at least one embodiment, the APU 1300 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., L2 cache 1428, L3 cache 1330, and L2 cache 1342), each of which may be private or shared among any number of components (e.g., cores 1320, core complex 1310, SIMD units 1352, compute units 1350, and graphics complex 1340).

[0167] Figure 14Shows a CPU 1400 according to at least one embodiment. In at least one embodiment, the CPU 1400 is developed by AMD Corporation of Santa Clara, California. In at least one embodiment, the CPU 1400 can be configured to execute application programs. In at least one embodiment, the CPU 1400 is configured to execute main control software, such as an operating system. In at least one embodiment, the CPU 1400 issues commands to control the operation of an external GPU (not shown). In at least one embodiment, the CPU 1400 can be configured to execute host-executable code derived from CUDA source code, and the external GPU can be configured to execute device-executable code derived from such CUDA source code. In at least one embodiment, the CPU 1400 includes, but is not limited to, any number of core complexes 1410, fabric 1460, I / O interfaces 1470, and memory controllers 1480.

[0168] In at least one embodiment, Figure 14 at least one component shown or described herein is used to implement the techniques and / or functions associated with Figures 1 to 6 the description. In at least one embodiment, the CPU 1400 performs one or more operations, at least in part based on the prefix sum associated with Figure 3 the description to convert vertex data to pixels, as described elsewhere herein.

[0169] In at least one embodiment, the core complex 1410 includes, but is not limited to, cores 1420(1)-1420(4) and an L3 cache 1430. In at least one embodiment, the core complex 1410 can include, but is not limited to, any number of cores 1420 and any combination of any number and type of caches. In at least one embodiment, the cores 1420 are configured to execute instructions of a specific ISA. In at least one embodiment, each core 1420 is a CPU core.

[0170] In at least one embodiment, each core 1420 includes, but is not limited to, an instruction fetch / decode unit 1422, an integer execution engine 1424, a floating-point execution engine 1426, and an L2 cache 1428. In at least one embodiment, the instruction fetch / decode unit 1422 fetches instructions, decodes these instructions, generates micro-operations, and dispatches individual micro-instructions to the integer execution engine 1424 and the floating-point execution engine 1426. In at least one embodiment, the instruction fetch / decode unit 1422 can dispatch one micro-instruction to the integer execution engine 1424 and another micro-instruction to the floating-point execution engine 1426 simultaneously. In at least one embodiment, the integer execution engine 1424 performs operations that are not limited to integer and memory operations. In at least one embodiment, the floating-point engine 1426 performs operations that are not limited to floating-point and vector operations. In at least one embodiment, the instruction fetch / decode unit 1422 dispatches micro-instructions to a single execution engine that replaces both the integer execution engine 1424 and the floating-point execution engine 1426.

[0171] In at least one embodiment, each core 1420(i) can access the L2 cache 1428(i) included in the core 1420(i), where i is an integer representing a specific instance of the core 1420. In at least one embodiment, each core 1420 included in a core complex 1410(j) is connected to other cores 1420 included in the core complex 1410(j) via an L3 cache 1430(j) included in the core complex 1410(j), where j is an integer representing a specific instance of the core complex 1410. In at least one embodiment, the core 1420 included in the core complex 1410(j) can access all L3 caches 1430(j) included in the core complex 1410(j), where j is an integer representing a specific instance of the core complex 1410. In at least one embodiment, the L3 cache 1430 can include, but is not limited to, any number of slices.

[0172] In at least one embodiment, the structure 1460 is a system interconnect that facilitates data and control transfer across the core complexes 1410(1)-1410(N) (where N is an integer greater than zero), the I / O interfaces 1470, and the memory controller 1480. In at least one embodiment, in addition to or instead of the structure 1460, the CPU 1400 may also include, but is not limited to, any number and type of system interconnects that facilitate data and control transfer across any number and type of directly or indirectly linked components that may be internal or external to the CPU 1400. In at least one embodiment, the I / O interface 1470 represents any number and type of I / O interfaces (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to the I / O interface 1470. In at least one embodiment, the peripheral devices coupled to the I / O interface 1470 may include, but are not limited to, a display, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc.

[0173] In at least one embodiment, the memory controller 1480 facilitates data transfer between the CPU 1400 and the system memory 1490. In at least one embodiment, the core complex 1410 and the graphics complex 1440 share the system memory 1490. In at least one embodiment, the CPU 1400 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 1480 and memory devices that may be dedicated to one component or shared among multiple components. In at least one embodiment, the CPU 1400 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., L2 cache 1428 and L3 cache 1430), each of which may be private to a component or shared among any number of components (e.g., cores 1420 and core complex 1410).

[0174] Figure 15An exemplary accelerator integrated slice 1590 according to at least one embodiment is shown. As used herein, a "slice" includes a designated portion of the processing resources of an accelerator integrated circuit. In at least one embodiment, the accelerator integrated circuit provides cache management, memory access, context management, and interrupt management services for multiple graphics processing engines among multiple graphics acceleration modules. Each graphics processing engine may include a separate GPU. Optionally, the graphics processing engine may include different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, the graphics acceleration module may be a GPU having multiple graphics processing engines. In at least one embodiment, the graphics processing engines may be individual GPUs integrated on a common package, line card, or chip.

[0175] In at least one embodiment, Figure 15 at least one component shown or described in is used to implement in conjunction with Figures 1 to 6 the techniques and / or functions described. In at least one embodiment, the acceleration integrated slice 1590 performs one or more operations based at least in part on a prefix sum described in conjunction with Figure 3 to convert vertex data to pixels, as described elsewhere herein.

[0176] The application program virtual address space 1582 within the system memory 1514 stores process elements 1583. In one embodiment, the process elements 1583 are stored in response to a GPU call 1581 from an application 1580 executing on the processor 1507. The process elements 1583 contain the processing state of the corresponding application 1580. The work descriptor ("WD") 1584 included in the process element 1583 may be a single job requested by the application or may contain a pointer to a job queue. In at least one embodiment, the WD 1584 is a pointer to a job request queue within the application program virtual address space 1582.

[0177] The graphics acceleration module 1546 and / or individual graphics processing engines may be shared by all or part of the processes in the system. In at least one embodiment, an infrastructure may be included for establishing a processing state and sending the WD 1584 to the graphics acceleration module 1546 to start a job in a virtualized environment.

[0178] In at least one embodiment, a dedicated process programming model is implemented. In this model, a single process owns the graphics acceleration module 1546 or an individual graphics processing engine. Since the graphics acceleration module 1546 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owned partition, and the operating system initializes the accelerator integrated circuit for the owned partition when the graphics acceleration module 1546 is allocated.

[0179] In operation, the WD fetch unit 1591 in the accelerator integrated slice 1590 fetches the next WD 1584, which includes an indication of work to be done by one or more graphics processing engines of the graphics acceleration module 1546. Data from the WD 1584 can be stored in the register 1545 and used by the memory management unit (“MMU”) 1539, the interrupt management circuit 1547, and / or the context management circuit 1548, as shown. For example, one embodiment of the MMU 1539 includes a segment / page walk circuit for accessing the segment / page table 1586 within the OS virtual address space 1585. The interrupt management circuit 1547 can process interrupt events (INT) 1592 received from the graphics acceleration module 1546. When performing a graphics operation, the virtual address 1593 generated by the graphics processing engine is translated to a physical address by the MMU 1539.

[0180] In one embodiment, the same register set 1545 is replicated for each graphics processing engine and / or the graphics acceleration module 1546 and can be initialized by the hypervisor or the operating system. Each of these replicated registers can be included in the accelerator integrated slice 1590. Exemplary registers that can be initialized by the hypervisor are shown in Table 1.

[0181] Table 1 – Hypervisor Initialized Registers

[0182]

[0183]

[0184] Exemplary registers that can be initialized by the operating system are shown in Table 2.

[0185] Table 2 – Operating System Initialized Registers

[0186] 1 Process and Thread Identification 2 Effective Address (EA) Context Save / Restore Pointer 3 Virtual Address (VA) Accelerator Utilization Record Pointer 4 Virtual Address (VA) Memory Segmentation Table Pointer 5 Authority Mask 6 Work Descriptor

[0187] In one embodiment, each WD 1584 is specific to a particular graphics acceleration module 1546 and / or a particular graphics processing engine. It contains all the information required for the graphics processing engine to do the work or the work to be done, or it can be a pointer to a memory location where the application has established a command queue of work to be done.

[0188] Figures 16A to 16BAn exemplary graphics processor in accordance with at least one embodiment of the present disclosure is shown. In at least one embodiment, any exemplary graphics processor may be fabricated using one or more IP cores. In addition to what is illustrated, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores. In at least one embodiment, the exemplary graphics processor is for use within a SoC.

[0189] Figure 16A An exemplary graphics processor 1610 of a SoC integrated circuit in accordance with at least one embodiment is shown, which may be fabricated using one or more IP cores. Figure 16B An additional exemplary graphics processor 1640 of a SoC integrated circuit in accordance with at least one embodiment is shown, which may be fabricated using one or more IP cores. In at least one embodiment, Figure 16A the graphics processor 1610 is a low-power graphics processor core. In at least one embodiment, Figure 16B the graphics processor 1640 is a higher-performance graphics processor core. In at least one embodiment, each graphics processor 1610, 1640 may be Figure 11 a variant of the graphics processor 1110.

[0190] In at least one embodiment, with respect to Figures 16A to 16B at least one of the components shown or described is used to implement the techniques and / or functions associated with Figures 1 to 6 those described. In at least one embodiment, the graphics processor 1640 performs one or more operations that convert vertex data to pixels based at least in part on a prefix sum associated with Figure 3 those described, as well as other operations described elsewhere herein.

[0191] In at least one embodiment, the graphics processor 1610 includes a vertex processor 1605 and one or more fragment processors 1615A - 1615N (e.g., 1615A, 1615B, 1615C, 1615D to 1615N - 1, and 1615N). In at least one embodiment, the graphics processor 1610 can execute different shader programs via separate logic such that the vertex processor 1605 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 1615A - 1615N perform fragment (e.g., pixel) shading operations for fragment or pixel or shader programs. In at least one embodiment, the vertex processor 1605 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 1615A - 1615N use the primitives and vertex data generated by the vertex processor 1605 to generate a frame buffer for display on a display device. In at least one embodiment, the fragment processors 1615A - 1615N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform operations similar to pixel shader programs provided in the Direct 3D API.

[0192] In at least one embodiment, the graphics processor 1610 additionally includes one or more MMUs 1620A - 1620B, caches 1625A - 1625B, and circuit interconnects 1630A - 1630B. In at least one embodiment, the one or more MMUs 1620A - 1620B provide virtual - to - physical address mapping for the graphics processor 1610, including for the vertex processor 1605 and / or the fragment processors 1615A - 1615N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in the one or more caches 1625A - 1625B. In at least one embodiment, the one or more MMUs 1620A - 1620B can be synchronized with other MMUs within the system, including one or more MMUs associated with Figure 11 one or more application processors 1105, image processors 1115, and / or video processors 1120 of the system such that each processor 1105 - 1120 can participate in a shared or unified virtual memory system. In at least one embodiment, the one or more circuit interconnects 1630A - 1630B enable the graphics processor 1610 to connect to other IP cores within the SoC via the internal bus of the SoC or via a direct connection.

[0193] In at least one embodiment, the graphics processor 1640 includes Figure 16AOne or more MMUs 1620A - 1620B, caches 1625A - 1625B, and circuit interconnects 1630A - 1630B of the graphics processor 1610. In at least one embodiment, the graphics processor 1640 includes one or more shader cores 1655A - 1655N (e.g., 1655A, 1655B, 1655C, 1655D, 1655E, 1655F, to 1655N - 1 and 1655N), which provide a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1640 includes an inter - core task manager 1645, which acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1655A - 1655N and a tiling unit 1658 to accelerate tiling operations for tile - based rendering, where the rendering operation of a scene is subdivided in image space, e.g., to take advantage of local spatial coherence within the scene or optimize the use of internal caches.

[0194] Figure 17A Shows a graphics core 1700 according to at least one embodiment. In at least one embodiment, the graphics core 1700 can be included within Figure 11 the graphics processor 1110. In at least one embodiment, the graphics core 1700 can be Figure 16BA unified shader core 1655A - 1655N. In at least one embodiment, the graphics core 1700 includes a shared instruction cache 1702, texture units 1718, and cache / shared memory 1720, which are shared by the execution resources within the graphics core 1700. In at least one embodiment, the graphics core 1700 may include multiple slices 1701A - 1701N or partitions per core, and the graphics processor may include multiple instances of the graphics core 1700. The slices 1701A - 1701N may include support logic, which includes local instruction caches 1704A - 1704N, thread schedulers 1706A - 1706N, thread dispatchers 1708A - 1708N, and a set of registers 1710A - 1710N. In at least one embodiment, the slices 1701A - 1701N may include a set of additional functional units (“AFU”) 1712A - 1712N, floating - point units (“FPU”) 1714A - 1714N, integer arithmetic logic units (“ALU”) 1716A - 1716N, address calculation units (“ACU”) 1713A - 1713N, double - precision floating - point units (“DPFPU”) 1715A - 1715N, and matrix processing units (“MPU”) 1717A - 1717N. In at least one embodiment, the graphics core 1700 is referred to as a compute unit or an arithmetic unit.

[0195] In one embodiment, the FPU 1714A - 1714N may perform single - precision (32 - bit) and half - precision (16 - bit) floating - point operations, while the DPFPU 1715A - 1715N may perform double - precision (64 - bit) floating - point operations. In at least one embodiment, the ALU 1716A - 1716N may perform variable - precision integer operations with 8 - bit, 16 - bit, and 32 - bit precision and may be configured for mixed - precision operations. In at least one embodiment, the MPU 1717A - 1717N may also be configured for mixed - precision matrix operations, including half - precision floating - point operations and 8 - bit integer operations. In at least one embodiment, the MPU 1717A - 1717N may perform various matrix operations to accelerate CUDA programs, including enabling accelerated general matrix - to - matrix multiplication (GEMM). In at least one embodiment, the AFU 1712A - 1712N may perform additional logical operations not supported by the floating - point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).

[0196] Figure 17BShows a general - purpose graphics processing unit (GPGPU) 1730 in at least one embodiment. In at least one embodiment, the GPGPU 1730 is highly parallel and suitable for deployment on a multi - chip module. In at least one embodiment, the GPGPU 1730 can be configured such that highly parallel computing operations can be performed by a GPU array. In at least one embodiment, the GPGPU 1730 can be directly linked to other instances of the GPGPU 1730 to create a multi - GPU cluster to improve the execution time for CUDA programs. In at least one embodiment, the GPGPU 1730 includes a host interface 1732 to enable connection to a host processor. In at least one embodiment, the host interface 1732 is a PCIe interface. In at least one embodiment, the host interface 1732 can be a vendor - specific communication interface or communication fabric. In at least one embodiment, the GPGPU 1730 receives commands from the host processor and uses a global scheduler 1734 to dispatch execution threads associated with those commands to a set of compute clusters 1736A - 1736H. In at least one embodiment, the compute clusters 1736A - 1736H share a cache memory 1738. In at least one embodiment, the cache memory 1738 can serve as a cache - of - caches for the cache memories within the compute clusters 1736A - 1736H.

[0197] In at least one embodiment, at least one component shown or described with respect to Figures 17A to 17B is used to implement techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, the GPGPU 1730 performs one or more operations that convert vertex data to pixels based at least in part on a prefix sum associated with Figure 3 described, and operations described elsewhere herein.

[0198] In at least one embodiment, the GPGPU 1730 includes memories 1744A - 1744B coupled to the compute clusters 1736A - 1736H via a set of memory controllers 1742A - 1742B. In at least one embodiment, the memories 1744A - 1744B can include various types of memory devices, including dynamic random - access memory (DRAM) or graphics random - access memory, such as synchronous graphics random - access memory (“SGRAM”), including graphics double - data rate (“GDDR”) memory.

[0199] In at least one embodiment, each of the compute clusters 1736A - 1736H includes a set of graphics cores, such as Figure 17AThe graphics core 1700, which can include various types of integer and floating-point logic units, can perform computational operations with various precisions, including those suitable for computations associated with CUDA programs. For example, in at least one embodiment, at least one subset of the floating-point units in each of the compute clusters 1736A - 1736H can be configured to perform 16-bit or 32-bit floating-point operations, while different subsets of floating-point units can be configured to perform 64-bit floating-point operations.

[0200] In at least one embodiment, multiple instances of the GPGPU 1730 can be configured to operate as compute clusters. The compute clusters 1736A - 1736H can implement any technically feasible communication technology for synchronization and data exchange. In at least one embodiment, multiple instances of the GPGPU 1730 communicate via the host interface 1732. In at least one embodiment, the GPGPU 1730 includes an I / O hub 1739 that couples the GPGPU 1730 to the GPU link 1740, enabling direct connection to other instances of the GPGPU 1730. In at least one embodiment, the GPU link 1740 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of the GPGPU 1730. In at least one embodiment, the GPU link 1740 is coupled to a high-speed interconnect to send and receive data to and from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of the GPGPU 1730 are located in separate data processing systems and communicate via network devices accessible via the host interface 1732. In at least one embodiment, the GPU link 1740 can be configured to be able to connect to the host processor, in addition to or in place of the host interface 1732. In at least one embodiment, the GPGPU 1730 can be configured to execute CUDA programs.

[0201] Figure 18A Shows a parallel processor 1800 according to at least one embodiment. In at least one embodiment, the various components of the parallel processor 1800 can be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (“ASICs”), or FPGAs.

[0202] In at least one embodiment, the parallel processor 1800 includes a parallel processing unit 1802. In at least one embodiment, the parallel processing unit 1802 includes an I / O unit 1804 that enables communication with other devices, including other instances of the parallel processing unit 1802. In at least one embodiment, the I / O unit 1804 can be directly connected to other devices. In at least one embodiment, the I / O unit 1804 is connected to other devices by using a hub or switch interface (e.g., memory hub 1805). In at least one embodiment, the connection between the memory hub 1805 and the I / O unit 1804 forms a communication link. In at least one embodiment, the I / O unit 1804 is connected to a host interface 1806 and a memory crossbar 1816, where the host interface 1806 receives commands for performing processing operations and the memory crossbar 1816 receives commands for performing memory operations.

[0203] In at least one embodiment, when the host interface 1806 receives a command buffer via the I / O unit 1804, the host interface 1806 can initiate work operations to execute those commands to the front end 1808. In at least one embodiment, the front end 1808 is coupled to a scheduler 1810 configured to allocate commands or other work items to a processing array 1812. In at least one embodiment, the scheduler 1810 ensures that the processing array 1812 is properly configured and in an active state before tasks are assigned to the processing array 1812 within the processing array 1812. In at least one embodiment, the scheduler 1810 is implemented by firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 1810 can be configured to perform complex scheduling and work allocation operations at both coarse-grained and fine-grained levels, enabling fast preemption and context switching of threads executing on the processing array 1812. In at least one embodiment, host software can demonstrate a workload for scheduling on the processing array 1812 via one of a plurality of graphics processing doorbells. In at least one embodiment, the workload can then be automatically allocated on the processing array 1812 by the scheduler 1810 logic within the microcontroller including the scheduler 1810.

[0204] In at least one embodiment, the processing array 1812 may include up to "N" processing clusters (e.g., cluster 1814A, cluster 1814B to cluster 1814N). In at least one embodiment, each of the clusters 1814A - 1814N of the processing array 1812 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 1810 may use various scheduling and / or work assignment algorithms to assign work to the clusters 1814A - 1814N of the processing array 1812, which may vary according to the workload generated by each type of program or computation. In at least one embodiment, the scheduling may be handled dynamically by the scheduler 1810, or may be assisted in part by compiler logic during the compilation of the program logic configured to be executed by the processing array 1812. In at least one embodiment, different clusters 1814A - 1814N of the processing array 1812 may be assigned to process different types of programs or to perform different types of computations.

[0205] In at least one embodiment, the processing array 1812 may be configured to perform various types of parallel processing operations. In at least one embodiment, the processing array 1812 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, the processing array 1812 may include logic for performing processing tasks that include filtering of video and / or audio data, performing modeling operations, including physical operations, and performing data transformations.

[0206] In at least one embodiment, the processing array 1812 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing array 1812 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, and tessellation logic and other vertex processing logic. In at least one embodiment, the processing array 1812 may be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 1802 may transfer data from the system memory via the I / O unit 1804 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 1822) during processing and then written back to the system memory.

[0207] In at least one embodiment, when the parallel processing unit 1802 is used to perform graphics processing, the scheduler 1810 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to the multiple clusters 1814A - 1814N of the processing array 1812. In at least one embodiment, portions of the processing array 1812 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen - space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of the clusters 1814A - 1814N can be stored in a buffer to allow the transfer of intermediate data between the clusters 1814A - 1814N for further processing.

[0208] In at least one embodiment, the processing array 1812 can receive a processing task to be executed via the scheduler 1810, which receives commands defining the processing task from the front - end 1808. In at least one embodiment, the processing task can include an index of the data to be processed, such as can include surface (patch) data, primitive data, vertex data, and / or pixel data, as well as status parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 1810 can be configured to obtain the index corresponding to the task, or can receive the index from the front - end 1808. In at least one embodiment, the front - end 1808 can be configured to ensure that the processing array 1812 is configured in a valid state before starting the workload specified by an incoming command buffer (e.g., batch - buffer, push - buffer, etc.).

[0209] In at least one embodiment, each of one or more instances of the parallel processing unit 1802 may be coupled to the parallel processor memory 1822. In at least one embodiment, the parallel processor memory 1822 may be accessed via a memory crossbar 1816 that may receive memory requests from the processing array 1812 as well as the I / O unit 1804. In at least one embodiment, the memory crossbar 1816 may access the parallel processor memory 1822 via a memory interface 1818. In at least one embodiment, the memory interface 1818 may include a plurality of partitioning units (e.g., partitioning unit 1820A, partitioning unit 1820B through partitioning unit 1820N), each of which may be coupled to a portion (e.g., a memory unit) of the parallel processor memory 1822. In at least one embodiment, the plurality of partitioning units 1820A-1820N are configured to be equal to the number of memory units such that the first partitioning unit 1820A has a corresponding first memory unit 1824A, the second partitioning unit 1820B has a corresponding memory unit 1824B, and the Nth partitioning unit 1820N has a corresponding Nth memory unit 1824N. In at least one embodiment, the number of partitioning units 1820A-1820N may not be equal to the number of memory devices.

[0210] In at least one embodiment, the memory units 1824A-1824N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, the memory units 1824A-1824N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps may be stored across the memory units 1824A-1824N, allowing the partitioning units 1820A-1820N to write portions of each rendering target in parallel to effectively utilize the available bandwidth of the parallel processor memory 1822. In at least one embodiment, local instances of the parallel processor memory 1822 may be excluded in favor of a unified memory design that utilizes system memory in combination with local cache memory.

[0211] In at least one embodiment, any one of clusters 1814A - 1814N of processing array 1812 can process data to be written into any of memory cells 1824A - 1824N within parallel processor memory 1822. In at least one embodiment, memory crossbar 1816 can be configured to transfer the output of each of clusters 1814A - 1814N to any of partitioning units 1820A - 1820N or to another cluster 1814A - 1814N, where the cluster 1814A - 1814N can perform other processing operations on the output. In at least one embodiment, each of clusters 1814A - 1814N can communicate with memory interface 1818 via memory crossbar 1816 to read from or write to various external storage devices. In at least one embodiment, memory crossbar 1816 has a connection to memory interface 1818 to communicate with I / O unit 1804 and a connection to a local instance of parallel processor memory 1822, thereby enabling processing units within different processing clusters 1814A - 1814N to communicate with system memory or other memories that are not local to parallel processing units 1802. In at least one embodiment, memory crossbar 1816 can use virtual channels to separate the traffic flow between clusters 1814A - 1814N and partitioning units 1820A - 1820N.

[0212] In at least one embodiment, multiple instances of parallel processing unit 1802 can be provided on a single insertion card, or multiple insertion cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 1802 can be configured to operate interoperably, even if the different instances have different numbers of processing cores, different numbers of local parallel processor memories, and / or other configuration differences. For example, in at least one embodiment, some instances of parallel processing unit 1802 can include floating - point units with higher precision relative to other instances. In at least one embodiment, systems incorporating one or more instances of parallel processing unit 1802 or parallel processor 1800 can be implemented in various configurations and form factors, including but not limited to desktop computers, laptop computers, or handheld personal computers, servers, workstations, gaming consoles, and / or embedded systems.

[0213] Figure 18BA processing cluster 1894 according to at least one embodiment is shown. In at least one embodiment, the processing cluster 1894 is included within a parallel processing unit. In at least one embodiment, the processing cluster 1894 is an instance of one of the processing clusters 1814A - 1814N of FIG. 18. In at least one embodiment, the processing cluster 1894 may be configured to execute a number of threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issue techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) techniques are used to support the parallel execution of a large number of generally synchronized threads, which uses a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster 1894.

[0214] In at least one embodiment, Figures 18A to 18B at least one component shown or described in is used to implement techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, the processing cluster 1894 performs one or more operations that are at least partially based on the prefix sums described in conjunction with Figure 3 and the prefix sums as described herein to transform vertex data into pixels.

[0215] In at least one embodiment, the operation of the processing cluster 1894 may be controlled by assigning processing tasks to the pipeline manager 1832 of the SIMT parallel processor. In at least one embodiment, the pipeline manager 1832 receives instructions from the scheduler 1810 of FIG. 18 and manages the execution of these instructions through the graphics multiprocessor 1834 and / or the texture unit 1836. In at least one embodiment, the graphics multiprocessor 1834 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 1894. In at least one embodiment, one or more instances of the graphics multiprocessor 1834 may be included within the processing cluster 1894. In at least one embodiment, the graphics multiprocessor 1834 may process data, and the data crossbar 1840 may be used to distribute the processed data to one of a number of possible destinations (including other shader units). In at least one embodiment, the pipeline manager 1832 may facilitate the distribution of the processed data by specifying the destination of the processed data to be distributed via the data crossbar 1840.

[0216] In at least one embodiment, each graphics multiprocessor 1834 within processing cluster 1894 may include the same set of functional execution logic (e.g., arithmetic logic unit, load store unit (LSU), etc.). In at least one embodiment, the functional execution logic may be configured in a pipeline manner, where new instructions may be issued before previous instructions are completed. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may exist.

[0217] In at least one embodiment, the instructions transmitted to processing cluster 1894 constitute a thread. In at least one embodiment, a set of threads executed across a group of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within the thread group may be assigned to a different processing engine within graphics multiprocessor 1834. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within graphics multiprocessor 1834. In at least one embodiment, when the number of threads included in the thread group is less than the number of processing engines, one or more processing engines may be idle during the cycle of processing the thread group. In at least one embodiment, the thread group may also include more threads than the number of processing engines within graphics multiprocessor 1834. In at least one embodiment, when the thread group includes more threads than the number of processing engines within graphics multiprocessor 1834, processing may be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups may be executed simultaneously on graphics multiprocessor 1834.

[0218] In at least one embodiment, graphics multiprocessor 1834 includes an internal cache memory to perform load and store operations. In at least one embodiment, graphics multiprocessor 1834 may forgo the internal cache and use the cache memory within processing cluster 1894 (e.g., L1 cache 1848). In at least one embodiment, each graphics multiprocessor 1834 may also access a partitioning unit (e.g., Figure 18AL2 caches within the partition units 1820A - 1820N), which are shared among all processing clusters 1894 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 1834 may also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 1802 may be used as global memory. In at least one embodiment, the processing cluster 1894 includes multiple instances of the graphics multiprocessor 1834, which may share common instructions and data that can be stored in the L1 cache 1848.

[0219] In at least one embodiment, each processing cluster 1894 may include an MMU 1845 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 1845 may reside within the memory interface 1818 of FIG. 18. In at least one embodiment, the MMU 1845 includes a set of page table entries (PTEs) that are used to map virtual addresses to the physical addresses of tiles (more information about tiles is discussed) and optionally to cache line indices. In at least one embodiment, the MMU 1845 may include a translation lookaside buffer (TLB) or a cache that may reside within the graphics multiprocessor 1834 or the L1 cache 1848 or the processing cluster 1894. In at least one embodiment, the physical addresses are processed to allocate surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index may be used to determine whether a request to a cache line is a hit or a miss.

[0220] In at least one embodiment, the processing cluster 1894 may be configured such that each graphics multiprocessor 1834 is coupled to a texture unit 1836 to perform texture mapping operations, e.g., which may involve determining texture sample locations, reading texture data, and filtering texture data. In at least one embodiment, texture data is read as needed from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 1834, and texture data is fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 1834 outputs the processed task to the data crossbar 1840 to provide the processed task to another processing cluster 1894 for further processing or to store the processed task in an L2 cache, local parallel processor memory, or system memory via the memory crossbar 1816. In at least one embodiment, a raster front operation unit (“preROP”) 1842 is configured to receive data from the graphics multiprocessor 1834 and direct the data to a ROP unit, which may be located together with the partitioning units described herein (e.g., the partitioning units 1820A - 1820N of FIG. 18). In at least one embodiment, the PreROP 1842 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.

[0221] Figure 18C A graphics multiprocessor 1896 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 1896 is Figure 18B the graphics multiprocessor 1834. In at least one embodiment, the graphics multiprocessor 1896 is coupled to the pipeline manager 1832 of the processing cluster 1894. In at least one embodiment, the graphics multiprocessor 1896 has an execution pipeline that includes, but is not limited to, an instruction cache 1852, an instruction unit 1854, an address mapping unit 1856, a register file 1858, one or more GPGPU cores 1862, and one or more LSUs 1866. The GPGPU cores 1862 and LSUs 1866 are coupled to a cache memory 1872 and a shared memory 1870 via a memory and cache interconnect 1868.

[0222] In at least one embodiment, the instruction cache 1852 receives a stream of instructions to be executed from the pipeline manager 1832. In at least one embodiment, the instructions are cached in the instruction cache 1852 and dispatched for execution by the instruction unit 1854. In one embodiment, the instruction unit 1854 may dispatch instructions as a thread group (e.g., a warp), and each thread of the thread group is assigned to a different execution unit within the GPGPU core 1862. In at least one embodiment, instructions may access any local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, the address mapping unit 1856 may be used to translate an address in the unified address space into a different memory address that can be accessed by the LSU 1866.

[0223] In at least one embodiment, the register file 1858 provides a set of registers for the functional units of the graphics multiprocessor 1896. In at least one embodiment, the register file 1858 provides temporary storage for the operands of the data paths of the functional units (e.g., the GPGPU core 1862, the LSU 1866) connected to the graphics multiprocessor 1896. In at least one embodiment, the register file 1858 is partitioned among each of the functional units such that a dedicated portion of the register file 1858 is assigned to each functional unit. In at least one embodiment, the register file 1858 is partitioned among different thread groups being executed by the graphics multiprocessor 1896.

[0224] In at least one embodiment, the GPGPU cores 1862 may each include an FPU and / or an ALU for executing the instructions of the graphics multiprocessor 1896. The GPGPU cores 1862 may be architecturally similar or the architectures may vary. In at least one embodiment, a first portion of the GPGPU core 1862 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU core 1862 includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 1896 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copy rectangle or pixel blend operations. In at least one embodiment, one or more of the GPGPU cores 1862 may also include fixed or special-function logic.

[0225] In at least one embodiment, the GPGPU core 1862 includes SIMD logic capable of executing a single instruction on multiple sets of data. In at least one embodiment, the GPGPU core 1862 can physically execute SIMD4, SIMD8, and SIMD9 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time or automatically generated when executing a program written and compiled for a single-program multiple-data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model can be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.

[0226] In at least one embodiment, the memory and cache interconnect 1868 is an interconnect network that connects each functional unit of the graphics multiprocessor 1896 to the register file 1858 and the shared memory 1870. In at least one embodiment, the memory and cache interconnect 1868 is a crossbar interconnect that allows the LSU 1866 to perform load and store operations between the shared memory 1870 and the register file 1858. In at least one embodiment, the register file 1858 can operate at the same frequency as the GPGPU core 1862, resulting in a very low latency for data transfer between the GPGPU core 1862 and the register file 1858. In at least one embodiment, the shared memory 1870 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 1896. In at least one embodiment, the cache memory 1872 can be used as, for example, a data cache to cache texture data communicated between the functional units and the texture unit 1836. In at least one embodiment, the shared memory 1870 can also be used as a program-managed cache. In at least one embodiment, in addition to the automatically cached data stored in the cache memory 1872, threads executing on the GPGPU core 1862 can also programmatically store data in the shared memory.

[0227] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated with the core on the same package or die and communicatively coupled to the core via an internal processor bus / interconnect (i.e., internal to the package or die). In at least one embodiment, regardless of how the GPU is connected, the processor core can allocate work to the GPU in the form of a sequence of commands / instructions contained in the WD. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0228] Figure 19 FIG. 1900 illustrates a graphics processor according to at least one embodiment. In at least one embodiment, the graphics processor 1900 includes a ring interconnect 1902, a pipeline front end 1904, a media engine 1937, and graphics cores 1980A - 1980N. In at least one embodiment, the ring interconnect 1902 couples the graphics processor 1900 to other processing units, including other graphics processors or one or more general processor cores. In at least one embodiment, the graphics processor 1900 is one of many processors integrated within a multi-core processing system.

[0229] In at least one embodiment, Figure 19 at least one component shown or described herein is used to implement the techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, the graphics processor 1900 performs one or more operations for converting vertex data to pixels that are at least partially based on a prefix sum as described in conjunction with Figure 3 described, and operations described elsewhere herein.

[0230] In at least one embodiment, the graphics processor 1900 receives multiple batches of commands via the ring interconnect 1902. In at least one embodiment, the input commands are interpreted by the command stream converter 1903 in the pipeline front end 1904. In at least one embodiment, the graphics processor 1900 includes scalable execution logic to perform 3D geometry processing and media processing via the graphics cores 1980A - 1980N. In at least one embodiment, for 3D geometry processing commands, the command stream converter 1903 provides the commands to the geometry pipeline 1936. In at least one embodiment, for at least some of the media processing commands, the command stream converter 1903 provides the commands to the video front end 1934, which is coupled to the media engine 1937. In at least one embodiment, the media engine 1937 includes a Video Quality Engine (VQE) 1930 for video and image post - processing, and a Multi - Format Encoding / Decoding (MFX) 1933 engine for providing hardware - accelerated media data encoding and decoding. In at least one embodiment, the geometry pipeline 1936 and the media engine 1937 each generate execution threads for the thread execution resources provided by at least one graphics core 1980A.

[0231] In at least one embodiment, the graphics processor 1900 includes scalable thread execution resources characterized by modular graphics cores 1980A - 1980N (sometimes referred to as core slices), each modular core having multiple sub - cores 1950A - 1950N, 1960A - 1960N (sometimes referred to as core sub - slices). In at least one embodiment, the graphics processor 1900 can have any number of graphics cores 1980A through 1980N. In at least one embodiment, the graphics processor 1900 includes a graphics core 1980A having at least a first sub - core 1950A and a second sub - core 1960A. In at least one embodiment, the graphics processor 1900 is a low - power processor having a single sub - core (e.g., 1950A). In at least one embodiment, the graphics processor 1900 includes multiple graphics cores 1980A - 1980N, each graphics core including a set of first sub - cores 1950A - 1950N and a set of second sub - cores 1960A - 1960N. In at least one embodiment, each of the first sub - cores 1950A - 1950N includes at least a first set of Execution Units (EU) 1952A - 1952N and media / texture samplers 1954A - 1954N. In at least one embodiment, each of the second sub - cores 1960A - 1960N includes at least a second set of execution units 1962A - 1962N and samplers 1964A - 1964N. In at least one embodiment, each of the sub - cores 1950A - 1950N, 1960A - 1960N shares a set of shared resources 1970A - 1970N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.

[0232] Figure 20 Illustrates a processor 2000 according to at least one embodiment. In at least one embodiment, the processor 2000 may include, but is not limited to, logic circuitry for executing instructions. In at least one embodiment, the processor 2000 may execute instructions, including x86 instructions, ARM instructions, special instructions for ASICs, etc. In at least one embodiment, the processor 2010 may include registers for storing packed data, such as 64-bit wide MMXTM registers in a microprocessor enabled with MMX technology by Intel Corporation in Santa Clara, California. In at least one embodiment, the MMX registers available in integer and floating-point forms may operate with packed data elements that accompany SIMD and Streaming SIMD Extensions (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or later versions (generally referred to as “SSEx” technologies) may hold such packed data operands. In at least one embodiment, the processor 2010 may execute instructions to accelerate CUAD programs.

[0233] In at least one embodiment, Figure 20 at least one component shown or described therein is used to implement the techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, the processor 2000 performs one or more operations for converting vertex data to pixels that are at least partially based on the prefix sum associated with Figure 3 described, and operations described elsewhere herein.

[0234] In at least one embodiment, the processor 2000 includes an in-order front end (“front end”) 2001 to fetch instructions to be executed and prepare the instructions for later use in the processor pipeline. In at least one embodiment, the front end 2001 may include several units. In at least one embodiment, the instruction prefetcher 2026 fetches instructions from memory and provides the instructions to the instruction decoder 2028, which in turn decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2028 decodes the received instructions into one or more operations of so-called “microinstructions” or “micro-operations” (also referred to as “uops” or “microinstructions”) for execution. In at least one embodiment, the instruction decoder 2028 parses the instructions into an opcode and corresponding data and control fields, which can be used by the microarchitecture to perform operations. In at least one embodiment, the trace cache 2030 may assemble the decoded microinstructions into a program-ordered sequence or trace in the microinstruction queue 2034 for execution. In at least one embodiment, when the trace cache 2030 encounters a complex instruction, the microcode ROM 2032 provides the microinstructions required to complete the operation.

[0235] In at least one embodiment, some instructions may be converted into a single micro-operation, while other instructions require several micro-operations to complete the entire operation. In at least one embodiment, if more than four microinstructions are required to complete an instruction, the instruction decoder 2028 may access the microcode ROM 2032 to execute the instruction. In at least one embodiment, the instructions may be decoded into a small number of microinstructions for processing at the instruction decoder 2028. In at least one embodiment, if multiple microinstructions are required to complete an operation, the instructions may be stored in the microcode ROM 2032. In at least one embodiment, the trace cache 2030 refers to an entry point programmable logic array (“PLA”) to determine the correct microinstruction pointer for reading a microcode sequence from the microcode ROM 2032 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after the microcode ROM 2032 finishes sequencing the micro-operations of the instruction, the front end 2001 of the machine may resume fetching micro-operations from the trace cache 2030.

[0236] In at least one embodiment, an out-of-order execution engine (“out-of-order engine”) 2003 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth and reorder the instruction stream to optimize performance as the instructions descend down the pipeline and are scheduled for execution. The out-of-order execution engine 2003 includes, but is not limited to, an allocator / register renamer 2040, a memory micro-instruction queue 2042, an integer / floating-point micro-instruction queue 2044, a memory scheduler 2046, a fast scheduler 2002, a slow / general floating-point scheduler (“slow / general FP scheduler”) 2004, and a simple floating-point scheduler (“simple FP scheduler”) 2006. In at least one embodiment, the fast scheduler 2002, the slow / general floating-point scheduler 2004, and the simple floating-point scheduler 2006 are also collectively referred to as “micro-instruction schedulers 2002, 2004, 2006”. The allocator / register renamer 2040 allocates the machine buffers and resources required for each micro-instruction to execute in order. In at least one embodiment, the allocator / register renamer 2040 renames logical registers to entries in the register file. In at least one embodiment, the allocator / register renamer 2040 also allocates entries for each micro-instruction in one of two micro-instruction queues, the memory micro-instruction queue 2042 for memory operations and the integer / floating-point micro-instruction queue 2044 for non-memory operations, in front of the memory scheduler 2046 and the micro-instruction schedulers 2002, 2004, 2006. In at least one embodiment, the micro-instruction schedulers 2002, 2004, 2006 determine when a micro-instruction is ready for execution based on the readiness of their dependent input register operand sources and the availability of execution resources micro-instructions that need to be completed. In at least one embodiment, the fast scheduler 2002 of at least one embodiment may be scheduled on each half of the main clock cycle, while the slow / general floating-point scheduler 2004 and the simple floating-point scheduler 2006 may be scheduled once per main processor clock cycle. In at least one embodiment, the micro-instruction schedulers 2002, 2004, 2006 arbitrate the scheduling ports to schedule micro-instructions for execution.

[0237] In at least one embodiment, execution block 2011 includes, but is not limited to, integer register file / branch network 2008, floating-point register file / branch network (“FP register file / branch network”) 2010, address generation units (“AGUs”) 2012 and 2014, fast arithmetic logic units (“fast ALUs”) 2016 and 2018, slow ALU 2020, floating-point ALU (“FP”) 2022, and floating-point move unit (“FP move”) 2024. In at least one embodiment, integer register file / branch network 2008 and floating-point register file / bypass network 2010 are also referred to herein as “register files 2008, 2010”. In at least one embodiment, AGUs 2012 and 2014, fast ALUs 2016 and 2018, slow ALU 2020, floating-point ALU 2022, and floating-point move unit 2024 are also referred to herein as “execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024”. In at least one embodiment, the execution block may include, but is not limited to, any number (including zero) and type of register files, branch networks, address generation units, and execution units (in any combination).

[0238] In at least one embodiment, register files 2008, 2010 may be arranged between microinstruction schedulers 2002, 2004, 2006 and execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024. In at least one embodiment, integer register file / branch network 2008 performs integer operations. In at least one embodiment, floating-point register file / branch network 2010 performs floating-point operations. In at least one embodiment, each of register files 2008, 2010 may include, but is not limited to, a branch network that may bypass or forward a just-completed result that has not yet been written to the register file to a new dependent. In at least one embodiment, register files 2008, 2010 may communicate data with each other. In at least one embodiment, integer register file / branch network 2008 may include, but is not limited to, two separate register files, one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, floating-point register file / branch network 2010 may include, but is not limited to, 128-bit-wide entries, since floating-point instructions typically have operands with widths of 64 to 128 bits.

[0239] In at least one embodiment, execution units 2012, 2014, 2016, 2018, 2020, 2022, 2024 may execute instructions. In at least one embodiment, register files 2008, 2010 store integer and floating-point data operand values that the microinstructions need to execute. In at least one embodiment, processor 2000 may include, but is not limited to, any number of execution units 2012, 2014, 2016, 2018, 2020, 2022, 2024 and combinations thereof. In at least one embodiment, floating-point ALU 2022 and floating-point move unit 2024 may execute floating-point, MMX, SIMD, AVX, and SSE or other operations, including specialized machine learning instructions. In at least one embodiment, floating-point ALU 2022 may include, but is not limited to, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro-operations. In at least one embodiment, instructions involving floating-point values may be processed with floating-point hardware. In at least one embodiment, ALU operations may be passed to fast ALUs 2016, 2018. In at least one embodiment, fast ALUs 2016, 2018 may execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, most complex integer operations enter slow ALU 2020 because slow ALU 2020 may include, but is not limited to, integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations may be performed by AGUs 2012, 2014. In at least one embodiment, fast ALU 2016, fast ALU 2018, and slow ALU 2020 may perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 2016, fast ALU 2018, and slow ALU 2020 may be implemented to support various data bit sizes including 16, 32, 128, 256, etc. In at least one embodiment, floating-point ALU 2022 and floating-point move unit 2024 may be implemented to support a certain range of operands with bits of various widths. In at least one embodiment, floating-point ALU 2022 and floating-point move unit 2024 may operate on 128-bit wide packed data operands in combination with SIMD and multimedia instructions.

[0240] In at least one embodiment, the microinstruction schedulers 2002, 2004, 2006 schedule dependent operations before the completion of the execution of the parent load. In at least one embodiment, since microinstructions can be scheduled and executed speculatively in the processor 2000, the processor 2000 may also include logic for handling memory misses. In at least one embodiment, if a data load miss occurs in the data cache, there may be dependent operations running in the pipeline that leave the scheduler temporarily without the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and independent operations may be allowed to complete. In at least one embodiment, the scheduler and the replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.

[0241] In at least one embodiment, the term "register" may refer to an on-board processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, the registers may be those that can be used from outside the processor (from the programmer's perspective). In at least one embodiment, the registers may not be limited to a particular type of circuit. Instead, in at least one embodiment, the registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented using a variety of different techniques by circuits within the processor, such as dedicated physical registers, physical registers dynamically allocated using register renaming, combinations of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also contains eight multimedia SIMD registers for packing data.

[0242] Figure 21Shows a processor 2100 according to at least one embodiment. In at least one embodiment, the processor 2100 includes, but is not limited to, one or more processor cores (cores) 2102A - 2102N, an integrated memory controller 2114, and an integrated graphics processor 2108. In at least one embodiment, the processor 2100 may include additional cores up to and including the additional core 2102N represented by the dashed box. In at least one embodiment, each processor core 2102A - 2102N includes one or more internal cache units 2104A - 2104N. In at least one embodiment, each processor core may also access one or more shared cache units 2106. In at least one embodiment, one or more processor cores 2102A - 2102N are referred to as one or more computing units or arithmetic units. In at least one embodiment, the processor 2100 includes one or more circuits for performing compiler and / or other operations, such as those described above in connection with Figures 1 to 6 the operations described. In at least one embodiment, the graphics processor 2108 performs operations, such as those described above in connection with Figures 1 to 6 the operations described.

[0243] In at least one embodiment, Figure 21 at least one component shown or described in Figures 1 to is used to implement the techniques and / or functions described in connection with ​ In at least one embodiment, the processor 2100 performs one or more operations that convert vertex data to pixels based at least in part on the prefix sum described in connection with

[0244] and the operations described elsewhere herein.

[0245] In at least one embodiment, the processor 2100 may further include a set of one or more bus controller units 2116 and a system agent core 2110. In at least one embodiment, one or more bus controller units 2116 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 2110 provides management functions for various processor components. In at least one embodiment, the system agent core 2110 includes one or more integrated memory controllers 2114 to manage access to various external memory devices (not shown).

[0246] In at least one embodiment, one or more processor cores 2102A - 2102N include support for simultaneous multi-threading. In at least one embodiment, the system agent core 2110 includes components for coordinating and operating the processor cores 2102A - 2102N during multi-threaded processing. In at least one embodiment, the system agent core 2110 may additionally include a power control unit (PCU) that includes logic and components to regulate one or more power states of the processor cores 2102A - 2102N and the graphics processor 2108.

[0247] In at least one embodiment, the processor 2100 additionally includes a graphics processor 2108 to perform graphics processing operations. In at least one embodiment, the graphics processor 2108 is coupled to the shared cache unit 2106 and the system agent core 2110 that includes one or more integrated memory controllers 2114. In at least one embodiment, the system agent core 2110 further includes a display controller 2111 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 2111 may also be a separate module coupled to the graphics processor 2108 via at least one interconnect, or may be integrated within the graphics processor 2108.

[0248] In at least one embodiment, a ring-based interconnect unit 2112 is used to couple the internal components of the processor 2100. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 2108 is coupled to the ring interconnect 2112 via an I / O link 2113.

[0249] In at least one embodiment, the I / O link 2113 represents at least one of a variety of I / O interconnections, including a package I / O interconnection that facilitates communication between various processor components and a high-performance embedded memory module 2118 (e.g., an eDRAM module). In at least one embodiment, each of the processor cores 2102A - 2102N and the graphics processor 2108 uses the embedded memory module 2118 as a shared LLC.

[0250] In at least one embodiment, the processor cores 2102A - 2102N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 2102A - 2102N are heterogeneous in terms of the ISA, where one or more of the processor cores 2102A - 2102N execute a common instruction set, while one or more other processor cores 2102A - 2102N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, in terms of the microarchitecture, the processor cores 2102A - 2102N are heterogeneous, where one or more cores with relatively high power consumption are coupled to one or more power cores with lower power consumption. In at least one embodiment, the processor 2100 can be implemented on one or more chips or be implemented as a SoC integrated circuit.

[0251] ​ A graphics processor core 2200 according to at least one of the described embodiments is shown. In at least one embodiment, the graphics processor core 2200 is included within a graphics core array. In at least one embodiment, the graphics processor core 2200 (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 2200 is an example of a graphics core slice, and the graphics processors described herein can include multiple graphics core slices based on a target power and performance envelope. In at least one embodiment, each graphics core 2200 can include a fixed - function block 2230, also referred to as a sub - slice, coupled to a plurality of sub - cores 2201A - 2201F, which includes modular blocks of general - purpose and fixed - function logic.

[0252] In at least one embodiment, ​ at least one of the components shown or described is used to implement the techniques and / or functions associated with ​ described. In at least one embodiment, the graphics core 2200 performs one or more operations that convert vertex data to pixels based at least in part on a prefix sum described in connection with ​ and the operations described elsewhere herein.

[0253] In at least one embodiment, the fixed function block 2230 includes a geometry / fixed function pipeline 2236, e.g., in a lower performance and / or lower power graphics processor implementation, the geometry / fixed function pipeline 2236 may be shared by all sub-cores in the graphics processor 2200. In at least one embodiment, the geometry / fixed function pipeline 2236 includes a 3D fixed function pipeline, a video front end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages a unified return buffer.

[0254] In at least one embodiment, the fixed function block 2230 further includes a graphics SoC interface 2237, a graphics microcontroller 2238, and a media pipeline 2239. The graphics SoC interface 2237 provides an interface between the graphics core 2200 and other processor cores in the SoC integrated circuit system. In at least one embodiment, the graphics microcontroller 2238 is a programmable sub-processor that can be configured to manage various functions of the graphics processor 2200, including thread dispatching, scheduling, and preemption. In at least one embodiment, the media pipeline 2239 includes logic that aids in decoding, encoding, preprocessing, and / or post-processing multimedia data including image and video data. In at least one embodiment, the media pipeline 2239 implements media operations via requests to the computation or sampling logic within the sub-cores 2201-2201F.

[0255] In at least one embodiment, the SoC interface 2237 enables the graphics core 2200 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared LLC memory, system RAM, and / or embedded on-chip or package DRAM. In at least one embodiment, the SoC interface 2237 can also enable communication with fixed function devices within the SoC (e.g., a camera imaging pipeline), and enable the use and / or implementation of global memory atoms that can be shared between the graphics core 2200 and the CPU within the SoC. In at least one embodiment, the SoC interface 2237 can also implement power management control for the graphics core 2200 and enable an interface between the clock domain of the graphics core 2200 and other clock domains within the SoC. In at least one embodiment, the SoC interface 2237 enables receipt of command buffers from a command stream converter and global thread dispatcher, which are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, when a media operation is to be performed, the commands and instructions can be dispatched to the media pipeline 2239, or when a graphics processing operation is to be performed, they can be allocated to the geometry and fixed function pipelines (e.g., geometry and fixed function pipeline 2236, geometry and fixed function pipeline 2214).

[0256] In at least one embodiment, the graphics microcontroller 2238 may be configured to perform various scheduling and management tasks for the graphics core 2200. In at least one embodiment, the graphics microcontroller 2238 may perform graphics and / or compute workload scheduling on various graphics parallel engines within the execution unit (EU) arrays 2202A - 2202F, 2204A - 2204F in the sub - cores 2201A - 2201F. In at least one embodiment, host software executing on the CPU core of the SoC including the graphics core 2200 may submit a workload to one of the multiple graphics processor doorbells, which invokes a scheduling operation on an appropriate graphics engine. In at least one embodiment, the scheduling operations include determining which workload to run next, submitting the workload to a command stream converter, preempting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, the graphics microcontroller 2238 may also facilitate a low - power or idle state for the graphics core 2200, thereby providing the graphics core 2200 with the ability to save and restore registers across low - power state transitions independent of the operating system and / or graphics driver software on the system within the graphics core 2200.

[0257] In at least one embodiment, the graphics core 2200 may have more or fewer sub - cores than the shown sub - cores 2201A - 2201F, up to N modular sub - cores. For each set of N sub - cores, in at least one embodiment, the graphics core 2200 may further include shared functional logic 2210, shared and / or cache memory 2212, a geometry / fixed - function pipeline 2214, and additional fixed - function logic 2216 to accelerate various graphics and compute processing operations. In at least one embodiment, the shared functional logic 2210 may include logic units (e.g., samplers, math, and / or inter - thread communication logic) that can be shared by each of the N sub - cores within the graphics core 2200. The shared and / or cache memory 2212 may be the LLC for the N sub - cores 2201A - 2201F within the graphics core 2200 and may also be used as shared memory accessible by multiple sub - cores. In at least one embodiment, a geometry / fixed - function pipeline 2214 may be included in place of the geometry / fixed - function pipeline 2236 within the fixed - function block 2230 and may include the same or similar logic units.

[0258] In at least one embodiment, the graphics core 2200 includes additional fixed function logic 2216, which may include various fixed function acceleration logics for use by the graphics core 2200. In at least one embodiment, the additional fixed function logic 2216 includes an additional geometry pipeline for use only in position shading. In position shading only, there are at least two geometry pipelines, and in the full geometry pipeline and culling pipeline within the geometry / fixed function pipelines 2216, 2236, it is the additional geometry pipeline that can be included in the additional fixed function logic 2216. In at least one embodiment, the culling pipeline is a trimmed version of the full geometry pipeline. In at least one embodiment, the full pipeline and the culling pipeline may execute different instances of an application, each instance having a separate environment. In at least one embodiment, position shading only may hide the long culling runs of discarded triangles, thus allowing shading to be completed earlier in some cases. For example, in at least one embodiment, the culling pipeline logic in the additional fixed function logic 2216 may execute the position shader in parallel with the main application and generally produce critical results faster than the full pipeline, because the culling pipeline fetches and masks the position attributes of vertices without performing rasterization and rendering pixels to the frame buffer. In at least one embodiment, the culling pipeline may use the generated critical results to calculate visibility information for all triangles, regardless of whether those triangles are culled. In at least one embodiment, the full pipeline (which may be referred to as a replay pipeline in this case) may consume the visibility information to skip culled triangles to only shade the visible triangles that are ultimately passed to the rasterization stage.

[0259] In at least one embodiment, the additional fixed function logic 2216 may also include general purpose processing acceleration logic, such as fixed function matrix multiplication logic, for implementing decelerated CUAD programs.

[0260] In at least one embodiment, a set of execution resources is included within each graphics sub-core 2201A - 2201F, which can be used to perform graphics, media, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader program. In at least one embodiment, the graphics sub-cores 2201A - 2201F include multiple EU arrays 2202A - 2202F, 2204A - 2204F, thread dispatch and inter-thread communication (TD / IC) logic 2203A - 2203F, 3D (e.g., texture) samplers 2205A - 2205F, media samplers 2206A - 2206F, shader processors 2207A - 2207F, and shared local memory (SLM) 2208A - 2208F. Each of the EU arrays 2202A - 2202F, 2204A - 2204F contains multiple execution units, which are GUGPUs capable of servicing graphics, media, or compute operations, performing floating-point and integer / fixed-point logical operations, including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 2203A - 2203F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between the threads executing on the execution units of the sub-core. In at least one embodiment, the 3D samplers 2205A - 2205F can read data related to textures or other 3D graphics into memory. In at least one embodiment, the 3D sampler can read texture data differently based on the configured sampling state and texture format associated with a given texture. In at least one embodiment, the media samplers 2206A - 2206F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics sub-core 2201A - 2201F can alternatively include a unified 3D and media sampler. In at least one embodiment, the threads executing on the execution units within each sub-core 2201A - 2201F can utilize the shared local memory 2208A - 2208F within each sub-core to enable the threads executing within a thread group to use a common pool of on-chip memory for execution.

[0261] ​FIG. 0 shows a parallel processing unit (“PPU”) 2300 in accordance with at least one embodiment. In at least one embodiment, the PPU 2300 is configured with machine-readable code that, if executed by the PPU 2300, causes the PPU 2300 to perform some or all of the processes and techniques described throughout this document. In at least one embodiment, the PPU 2300 is a multi-threaded processor implemented on one or more integrated circuit devices and utilizes multi-threading as a latency hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) that are executed in parallel on multiple threads. In at least one embodiment, a thread refers to an execution thread and is an instance of a set of instructions configured to be executed by the PPU 2300. In at least one embodiment, the PPU 2300 is a graphics processing unit (“GPU”) configured to implement a graphics rendering pipeline for processing three-dimensional (“3D”) graphics data to generate two-dimensional (“2D”) image data for display on a display device (such as an LCD device). In at least one embodiment, the PPU 2300 is used to perform computations such as linear algebra operations and machine learning operations. ​ An example parallel processor is shown for illustrative purposes only and should be construed as a non-limiting example of a processor architecture implemented in at least one embodiment.

[0262] In at least one embodiment, ​ at least one component shown or described herein is used to implement the techniques and / or functions associated with ​ the description. In at least one embodiment, the PPU 2300 performs one or more operations for converting vertex data to pixels that are at least partially based on a prefix sum described in connection with ​ the description, as well as operations described elsewhere herein.

[0263] In at least one embodiment, one or more PPU 2300s are configured to accelerate high-performance computing (“HPC”), data center, and machine learning applications. In at least one embodiment, one or more PPU 2300s are configured to accelerate CUDA programs. In at least one embodiment, PPU 2300 includes, but is not limited to, I / O unit 2306, front-end unit 2310, scheduler unit 2312, work distribution unit 2314, hub 2316, crossbar (“Xbar”) 2320, one or more general processing clusters (“GPC”) 2318, and one or more partition units (“memory partition units”) 2322. In at least one embodiment, PPU 2300 is connected to a host processor or other PPU 2300 via one or more high-speed GPU interconnects (“GPU interconnect”) 2308. In at least one embodiment, PPU 2300 is connected to a host processor or other peripheral devices via a system bus or interconnect 2302. In one embodiment, PPU 2300 is connected to a local memory including one or more memory devices (“memory”) 2304. In at least one embodiment, memory device 2304 includes, but is not limited to, one or more dynamic random access memory (“DRAM”) devices. In at least one embodiment, one or more DRAM devices are configured and / or configurable as a high-bandwidth memory (“HBM”) subsystem, and multiple DRAM dies are stacked within each device.

[0264] In at least one embodiment, high-speed GPU interconnect 2308 may refer to a wire-based multi-channel communication link used by the system for scaling, and includes one or more PPU 2300s in combination with one or more CPUs (“CPU”), supporting cache coherence between the PPU 2300 and the CPU and CPU master control. In at least one embodiment, high-speed GPU interconnect 2308 transfers data and / or commands to other units of PPU 2300 via hub 2316, such as one or more copy engines, video encoders, video decoders, power management units, and / or other components that may not be explicitly shown in ​ it.

[0265] In at least one embodiment, I / O unit 2306 is configured to receive from a host processor via system bus 2302 ( ​send and receive communications (e.g., commands, data) (not shown in the figure). In at least one embodiment, I / O unit 2306 communicates directly with the host processor via system bus 2302 or through one or more intermediate devices (such as a memory bridge). In at least one embodiment, I / O unit 2306 may communicate with one or more other processors (such as one or more PPU 2300) via system bus 2302. In at least one embodiment, I / O unit 2306 implements a PCIe interface for communication via the PCIe bus. In at least one embodiment, I / O unit 2306 implements an interface for communicating with external devices.

[0266] In at least one embodiment, I / O unit 2306 decodes packets received via system bus 2302. In at least one embodiment, at least some of the packets represent commands configured to cause PPU 2300 to perform various operations. In at least one embodiment, I / O unit 2306 sends the decoded commands to various other units of PPU 2300 as specified by the commands. In at least one embodiment, the commands are sent to the front-end unit 2310 and / or sent to the hub 2316 or other units of PPU 2300, such as one or more copy engines, video encoders, video decoders, power management units, etc. ( ​ not explicitly shown in the figure). In at least one embodiment, I / O unit 2306 is configured to route communications between various logic units of PPU 2300.

[0267] In at least one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload to PPU 2300 for processing. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in a memory that can be accessed (e.g., read / written) by both the host processor and PPU2300 - the host interface unit may be configured to access the buffer in the system memory connected to system bus 2302 via a memory request transmitted through I / O unit 2306 via system bus 2302. In at least one embodiment, the host processor writes the command stream to the buffer and then sends a pointer indicating the start of the command stream to PPU 2300, such that the front-end unit 2310 receives one or more command stream pointers and manages one or more command streams, reads commands from the command stream, and forwards the commands to various units of PPU 2300.

[0268] In at least one embodiment, the front-end unit 2310 is coupled to a scheduler unit 2312 that configures various GPCs 2318 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 2312 is configured to track state information related to the various tasks managed by the scheduler unit 2312, where the state information may indicate which GPC 2318 a task is assigned to, whether the task is active or inactive, the priority associated with the task, and so on. In at least one embodiment, the scheduler unit 2312 manages multiple tasks executed on one or more GPCs 2318.

[0269] In at least one embodiment, the scheduler unit 2312 is coupled to a work distribution unit 2314 that is configured to dispatch tasks for execution on the GPCs 2318. In at least one embodiment, the work distribution unit 2314 tracks the multiple scheduled tasks received from the scheduler unit 2312 and the work distribution unit 2314 manages the pool of pending tasks and the pool of active tasks for each GPC 2318. In at least one embodiment, the pool of pending tasks includes multiple time slots (e.g., 32 time slots) that contain tasks assigned to be processed by a particular GPC 2318; the pool of active tasks may include multiple time slots (e.g., 4 time slots) for tasks to be actively processed by the GPC 2318 such that as one of the GPCs 2318 finishes executing a task, that task is evicted from the active task pool of the GPC 2318, and one of the other tasks from the pool of pending tasks is selected and scheduled for execution on the GPC 2318. In at least one embodiment, if an active task is idle on a GPC 2318, e.g., while waiting for a data dependency to be resolved, the active task is evicted from the GPC 2318 and returned to the pool of pending tasks while another task from the pool of pending tasks is selected and scheduled for execution on the GPC 2318.

[0270] In at least one embodiment, the work distribution unit 2314 communicates with one or more GPCs 2318 via an XBar 2320. In at least one embodiment, the XBar 2320 is an interconnection network that couples many of the units of the PPU 2300 to other units of the PPU 2300 and may be configured to couple the work distribution unit 2314 to a particular GPC 2318. In at least one embodiment, one or more other units of the PPU 2300 may also be connected to the XBar 2320 via a hub 2316.

[0271] In at least one embodiment, tasks are managed by a scheduler unit 2312 and assigned to one of the GPCs 2318 by a work distribution unit 2314. The GPCs 2318 are configured to process the tasks and produce results. In at least one embodiment, the results can be consumed by other tasks in the GPC 2318, routed to a different GPC 2318 via the XBar 2320, or stored in the memory 2304. In at least one embodiment, the results can be written to the memory 2304 by a partitioning unit 2322, which implements a memory interface for writing data to or reading data from the memory 2304. In at least one embodiment, the results can be transmitted to another PPU 2300 or CPU via the high-speed GPU interconnect 2308. In at least one embodiment, the PPU 2300 includes, but is not limited to, U partitioning units 2322, which is equal to the number of separate and distinct memory devices 2304 coupled to the PPU 2300.

[0272] In at least one embodiment, a host processor executes a driver core that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 2300. In one embodiment, multiple compute applications are executed simultaneously by the PPU 2300, and the PPU 2300 provides isolation, quality of service ("QoS"), and separate address spaces for the multiple compute applications. In at least one embodiment, an application generates instructions (e.g., in the form of API calls) that cause the driver core to generate one or more tasks for execution by the PPU 2300, and the driver core outputs the tasks to one or more streams processed by the PPU 2300. In at least one embodiment, each task includes one or more related thread groups, which may be referred to as warps. In at least one embodiment, a warp includes multiple related threads (e.g., 32 threads) that can be executed in parallel. In at least one embodiment, cooperating threads may refer to multiple threads, including instructions for executing a task and exchanging data via shared memory.

[0273] ​ A GPC 2400 is shown in accordance with at least one embodiment. In at least one embodiment, the GPC 2400 is ​GPC 2318. In at least one embodiment, each GPC 2400 includes, but is not limited to, a plurality of hardware units for processing tasks, and each GPC 2400 includes, but is not limited to, a pipeline manager 2402, a pre-raster operation unit (“PROP”) 2404, a raster engine 2408, a work distribution crossbar (“WDX”) 2416, a memory management unit (“MMU”) 2418, one or more data processing clusters (“DPC”) 2406, and any suitable combination of components.

[0274] In at least one embodiment, regarding ​ at least one of the components shown or described is used to implement the techniques and / or functions associated with ​ those described. In at least one embodiment, GPC 2400 performs one or more operations that convert vertex data to pixels based at least in part on the prefixes associated with ​ those described, as well as the operations described elsewhere herein.

[0275] In at least one embodiment, the operation of GPC 2400 is controlled by the pipeline manager 2402. In at least one embodiment, the pipeline manager 2402 manages the configuration of one or more DPC 2406 to process the tasks assigned to GPC 2400. In at least one embodiment, the pipeline manager 2402 configures at least one of one or more DPC 2406 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, DPC 2406 is configured to execute a vertex shader program on a programmable streaming multiprocessor (“SM”) 2414. In at least one embodiment, the pipeline manager 2402 is configured to route data packets received from a work distribution unit to the appropriate logic units within GPC 2400, and in at least one embodiment, some data packets can be routed to fixed function hardware units in PROP 2404 and / or raster engine 2408, while other data packets can be routed to DPC 2406 for processing by primitive engine 2412 or SM 2414. In at least one embodiment, the pipeline manager 2402 configures at least one of DPC 2406 to implement a neural network model and / or a computing pipeline. In at least one embodiment, the pipeline manager 2402 configures at least one of DPC 2406 to execute at least a portion of a CUDA program.

[0276] In at least one embodiment, the PROP unit 2404 is configured to route data generated by the raster engine 2408 and DPC 2406 to a raster operation (“ROP”) unit in a partition unit, such as those associated with ​Memory partition units 2322 and the like described in more detail. In at least one embodiment, the PROP unit 2404 is configured to perform optimizations for color mixing, organize pixel data, perform address translation, and so on. In at least one embodiment, the raster engine 2408 includes, but is not limited to, a plurality of fixed-function hardware units configured to perform various raster operations, and in at least one embodiment, the raster engine 2408 includes, but is not limited to, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile aggregation engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives the transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices; the plane equations are transmitted to the coarse raster engine to generate coverage information for the primitive (e.g., the x, y coverage range masks of the tile); the output of the coarse raster engine is transmitted to the culling engine, where fragments associated with primitives that fail the z-test are culled and transmitted to the clipping engine, where fragments located outside the view frustum are clipped. In at least one embodiment, the clipped and culled fragments are passed to the fine raster engine to generate the attributes of the pixel fragments based on the plane equations generated by the setup engine. In at least one embodiment, the output of the raster engine 2408 includes fragments that will be processed by any appropriate entity (e.g., a fragment shader implemented within the DPC 2406).

[0277] In at least one embodiment, each DPC 2406 included in the GPC 2400 includes, but is not limited to, an M pipeline controller ("MPC") 2410; a primitive engine 2412; one or more SMs 2414; and any suitable combination thereof. In at least one embodiment, the MPC 2410 controls the operation of the DPC 2406 and routes the packets received from the pipeline manager 2402 to the appropriate units within the DPC 2406. In at least one embodiment, packets associated with vertices are routed to the primitive engine 2412, which is configured to fetch vertex attributes associated with the vertices from memory; conversely, packets associated with shader programs can be sent to the SM 2414.

[0278] In at least one embodiment, SM 2414 includes, but is not limited to, programmable stream processors configured to process tasks represented by multiple threads. In at least one embodiment, SM 2414 is multithreaded and configured to execute multiple threads (e.g., 32 threads) from a particular thread group simultaneously, and implements a single instruction, multiple data (“SIMD”) architecture, where each thread in a group of threads (e.g., a warp) is configured to process different data sets based on the same instruction set. In at least one embodiment, all threads in a thread group execute the same instruction. In at least one embodiment, SM 2414 implements a single instruction, multiple thread (“SIMT”) architecture, where each thread in a group of threads is configured to process different data sets based on the same instruction set, but where individual threads in the thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, such that when threads in a warp diverge, concurrency between the warp and serial execution within the warp is achieved. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, such that equal concurrency is achieved between all threads within and between warps. In at least one embodiment, an execution state is maintained for each individual thread, and threads executing the same instruction can converge and execute in parallel to improve efficiency. The following describes at least one embodiment of SM 2414 in conjunction with ​ At least one embodiment of SM 2414 is described in more detail.

[0279] In at least one embodiment, MMU 2418 provides an interface between GPC 2400 and a memory partitioning unit (e.g., ​ partitioning unit 2322), and MMU 2418 provides virtual address to physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, MMU 2418 provides one or more translation lookaside buffers (“TLBs”) for performing virtual address to physical address translation in memory.

[0280] ​ A stream multiprocessor (“SM”) 2500 according to at least one embodiment is shown. In at least one embodiment, SM 2500 is ​of SM 2414. In at least one embodiment, SM 2500 includes, but is not limited to, an instruction cache 2502; one or more scheduler units 2504; a register file 2508; one or more processing cores ("cores") 2510; one or more special function units ("SFUs") 2512; one or more load / store units ("LSUs") 2514; an interconnect network 2516; a shared memory / level 1 ("L1") cache 2518; and any suitable combination thereof. In at least one embodiment, a work distribution unit schedules tasks to be executed on a general processing cluster ("GPC") of a parallel processing unit ("PPU"), and each task is assigned to a specific data processing cluster ("DPC") within the GPC, and if the task is associated with a shader program, the task is assigned to one of the SM 2500s. In at least one embodiment, the scheduler unit 2504 receives tasks from the work distribution unit and manages the instruction scheduling for one or more thread blocks assigned to the SM 2500. In at least one embodiment, the scheduler unit 2504 schedules thread blocks to execute as warps of parallel threads, where each thread block is assigned at least one warp. In at least one embodiment, each warp executes threads. In at least one embodiment, the scheduler unit 2504 manages multiple different thread blocks, assigns warps to different thread blocks, and then dispatches instructions from multiple different cooperating groups to various functional units (e.g., processing cores 2510, SFUs 2512, and LSUs 2514) within each clock cycle. In at least one embodiment, SM 2500 includes one or more thread block clusters, where the thread block clusters can implement programmatic control of locality at a granularity greater than that of a single thread block of a single streaming multiprocessor (SM). In at least one embodiment, the thread block clusters (also referred to as "clusters") enable multiple thread blocks running concurrently across streaming multiprocessors to synchronize and cooperatively acquire, exchange, or otherwise use data.

[0281] In at least one embodiment, ​ at least one component shown or described herein is used to implement the techniques and / or functions associated with ​ described. In at least one embodiment, SM 2500 performs one or more operations for converting vertex data to pixels based at least in part on a prefix sum associated with ​ described, as well as operations described elsewhere herein.

[0282] In at least one embodiment, a "cooperative group" may refer to a programming model for organizing groups of communicating threads that allows developers to express the granularity at which threads are communicating, enabling a richer and more efficient parallel decomposition. In at least one embodiment, a cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. In at least one embodiment, the API of a conventional programming model provides a single, simple construct for synchronizing cooperative threads: a barrier for all threads across a thread block (e.g., the syncthreads() function). However, in at least one embodiment, a programmer can define groups of threads at a granularity smaller than that of a thread block and synchronize within the defined groups to achieve higher performance, design flexibility, and software reuse in the form of collective group-wide functional interfaces. In at least one embodiment, cooperative groups enable a programmer to explicitly define groups of threads at sub-block and multi-block granularities and perform collective operations such as synchronizing the threads within a cooperative group. In at least one embodiment, the sub-block granularity is as small as a single thread. In at least one embodiment, the programming model supports clean composition across software boundaries so that library and utility functions can synchronize safely within their local environments without having to make assumptions about convergence. In at least one embodiment, cooperative group primitives enable new patterns of cooperative parallelism, including but not limited to producer-consumer parallelism, opportunistic parallelism, and global synchronization across an entire thread block grid.

[0283] In at least one embodiment, a dispatch unit 2506 is configured to send instructions to one or more of the functional units, and a scheduler unit 2504 includes, but is not limited to, two dispatch units 2506 that enable two different instructions from the same warp to be dispatched each clock cycle. In at least one embodiment, each scheduler unit 2504 includes a single dispatch unit 2506 or additional dispatch units 2506.

[0284] In at least one embodiment, each SM 2500 includes, but is not limited to, a register file 2508 in at least one embodiment, and this register file 2508 provides a set of registers for the functional units of the SM 2500. In at least one embodiment, the register file 2508 is divided among each functional unit, thereby allocating a dedicated portion of the register file 2508 to each functional unit. In at least one embodiment, the register file 2508 is divided among different warps executed by the SM 2500, and the register file 2508 provides temporary storage for the operands of the data paths connected to the functional units. In at least one embodiment, each SM 2500 includes, but is not limited to, a plurality of L processing cores 2510. In at least one embodiment, the SM2500 includes, but is not limited to, a large number (e.g., 128 or more) of different processing cores 2510. In at least one embodiment, each processing core 2510 includes, but is not limited to, a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit, which includes, but is not limited to, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In at least one embodiment, the processing core 2510 includes, but is not limited to, 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0285] In at least one embodiment, the tensor cores are configured to perform matrix operations. In at least one embodiment, one or more tensor cores are included in the processing core 2510. In at least one embodiment, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4×4 matrix and performs matrix multiplication and accumulation operations D = A×B + C, where A, B, C, and D are 4×4 matrices.

[0286] In at least one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, and the accumulation matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor core performs 32-bit floating-point accumulation operations on 16-bit floating-point input data. In at least one embodiment, the 16-bit floating-point multiplication uses 64 operations and results in a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point addition for 4x4x4 matrix multiplication. In at least one embodiment, the tensor core is used to perform larger two-dimensional or higher-dimensional matrix operations composed of these smaller elements. In at least one embodiment, APIs (such as the CUDA-C++ API) expose specialized matrix load, matrix multiplication and accumulation, and matrix store operations to effectively use the tensor core from a CUDA-C++ program. In at least one embodiment, at the CUDA level, the warp-level interface assumes a 16×16 size matrix across all 32 warp threads.

[0287] In at least one embodiment, each SM 2500 includes, but is not limited to, M SFUs 2512 that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In at least one embodiment, the SFU 2512 includes, but is not limited to, a tree traversal unit configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFU 2512 includes, but is not limited to, a texture unit configured to perform texture mapping filtering operations. In at least one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texture pixels) from memory and sample the texture map to produce a sampled texture value for use by a shader program executed by the SM 2500. In at least one embodiment, the texture map is stored in the shared memory / L1 cache 2518. In at least one embodiment, the texture unit uses mip maps (e.g., texture maps with different levels of detail) to implement texture operations (such as filtering operations). In at least one embodiment, each SM 2500 includes, but is not limited to, two texture units.

[0288] In at least one embodiment, each SM 2500 includes, but is not limited to, N LSUs 2514 that implement load and store operations between the shared memory / L1 cache 2518 and the register file 2508. In at least one embodiment, each SM 2500 includes, but is not limited to, an interconnect network 2516 that connects each functional unit to the register file 2508 and the LSU 2514 to the register file 2508 and the shared memory / L1 cache 2518. In at least one embodiment, the interconnect network 2516 is a crossbar switch that can be configured to connect any functional unit to any register in the register file 2508 and the LSU 2514 to memory locations in the register file 2508 and the shared memory / L1 cache 2518.

[0289] In at least one embodiment, the shared memory / L1 cache 2518 is an array of on-chip memory that, in at least one embodiment, allows data storage and communication between the SM 2500 and the primitive engine and between threads in the SM 2500. In at least one embodiment, the shared memory / L1 cache 2518 includes, but is not limited to, a storage capacity of 128 KB and is located in the path from the SM 2500 to the partition unit. In at least one embodiment, the shared memory / L1 cache 2518 is used for caching reads and writes in at least one embodiment. In at least one embodiment, one or more of the shared memory / L1 cache 2518, the L2 cache, and the memory are a backing store.

[0290] In at least one embodiment, combining data cache and shared memory functionality into a single memory block provides improved performance for both types of memory access. In at least one embodiment, the capacity is used by programs that do not use the shared memory or used as a cache, e.g., if the shared memory is configured to use half of the capacity, texture and load / store operations can use the remaining capacity. According to at least one embodiment, the integration within the shared memory / L1 cache 2518 enables the shared memory / L1 cache 2518 to function as a high-throughput pipeline for streaming data while providing high-bandwidth and low-latency access to frequently reused data. In at least one embodiment, when configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function GPU is bypassed, thus creating a simpler programming model. In at least one embodiment, in a general-purpose parallel computing configuration, the work distribution unit directly assigns and distributes blocks of threads to the DPC. In at least one embodiment, the threads within a block execute the same program, use a unique thread ID in the computation to ensure that each thread generates a unique result, execute the program and perform computations using the SM 2500, communicate between threads using the shared memory / L1 cache 2518, and read and write global memory using the LSU 2514 via the shared memory / L1 cache 2518 and the memory partition unit. In at least one embodiment, when configured for general-purpose parallel computing, the SM 2500 writes commands to the scheduler unit 2504 that can be used to start new work on the DPC. In at least one embodiment, the SM 2500 includes one or more distributed shared memories (or distributed shared memory) that support direct SM-to-SM operations, such as loading, storing, and performing atomic operations across multiple SM shared memory blocks.

[0291] In at least one embodiment, the SM 2500 includes one or more asynchronous execution functions, which include a Tensor Memory Accelerator (TMA) unit that can transfer data blocks between global memory and shared memory. In at least one embodiment, one or more processors use or access one or more TMAs to perform bidirectional copy operations, such as from global memory to shared memory and vice versa. In at least one embodiment, the SM 2500 includes one or more TMAs to asynchronously copy between thread blocks in a cluster. In at least one embodiment, the SM 2500 includes one or more asynchronous transaction barriers to perform atomic data movement and synchronization. In at least one embodiment, the SM 2500 includes a Tensor Core Transformer Engine, which includes software and one or more cores to accelerate transformer model training and inference. In at least one embodiment, the transformer (one or more processor cores) that executes one or more Tensor Core Transformer Engines manages and dynamically selects FP8 and 16-bit computations by re-transforming and scaling between FP8 and 16 bits in each layer of one or more neural networks.

[0292] In at least one embodiment, the PPU is included in or coupled to a desktop computer, laptop computer, tablet computer, server, supercomputer, smart phone (e.g., wireless, handheld device), PDA, digital camera, vehicle, head-mounted display, handheld electronic device, etc. In at least one embodiment, the PPU is implemented on a single semiconductor substrate. In at least one embodiment, the PPU is included in a System-on-Chip (SoC) together with one or more other devices (such as an additional PPU, memory, RISC CPU, MMU, digital-to-analog converter (“DAC”), etc.).

[0293] In at least one embodiment, the PPU may be included on a graphics card that includes one or more storage devices. The graphics card may be configured to connect to a PCIe slot on a desktop computer motherboard. In at least one embodiment, the PPU may be an integrated GPU (“iGPU”) included in the chipset of a motherboard.

[0294] Software Fabric for General Computing

[0295] The following figures illustrate, but are not limited to, exemplary software fabrics for implementing at least one embodiment.

[0296] ​Shows the software stack of a programming platform according to at least one embodiment. In at least one embodiment, the programming platform is a platform for accelerating computing tasks using the hardware on a computing system. In at least one embodiment, software developers can access the programming platform through libraries, compiler directives, and / or extensions to the programming language. In at least one embodiment, the programming platform can be, but is not limited to, CUDA, Radeon Open Compute Platform (“ROCm”), OpenCL (OpenCL developed by the Khronos group) TM ), SYCL, or Intel One API.

[0297] In at least one embodiment, ​ at least one component shown or described in is used to implement the techniques and / or functions described in conjunction with ​ description. In at least one embodiment, the software stack 2600 performs one or more operations for converting vertex data to pixels based at least in part on the prefix sum described in conjunction with ​ description, as well as operations described elsewhere herein.

[0298] In at least one embodiment, the software stack 2600 of the programming platform provides an execution environment for the application 2601. In at least one embodiment, the application 2601 can include any computer software capable of being launched on the software stack 2600. In at least one embodiment, the application 2601 can include, but is not limited to, artificial intelligence (“AI”) / machine learning (“ML”) applications, high-performance computing (“HPC”) applications, virtual desktop infrastructure (“VDI”), or data center workloads.

[0299] In at least one embodiment, the application 2601 and the software stack 2600 run on the hardware 2607. In at least one embodiment, the hardware 2607 can include one or more GPUs, CPUs, FPGAs, AI engines, and / or other types of computing devices that support the programming platform. In at least one embodiment, for example, using CUDA, the software stack 2600 can be vendor-specific and only compatible with devices from a specific vendor. In at least one embodiment, for example, in using OpenCL, the software stack 2600 can be used with devices from different vendors. In at least one embodiment, the hardware 2607 includes a host connected to one or more devices that can be accessed via application programming interface (API) calls to perform computing tasks. In at least one embodiment, compared to the host within the hardware 2607, which can include, but is not limited to, a CPU (but can also include computing devices) and its memory, the devices within the hardware 2607 can include, but are not limited to, GPUs, FPGAs, AI engines, or other computing devices (but can also include CPUs) and their memory.

[0300] In at least one embodiment, the software stack 2600 of the programming platform includes, but is not limited to, a plurality of libraries 2603, a runtime 2605, and a device kernel driver 2606. In at least one embodiment, each of the libraries 2603 may include data and programming code that can be used by a computer program and utilized during software development. In at least one embodiment, the libraries 2603 may include, but are not limited to, pre-written code and subroutines, classes, values, type specifications, configuration data, documentation, help data, and / or message templates. In at least one embodiment, the libraries 2603 include functions that are optimized for execution on one or more types of devices. In at least one embodiment, the libraries 2603 may include, but are not limited to, functions for performing mathematical, deep learning, and / or other types of operations on the device. In at least one embodiment, the libraries 2603 are associated with corresponding APIs 2602, which may include one or more APIs that expose the functions implemented in the libraries 2603. In at least one embodiment, a processor (e.g., a CPU, a GPU) executes, invokes, or otherwise uses one or more APIs to prioritize kernels. For example, a first kernel (e.g., a parent kernel) may start a second kernel (e.g., a child kernel), and the processor may use the second kernel to start additional kernels (e.g., grandchild kernels) independent of the first kernel. In at least one embodiment, the processor executes an API or calls an API to be executed from memory to support dynamic stream prioritization (e.g., updating priorities when performing operations using streams). For example, when the processor executes the API, it allows a programmer to copy the stream priority from one stream to one or more other streams.

[0301] In at least one embodiment, the software stack 2600 includes APIs to support dynamic stream prioritization (e.g., updating the priority while an operation is being performed using the stream), which allows a programmer to set the priority of a stream at any time after the stream has been created. In at least one embodiment, the software stack 2600 includes APIs to support dynamic stream prioritization (e.g., updating the priority while an operation is being performed using the stream), which allows a programmer to obtain the current priority of a stream, where the priority is one of the multiple attributes of the stream. In at least one embodiment, the software stack 2600 includes APIs to support dynamic stream prioritization (e.g., updating the priority while an operation is being performed using the stream), which allows a programmer to obtain the current priority of a stream as a single attribute. In at least one embodiment, the software stack 2600 includes APIs to support dynamic stream prioritization (e.g., updating the priority when the stream is used to perform an operation), which allows a programmer to launch a kernel to perform an operation on a stream at a set priority, which may be different from the stream priority. In at least one embodiment, the software stack 2600 includes APIs to indicate whether an object (e.g., a thread synchronization object, such as a barrier) tracks whether all data movement operations of a set of threads running on a GPU have a specified state after a specified period of time, where the specified state can be a state indicating that the data has been moved and is ready to be used, and an expected parity value is used as an input to the API to specify it.

[0302] In at least one embodiment, the software stack 2600 includes one or more APIs for updating a kernel. In at least one embodiment, a processor executes an API or calls an API from memory to update an existing API to support a context-free kernel, which allows a programmer to add a kernel node to a graph without a graphical context so that the graphical context can be dynamically associated with the kernel at runtime. In at least one embodiment, the software stack 2600 includes one or more APIs to allow a programmer to obtain a kernel identifier and a graphical context as separate parameters from a kernel node in order to obtain parameters from a kernel and a context-free kernel. In at least one embodiment, the software stack 2600 includes one or more APIs to use a parallel processor (e.g., one or more graphics processing units) to launch a task graph (e.g., a task graph) and execute one or more task graphs (e.g., including one or more programs).

[0303] In at least one embodiment, the software stack 2600 includes one or more APIs to associate one or more instructions with one or more memory ordering operations (such as fence or membar operations). In at least one embodiment, instructions are associated with one or more domains such that the memory ordering operations are performed in association with one or more specific domains without interfering with instructions of other domains. The API indicates that a thread has arrived (e.g., at a thread synchronization barrier), or has completed a work phase related to an asynchronous data movement operation on the GPU. In at least one embodiment, the software stack 2600 includes one or more to allow a programmer to manually indicate an expected transaction count when a thread completes a work phase, and this transaction count is used to update an object that tracks whether all data movement operations of a group of threads have been completed.

[0304] In at least one embodiment, the application 2601 is written as source code, which is compiled into executable code, as discussed in more detail below in conjunction with ​ In at least one embodiment, the executable code of the application 2601 can run at least in part on an execution environment provided by the software stack 2600. In at least one embodiment, during the execution of the application 2601, code that needs to run on the device (as compared to the host) can be obtained. In such a case, in at least one embodiment, the runtime 2605 can be called to load and start the required code on the device. In at least one embodiment, the runtime 2605 can include any technically feasible runtime system capable of supporting the execution of the application 2601.

[0305] In at least one embodiment, the runtime 2605 is implemented as one or more runtime libraries associated with a corresponding API (which is shown as API 2604). In at least one embodiment, one or more such runtime libraries can include, but are not limited to, functions for memory management, execution control, device management, error handling, and / or synchronization, etc. In at least one embodiment, the memory management functions can include, but are not limited to, functions for allocating, deallocating, and copying device memory, and transferring data between host memory and device memory. In at least one embodiment, the execution control functions can include, but are not limited to, functions for starting a function on the device (sometimes called a "kernel" when the function is a global function callable from the host), and functions for setting property values in a buffer maintained by the runtime library for a given function to be executed on the device.

[0306] In at least one embodiment, the runtime library and corresponding API 2604 can be implemented in any technically feasible manner. In at least one embodiment, one (or any number of) APIs can expose a set of low-level functions for fine-grained control of the device, while another (or any number of) APIs can expose such a set of higher-level functions. In at least one embodiment, a high-level runtime API can be built on top of the low-level API. In at least one embodiment, one or more runtime APIs can be language-specific APIs layered on top of a language-independent runtime API.

[0307] In at least one embodiment, one or more processors disclosed in the "processing system" can execute, access, or otherwise use the software stack 2600. For example, the APU 1300, CPU 1400, 16A - 16B exemplary graphics processors, general-purpose graphics processing unit ("GPGPU") 1730, parallel processor 1800, processing cluster 1894, graphics multiprocessor 1834, graphics multiprocessor 1896, graphics processor 1900, processor 2000, processor 2100, parallel processing unit ("PPU") 2300, GPC 2400, and / or streaming multiprocessor ("SM") 2500 can execute, use, call, or otherwise implement (e.g., by accessing memory) one or more APIs included in the software stack 2600.

[0308] In at least one embodiment, the device kernel driver 2606 is configured to facilitate communication with the underlying device. In at least one embodiment, the device kernel driver 2606 can provide APIs such as API 2604 and / or low-level functions on which other software depends. In at least one embodiment, the device kernel driver 2606 can be configured to compile intermediate representation ("IR") code into binary code at runtime. In at least one embodiment, for CUDA, the device kernel driver 2606 can compile non-hardware-specific parallel thread execution ("PTX") IR code into binary code for a specific target device (caching the compiled binary code) at runtime, which is sometimes also referred to as "final" code. In at least one embodiment, doing so can allow the final code to run on the target device, which may not have existed when the source code was initially compiled into PTX code. Alternatively, in at least one embodiment, the device source code can be compiled into binary code offline, without the device kernel driver 2606 having to compile the IR code at runtime.

[0309] ​ Shown in accordance with at least one embodiment ​CUDA implementation of the software stack 2600. In at least one embodiment, the CUDA software stack 2700 on which the application 2701 can be launched includes a CUDA library 2703, a CUDA runtime 2705, a CUDA driver 2707, and a device kernel driver 2708. In at least one embodiment, the CUDA software stack 2700 executes on the hardware 2709, which may include a CUDA-enabled GPU developed by NVIDIA Corporation of Santa Clara, California.

[0310] In at least one embodiment, Figure 27 at least one component shown or described herein is used to implement the techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, the application 2701 performs at least in part one or more operations for converting vertex data to pixels based on the prefix sum described in conjunction with Figure 3 and the operations described elsewhere herein.

[0311] In at least one embodiment, the application 2701, the CUDA runtime 2705, and the device kernel driver 2708 may perform functions similar to those of the application 2601, the runtime 2605, and the device kernel driver 2606, respectively, as described above in connection with Figure 26A description thereof is provided. In at least one embodiment, the CUDA driver 2707 includes a library (libcuda.so) that implements the CUDA driver API 2706. In at least one embodiment, similar to the CUDA runtime API 2704 implemented by the CUDA runtime library (cudart), the CUDA driver API 2706 may expose, but is not limited to, functions for memory management, execution control, device management, error handling, synchronization, and / or graphics interoperability, etc. In at least one embodiment, the CUDA driver API 2706 differs from the CUDA runtime API 2704 in that the CUDA runtime API 2704 simplifies device code management by providing implicit initialization, context (similar to a process) management, and module (similar to a dynamically loaded library) management. In contrast to the high-level CUDA runtime API 2704, in at least one embodiment, the CUDA driver API 2706 is a low-level API that provides more fine-grained control of the device, particularly with respect to context and module loading. In at least one embodiment, the CUDA driver API 2706 may expose functions for context management that are not exposed by the CUDA runtime API 2704. In at least one embodiment, the CUDA driver API 2706 is also language-independent and, in addition to supporting the CUDA runtime API 2704, also supports, for example, OpenCL. Further, in at least one embodiment, the development libraries, including the CUDA runtime 2705, can be considered separate from the driver components, including the user-mode CUDA driver 2707 and the kernel-mode device driver 2708 (sometimes also referred to as the "display" driver).

[0312] In at least one embodiment, the CUDA library 2703 may include, but is not limited to, math libraries, deep learning libraries, parallel algorithm libraries, and / or signal / image / video processing libraries that can be utilized by parallel computing applications (such as application 2701). In at least one embodiment, the CUDA library 2703 may include math libraries, such as the cuBLAS library, which is an implementation of the basic linear algebra subprograms ("BLAS") for performing linear algebra operations; the cuFFT library for computing the fast Fourier transform ("FFT"), and the cuRAND library for generating random numbers, etc. In at least one embodiment, the CUDA library 2703 may include deep learning libraries, such as the cuDNN library for primitives of deep neural networks and the TensorRT platform for high-performance deep learning inference, etc.

[0313] Figure 28 Shows according to at least one embodiment of Figure 26ROCm implementation of the software stack 2600. In at least one embodiment, the ROCm software stack 2800 on which the application 2801 can be launched includes a language runtime 2803, a system runtime 2805, a thunk 2807, and a ROCm kernel driver 2808. In at least one embodiment, the ROCm software stack 2800 executes on the hardware 2809, which may include a ROCm-enabled GPU developed by AMD Corporation of Santa Clara, California.

[0314] In at least one embodiment, Figure 28 at least one component shown or described in is used to implement the techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, the ROCm software stack @22200 performs at least in part one or more operations for converting vertex data to pixels based on the prefix sums described in conjunction with Figure 3 and prefix sums as described elsewhere herein.

[0315] In at least one embodiment, the application 2801 may perform functions similar to those of the application 2601 discussed above in conjunction with Figure 26 . Additionally, in at least one embodiment, the language runtime 2803 and the system runtime 2805 may perform functions similar to those of the runtime 2605 discussed above in conjunction with Figure 26 . In at least one embodiment, the language runtime 2803 and the system runtime 2805 differ in that the system runtime 2805 is a language-agnostic runtime that implements the ROCr system runtime API 2804 and utilizes the Heterogeneous System Architecture (“HSA”) runtime API. In at least one embodiment, the HSA runtime API is a thin user-mode API that exposes interfaces for accessing and interacting with the AMD GPU, including functions for memory management, execution control for dispatching kernels through the architecture, error handling, system and agent information, and runtime initialization and shutdown, etc. In at least one embodiment, compared to the system runtime 2805, the language runtime 2803 is an implementation of a language-specific runtime API 2802 layered on top of the ROCr system runtime API 2804. In at least one embodiment, the language runtime API may include, but is not limited to, the Portable Heterogeneous Computing Interface (“HIP”) language runtime API, the Heterogeneous Computing Compiler (“HCC”) language runtime API, or the OpenCL API, etc. In particular, the HIP language is an extension of the C++ programming language with a functionally similar version of the CUDA mechanism, and in at least one embodiment, the HIP language runtime API includes those associated with the above in conjunction with Figure 27Functions similar to the CUDA runtime API 2704 being discussed, such as functions for memory management, execution control, device management, error handling, and synchronization, etc.

[0316] In at least one embodiment, thunk (ROCt) 2807 is an interface 2806 that can be used to interact with the underlying ROCm driver 2808. In at least one embodiment, the ROCm driver 2808 is a ROCk driver, which is a combination of an AMDGPU driver and an HSA kernel driver (amdkfd). In at least one embodiment, the AMDGPU driver is a device kernel driver for GPUs developed by AMD, which performs functions similar to the device kernel driver 2606 discussed above in conjunction with Figure 26 the device kernel driver 2606. In at least one embodiment, the HSA kernel driver is a driver that allows different types of processors to more effectively share system resources via hardware features.

[0317] In at least one embodiment, various libraries (not shown) can be included in the ROCm software stack 2800 above the language runtime 2803, and provide functions similar to the CUDA libraries 2703 discussed above in conjunction with Figure 27 the CUDA libraries 2703. In at least one embodiment, the various libraries can include but are not limited to math, deep learning, and / or other libraries, such as the hipBLAS library that implements functions similar to CUDA cuBLAS, the rocFFT library similar to CUDA cuFFT for computing FFT, etc.

[0318] Figure 29 An OpenCL implementation of the software stack 2600 according to at least one embodiment is shown. Figure 26 In at least one embodiment, the OpenCL software stack 2900 on which an application 2901 can be launched includes an OpenCL framework 2910, an OpenCL runtime 2906, and a driver 2907. In at least one embodiment, the OpenCL software stack 2900 executes on hardware 2709 that is not vendor - specific. In at least one embodiment, since devices developed by different vendors support OpenCL, specific OpenCL drivers may be required to interoperate with hardware from such vendors.

[0319] In at least one embodiment, Figure 29 at least one component shown or described in Figures 1 to 6 is used to implement the techniques and / or functions described in conjunction with Figure 3 the prefix sum described above and convert vertex data to pixels, as well as operations described elsewhere herein.

[0320] In at least one embodiment, the application 2901, the OpenCL runtime 2906, the device kernel driver 2907, and the hardware 2908 may respectively perform functions similar to the application 2601, the runtime 2605, the device kernel driver 2606, and the hardware 2607 discussed above. In at least one embodiment, the application 2901 further includes an OpenCL kernel 2902 having code to be executed on the device. Figure 26 In at least one embodiment, OpenCL defines a "platform" that allows a host to control devices connected to the host. In at least one embodiment, the OpenCL framework provides a platform layer API and a runtime API, shown as platform API 2903 and runtime API 2905. In at least one embodiment, the runtime API 2905 uses a context to manage the execution of kernels on a device. In at least one embodiment, each identified device may be associated with a respective context, and the runtime API 2905 may use this context to manage the command queue, program objects, and kernel objects, shared memory objects, etc. of the device. In at least one embodiment, the platform API 2903 exposes functions that allow a device context to be used to select and initialize a device, submit work to the device via a command queue, and enable data transfer to and from the device, etc. Additionally, in at least one embodiment, the OpenCL framework provides various built-in functions (not shown), including mathematical functions, relational functions, and image processing functions, etc.

[0321] In at least one embodiment, the compiler 2904 is also included in the OpenCL framework 2910. In at least one embodiment, the source code may be compiled offline before executing the application or compiled online during the execution of the application. Contrary to CUDA and ROCm, the OpenCL application in at least one embodiment may be compiled online by the compiler 2904, and the compiler 2904 is included to represent any number of compilers that may be used to compile source code and / or IR code (e.g., standard portable intermediate representation ("SPIR-V") code) into binary code. Alternatively, in at least one embodiment, the OpenCL application may be compiled offline before executing such an application.

[0322]

[0323] Figure 30 ​Shows software supported by a programming platform according to at least one embodiment. In at least one embodiment, the programming platform 3004 is configured to support various programming models 3003, middleware and / or libraries 3002, and frameworks 3001 that an application 3000 can rely on. In at least one embodiment, the application 3000 can be an AI / ML application implemented using, for example, a deep learning framework (e.g., MXNet, PyTorch, or TensorFlow), which can rely on libraries such as cuDNN, NVIDIA Collective Communications Library (“NCCL”), and / or NVIDIA Developer Data Loading Library (“DALI”) CUDA libraries to provide accelerated computing on the underlying hardware.

[0324] In at least one embodiment, Figure 30 at least one of the components shown or described is used to implement the techniques and / or functions described in connection with Figures 1 to 6 In at least one embodiment, the application 3002 performs one or more operations of converting vertex data to pixels based at least in part on the prefix sum described in connection with Figure 3 and the operations described elsewhere herein.

[0325] In at least one embodiment, the programming platform 3004 can be one of the CUDA, ROCm, or OpenCL platforms described above in connection with Figure 27 , Figure 28 and Figure 29 respectively. In at least one embodiment, the programming platform 3004 supports multiple programming models 3003, which are abstractions of the underlying computing system that allow the expression of algorithms and data structures. In at least one embodiment, the programming model 3003 can expose the characteristics of the underlying hardware to improve performance. In at least one embodiment, the programming model 3003 can include, but is not limited to, CUDA, HIP, OpenCL, C++ Accelerated Massive Parallelism (“C++AMP”), Open Multi-Processing (“OpenMP”), Open Accelerator (“OpenACC”), and / or Vulcan Compute.

[0326] In at least one embodiment, the library and / or middleware 3002 provides an implementation of the abstraction of the programming model 3004. In at least one embodiment, such a library includes data and programming code that can be used by a computer program and utilized during software development. In at least one embodiment, in addition to those that can be obtained from the programming platform 3004, such middleware also includes software that provides services to applications. In at least one embodiment, the library and / or middleware 3002 may include, but is not limited to, cuBLAS, cuFFT, cuRAND, and other CUDA libraries, or rocBLAS, rocFFT, rocRAND, and other ROCm libraries. Additionally, in at least one embodiment, the library and / or middleware 3002 may include the NCCL and the ROCm Communication Collective Library ("RCCL") libraries, which provide communication routines for GPUs, the MIOpen library for deep learning acceleration, and / or the Eigen library for linear algebra, matrix and vector operations, geometric transformations, numerical solvers, and related algorithms.

[0327] In at least one embodiment, the application framework 3001 depends on the library and / or middleware 3002. In at least one embodiment, each application framework 3001 is a software framework for implementing the standard structure of application software. Returning to the AI / ML example discussed above, in at least one embodiment, a framework (such as the Caffe, Caffe2, TensorFlow, Keras, PyTorch, or MxNet deep learning frameworks) can be used to implement AI / ML applications.

[0328] Figure 31 Shown is compiled code according to at least one embodiment to execute on Figures 26 - 29 one of the programming platforms. In at least one embodiment, the compiler 3101 receives the source code 3100, which includes both host code and device code. In at least one embodiment, the compiler 3101 is configured to convert the source code 3100 into host-executable code 3102 for execution on the host and device-executable code 3103 for execution on the device. In at least one embodiment, the source code 3100 can be compiled offline before executing the application or online during the execution of the application. In at least one embodiment, the compiler 3101 includes or can access one or more libraries to identify a series of API calls to execute a single fused API, where the single fused API is a combined API of two or more APIs.

[0329] In at least one embodiment, Figure 31 at least one component shown or described in Figures 1 to 6The described technology and / or functionality. In at least one embodiment, compiler 3101 compiles code for causing one or more processors to perform one or more operations, the operations being at least partially based on prefix sums described in conjunction with Figure 3 to convert vertex data to pixels and operations described elsewhere herein.

[0330] In at least one embodiment, source code 3100 may include code in any programming language supported by compiler 3101, such as C++, C, Fortran, etc. In at least one embodiment, source code 3100 may be included in a single-source file that has a mixture of host code and device code and indicates the location of the device code therein. In at least one embodiment, the single-source file may be a.cu file that includes CUDA code or a.hip.cpp file that includes HIP code. Alternatively, in at least one embodiment, source code 3100 may include multiple source code files instead of a single-source file, in which the host code and device code are separated.

[0331] In at least one embodiment, compiler 3101 is configured to compile source code 3100 into host-executable code 3102 for execution on a host and device-executable code 3103 for execution on a device. In at least one embodiment, compiler 3101 performs operations including parsing source code 3100 into an abstract syntax tree (AST), performing optimizations, and generating executable code. In at least one embodiment where source code 3100 includes a single-source file, compiler 3101 may separate the device code from the host code in such single-source file, compile the device code and the host code into device-executable code 3103 and host-executable code 3102 respectively, and link the device-executable code 3103 and the host-executable code 3102 together in a single file, as discussed in more detail below with respect to Figure 32 More detailed discussion.

[0332] In at least one embodiment, host-executable code 3102 and device-executable code 3103 may be in any suitable format, such as binary code and / or IR code. In the case of CUDA, in at least one embodiment, host-executable code 3102 may include native object code, while device-executable code 3103 may include code in PTX intermediate representation. In at least one embodiment, in the case of ROCm, both host-executable code 3102 and device-executable code 3103 may include target binary code.

[0333] Figure 32 is compiled code according to at least one embodiment for execution on Figures 26 - 29A more detailed illustration executed on one of the programming platforms. In at least one embodiment, the compiler 3201 is configured to receive the source code 3200, compile the source code 3200, and output the executable file 3210. In at least one embodiment, the source code 3200 is a single source file, such as a.cu file,.hip.cpp file, or a file in other formats, which includes both host code and device code. In at least one embodiment, the compiler 3201 can be, but is not limited to, the NVIDIA CUDA compiler ("NVCC") for compiling CUDA code in a.cu file, or the HCC compiler for compiling HIP code in a.hip.cpp file.

[0334] In at least one embodiment, Figure 32 at least one component shown or described in is used to implement the techniques and / or functions associated with Figures 1 to 6 described. In at least one embodiment, the compiler 3201 compiles code that causes one or more processors to perform one or more operations that convert vertex data to pixels based at least in part on the prefix sums described in connection with Figure 3 and the prefix sums described elsewhere herein.

[0335] In at least one embodiment, the compiler 3201 includes a compiler front-end 3202, a host compiler 3205, a device compiler 3206, and a linker 3209. In at least one embodiment, the compiler front-end 3202 is configured to separate the device code 3204 from the host code 3203 in the source code 3200. In at least one embodiment, the device code 3204 is compiled by the device compiler 3206 into device executable code 3208, which, as described, can include binary code or IR code. In at least one embodiment, the host code 3203 is separately compiled by the host compiler 3205 into host executable code 3207. In at least one embodiment, for NVCC, the host compiler 3205 can be, but is not limited to, a general C / C++ compiler that outputs native object code, while the device compiler 3206 can be, but is not limited to, a compiler based on the low-level virtual machine ("LLVM") that forks the LLVM compiler infrastructure and outputs PTX code or binary code. In at least one embodiment, for HCC, both the host compiler 3205 and the device compiler 3206 can be, but is not limited to, LLVM-based compilers that output target binary code.

[0336] In at least one embodiment, after compiling source code 3200 into host-executable code 3207 and device-executable code 3208, linker 3209 links the host and device-executable codes 3207 and 3208 together in executable file 3210. In at least one embodiment, the native object code of the host and PTX or the binary code of the device can be linked together in an Executable and Linkable Format (“ELF”) file, which is a container format for storing object code.

[0337] Figure 33 Illustrates the transformation of source code prior to compilation according to at least one embodiment. In at least one embodiment, source code 3300 is passed through transformation tool 3301, which transforms source code 3300 into transformed source code 3302. In at least one embodiment, compiler 3303 is used to compile the transformed source code 3302 into host-executable code 3304 and device-executable code 3305, the process of which is similar to the process of compiler 3101 compiling source code 3100 into host-executable code 3102 and device-executable code 3103, as discussed above in connection with Figure 31 discussed.

[0338] In at least one embodiment, Figure 33 at least one component shown or described in Figures 1 to 6 is used to implement the techniques and / or functions described in connection with Figure 3 described prefix sums and prefix sums as described elsewhere herein to perform one or more operations that transform vertex data into pixels.

[0339] In at least one embodiment, the transformation performed by transformation tool 3301 is used to port source code 3300 to execute in a different environment than originally intended to run on. In at least one embodiment, transformation tool 3301 may include, but is not limited to, a HIP converter, which is used to “hipify” CUDA code for the CUDA platform into HIP code that can be compiled and executed on the ROCm platform. In at least one embodiment, the transformation of source code 3300 may include: parsing source code 3300 and converting calls to APIs provided by one programming model (e.g., CUDA) into corresponding calls to APIs provided by another programming model (e.g., HIP), as described below in connection with Figure 34A and Figure 35This will be discussed in more detail. Returning to the example of porting CUDA code, in at least one embodiment, calls to the CUDA runtime API, the CUDA driver API, and / or the CUDA libraries can be converted to corresponding HIP API calls. In at least one embodiment, the automatic conversion performed by the conversion tool 3301 may sometimes be incomplete and additional manual effort may be required to fully port the source code 3300.

[0340] Configuring the GPU for general computing

[0341] The following figures illustrate, but are not limited to, an exemplary architecture for compiling and executing computational source code according to at least one embodiment.

[0342] Figure 34A A system 3400 is shown that is configured to compile and execute CUDA source code 3410 using different types of processing units according to at least one embodiment. In at least one embodiment, the system 3400 includes, but is not limited to, CUDA source code 3410, a CUDA compiler 3450, host executable code 3470(1), host executable code 3470(2), CUDA device executable code 3484, a CPU 3490, a CUDA-enabled GPU 3494, a GPU 3492, a CUDA-to-HIP conversion tool 3420, HIP source code 3430, a HIP compiler driver 3440, an HCC 3460, and HCC device executable code 3482.

[0343] In at least one embodiment, the CUDA source code 3410 is a collection of human-readable code in the CUDA programming language. In at least one embodiment, the CUDA code is human-readable code in the CUDA programming language. In at least one embodiment, the CUDA programming language is an extension of the C++ programming language that includes, but is not limited to, mechanisms for defining device code and differentiating between device code and host code. In at least one embodiment, the device code is source code that can be executed in parallel on a device after compilation. In at least one embodiment, the device can be a processor optimized for parallel instruction processing, such as a CUDA-enabled GPU 3490, a GPU 3492, or another GPGPU, etc. In at least one embodiment, the host code is source code that can be executed on a host after compilation. In at least one embodiment, the host is a processor optimized for sequential instruction processing, such as a CPU 3490.

[0344] In at least one embodiment, the CUDA source code 3410 includes, but is not limited to, any number (including zero) of global functions 3412, any number (including zero) of device functions 3414, any number (including zero) of host functions 3416, and any number (including zero) of host / device functions 3418. In at least one embodiment, the global functions 3412, device functions 3414, host functions 3416, and host / device functions 3418 may be mixed in the CUDA source code 3410. In at least one embodiment, each global function 3412 can be executed on the device and can be called from the host. Thus, in at least one embodiment, one or more of the global functions 3412 can serve as an entry point for the device. In at least one embodiment, each global function 3412 is a kernel. In at least one embodiment and in a technique called dynamic parallelism, one or more of the global functions 3412 define a kernel that can be executed on the device and can be called from such a device. In at least one embodiment, the kernel is executed N times in parallel by N different threads on the device during execution (where N is any positive integer).

[0345] In at least one embodiment, each device function 3414 is executed on the device and can only be called from such a device. In at least one embodiment, each host function 3416 is executed on the host and can only be called from such a host. In at least one embodiment, each host / device function 3416 defines both a host version of the function that is executable on the host and can only be called from such a host, and a device version of the function that is executable on the device and can only be called from such a device.

[0346] In at least one embodiment, the CUDA source code 3410 may further include, but is not limited to, any number of calls to any number of functions defined by the CUDA runtime API 3402. In at least one embodiment, the CUDA runtime API 3402 may include, but is not limited to, any number of functions executed on the host for allocating and deallocating device memory, transferring data between host memory and device memory, managing a system with multiple devices, etc. In at least one embodiment, the CUDA source code 3410 may further include any number of calls to any number of functions specified in any number of other CUDA APIs. In at least one embodiment, a CUDA API may be any API designed to be used by CUDA code. In at least one embodiment, CUDA APIs include, but are not limited to, the CUDA runtime API 3402, the CUDA driver API, APIs for any number of CUDA libraries, etc. In at least one embodiment and with respect to the CUDA runtime API 3402, the CUDA driver API is a lower-level API but may provide more fine-grained control of the device. In at least one embodiment, examples of CUDA libraries include, but are not limited to, cuBLAS, cuFFT, cuRAND, cuDNN, etc.

[0347] In at least one embodiment, the CUDA compiler 3450 compiles the input CUDA code (e.g., the CUDA source code 3410) to generate host-executable code 3470(1) and CUDA device-executable code 3484. In at least one embodiment, the CUDA compiler 3450 is NVCC. In at least one embodiment, the host-executable code 3470(1) is a compiled version of the host code included in the input source code that is executable on the CPU 3490. In at least one embodiment, the CPU 3490 may be any processor optimized for sequential instruction processing.

[0348] In at least one embodiment, the CUDA device executable code 3484 is a compiled version of the device code included in the input source code that is executable on a CUDA-enabled GPU 3494. In at least one embodiment, the CUDA device executable code 3484 includes, but is not limited to, binary code. In at least one embodiment, the CUDA device executable code 3484 includes, but is not limited to, IR code, such as PTX code, which is further compiled by the device driver at runtime into binary code for a specific target device (e.g., a CUDA-enabled GPU 3494). In at least one embodiment, a CUDA-enabled GPU 3494 can be any processor that is optimized for parallel instruction processing and supports CUDA. In at least one embodiment, the CUDA-enabled GPU 3494 is developed by NVIDIA Corporation of Santa Clara, California.

[0349] In at least one embodiment, the CUDA-to-HIP conversion tool 3420 is configured to convert CUDA source code 3410 into functionally similar HIP source code 3430. In at least one embodiment, the HIP source code 3430 is a collection of human-readable code of the HIP programming language. In at least one embodiment, HIP code is human-readable code of the HIP programming language. In at least one embodiment, the HIP programming language is an extension of the C++ programming language that includes, but is not limited to, a functionally similar version of the CUDA mechanism for defining device code and differentiating device code from host code. In at least one embodiment, the HIP programming language may include a subset of the functionality of the CUDA programming language. In at least one embodiment, for example, the HIP programming language includes, but is not limited to, a mechanism for defining global functions 3412, but such a HIP programming language may lack support for dynamic parallelism, and thus, the global functions 3412 defined in the HIP code can only be called from the host.

[0350] In at least one embodiment, the HIP source code 3430 includes, but is not limited to, any number (including zero) of global functions 3412, any number (including zero) of device functions 3414, any number (including zero) of host functions 3416, and any number (including zero) of host / device functions 3418. In at least one embodiment, the HIP source code 3430 may also include any number of calls to any number of functions specified in the HIP runtime API 3432. In one embodiment, the HIP runtime API 3432 includes, but is not limited to, functionally similar versions of a subset of the functions included in the CUDA runtime API 3402. In at least one embodiment, the HIP source code 3430 may also include any number of calls to any number of functions specified in any number of other HIP APIs. In at least one embodiment, a HIP API may be any API designed for use with HIP code and / or ROCm. In at least one embodiment, HIP APIs include, but are not limited to, the HIP runtime API 3432, the HIP driver API, APIs for any number of HIP libraries, APIs for any number of ROCm libraries, and the like.

[0351] In at least one embodiment, the CUDA-to-HIP conversion tool 3420 converts each kernel call in the CUDA code from CUDA syntax to HIP syntax and converts any number of other CUDA calls in the CUDA code to any number of other functionally similar HIP calls. In at least one embodiment, a CUDA call is a call to a function specified in the CUDA API, and a HIP call is a call to a function specified in the HIP API. In at least one embodiment, the CUDA-to-HIP conversion tool 3420 converts any number of calls to functions specified in the CUDA runtime API 3402 to any number of calls to functions specified in the HIP runtime API 3432.

[0352] In at least one embodiment, the CUDA-to-HIP conversion tool 3420 is a tool called hipify-perl that performs a text-based conversion process. In at least one embodiment, the CUDA-to-HIP conversion tool 3420 is a tool called hipify-clang that performs a more complex and robust conversion process relative to hipify-perl, which involves using clang (a compiler front-end) to parse the CUDA code and then transforming the resulting symbols. In at least one embodiment, in addition to the modifications performed by the CUDA-to-HIP conversion tool 3420, correctly converting the CUDA code to HIP code may also require modifications (e.g., manual editing).

[0353] In at least one embodiment, the HIP compiler driver 3440 is a front end that determines the target device 3446 and then configures a compiler compatible with the target device 3446 to compile the HIP source code 3430. In at least one embodiment, the target device 3446 is a processor optimized for parallel instruction processing. In at least one embodiment, the HIP compiler driver 3440 can determine the target device 3446 in any technically feasible manner.

[0354] In at least one embodiment, if the target device 3446 is CUDA-compatible (e.g., a CUDA-enabled GPU 3494), the HIP compiler driver 3440 generates a HIP / NVCC compilation command 3442. In at least one embodiment and as described in Figure 34B more detail, the HIP / NVCC compilation command 3442 configures the CUDA compiler 3450 to compile the HIP source code 3430 using, but not limited to, HIP-to-CUDA conversion headers and the CUDA runtime library. In at least one embodiment and in response to the HIP / NVCC compilation command 3442, the CUDA compiler 3450 generates host-executable code 3470(1) and CUDA device-executable code 3484.

[0355] In at least one embodiment, if the target device 3446 is not CUDA-compatible, the HIP compiler driver 3440 generates a HIP / HCC compilation command 3444. In at least one embodiment and as described in Figure 34C more detail, the HIP / HCC compilation command 3444 configures the HCC 3460 to compile the HIP source code 3430 using HCC headers and the HIP / HCC runtime library. In at least one embodiment and in response to the HIP / HCC compilation command 3444, the HCC 3460 generates host-executable code 3470(2) and HCC device-executable code 3482. In at least one embodiment, the HCC device-executable code 3482 is a compiled version of the device code included in the HIP source code 3430 that can be executed on the GPU 3492. In at least one embodiment, the GPU 3492 can be any processor optimized for parallel instruction processing that is not CUDA-compatible and is HCC-compatible. In at least one embodiment, the GPU 3492 is developed by AMD Corporation of Santa Clara, California. In at least one embodiment, the GPU 3492 is a non-CUDA-enabled GPU 3492.

[0356] For illustrative purposes only, in Figure 34ADepicted are three different processes that can be implemented in at least one embodiment to compile CUDA source code 3410 for execution on a CPU 3490 and different devices. In at least one embodiment, the direct CUDA process compiles CUDA source code 3410 for execution on a CPU 3490 and a CUDA-enabled GPU 3494 without converting the CUDA source code 3410 to HIP source code 3430. In at least one embodiment, the indirect CUDA process converts CUDA source code 3410 to HIP source code 3430 and then compiles the HIP source code 3430 for execution on a CPU 3490 and a CUDA-enabled GPU 3494. In at least one embodiment, the CUDA / HCC process converts CUDA source code 3410 to HIP source code 3430 and then compiles the HIP source code 3430 for execution on a CPU 3490 and a GPU 3492.

[0357] The direct CUDA process that can be implemented in at least one embodiment can be depicted by a dashed line and a series of bubble annotations A1 - A3. In at least one embodiment, and as shown by bubble annotation A1, a CUDA compiler 3450 receives CUDA source code 3410 and a CUDA compilation command 3448 that configures the CUDA compiler 3450 to compile the CUDA source code 3410. In at least one embodiment, the CUDA source code 3410 used in the direct CUDA process is written in the CUDA programming language, which is based on other programming languages in addition to C++ (such as C, Fortran, Python, Java, etc.). In at least one embodiment, and in response to the CUDA compilation command 3448, the CUDA compiler 3450 generates host executable code 3470(1) and CUDA device executable code 3484 (denoted by bubble annotation A2). In at least one embodiment and as shown by bubble annotation A3, the host executable code 3470(1) and the CUDA device executable code 3484 can be executed on a CPU 3490 and a CUDA-enabled GPU 3494, respectively. In at least one embodiment, the CUDA device executable code 3484 includes, but is not limited to, binary code. In at least one embodiment, the CUDA device executable code 3484 includes, but is not limited to, PTX code and is further compiled into binary code for a specific target device at runtime.

[0358] An indirect CUDA flow that can be implemented in at least one embodiment can be described by dashed lines and a series of bubble annotations B1 - B6. In at least one embodiment and as shown by bubble annotation B1, the CUDA-to-HIP conversion tool 3420 receives the CUDA source code 3410. In at least one embodiment and as shown by bubble annotation B2, the CUDA-to-HIP conversion tool 3420 converts the CUDA source code 3410 into HIP source code 3430. In at least one embodiment and as shown by bubble annotation B3, the HIP compiler driver 3440 receives the HIP source code 3430 and determines whether the target device 3446 has CUDA enabled.

[0359] In at least one embodiment and as shown by bubble annotation B4, the HIP compiler driver 3440 generates a HIP / NVCC compilation command 3442 and sends both the HIP / NVCC compilation command 3442 and the HIP source code 3430 to the CUDA compiler 3450. In at least one embodiment and as described in more detail in conjunction with Figure 34B the HIP / NVCC compilation command 3442 configures the CUDA compiler 3450 to compile the HIP source code 3430 using, but not limited to, HIP-to-CUDA conversion headers and the CUDA runtime library. In at least one embodiment and in response to the HIP / NVCC compilation command 3442, the CUDA compiler 3450 generates host executable code 3470(1) and CUDA device executable code 3484 (represented by bubble annotation B5). In at least one embodiment and as shown by bubble annotation B6, the host executable code 3470(1) and the CUDA device executable code 3484 can be executed on the CPU 3490 and the CUDA-enabled GPU 3494, respectively. In at least one embodiment, the CUDA device executable code 3484 includes, but is not limited to, binary code. In at least one embodiment, the CUDA device executable code 3484 includes, but is not limited to, PTX code and is further compiled into binary code for a specific target device at runtime.

[0360] A CUDA / HCC flow that can be implemented in at least one embodiment can be described by solid lines and a series of bubble annotations C1 - C6. In at least one embodiment and as shown by bubble annotation C1, the CUDA-to-HIP conversion tool 3420 receives the CUDA source code 3410. In at least one embodiment and as shown by bubble annotation C2, the CUDA-to-HIP conversion tool 3420 converts the CUDA source code 3410 into HIP source code 3430. In at least one embodiment and as shown by bubble annotation C3, the HIP compiler driver 3440 receives the HIP source code 3430 and determines that the target device 3446 does not have CUDA enabled.

[0361] In at least one embodiment, the HIP compiler driver 3440 generates a HIP / HCC compilation command 3444 and sends both the HIP / HCC compilation command 3444 and the HIP source code 3430 to the HCC 3460 (denoted by bubble annotation C4). In at least one embodiment and as described in more detail in conjunction with Figure 34C and as described in more detail, the HIP / HCC compilation command 3444 configures the HCC 3460 to compile the HIP source code 3430 using, but not limited to, the HCC headers and the HIP / HCC runtime libraries. In at least one embodiment and in response to the HIP / HCC compilation command 3444, the HCC 3460 generates host executable code 3470(2) and HCC device executable code 3482 (denoted by bubble annotation C5). In at least one embodiment and as shown by bubble annotation C6, the host executable code 3470(2) and the HCC device executable code 3482 can be executed on the CPU 3490 and the GPU 3492, respectively.

[0362] In at least one embodiment, after converting the CUDA source code 3410 to the HIP source code 3430, the HIP compiler driver 3440 can subsequently be used to generate executable code for the CUDA-enabled GPU 3494 or the GPU 3492 without re-running CUDA as the HIP conversion tool 3420. In at least one embodiment, the CUDA-to-HIP conversion tool 3420 converts the CUDA source code 3410 to the HIP source code 3430 and then stores it in memory. In at least one embodiment, the HIP compiler driver 3440 then configures the HCC 3460 to generate the host executable code 3470(2) and the HCC device executable code 3482 based on the HIP source code 3430. In at least one embodiment, the HIP compiler driver 3440 subsequently configures the CUDA compiler 3450 to generate the host executable code 3470(1) and the CUDA device executable code 3484 based on the stored HIP source code 3430.

[0363] Figure 34B A system 3404 configured to compile and execute Figure 34A the CUDA source code 3410 using the CPU 3490 and the CUDA-enabled GPU 3494 is shown in accordance with at least one embodiment. In at least one embodiment, the system 3404 includes, but is not limited to, the CUDA source code 3410, the CUDA-to-HIP conversion tool 3420, the HIP source code 3430, the HIP compiler driver 3440, the CUDA compiler 3450, the host executable code 3470(1), the CUDA device executable code 3484, the CPU 3490, and the CUDA-enabled GPU 3494.

[0364] In at least one embodiment and as previously described herein in connection with Figure 34A what has been described, the CUDA source code 3410 includes, but is not limited to, any number (including zero) of global functions 3412, any number (including zero) of device functions 3414, any number (including zero) of host functions 3416, and any number (including zero) of host / device functions 3418. In at least one embodiment, the CUDA source code 3410 also includes, but is not limited to, any number of calls to any number of functions specified in any number of CUDA APIs.

[0365] In at least one embodiment, the CUDA-to-HIP conversion tool 3420 converts the CUDA source code 3410 into HIP source code 3430. In at least one embodiment, the CUDA-to-HIP conversion tool 3420 converts each kernel call in the CUDA source code 3410 from CUDA syntax to HIP syntax and converts any number of other CUDA calls in the CUDA source code 3410 into any number of other functionally similar HIP calls.

[0366] In at least one embodiment, the HIP compiler driver 3440 determines that the target device 3446 is CUDA-enabled and generates HIP / NVCC compile commands 3442. In at least one embodiment, the HIP compiler driver 3440 then configures the CUDA compiler 3450 via the HIP / NVCC compile commands 3442 to compile the HIP source code 3430. In at least one embodiment, as part of configuring the CUDA compiler 3450, the HIP compiler driver 3440 provides access to a HIP to CUDA translation header 3452. In at least one embodiment, the HIP to CUDA translation header 3452 translates any number of mechanisms (e.g., functions) specified in any number of HIP APIs to any number of mechanisms specified in any number of CUDA APIs. In at least one embodiment, the CUDA compiler 3450 uses the HIP to CUDA translation header 3452 in conjunction with a CUDA runtime library 3454 corresponding to the CUDA runtime API 3402 to generate host executable code 3470 (1) and CUDA device executable code 3484. In at least one embodiment, the host executable code 3470(1) and the CUDA device executable code 3484 can then be executed on the CPU 3490 and the CUDA-enabled GPU 3494, respectively. In at least one embodiment, the CUDA device executable code 3484 includes, but is not limited to, binary code. In at least one embodiment, the CUDA device executable code 3484 includes, but is not limited to, PTX code and is further compiled into binary code for a specific target device at runtime.

[0367] Figure 34C A system 3406 is shown that is configured to compile and execute using a CPU 3490 and a non-CUDA enabled GPU 3492 according to at least one embodiment. Figure 34A CUDA source code 3410. In at least one embodiment, system 3406 includes, but is not limited to, CUDA source code 3410, CUDA to HIP conversion tool 3420, HIP source code 3430, HIP compiler driver 3440, HCC 3460, host executable code 3470(2), HCC device executable code 3482, CPU 3490, and GPU 3492.

[0368] In at least one embodiment, Figures 34A to 34C At least one component shown or described in the Figures 1 to 6 In at least one embodiment, system 3406 performs at least in part based on the combination of Figure 3 The prefix described and one or more operations that convert vertex data into pixels, as well as operations described elsewhere in this document.

[0369] In at least one embodiment, and as previously described herein in connection with Figure 34A what is described, the CUDA source code 3410 includes, but is not limited to, any number (including zero) of global functions 3412, any number (including zero) of device functions 3414, any number (including zero) of host functions 3416, and any number (including zero) of host / device functions 3418. In at least one embodiment, the CUDA source code 3410 also includes, but is not limited to, any number of calls to any number of functions specified in any number of CUDA APIs.

[0370] In at least one embodiment, the CUDA-to-HIP conversion tool 3420 converts the CUDA source code 3410 into HIP source code 3430. In at least one embodiment, the CUDA-to-HIP conversion tool 3420 converts each kernel call in the CUDA source code 3410 from CUDA syntax to HIP syntax and converts any number of other CUDA calls in the source code 3410 into any number of other functionally similar HIP calls.

[0371] In at least one embodiment, the HIP compiler driver 3440 then determines that the target device 3446 is not CUDA-enabled and generates a HIP / HCC compilation command 3444. In at least one embodiment, the HIP compiler driver 3440 then configures the HCC 3460 to execute the HIP / HCC compilation command 3444, thereby compiling the HIP source code 3430. In at least one embodiment, the HIP / HCC compilation command 3444 configures the HCC 3460 to use, but is not limited to, the HIP / HCC runtime library 3458 and the HCC headers 3456 to generate host-executable code 3470(2) and HCC device-executable code 3482. In at least one embodiment, the HIP / HCC runtime library 3458 corresponds to the HIP runtime API 3432. In at least one embodiment, the HCC headers 3456 include, but are not limited to, any number and type of interoperability mechanisms for HIP and HCC. In at least one embodiment, the host-executable code 3470(2) and the HCC device-executable code 3482 can be executed on the CPU 3490 and the GPU 3492, respectively.

[0372] Figure 35 Illustrated is a diagram according to at least one embodiment by Figure 34CExemplary kernels converted by the CUDA-to-HIP conversion tool 3420. In at least one embodiment, the CUDA source code 3410 divides the overall problem that a given kernel is designed to solve into relatively coarse sub-problems that can be solved independently using thread blocks. In at least one embodiment, each thread block includes, but is not limited to, any number of threads. In at least one embodiment, each sub-problem is divided into relatively fine pieces that can be solved in parallel by the threads ...

Claims

1. A processor, comprising: One or more circuits for identifying one or more pixels within one or more polygons, which are identified at least in part based on whether one or more pixels adjacent to the one or more pixels are partially outside the one or more polygons.

2. The processor according to claim 1, wherein the one or more circuits are for identifying the one or more pixels within the one or more polygons at least in part based on gradients of corners of one or more adjacent pixels.

3. The processor according to claim 1, wherein the one or more circuits are for identifying the one or more pixels within the one or more polygons at least in part based on one or more prefix sums, the prefix sums using information indicating whether one or more adjacent pixels in a row are partially outside the one or more polygons.

4. The processor according to claim 1, wherein the one or more circuits are for generating information indicating whether one or more adjacent pixels are partially outside the one or more polygons at least in part based on one or more directions of one or more sides of the one or more polygons.

5. The processor according to claim 1, wherein the one or more circuits are for identifying the one or more pixels at least in part based on one or more winding numbers generated using gradients assigned to portions of pixels and one or more prefix sums.

6. The processor according to claim 1, wherein the one or more circuits are for identifying the one or more pixels within the one or more polygons by at least identifying portions of the one or more pixels within the one or more polygons.

7. The processor according to claim 1, wherein the one or more circuits are for identifying pixels using different combinations of hardware resources for different pixels.

8. A system, comprising: One or more processors for identifying one or more pixels within one or more polygons, which are identified at least in part based on whether one or more pixels adjacent to the one or more pixels are partially outside the one or more polygons.

9. The system according to claim 8, wherein the one or more processors are for identifying the one or more pixels within the one or more polygons at least in part based on gradients of pixel points in one or more adjacent pixels.

10. The system according to claim 8, wherein the one or more processors are for identifying the one or more pixels within the one or more polygons at least in part based on one or more inclusive prefix sums, the inclusive prefix sums using information indicating whether one or more adjacent pixels in a column are partially outside the one or more polygons.

11. The system according to claim 8, wherein the one or more processors are configured to generate information indicating whether one or more adjacent pixels are at least partially outside the one or more polygons, at least in part based on one or more directions of one or more edges of the one or more polygons.

12. The system according to claim 8, wherein the one or more processors are configured to use one or more adjacent pixels and perform one or more prefix sums to generate a domain of one or more winding numbers to indicate whether the one or more pixels are within the one or more polygons.

13. The system according to claim 8, wherein the one or more processors are configured to identify the one or more pixels within the one or more polygons at least in part based on summing gradients of corners in one or more adjacent pixels.

14. The system according to claim 8, wherein the one or more processors are configured to identify pixels at least in part based on a variable rate shading process.

15. A method comprising: identifying one or more pixels within one or more polygons, the identifying being at least in part based on whether one or more pixels adjacent to the one or more pixels are partially outside the one or more polygons.

16. The method according to claim 15, further comprising: identifying the one or more pixels within the one or more polygons at least in part based on gradients of corners of one or more adjacent pixels.

17. The method according to claim 15 further comprises: identifying the one or more pixels within the one or more polygons at least in part based on one or more prefix sum functions that use gradients of corners of one or more adjacent pixels as inputs.

18. The method according to claim 15, further comprising: generating information indicating whether one or more adjacent pixels are partially outside the one or more polygons at least in part based on positions of corners in one or more adjacent pixels identified relative to one or more edges of the one or more polygons.

19. The method according to claim 15 further comprises: generating a domain of one or more winding numbers to indicate whether the one or more pixels should be activated.

20. The method according to claim 15 further comprises: identifying the number of corners within the one or more polygons in one or more adjacent pixels.