Scalable parallel tessellation
By identifying the mosaic factor of the patch and assigning mosaic examples to multiple parallel mosaic pipelines, the problem of inadequate sorting requirements in the mosaic process in the prior art is solved, and the processing efficiency and storage requirements of the graphics processing system are improved.
Patent Information
- Application Number
- CN201910630713.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-13
- Filing Date
- 2019-07-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-07-12
AI Technical Summary
During the parallel inlay process, the existing graphics processing systems are not adapted to the sorting requirements due to the relative size changes of the patches, which increases the waiting time, and the memory size requirement is large, which affects the system efficiency.
By identifying the mosaic factor of the patch, determine the mosaic examples used in the mosaic process and assign these mosaic examples in multiple parallel mosaic pipelines to achieve parallel operation and improve system efficiency.
Reduces idle time of mosaic pipelines, improves processing efficiency, reduces memory size requirements, and maintains the correct order of output primitives.
Smart Images

Figure CN110717989B_ABST
Abstract
Description
Background Art
[0001] In a graphics processing system, complex geometric surfaces can be represented by tiles using geometry data. The geometry data can be in the form of control points that define the surface as a curve (e.g., a Bezier curve). Typically, such surfaces are processed in a graphics processing system by performing tessellation of the surface to divide the surface into a mesh of primitives, which are typically in the form of triangles, as defined according to a graphics processing API used to render graphics, such as OpenGL and Direct3D.
[0002] Graphics processing systems are generally efficient due to their ability to perform parallel processing, in which large amounts of data are processed in parallel to reduce latency. However, one of the requirements of the tessellation process defined by several APIs is that the order in which tiles are submitted to the tessellator is maintained in the order in which the tessellator issues primitives. In other words, the primitives of the first tile received must be issued before the primitives of the second tile received. This ordering requirement can be problematic for graphics processing systems because the relative sizes of tiles can vary greatly.
[0003] Figure 1 An example tessellation system 100 is shown that includes several parallel tessellation units 110, 120, 130, each configured to tessellate a tile. In this example, three tiles 101-103 are received in sequence and dispatched for parallel processing. Figure 1 In the example of , the first received tile 101 is sent to the tessellation unit 110, the second received tile 102 is sent to the tessellation unit 120, and the third received tile 103 is sent to the tessellation unit 130. In this example, the first received tile 110 will be tessellated into a greater number of primitives 111 than the number of primitives 112, 113 to be tessellated for the tiles 102, 103, respectively (e.g., because the subsequently received tile requires a lower level of detail or is a simpler or smaller tile).
[0004] Processing tiles in parallel provides increased throughput in many cases. However, since the order of received tiles must be maintained in the order in which primitives were issued, latency may increase if the relative amount of processing required for each tile is significantly different. Figure 1In the example of , the amount of processing required to process tile 101 to produce primitive 111 is much greater than the amount of processing required to process tiles 102 and 103, and thus the amount of time required to process tile 102 may be less than that required to process tile 101. Therefore, primitives 112 and 113 may be produced before primitive 111, which is contrary to the requirements of many APIs. The in-order requirement forces each parallel tessellation unit to effectively serialize with the surrounding units, and to alleviate this serialization, large memories may be placed at the output to the tessellation unit to allow the output to be buffered. Memory 140 may be written to in any order as each tessellation unit outputs primitives, and then may be read in this order in order to maintain the correct primitive order required by the API.
[0005] However, the required size of memory 140 may be significant, and may scale with the number of parallel processors in operation. The maximum number of vertices generated by the tessellation of a single tile may be determined by the API, and may be, for example, approximately 4096 vertices, with a typical vertex size of 64 to 96 bytes. In a system with multiple tessellation units, memory 140 may need to be sized so that it can store at least the worst case output (e.g., 4096 vertices) vertices from each tessellation unit. It can be seen that with these example values and a relatively small number of tessellation units, such as four tessellation units, the size of memory 140 may be approximately 1 MB.
[0006] For example, memory 140 may be made larger if, for example, additional buffering is needed, or memory 140 may be made smaller, for example, to target a typical expected number of vertices per tile rather than a worst-case number. However, if memory 140 is not large enough to contain the output from the tiles being processed in parallel at any particular time, the tessellation unit may need to be stopped (i.e., paused) to ensure proper primitive ordering. This may reduce throughput and / or increase latency. Summary of the invention
[0007] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0008] A method of tessellizing tiles to generate tessellation geometry data representing the tessellated tiles is provided, the method comprising: processing received geometry data representing the tiles to identify a tessellation factor for the tiles; determining tessellation instances to be used in tessellizing the tiles based on the identified tessellation factor for the tiles; and distributing the tessellation instances among a plurality of tessellation pipelines operating in parallel, wherein a respective set of one or more of the tessellation instances is allocated to each of the tessellation pipelines, and wherein each of the tessellation pipelines generates tessellation geometry data associated with the respective allocated set of one or more of the tessellation instances.
[0009] A tessellation module is provided, which is configured to tessellate tiles to generate tessellation geometry data representing the tessellated tiles, the tessellation module comprising: tessellation factor logic, which is configured to process received geometry data representing the tiles to identify a tessellation factor of the tiles; a plurality of tessellation pipelines, which are arranged to operate in parallel; and a controller, which is configured to: determine tessellation instances to be used in the process of tessellizing the tiles based on the identified tessellation factors of the tiles; and distribute the tessellation instances among the plurality of tessellation pipelines, thereby distributing a respective set of one or more of the tessellation instances to each of the tessellation pipelines, wherein each of the tessellation pipelines is configured to generate tessellation geometry data associated with the distributed set of one or more of the tessellation instances.
[0010] A tessellation module is provided, which is configured to tessellate tiles to generate tessellation geometry data representing the tessellated tiles, the tessellation module comprising: a plurality of cores, each core comprising a plurality of tessellation pipelines and a controller; and a tile dispatcher configured to replicate a set of tiles and pass the set of tiles to each of the plurality of cores; wherein each of the cores is configured to: process a respective tile of the set at a respective tessellation pipeline to identify a tessellation factor for the tile of the set; determine, at the controller of the core, a tessellation instance to be used in tessellizing the set of tiles based on the identified tessellation factor for the set of tiles; determine, at the controller of the core, an allocation of the tessellation instance among the tessellation pipelines of the core; and process the tessellation instance at the allocated tessellation pipelines to generate tessellation geometry data associated with the respective allocated tessellation instance, wherein the controllers of the plurality of cores are configured to cause a subset of the tessellation instances of a tile to be allocated to the tessellation pipelines of the cores, and cause all tessellation instances of a tile to be processed collectively on all cores.
[0011] The inlay module may be embodied in hardware on an integrated circuit. A method of manufacturing an inlay module at an integrated circuit manufacturing system may be provided. An integrated circuit definition data set may be provided that, when processed in the integrated circuit manufacturing system, configures the system to manufacture the inlay module. A non-transitory computer-readable storage medium may be provided on which a computer-readable description of an integrated circuit is stored, which, when processed, causes a layout processing system to generate a circuit layout description used in the integrated circuit manufacturing system to manufacture the inlay module.
[0012] An integrated circuit manufacturing system may be provided, the integrated circuit manufacturing system comprising: a non-transitory computer-readable storage medium having stored thereon a computer-readable integrated circuit description describing a damascene module; a layout processing system configured to process the integrated circuit description so as to generate a circuit layout description of an integrated circuit embodying the damascene module; and an integrated circuit production system configured to manufacture the damascene module based on the circuit layout description.
[0013] Computer program code for executing any of the methods described herein may be provided.A non-transitory computer-readable storage medium having computer-readable instructions stored thereon may be provided, which when executed at a computer system causes the computer system to perform any of the methods as described herein.
[0014] The above features may be combined as desired, as will be apparent to the skilled person, and may be combined with any aspects of the examples described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Examples will now be described in detail with reference to the accompanying drawings, in which:
[0016] Figure 1 A block diagram of a mosaic system is shown;
[0017] FIG. 2( a ) shows an example mosaic module according to the present disclosure;
[0018] FIG2( b ) is a flow chart illustrating an example method of distributing tessellation instances among a tessellation pipeline according to the present disclosure;
[0019] Figure 3 shows an example process for writing data to a buffer;
[0020] Figure 4 shows an example process for reading data from a buffer;
[0021] Figure 5 Another example mosaic module according to the present disclosure is shown;
[0022] Figure 6 shows an example of data organization within a buffer;
[0023] Figure 7 (a) through 7(e) illustrate an example sequence of steps taken to process multiple mosaic instances;
[0024] Figure 8 (a) through 8(e) illustrate another example sequence of steps taken to process multiple mosaic instances;
[0025] Fig. 9 Another example mosaic module according to the present disclosure is shown;
[0026] Fig.10 is used Fig. 9 A flowchart of a method for inlaying patches with an inlay module is shown;
[0027] Fig.11 A computer system is shown in which a mosaic module is implemented; and
[0028] Fig.12 An integrated circuit manufacturing system for producing an integrated circuit embodying a damascene module is shown.
[0029] The accompanying drawings illustrate various examples. A skilled person will appreciate that an element boundary (e.g., a box, a group of boxes, or other shapes) shown in the drawings represents an example of a boundary. In some examples, an element may be designed as multiple elements, or multiple elements may be designed as one element. Where appropriate, common reference numerals are used throughout the figures to indicate similar features. DETAILED DESCRIPTION
[0030] The following description is presented by way of example to enable those skilled in the art to make and use the invention.The present invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to those skilled in the art.
[0031] The arrangement described herein provides an improved tessellation method, wherein the operations required to tessellate tiles can be divided into smaller workloads that can be allocated (or dispatched) among multiple tessellation pipelines for parallel operation. By providing the arrangement described herein, parallel tessellation of tiles of significantly different sizes can be performed by multiple tessellation pipelines on one or more processing cores without causing a reduction in processing that occurs in the prior art system described above due to the serialization of tile processing. Specifically, the tessellation work can be decomposed into different parts (or "tessellation instances") and dispatched on multiple tessellation pipelines. This reduces the amount of time that a tessellation pipeline is idle waiting for other tessellation pipelines to complete their work while maintaining the order in which the tessellated primitives are output. The term "tessellation pipeline" is used herein to refer to hardware for performing a sequence of processing stages, wherein the output of one processing stage provides input for a subsequent processing stage. A "tessellation pipeline" may or may not be dedicated to performing tessellation tasks. For example, a “tessellation pipeline” described herein may be a general processing pipeline that can perform several different types of processing tasks, such as executing programmable shading programs for a tessellation level, as well as performing other processing tasks such as vertex processing tasks and pixel processing tasks (e.g., texturing and shading), to name just a few examples.
[0032] Embodiments will now be described by way of example only.
[0033] FIG2( a ) shows a tessellation module 200 according to an example of the present disclosure. The tessellation module 200 includes a tessellation factor module 210 , a scheduler 220 , a plurality of tessellation pipelines 230 - 1 , 230 - 2 , 230 - 3 , and an optional memory 240 .
[0034] The tessellation factor module 210 is configured to receive the geometry data of the tile and process the geometry data of the tile to determine the tessellation factor to be used for tessellation of the tile. The tessellation factor is a value that defines the level of granularity at which the tile will be tessellated (usually defined per edge or per vertex). The tessellation factor thus defines the amount of tessellation to be performed on the tile, and therefore defines the number of primitives to be generated during tessellation. Therefore, from the tessellation factor it is possible to determine the amount of geometry data to be generated during tessellation of the tile. The tessellation factor module 210 may be referred to herein as "tessellation factor logic". In some examples (but not all examples), the tessellation factor logic may share processing resources with the tessellation pipeline 230, such as so that the tessellation factor logic and the tessellation pipeline are implemented using the same processing hardware, but it is shown as a separate component in FIG. 2 to illustrate the process flow of the tiles being processed in the tessellation module 200.
[0035] The scheduler 220 is configured to generate one or more tessellation instances for a given tile based on the determined tessellation factor for the tile. The scheduler 220 may be referred to herein as a controller. Each tessellation instance is associated with at least a portion of the tessellation geometry data for the tile, such that the geometries associated with all tessellation instances for the tile collectively define the tessellation geometry data for the tile. Thus, the tessellation instance may be considered to identify the geometry workload that will be performed to tessellate all or a portion of a tile.
[0036] By generating mosaic instances, the total amount of work required for mosaicking is divided into one or more batches of work that can be executed independently of each other. The mosaic instance therefore represents at least a portion of the data to be mosaicked. The scheduler is configured to dispatch mosaic instances to be processed by mosaic pipelines 230-1 to 230-3. Mosaic instances can be defined as having the same size, as will be explained in more detail later. The scheduler 220 can be configured to queue mosaic instances and dispatch mosaic instances to mosaic pipelines in a first-in, first-out order. In a simple example, the next mosaic instance that is not processed by the mosaic pipeline is passed for processing. This mosaic instance is passed to the next mosaic pipeline that becomes available or idle for processing, which occurs when the mosaic pipeline has completed processing of the previously received mosaic instance. However, in other examples, multiple mosaic instances can be submitted simultaneously to be processed by the mosaic pipeline. The mosaic pipeline runs tasks from one or more mosaic instances in any suitable order to process the mosaic instance. By submitting multiple tessellation instances to the pipeline at a given time, if one tessellation instance stalls for some reason, the pipeline can work on another tessellation instance so that the pipeline does not become idle. In addition, as mentioned above, the pipeline can process other types of work along with the tasks of the tessellation instance so that the pipeline does not become idle without tessellation work.
[0037] Each tessellation pipeline 230-1 to 230-3 includes a processing element configured to perform at least a portion of the tessellation process. In this way, tessellation occurs in each tessellation pipeline on a subset of the geometric shapes defined by the tiles. As will be appreciated, during identification of the tessellation factor, one or more steps of the tessellation process may need to be performed before scheduling the tessellation instance for processing. In some arrangements, this work is re-executed in the tessellation pipeline. However, in some other arrangements, this work is not re-executed in the tessellation pipeline. Instead, the scheduler 220 may store any data generated by the operations performed by the tessellation factor module 210 in the process of generating the tessellation factor and passed to the tessellation pipeline to avoid re-executing the operations required to generate this data. Therefore, the work performed by the tessellation pipeline may be a pared-down version of the work performed in a conventional single-phase tessellation pipeline. The tessellation pipelines 230-1 to 230-3 tessellate the received geometric shape data associated with a specific tessellation instance, which is assigned to the tessellation pipeline to generate primitive data defining the tessellated primitives generated during the tessellation. The geometric shape data is stored in the memory 240. Memory 240 is a memory configured to store primitive data generated by the tessellation pipelines 230-1 to 230-3 and to emit geometry in the correct order for further processing. The geometry, after tessellation, is typically emitted to a subsequent geometry processing pipeline stage (e.g., for performing clipping, viewport transformation, or projection, etc.), which may be performed by fixed function circuits in separate hardware units or may be performed by executing appropriate instructions on a processing unit that may or may not be part of the tessellation pipeline 230.
[0038] As previously mentioned, the tessellation pipeline may generate only a subset of the primitive data for a tile. The primitive data stored in the memory 240 is stored so that the primitive data can be combined to form a complete set of primitive data for the tile. For example, a tile may be defined by geometry data formed by four separate tessellation instances. Primitive data derived from the four tessellation instances may be stored, for example, in contiguous memory locations in the memory 240 so that the memory region of the primitive data spanning the four tessellation instances collectively defines the primitive data for the tile.
[0039] Figure 3 and 4 20. An example operation of the tessellation module is shown in FIG. As stated above, a buffer (which may be implemented in scheduler 220) is configured to hold geometry data to be dispatched among tessellation instances. The buffer may be a first-in, first-out buffer. In an example embodiment, there are two separate processes that operate to control the flow of data through the buffer. Specifically, a first process controls the writing of data to the buffer, and a second process controls the reading of data from the buffer. In this way, the buffer may be implemented as a circular buffer, where pointers may be used to handle writing to and reading from the buffer.
[0040] FIG. 2( b) shows a method 250 performed by the tessellation module 200 to tessellate a tile according to an example. The method 250 begins by identifying a tessellation factor for a tile at step 260. At step 270, the number of tessellation instances to be used to tessellate the tile is determined. For example, the number of tessellation instances to be used will depend on the tessellation factor. In one approach, it is possible to determine the number of primitives to be generated for tessellation of the tile based on the tessellation factor. Determining the number of tessellation instances may involve dividing the total number of primitives to be generated to represent the tile into predefined primitive batches to be allocated to different tessellation instances. At step 280, tessellation instances are allocated among the tessellation pipelines to tessellate corresponding portions of the tile in parallel. In other words, in step 280, the allocation of tessellation instances among the tessellation pipelines is determined.
[0041] Figure 3 An example method 300 for identifying a mosaic factor for a mosaic tile is shown. Specifically, Figure 3 The method 300 of begins at step 310, where input geometry data for the next tile to be processed is received according to the application and according to a predetermined order of the API. The input geometry data for the received tile may be defined by a number of control points. At step 320, the received geometry data is processed to determine the tessellation factor of the tile. Processing the received tile to determine the tessellation factor may include performing vertex shading and hull shading (or at least a portion of these shading processes). Vertex shading is a user-defined shading process that operates on a per-vertex (or per-control point) basis. Hull shading occurs together based on multiple control points. These processes will be described in more detail below. After the tessellation factor of the tile has been determined, at step 330, the number of tessellation instances to be used to tessellate the tile is determined, and the tessellation factor is written to a buffer. At step 340, if the buffer full threshold is not met, then the subsequent tile is retrieved at step 310. If the buffer full threshold is met, then the method 300 waits until the buffer is clear enough to store the tessellation factor of the subsequent tile before receiving the subsequent tile.
[0042] In an example method of determining the number of tessellation instances, the number of vertices that would be produced by tessellizing the tile using the determined tessellation factor is determined. The determination of the tessellation instances is less expensive to perform than a full tessellation process because only the input geometry data of the tile (rather than all data produced by the tile) needs to be processed, and furthermore the shading process within the hull shading stage that needs to be performed to determine the tessellation factor may only process the subset required to fully tessellate the tile. In this way, only the data required to determine the count of the number of primitives produced is determined and recorded.
[0043] Example equations for determining the number of tessellation instances generated from the geometry data (eg, control points) of a tile are set forth below.
[0044]
[0045] where J is the number of generated tessellation instances for a given patch, and N verts is the number of vertices that will be generated by performing tessellation of the tile according to the tessellation factor defined for the tile, and L is the number of vertices that should be batched at each tessellation pipeline. In one example, N may be determined based on the tessellation factor identified for the tile. verts L may be statically determined (eg, predetermined) based on the amount of memory storage available to store vertex data generated by each tessellation pipeline. In other words, L is the maximum number of vertices that can be assigned to a single pipeline so that processing does not stall due to lack of storage space.
[0046] For example, if each tessellation pipeline has an allocated memory size L of 1,000 vertices and the tile will generate 5,000 vertices (N verts ), then J=5 and five different tessellation instances are generated. Each tessellation instance is scheduled for processing in the tessellation pipeline.
[0047] Figure 4 A method 350 for assigning tessellation instances to a tessellation pipeline for processing is shown. At step 360, it is determined whether the buffer from which data is to be retrieved is empty. If the buffer is empty, there is currently no tile data to be processed, and the method waits for the data for the tile to be ready. If the buffer is not empty (i.e., there is some data to be processed), it is determined at step 370 whether the tessellation pipeline is available to process the data for the tile. A tessellation pipeline is available when at least the first stage of the pipeline is no longer processing a previously received tessellation instance. In some arrangements, because the tessellation pipeline can be configured to implement a pipelined process, processing of a current tessellation instance at a first pipeline stage can begin while the pipeline is simultaneously processing a previous tessellation instance at a later pipeline stage. In an example, the tessellation pipeline can generate an "available" signal at an appropriate stage during its processing of a tessellation instance to indicate that the tessellation pipeline is available to receive and begin processing the next tessellation instance. At step 380, the next tessellation instance is assigned to an available tessellation pipeline for processing, and the method returns to step 360 where it is determined whether there is more data to be sent to the tessellation pipeline for processing. The assignment of tessellation instances to tessellation pipelines may involve transmitting to the tessellation pipelines the input geometry data (e.g., control points) of the tile and the tessellation factor to be used for tessellizing the tile, as well as any edge data generated in determining the tessellation factor.
[0048] Figure 3 and 4The methods may be run in parallel in the tessellation module such that method 300 operates to fill a buffer with data containing tessellation factors for one or more tiles, and method 350 operates to read data from the buffer when a tessellation instance is allocated for processing one or more tiles.
[0049] Each pipeline may be configured to process more than one tessellation instance at a time (eg, from more than one tile), which may allow the pipeline to avoid becoming idle, or at least reduce the amount of time that the pipeline remains idle.
[0050] In one example, the geometry data associated with the tessellation instance is formed by dividing the tile into separate batches of output geometry data to be processed that will each produce a maximum number of vertices, which maximum number can be determined based on the identified tessellation factor. The next tessellation instance is determined based on the data produced by processing the current tile. As specified by the API, the geometry data produced from each tile will be output from the tessellation system 200 in the order in which the tile input data is received. Therefore, control logic coupled to each of the tessellation pipelines can be used to ensure that the order of primitives / vertices is maintained when processed primitives / vertices are issued or read from the memory 240 of the tessellation system. For example, the tessellation system can communicate with subsequent pipeline stages to indicate the availability of processed primitives / vertices by sending a signal, setting a flag, or incrementing a counter, and the subsequent stage can receive the signal or test the flag or counter to determine when the processed primitives / vertices associated with a particular tessellation instance can be read from the memory.
[0051] A tessellation instance may be associated with a predetermined maximum number L of vertices. Given a tile to be processed, it may be determined how many tessellation instances will need to be used. For example, based on a tessellation factor identified for the tile, it may be determined how many vertices will be generated during tessellation, given by N. verts Given. According to N verts With the determination of , it is possible to calculate the number of tessellation instances that need to be generated - i.e. Vertex. Among them N verts =4,500 and L=1,000, the first tessellation instance may involve generating the first 1,000 vertices (e.g., with indices 0 to 999), the second tessellation may involve generating the next 1,000 vertices (e.g., with indices 1,000 to 1,999), and so on. The final fifth tessellation instance may include the final 500 vertices (e.g., with indices 4,000 to 4,499). Alternatively, the vertices may be distributed more evenly among the tessellation instances. For example, 4,500 vertices may be distributed to 5 instances by associating 900 vertices with each tessellation instance.
[0052] As will be appreciated from the above, a tessellation instance therefore involves a subset of the tessellation work required to tessellate a tile. The data required for each tessellation instance includes the necessary data for the vertices to be processed in order to generate the primitive associated with the tessellation instance. The data includes all tile control data and tessellation factors, together with data indicating where in the tile tessellation should begin for a given instance. It will also be appreciated that the data may depend on the position of the vertices associated with the tessellation instance within the tile being tessellated. For example, for a high index vertex, a subset of the tessellation operations for a lower index vertex may have to be performed in order to allow a complete primitive to be formed.
[0053] Figure 5 An example tessellation module 500 according to the present disclosure is shown. The tessellation module 500 includes a first tessellation stage 510, a controller 520, a second tessellation stage 530, and an optional memory 540 (although the memory may be external to the tessellation module 500). The tessellation module 500 is similar to the tessellation module 200 described above and shown in FIG. 2 . In the tessellation module 500, the tessellation factor logic is implemented as the first tessellation stage 510, the tessellation pipeline is implemented in the second tessellation stage 530; and the controller 520 includes a scheduler 521 and other components as described below. The first and second tessellation stages (510 and 530) may share processing resources, for example, such that they are implemented using the same processing hardware, but they are Figure 5 is shown as a separate stage to illustrate the functionality of the way in which tiles are processed in a pipelined fashion.
[0054] The tessellation module 500 is provided with geometry data for one or more tiles from a geometry source 300 configured to provide the geometry data for the tiles in an order defined by an external operating application. The geometry source 300 receives control signals from the tessellation module 500 that control the transmission of the geometry data to the tessellation module 500. The geometry data for a particular tile may include untransformed vertex inputs in the form of control points that define the surface geometry of the tile. The geometry data for the tile is received from the geometry source 300 at a first tessellation stage 510.
[0055] The first tessellation stage 510 is configured to process the input geometry data of the tile to determine the tessellation factor of the tile so that it can be determined how many tessellation instances are to be instantiated by the controller 520 to tessellate the tile. The amount of processing required by the first tessellation stage to determine the tessellation factor may depend on the application being run. For example, the tessellation factor may be provided directly, i.e., the tessellation factor may be hard-coded. If this is the case, then no geometry processing is required by the first tessellation stage. For some applications, the tessellation factor may be determined programmatically, for example, based on the distance of the tile from the screen and / or based on the size of the tile. For such applications, it may be necessary to process untransformed vertex data (e.g., the control points of the tile) to determine the tessellation factor.
[0056] In one example, the first tessellation level may include one or more instances of a first vertex shader 511. The one or more first vertex shaders 511 may be configured to perform programmed per-vertex operations on the received untransformed vertex data. For example, the one or more first vertex shaders may be configured to perform at least a subset of the functions performed by a vertex shader as defined in the Direct3D or OpenGL standards. Because the tessellation module 500 may include one or more first vertex shaders, per-vertex shading operations may be performed in parallel on control points for a given tile, with each first vertex shader performing a subset of the per-vertex operations for the tile.
[0057] The processed vertex data output from the one or more first vertex shaders 511 is passed to one or more first tile shaders 512, which are configured to perform operations on multiple vertices by receiving one or more processed vertices and process the received vertices together. For example, the one or more tile shaders 512 may be configured to perform at least a subset of the functions performed by a hull shader as defined in the Direct3D standard or a tessellation control shader as defined in the OpenGL standard. The one or more first tile shaders 512 are configured to perform the minimum amount of processing required to generate a tessellation factor. Accordingly, the vertex shader and tile shader may have a reduced size and / or complexity compared to a complete vertex / hull shader required to fully implement operations as defined by the application programmer for these levels (as defined by the Direct3D and / or OpenGL standards).
[0058] The one or more first tile shaders 512 are configured to pass the identified tessellation factor of the tile and optionally any edge data resulting from the processing to the controller 520. The edge data may, for example, include coefficients of the tile. The controller 520 includes a buffer 522 configured to store data about the processed tile. The controller 520 further includes a scheduler 521 and a tessellation instance dispatcher 523.
[0059] The buffer 522 is configured to store data generated by the first mosaic stage 510 for each of the plurality of tiles. An example of the organization of data within the buffer 522 is shown in Figure 6 The buffer 600 is described in detail. Figure 6As shown in , data associated with each tile may be stored together. For example, for each tile, the buffer 600 may store a tile identifier 610 that identifies the specific tile to be processed. The buffer 600 may also store an execution address 620 for each tile, which identifies the memory address of the instruction to be executed during the second tessellation level 530 tessellation of the tile. For example, this may include vertex shading instructions, shell shading instructions, and / or domain shading instructions. For each tile, the buffer 600 may also store the tessellation factor 630 determined in the first tessellation level 510. The buffer 600 may also optionally store edge data for each tile generated as a result of processing data in the first tessellation level 510. The edge data may include some or all data generated as a result of processing performed during the first tessellation level and may be reused during the second tessellation level. By storing this data, the edge data does not have to be regenerated during the second tessellation level, which can reduce the amount of repeated processing in the second tessellation level caused by splitting the tessellation into multiple levels.
[0060] Buffer 522 stores data including the tessellation factor for each tile to be processed. Figure 5 In the embodiment, the controller 520 is configured to identify the number of tessellation instances to be used for processing each tile from the tessellation factor. The number of tessellation instances may be stored, for example, in a tessellation instance dispatcher 523 or in a buffer 522. The tessellation instance dispatcher 523 is configured to distribute (e.g., dispatch) the tessellation instances among the tessellation pipelines in the second tessellation stage 530. The tessellation instance dispatcher 523 may be configured in an example to implement Figure 4 Specifically, the tessellation instance dispatcher 523 may be configured to determine whether the buffer 522 is empty. If the buffer is not empty, there is at least one tessellation instance of the tile to be processed.
[0061] As mentioned above, the tessellation instance dispatcher 523 may be configured to determine the number of tessellation instances to be generated to process the tile based on the tessellation factor of the tile. The tessellation instance dispatcher 523 then determines whether there is a tessellation pipeline available to process the next tessellation instance to be processed. For example, the tessellation instance dispatcher 523 may receive a signal from the scheduler 521 indicating the availability status of one or more tessellation pipelines. If the tessellation pipeline is identified as available, the tessellation instance dispatcher 523 provides the next tessellation instance to the available tessellation pipeline. Even if the tessellation pipeline is not currently idle, the tessellation pipeline may be "available" when it is ready to receive a tessellation instance. The tessellation instance provided to the tessellation pipeline may be queued at the tessellation pipeline for processing (e.g., in a FIFO). The execution address, tessellation factor, and optional edge data of a particular tile are passed to the particular tessellation pipeline for processing. The dispatcher 523 also provides an indication to the particular tessellation pipeline of which portion of the tile the particular tessellation instance relates to. Tessellation instance dispatcher 523 can keep track of tessellation instances to be dispatched for a particular tile. For example, for each tile, dispatcher 523 can maintain a count of the number of tessellation instances required to process the tile and maintain an indication of which tessellation instances have been sent for processing. A flag can be used to maintain the processing status of each tessellation instance.
[0062] The scheduler 521 is configured to control reading from and writing to the buffer 522 to ensure that the buffer does not overflow, while also attempting to minimize the amount of time that the buffer is empty. This allows the tessellation module 500 to maximize the amount of time that the first and second tessellation stages 510 and 530 are operating to optimize the amount of processing. Specifically, the scheduler 521 monitors the number of entries currently in the buffer. If the buffer is not full (e.g., the buffer threshold is not met), the scheduler 521 sends a signal to the geometry source 300 to send another tile of data for processing by the first tessellation stage 510. In addition, the scheduler 521 is configured to control the tessellation instance dispatcher 523 by sending a control signal to send the data of the tessellation instance to the tessellation pipeline in the second tessellation stage 530. The scheduler 521 controls the tessellation instance dispatcher 523 based on the availability of the tessellation pipeline received from the second tessellation stage 530 as status information.
[0063] exist Figure 5 In the example of , the second tessellation stage 530 includes a plurality of tessellation pipelines, each of which includes a second vertex shader 531, a second tile shader, and a domain shader 533. The tessellation pipelines may also include a fixed function tessellation block (not shown) that performs a tessellation process as defined in more detail below. The tessellation pipelines may also include a geometry shader that is configured to apply geometry shading to the output of the domain shader 533.
[0064] The second vertex shaders 531 are each configured to perform tessellation pipeline operations on a per-vertex basis (e.g., on a control point basis for a tile). Specifically, the second vertex shaders 531 may be configured to perform at least a subset of the functions performed by a vertex shader as defined in the Direct3D or OpenGL standard. Because some vertex shading required for tessellation of tiles is performed by the one or more first vertex shaders 511 in the first tessellation level 510, the processing may optionally be skipped in the second tessellation level 530. For example, in the case where edge data 640 regarding output from the first vertex shader 511 is stored in the buffer 522, it is possible to skip the processing during the second tessellation level. For example, the first and second vertex shaders may collectively define a vertex shader as defined in the Direct3D or OpenGL standard, wherein each of the first and second vertex shaders performs a respective subset of the defined functionality. For example, the first vertex shader 511 may perform the geometry processing necessary to provide the first tile shader 512 with the required geometry data to identify the tessellation factor, while the second vertex shader 531 may perform other types of data processing (e.g., the second vertex shader 531 may change the basis function of the tile (e.g., Bezier to Catmul-Rom)). Alternatively, it is possible to reduce the storage requirements in the buffer 522 by not storing the output of the first vertex shader between tessellation levels. In this way, the second vertex shader 531 may be required to repeat some of the processing that has already been performed by the first vertex shader 511. For example, Figure 5 As shown in , the second vertex shader may be configured to receive untransformed geometry data from the geometry source 300. As a result, the first vertex shader output does not have to be stored in the buffer 522.
[0065] The second tile shader 532 may be configured to perform at least a subset of the functions performed by a hull shader as defined in the Direct3D standard or a tessellation control shader as defined in the OpenGL standard. In this example, the second tile shader 532 is stripped of any processing involved in generating tessellation factors and optionally any edge data. This is because this data has been determined during the first tessellation level and is held in the buffer 522, so it is not necessary to generate this data again. The results generated by the second tile shader (along with the pre-generated tessellation factors and edge data) are passed to a fixed function tessellation module (not shown), which performs a predefined process of tessellizing the geometry of the tessellation instance based on the tessellation factors and edge data to generate output data defining domain indices and coordinates for sub-tiling. For example, the outputs of the second tile shader 532 and the fixed function tessellation are tessellated primitives and domain indices and UV coordinates. Alternatively, domain points may be pre-generated by the fixed function tessellation unit within the tessellation instance dispatcher and dispatched directly with the tile instance. Like the first and second vertex shaders, the first and second tile shaders may collectively define a shell shader or a tessellation control shader, wherein each of the first and second tile shaders performs a respective subset of the defined functionality. Alternatively, the second tile shader may repeat at least a portion of the processing performed by the first tile shader in order to reduce the amount of storage required by the buffer 522.
[0066] The one or more domain shaders 533 may be configured according to domain shaders as defined in the Direct3D standard and tessellation evaluation shaders as defined in the OpenGL standard. Specifically, the domain shader 533 is configured to consume the output domain coordinates from the fixed function tessellation unit and the output control points from the second tile shader 532, and generate the positions (and other data) of one or more vertices of the tessellated geometry. For a tessellation instance, the vertices of the tessellation instance are generated and passed to the memory 540. The vertex data for each tile may be provided from the memory 540 for further processing. For example, the tessellated geometry may be further processed using a geometry shader and then passed to a culling module configured to cull vertices that are not visible in the scene (e.g., using backface culling or small object culling), and then passed to the clipping, viewport transformation, and projection modules.
[0067] As mentioned previously, the memory 540 may be one or more physical memories configured to store the results of each tessellation pipeline. For example, the one or more physical memories may form a plurality of logical memories, wherein each logical memory is configured to store the combined geometry from each of a plurality of tessellation instances that collectively define the tessellated vertices of the tile. In this way, the tessellated vertex data of the tile may be reassembled in the memory 540. This will be relative to Figure 7 Explain in more detail.
[0068] Figure 7 (a) to 7(e) illustrate a simple example in which a sequence of steps is taken to process multiple tessellation instances using four tessellation pipelines (ie, pipelines 230-1 to 230-4). Figure 7 (a) shows the first step, in which nine tessellation instances are identified. In this example, there are three individually tessellated tiles, each separated into three tessellation instances such that each tile contains a first (denoted as tile "x" TI 0), a second (denoted as tile "x" TI 1), and a third (tile "x" TI 2) tessellation instance. As can be seen from Figure 7 It can be seen that the resulting vertex data is to be stored in the memory 700. Figure 7 In the example of , a single physical memory is used. The single physical memory is separated into three logic blocks, wherein each logic block is configured to store vertex data generated for a tile. For example, the first logic block 710 is configured to store vertex data of a first tile, the second logic block 720 is configured to store vertex data of a second tile, and the third logic block 730 is configured to store vertex data of a third tile.
[0069] exist Figure 7 In (b), it is determined that all four tessellation pipelines 230-1 to 230-4 are available for processing because tessellation has just begun in this example. Accordingly, the first tessellation instance (tile OT1 0) is passed by the tessellation instance dispatcher to the first tessellation pipeline 230-1 for processing. Similarly, the next tessellation instance (tile OT1 1) is passed to the next tessellation pipeline 230-2, and so on, until the first four identified tessellation instances have been passed to the four tessellation pipelines for processing. Thus, there are five tessellation instances yet to be allocated for processing by the tessellation pipelines. No further allocation of tessellation instances to tessellation pipelines can occur at this point because there are no more available tessellation pipelines. In Figure 7 In the simplified examples shown in (a) through 7(e), the pipeline contains a single instance at a time. However, in other examples, the pipeline may not be constrained to contain only a single instance at a time. The vertex shading, tile shading, and domain shading stages are inherently programmable, so it may be beneficial for the pipeline to process multiple instances in parallel, allowing the pipeline to hide (a) internal pipeline latencies and (b) any latencies associated with external memory fetches. In these examples, memory 700 has (at least) enough space to consume enough parallel instances to hide at least the internal latencies.
[0070] exist Figure 7 At (c), the tessellation pipelines have each completed processing of the first batch of received tessellation instances and have provided the resulting vertex data for the first batch of tessellation instances to memory 700. Figure 7(c) As can be seen, the vertex data generated from the tessellation instance of the first tile is stored sequentially in the logical memory 710. Similarly, the first tessellation instance of the second tile (P1TI0) is stored in a logical memory configured to store vertex data for the second tessellation pipeline. In other examples, the memory 700 may not be divided into separate logical blocks, and the storage of vertex data generated from the tessellation instances may be stored out of order in the logical memory or in a single memory 700. The allocation of storage space from the memory may be managed by any memory management technique, such as using pointers, indexes, or linked lists, which allows the generated vertex data to be located and read out to subsequent pipeline stages in sequence. Figure 7 In the example of (c), all vertex data generated from the tessellation instance of the first tile is available in memory. The availability of data can be indicated to subsequent pipeline stages, and the data can then be read from memory 700. Figure 7 7(e) reads data from memory simultaneously, and memory may then be freed up for storing vertex data generated from tessellation instances of more tiles. In another example, the availability of data may be indicated to subsequent pipeline stages for vertex data from each of the tessellation instances individually, rather than waiting until vertex data for a complete tile is available. The order in which vertices arrive at subsequent pipeline stages may be maintained by communicating between the tessellation system and subsequent pipeline stages, such as by sending signals, setting flags, or incrementing counters as described above, so that subsequent pipeline stages read each item of vertex data generated in sequence, and not before it becomes available in memory 700.
[0071] As described earlier, the tessellation pipeline can identify when it is available to receive a tessellation instance. For example, where the tessellation pipeline is a pipelined process, it is possible to receive the next tessellation instance before the previous tessellation instance is completed. Once it has been identified that the tessellation pipeline is available to receive a tessellation instance, the next tessellation instance to be processed is passed to the tessellation pipeline for processing. Figure 7 As can be seen in (c), a second batch formed by the next four tessellation instances from the list of tessellation instances to be processed are respectively passed to the tessellation pipeline for processing.
[0072] exist Figure 7 In (d), vertex data for each tessellation instance of the second batch is generated and stored in the appropriate portion of memory 700. As can be seen, vertex data for the second tile has been stored in logical memory location 720. Vertex data for the first and second tessellation instances (P2TI 0 and TI 1) of the third tile are stored in logical memory 730 of that tile. Figure 7 (d), the remaining third tessellation instance (P 2TI 2) of the third tile is passed to the first tessellation pipeline and processed and stored in the logic memory 730, as shown in FIG. Figure 7 As shown in (e).
[0073] Figure 8 A similar arrangement is shown, where three different tiles with different numbers of vertices are to be tessellated. Figure 7 In , each tile generates three tessellation instances when processed in the first tessellation level. In contrast, in Figure 8 The first tile (Tile 0) forms a single tessellation instance, the second tile (Tile 1) forms five tessellation instances, and the third tile (Tile 2) forms three tessellation instances. Fill at a rate that depends on the number of tessellation pipelines present in the tessellation module Figure 8 Memory 800.
[0074] Similar to Figure 7 The example shown in Figure 8 In the example shown, the pipeline contains a single instance at a time. However, as described above, in other examples, the pipeline may not be constrained to contain only a single instance at a time, and in fact the pipeline may process multiple instances in parallel.
[0075] Figure 7 and 8 An example of a system is shown in which the memory 700 or 800 is large enough to contain all vertex data generated by the tessellation instance. Figure 7 and 8 and more specifically compared to Figure 1 In systems with tessellation, scheduling of tessellation instances into tessellation pipelines allows the amount of required memory to be reduced significantly further. Figure 7 In the example of , it can be seen that the tessellation instances are dispatched across the four tessellation pipelines, such that the tessellation instance for tile 0 is scheduled before the tessellation instance for tile 1, and the tessellation instance for tile 1 is scheduled before the tessellation instance for tile 2. This is consistent with Figure 1 In contrast to the example Figure 1 In the example of , each tile is scheduled to be fully tessellated on a specific tessellation unit. Figure 7 As can be seen in (c), the first four sets of generated vertex data written to logical memories 710 and 720 are the first four sets that must be read from memory 700 when reading out the vertices in the correct order. Figure 7 In (d), the next four sets of generated vertex data written to logical memories 720 and 730 are the next four sets that must be read from memory in sequence following the vertex data from the previous step. Figure 7(e), the final set of generated vertex data written to logical memory 730 is the last set that must be read from memory. The requirement for reordering vertex data sets is therefore limited to the number of vertex data sets that can be generated by four pipelines. In theory, a memory large enough to store four sets of generated vertex data (or T sets of generated vertex data in a system with T tessellation units) is sufficient. If each tessellation pipeline can contain more than one tessellation instance at a time, the memory requirements can be increased. For example, a system with four tessellation pipelines, each of which can contain two tessellation instances, can generate up to eight vertex data sets in any order. A memory capable of storing eight vertex data sets can therefore be used to allow reordering. If additional buffering is desired, the memory size can also be increased beyond the size calculated in this way. For example, double buffering can be used so that the tessellation pipeline can write to the memory while the subsequent pipeline stage is reading out. Additional buffering can be used, for example as a FIFO buffer, to smooth out data streams in which the rate at which the tessellation unit generates vertex data or the rate at which the subsequent pipeline stage consumes is uneven. The size of the tessellation instance can be selected in order to target a specific memory size. In an example where a tessellation instance is associated with up to 1000 vertices, the visible memory is approximately Figure 1 The system requires one-fourth the memory size, Figure 1 In a system with tessellation pipelines, a tile may generate up to 4096 vertices. The total number of vertices that may be generated from a tile may not be under the control of the tessellation system designer, but the size of the tessellation instance is under the control of the tessellation system designer. The number of vertices associated with a tessellation instance may be made much smaller, such as 16 vertices, in which case the amount of memory required is reduced to approximately 6 kilobytes (for a system with four tessellation pipelines).
[0076] In the arrangement described above, tessellation instances are defined based on a predetermined number of tessellated vertices (i.e., the vertex count), and are related to the amount of memory allocated to each tessellation pipeline. In the arrangement described above, some tessellation instances may be associated with fewer vertices than the vertex count. For example, if the vertex count is 1,000 and the tile will produce 2,225 tessellated vertices, the first and second tessellation instances may each be associated with 1,000 vertices, but the third tessellation instance may be associated with only 225 vertices. It will be appreciated that this may result in a reduction in processing throughput, because if a tessellation pipeline is processing a tessellation instance that will produce a number of vertices less than the vertex count, the tessellation pipeline may not be operating at full capacity.
[0077] To account for this reduction in processing volume, in some arrangements it is possible to combine tessellation instances from different tiles that, when combined, produce a number of vertices less than or equal to the vertex count. For example, vertices from the first tessellation instance of a tile may be included in the final tessellation instance of the previous tile. While this approach may mean that some tessellation instances have a better number of vertices to produce, there may be additional complexity in processing these tessellation instances because data about more than one tile may need to be provided to the tessellation pipeline used to process a particular tessellation instance, and because more than one tessellation operation may be required to process a particular tessellation instance.
[0078] Fig. 9 Another example tessellation module 900 according to the present disclosure is shown. The tessellation module 900 includes three processing cores: core 0 (9020), core 1 (9021), and core 2 (9022). Each core includes a controller 904; four tessellation pipelines 906, 907, 908, and 909; and a memory 910. The tessellation module 900 also includes a tile dispatcher 912.
[0079] The tessellation module 900 is provided with geometry data for one or more tiles from a geometry source 300 configured to provide the geometry data for the tiles in an order defined by an external operating application. The geometry data for a particular tile may include untransformed vertex input in the form of control points defining the surface geometry of the tile.
[0080] refer to Fig.10 The flowchart shown in describes the operation of the tessellation module 900. In step S1002, geometry data for a set of one or more tiles is received from the geometry source 300 at the tile dispatcher 912.
[0081] In step S1004, tile dispatcher 912 copies a set of tiles and passes the set of tiles to each core. The number of tiles included in the set can be chosen to match the number of mosaic pipelines in each of cores 902. Fig. 9 In the example shown in , the tile set includes four tiles, and this set of four tiles is provided to each of the cores 9020 , 9021 , and 9022 .
[0082] In step S1006, each of the cores operates independently to determine the tessellation factor for the tiles of the collection. As described in the example above, the tessellation factor is determined by executing the vertex shader and the tile shader. This can be described as the first execution stage. Step S1006 involves running the vertex and tile shaders for the four tiles of the collection at each of the cores 902. Because each core 902 contains four pipelines (i.e., the number of pipelines in the core is the same as the number of tiles in the collection), each pipeline in the core performs vertex shading and tile shading for a corresponding tile of the collection. By matching the number of tiles in the collection with the number of tessellation pipelines in the core, optimal utilization of the hardware can be achieved.
[0083] For example, a set of tiles assigned to four cores includes four tiles: tile 0, tile 1, tile 2, and tile 3. In core 0 9020, pipeline 0 9060 performs vertex shading and tile shading (e.g., including hull shading) for tile 0; pipeline 1 9070 performs vertex shading and tile shading (e.g., including hull shading) for tile 1; pipeline 2 9080 performs vertex shading and patch shading (e.g., including hull shading) for tile 2; and pipeline 3 9090 performs vertex shading and patch shading (e.g., including hull shading) for tile 3. Similarly, in core 1 9021, pipeline 0 9061 performs vertex shading and patch shading (e.g., including hull shading) for tile 0; pipeline 1 9071 performs vertex shading and patch shading (e.g., including hull shading) for tile 1; pipeline 2 9081 performs vertex shading and patch shading (e.g., including hull shading) for tile 2; and pipeline 3 9091 performs vertex shading and patch shading (e.g., including hull shading) for tile 3. In addition, in core 2 9022, pipeline 0 9062 performs vertex shading and patch shading (e.g., including hull shading) for tile 0; pipeline 1 9072 performs vertex shading and patch shading (e.g., including hull shading) for tile 1; pipeline 2 9082 performs vertex shading and patch shading (e.g., including hull shading) for tile 2; and pipeline 3 9092 performs vertex shading and patch shading (e.g., including hull shading) for tile 3.
[0084] Thus, after step S1006, each core has determined the tessellation factor for each tile of the set. In step S1008, for each of the cores 902, the controller 904 determines the tessellation instances to be processed at that particular core. In other words, in step S1008, for each of the cores 902, the controller 904 determines the distribution of the tessellation instances to be processed across the tessellation pipelines of that core. The controller 904 of each core 902 has all the information it needs to figure out which tessellation instances of the tile to be processed at that core. For example, the controller 904 of each core 902 may know: (i) the number of cores 902 and / or the number of tessellation pipelines 906-909 in the tessellation module 900, (ii) the functional location of the core 902 within the tessellation module 900, and (iii) the available output storage of the memory 910 in the core 902. Based on this information, the core 902 x Controller 904 x Identifiable core 902 x Which tessellation instances of the tile are to be processed. This information may be predetermined and stored locally in the controller 904 of the core 902, or some or all of this information may be provided to the core 902 from the tile dispatcher 912. In this way, the cores 902 operate collectively to process all tessellation instances of the tile. In other words, a subset of the tessellation instances of the tile are assigned to the tessellation pipelines of the cores, where collectively on all cores, all tessellation instances of the tile are processed. The vertex and tile shading operations of the first execution stage are repeated across different cores, but the domain shading operations (of the tessellation instances) are not repeated across different cores. The controller 904 passes the appropriate tessellation instances to the corresponding tessellation pipelines 906-909 within the core 902.
[0085] The distribution of tessellation instances across the tessellation pipelines of multiple cores is preferably such that tessellation instances of one tile are processed in parallel in as many tessellation pipelines as possible, with tessellation instances of a first tile being scheduled before instances of a second tile. In this way, tessellation instances are also achieved in systems with multiple processing cores. Figure 7 and 8 The advantages of scheduling tessellation instances demonstrated in the description of FIG. 1006 are demonstrated. For example, there is some duplication of workload at S1006, where the tessellation factors for each tile are calculated at each core. However, this is a relatively small amount of computation, and it permits each core to perform the allocation of tessellation instances to its own tessellation pipeline without the need to communicate with other cores. Avoiding the need for cores to communicate with each other avoids the need for a central control unit (which can become a bottleneck in either processing or silicon layout), and enables a more scalable parallel tessellation system.
[0086] In step S1010, the tessellation pipelines 906-909 process the tessellation instance in the second execution stage to generate the tessellation geometry of the tile. As described above, the processing of the tessellation instance involves performing domain shading operations. Because vertex shading and tile shading operations are performed for each tile in each core, each core is able to access the results of the vertex and tile shading operations performed during the first execution stage. Domain shading may include consuming output domain coordinates from the fixed function tessellation unit and output control points from the tile shader, as well as generating the positions (and other data) of one or more vertices of the tessellated geometry. For the tessellation instance, the vertices of the tessellation instance are generated and passed to the memory 910 of the core 902.
[0087] In step S1011, the tessellated vertex data for each tile may be provided from the memory 910 of each of the cores 902 for further processing. As part of step S1011, control logic (e.g., controller 904) controls the emission of the tessellated vertex data for the tiles to ensure that the correct vertex ordering (according to the submission order of the geometry from the geometry source 300) is maintained. For example, the emission of processed vertices may be blocked for a tessellation instance until the processed vertices have been emitted for all previous tessellation instances. For example, the emitted tessellated geometry may be further processed using a geometry shader and then passed to a culling module configured to cull vertices that are not visible in the scene (e.g., using backface culling or small object culling), and then passed onto the clipping, viewport transformation, and projection modules.
[0088] In step S1012, the tessellation module 900 determines whether there are more tile sets to be tessellated. If there are more tiles to be tessellated, the method returns to step S1004 so that another tile set is copied and passed to each core. If necessary, a signal is sent to the geometry source to send more geometry data to the tile dispatcher 912. If it is determined in step S1012 that there are no more tile sets to be tessellated, the method proceeds to S1014, where the method ends.
[0089] References Fig. 9 and 10 The described scheme may avoid performing vertex shading and tile shading stages in the second execution stage (i.e., after the tessellation instance has been determined). Repeating the vertex shading and tile shading stages across all cores ensures that each core has the results of vertex shading and tile shading operations for any tile for which the tessellation instance may be processed at that core. The controller 904 may include buffering to store data generated during the first execution stage so that it may be reused during the second execution stage. Alternatively, the second execution stage may repeat at least a portion of the processing performed by the first execution stage in order to reduce the amount of storage required for buffering in the controller 904.
[0090] In one example, the memory 910 of each of the cores 902 has capacity for 16 output (i.e., tessellated) vertices. It should be noted that this number may vary based on vertex size, but for this simple example, it is assumed that vertex data for 16 vertices can be stored in each memory 910 at a given time. Thus, each tessellation instance is associated with four tessellated vertices for a tile, so that a tessellation instance can be provided to each of the four pipelines 906-909 within the core at a given time. Four tiles (Tile 0, Tile 1, Tile 2, and Tile 3) are included in the set.
[0091] In this example, first, on each core 902, tessellation pipeline 0 906 performs vertex shading and tile shading on tile 0; tessellation pipeline 1 907 performs vertex shading and tile shading on tile 1; tessellation pipeline 2 908 performs vertex shading and tile shading on tile 2; and tessellation pipeline 3 909 performs vertex shading and tile shading on tile 3. Tile 0 generates 384 vertices, tile 1 generates 96 vertices, tile 2 generates 40 vertices, and tile 3 generates 180 vertices.
[0092] Each of the controllers 904 determines that tile 0 is to be processed as 96 tessellation instances; tile 1 is to be processed as 24 tessellation instances; tile 2 is to be processed as 10 tessellation instances; and tile 3 is to be processed as 45 tessellation instances. These tessellation instances are assigned for execution by the pipelines of the core 902. The following table shows how the tessellation instances (which may each be associated with up to four tessellated vertices) are dispatched across the different pipelines of the different cores for the four tiles:
[0093] core Pipeline Patches vertex 0 0 0 0-3 0 1 0 4-7 0 2 0 8-11 0 3 0 12-15 1 0 0 16-19 1 1 0 20-23 1 2 0 24-27 1 3 0 28-31 2 0 0 32-35 2 1 0 36-39 2 2 0 40-43 2 3 0 44-47 0 0 0 48-51 0 1 0 52-55 0 2 0 56-59 0 3 0 60-63 1 0 0 64-67 1 1 0 68-71 1 2 0 72-75 1 3 0 76-79 2 0 0 80-83 2 1 0 84-87 2 2 0 88-91 2 3 0 92-95 0 0 1 0-3 0 1 1 4-7 0 2 1 8-11 0 3 1 12-15 1 0 1 16-19 1 1 1 20-23 2 0 2 0-3 2 1 2 4-7 2 2 2 8-9 0 0 3 0-3 0 1 3 4-7 0 2 3 8-11 0 3 3 12-15 1 0 3 16-19 1 1 3 20-23 1 2 3 24-27 1 3 3 28-31 2 0 3 32-35 2 1 3 36-39 2 2 3 40-43 2 3 3 44
[0094] Each row of the table shown above relates to a tessellation instance and indicates which pipeline of which core processes the tessellation instance and also indicates which vertex of which tile is produced by processing the tessellation instance. Different cores and different pipelines of a core operate in parallel.
[0095] Fig.11 1 shows a computer system in which the graphics processing system and tessellation modules described herein may be implemented. The computer system includes a CPU 1102, a GPU 1104, a memory 1106, and other devices 1112, such as a display 1116, a speaker 1118, and a camera 1114. A tessellation module 1110 (e.g., tessellation modules 200, 500, and 900) is implemented on the GPU 1104. The components of the computer system may communicate with each other via a communication bus 1120.
[0096] refer to Figures 1 to 10The described mosaic modules are shown as including several functional blocks. This is merely illustrative and is not intended to define a strict division between different logical elements of such entities. Each functional block may be provided in any suitable manner. It should be understood that intermediate values described herein as being formed by the mosaic modules need not be physically generated by the mosaic modules at any point and may simply represent logical values between their inputs and outputs that conveniently describe the processing performed by the mosaic modules.
[0097] The inlay module described herein may be embodied in hardware on an integrated circuit. The inlay module described herein may be configured to perform any of the methods described herein. Typically, any of the functions, methods, techniques, or components described above may be implemented in software, firmware, hardware (e.g., fixed logic circuits) or any combination thereof. The terms "module," "functionality," "component," "element," "unit," "block," and "logic" may be used herein to generally represent software, firmware, hardware, or any combination thereof. In the case of software implementation, a module, functionality, component, element, unit, block, or logic represents a program code that performs a specified task when executed on a processor. The algorithms and methods described herein may be performed by one or more processors that execute code, and the code causes the processor to perform the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disks, flash memory, hard disk storage, and other memory devices that can use magnetic, optical, and other technologies to store instructions or other data and can be accessed by a machine.
[0098] As used herein, the terms computer program code and computer readable instructions refer to any kind of executable code for execution by a processor, including code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, byte code, code defining an integrated circuit (e.g., hardware description language or netlist), and code expressed in programming language code such as C, Java, or OpenCL. Executable code can be, for example, any kind of software, firmware, script, module, or library that, when properly executed, processed, interpreted, compiled, run in a virtual machine or other software environment, causes a processor of a computer system supporting the executable code to perform the tasks specified by the code.
[0099] A processor, computer or computer system may be any kind of device, machine or special purpose circuit, or a collection or part thereof, which has processing capabilities so that instructions can be executed. A processor may be any kind of general or special purpose processor, such as a CPU, GPU, system on a chip, state machine, media processor, application specific integrated circuit (ASIC), programmable logic array, field programmable gate array (FPGA), etc. A computer or computer system may include one or more processors.
[0100] The present invention is also intended to cover software that defines the hardware configuration as described herein, such as HDL (Hardware Description Language) software, for designing integrated circuits or for configuring programmable chips to perform desired functions. That is, a computer-readable storage medium may be provided having encoded thereon computer-readable program code in the form of an integrated circuit definition data set that, when processed (i.e., run) in an integrated circuit manufacturing system, configures the system to manufacture a mosaic module configured to perform any of the methods described herein, or to manufacture a mosaic module including any of the devices described herein. The integrated circuit definition data set may be, for example, an integrated circuit description.
[0101] Thus, a method of manufacturing a damascene module as described herein at an integrated circuit manufacturing system may be provided. Furthermore, an integrated circuit definition data set may be provided which, when processed in an integrated circuit manufacturing system, causes the method of manufacturing a damascene module to be performed.
[0102] The integrated circuit definition data set may be in the form of computer code, for example, as a netlist, code for configuring a programmable chip, as a hardware description language that defines an integrated circuit at any level, including as register transfer level (RTL) code, as a high-level circuit representation such as Verilog or VHDL, and as a low-level circuit representation such as OASIS (RTM) and GDSII. The higher-level representation (e.g., RTL) that logically defines the integrated circuit may be processed at a computer system configured to generate a manufacturing definition of the integrated circuit in the context of a software environment that includes definitions of circuit elements and rules for combining those elements to generate a manufacturing definition of the integrated circuit so defined by the representation. As is typically the case with software executing on a computer system to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to configure the computer system for generating a manufacturing definition of the integrated circuit to execute code that defines the integrated circuit to generate the manufacturing definition of the integrated circuit.
[0103] Now refer to Fig.12 An example of processing an integrated circuit definition data set at an integrated circuit manufacturing system to configure the system to manufacture a damascene module is described.
[0104] Fig.12 An example of an integrated circuit (IC) manufacturing system 1202 is shown, which is configured to manufacture a damascene module as described in any of the examples herein. Specifically, the IC manufacturing system 1202 includes a layout processing system 1204 and an integrated circuit generation system 1206. The IC manufacturing system 1202 is configured to receive an IC definition data set (e.g., defining a damascene module as described in any of the examples herein), process the IC definition data set, and generate an IC (e.g., which embodies a damascene module as described in any of the examples herein) based on the IC definition data set. The processing of the IC definition data set configures the IC manufacturing system 1202 to manufacture an integrated circuit embodying a damascene module as described in any of the examples herein.
[0105] The layout processing system 1204 is configured to receive and process an IC definition data set to determine a circuit layout. Methods for determining a circuit layout based on an IC definition data set are known in the art and may, for example, involve synthesizing RTL code to determine a gate-level representation of a circuit to be generated, for example, in terms of logic components (e.g., NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). By determining location information for the logic components, the circuit layout may be determined based on the gate-level representation of the circuit. This may be done automatically or with user involvement to optimize the circuit layout. When the layout processing system 1204 has determined the circuit layout, it may output the circuit layout definition to the IC generation system 1206. The circuit layout definition may be, for example, a circuit layout description.
[0106] As is known in the art, IC production system 1206 produces ICs according to the circuit layout definition. For example, IC production system 1206 can implement a semiconductor device manufacturing process to produce ICs, which can involve a multi-step sequence of photolithography and chemical processing steps during which electronic circuits are gradually formed on a wafer made of semiconductor material. The circuit layout definition can be in the form of a mask that can be used in a photolithography process to produce the IC according to the circuit definition. Alternatively, the circuit layout definition provided to IC production system 1206 can be in the form of a computer readable code that IC production system 1206 can use to form a suitable mask for producing the IC.
[0107] The different processes performed by the IC manufacturing system 1202 can all be performed at one location, such as by one party. Alternatively, the IC manufacturing system 1202 can be a distributed system such that some processes can be performed at different locations and can be performed by different parties. For example, some of the following stages can be performed at different locations and / or by different parties: (i) synthesizing RTL code representing an IC definition data set to form a gate-level representation of a circuit to be produced, (ii) generating a circuit layout based on the gate-level representation, (iii) forming a mask based on the circuit layout, and (iv) using the mask to manufacture the integrated circuit.
[0108] In other examples, processing an integrated circuit definition data set at an integrated circuit manufacturing system may configure the system to manufacture a mosaic module without processing the IC definition data set to determine the circuit layout. For example, the integrated circuit definition data set may define a configuration of a reconfigurable processor (e.g., an FPGA), and processing the data set may configure the IC manufacturing system to produce a reconfigurable processor having the defined configuration (e.g., by loading the configuration data into the FPGA).
[0109] In some embodiments, when processed in an integrated circuit manufacturing system, the integrated circuit manufacturing definition data set may cause the integrated circuit manufacturing system to produce an apparatus as described herein. Fig.12 The integrated circuit manufacturing system can be configured in the described manner to manufacture the device described herein.
[0110] In some examples, an integrated circuit definition data set may include software that runs on, or in combination with, the hardware defined at the data set. Fig.12 In the example shown, the IC production system can be further configured by the integrated circuit definition dataset to load firmware onto the integrated circuit according to the program code defined in the integrated circuit definition dataset when manufacturing the integrated circuit, or otherwise provide the integrated circuit with program code for use with the integrated circuit.
[0111] Compared to known implementations, the implementation of the concepts set forth in this application in devices, equipment, modules and / or systems (and in the methods implemented herein) can cause performance improvements. The performance improvement may include one or more of improved computing performance, reduced waiting time, increased throughput and / or reduced power consumption. During the manufacture of such devices, equipment, modules and systems (e.g., in integrated circuits), a trade-off can be made between performance improvements and physical implementations to improve the manufacturing method. For example, a trade-off can be made between performance improvements and layout area to match the performance of known implementations, but using less silicon. For example, this can be accomplished by reusing functional blocks in a serial manner or sharing functional blocks between elements of devices, equipment, modules and / or systems. In contrast, the concepts set forth in this application that cause improvements in the physical implementations of devices, equipment, modules and systems (e.g., reduced silicon area) can be weighed against performance improvements. For example, this can be accomplished by manufacturing multiple instances of a module within a predefined area budget.
[0112] Applicants hereby independently disclose each individual feature described herein and any combination of two or more such features, to the extent that such feature or combination can be implemented according to the common general knowledge of a person skilled in the art based on the entirety of this specification, regardless of whether such feature or combination of features solves any problem disclosed herein. In view of the foregoing description, it will be apparent to those skilled in the art that various modifications may be made within the scope of the invention.
Claims
1. A method for tessellizing tiles to generate tessellated geometric data representing the tessellated tiles, the method comprising: processing received geometry data representing a tile to identify a tessellation factor of the tile; determining tessellation instances to be used for tessellizing the tile based on the identified tessellation factor for the tile, wherein each of the tessellation instances determined for the tile is associated with a portion of tessellated geometry that will result when tessellizing the tile, such that the tessellated geometry associated with all of the tessellation instances for the tile collectively defines the tessellated geometry data for the tile; and distributing the tessellation instances among a plurality of tessellation pipelines operating in parallel, wherein a respective set of one or more of the tessellation instances is allocated to each of the tessellation pipelines, and wherein each of the tessellation pipelines produces the tessellated geometry data associated with the respective allocated set of one or more of the tessellation instances; The determining of the tessellation instances to be used for tessellation of the tile comprises determining the number of tessellation instances to be used for tessellation of the tile by: Determine the number of vertices (N) to be generated for the tile during tessellation based on the determined tessellation factor for the tile verts );as well as The number of vertices (N verts ) divided by a predetermined number (L).
2. The method of claim 1, wherein the predetermined number (L) represents the number of vertices to be batched at each tessellation pipeline.
3. The method according to claim 1 or 2, wherein the predetermined number (L) is the maximum number of vertices that can be assigned to a single pipeline so that processing does not stall due to lack of storage space.
4. The method of claim 1 or 2, wherein the predetermined number (L) is determined based on an amount of memory storage available for storing vertex data generated by each of the tessellation pipelines.
5. The method of claim 1 or 2, wherein said processing received geometry data representing a tile to identify a tessellation factor for the tile comprises determining the tessellation factor for the tile.
6. The method according to claim 1 or 2, further comprising: replicating a set of tiles and delivering the set of tiles to each of a plurality of cores, each core comprising a plurality of tessellation pipelines; wherein the method comprises at each of the cores: processing respective tiles of the set at respective tessellation pipelines to identify tessellation factors for the tiles of the set, wherein the tessellation instances to be used for tessellation of the tiles are determined based on the identified tessellation factors; determining the distribution of the tessellation instances among the tessellation pipelines of the core; and processing the tessellation instances at the allocated tessellation pipelines to generate tessellated geometry data associated with the respective allocated tessellation instances, Wherein a subset of the tessellation instances of a tile are assigned to the tessellation pipelines of cores, wherein collectively on all the cores, all the tessellation instances of the tile are processed.
7. A tessellation module configured to tessellate tiles to generate tessellated geometric shape data representing the tessellated tiles, the tessellation module comprising: tessellation factor logic configured to process received geometry data representing a tile to identify a tessellation factor for the tile; a plurality of tessellation pipelines arranged to operate in parallel; and A controller configured to: determining tessellation instances to be used for tessellizing the tile based on the identified tessellation factor for the tile, wherein each of the tessellation instances determined for the tile is associated with a portion of tessellated geometry that will result when tessellizing the tile, such that the tessellated geometry associated with all of the tessellation instances for the tile collectively defines the tessellated geometry data for the tile; and distributing the tessellation instances among the plurality of tessellation pipelines, whereby a respective set of one or more of the tessellation instances is distributed to each of the tessellation pipelines, wherein each of the tessellation pipelines is configured to generate the tessellated geometry data associated with the distributed set of one or more of the tessellation instances; and Wherein, in order to determine the tiling instance to be used for tiling the tile, the controller is configured to: Determine the number of vertices (N) to be generated for the tile during tessellation based on the determined tessellation factor for the tile verts );as well as The number of vertices (N verts ) divided by a predetermined number (L). 8 . The tessellation module of claim 7 , wherein the controller comprises a tessellation instance dispatcher configured to distribute the tessellation instances among the plurality of tessellation pipelines.
9. A tessellation module according to claim 7 or 8, wherein each of the tessellation instances is associated with a different portion of the tessellated geometric shape to be produced when tessellizing the tile.
10. The inlay module according to claim 7 or 8, wherein the controller is configured to: determining the tessellation instance by determining a first tessellation instance associated with a first portion of the tessellated geometric data and a second tessellation instance associated with a second, different portion of the tessellated geometric data; and The tessellation instances are distributed among the plurality of tessellation pipelines by distributing the first tessellation instance to a first tessellation pipeline and distributing the second tessellation instance to a second, different tessellation pipeline.
11. The inlay module according to claim 7 or 8, wherein: The tessellation factor logic is configured to receive geometry data representing a subsequent tile and process the received geometry data representing the subsequent tile to identify a subsequent tessellation factor for the subsequent tile; as well as The controller is configured to determine a subsequent tessellation instance to be used for tessellation of the subsequent tile based on the identified subsequent tessellation factor of the subsequent tile, and to distribute each of the subsequent tessellation instances among the tessellation pipeline after all tessellation instances of previously received tiles have been distributed to the tessellation pipeline.
12. The tessellation module of claim 11, wherein each of the tessellation pipelines is associated with a memory, and wherein a size of the predetermined number depends on a size of the memory.
13. The tessellation module of claim 7 or 8, wherein the controller is configured to provide edge data to each of the tessellation pipelines, the edge data comprising data resulting from identifying the tessellation factors, and wherein the tessellation pipelines are configured to process the tessellation instances using the edge data.
14. The tessellation module of claim 7 or 8, wherein the controller is configured to provide the identified tessellation factor of the tile and the received geometry data representing the tile to each of the tessellation pipelines, wherein the controller further comprises a buffer configured to store the identified tessellation factor; and wherein the controller is configured to retrieve the tessellation factor from the buffer so as to provide the identified tessellation factor to each of the tessellation pipelines.
15. The tessellation module according to claim 7 or 8, wherein the controller is further configured to control the issuing of tessellated geometry data such that tessellated geometry data for a tessellation instance is not issued until tessellated geometry data has been issued for all previous tessellation instances.
16. A tessellation module configured to tessellate tiles to generate tessellated geometric shape data representing the tessellated tiles, the tessellation module comprising: a plurality of cores, each core comprising a plurality of tessellated pipelines and controllers arranged to operate in parallel; as well as a tile dispatcher configured to replicate a set of tiles and deliver the set of tiles to each of the plurality of cores; Each of the cores is configured to: processing a respective tile of the set at a respective tessellation pipeline to identify a tessellation factor for the tile of the set; determining, at the controller of the core, tessellation instances to be used for tessellizing the tiles of the set based on the identified tessellation factors of the tiles of the set, wherein each of the tessellation instances determined for a tile is associated with a portion of tessellated geometry that will be produced when tessellizing the tile, such that the tessellated geometry associated with all of the tessellation instances of the tile collectively defines the tessellated geometry data for the tile; determining, at the controller of the core, a distribution of the tessellation instances among the tessellation pipeline of the core; and processing the tessellation instances at the allocated tessellation pipelines to generate tessellated geometry data associated with the respective allocated tessellation instances, Wherein the controllers of the plurality of cores are configured to cause subsets of the tessellation instances of a tile to be assigned to the tessellation pipelines of a core and to cause all of the tessellation instances of the tile to be processed collectively on all of the cores.
17. The tessellation module of claim 16, wherein each core comprises a number of tessellation pipelines equal to the number of tiles in the set.
18. A tessellation module according to claim 16 or 17, wherein each core includes a memory, and wherein the controller of a particular core is configured to determine the allocation of the tessellation instances among the tessellation pipelines of the particular core based on: (i) the number of cores in the tessellation module and / or the number of tessellation pipelines in the cores in the tessellation module, (ii) the functional location of the particular core within the plurality of cores of the tessellation module, and (iii) the available output storage of the memory in the core.
19. A tessellation module according to claim 16 or 17, wherein the controller of the core is configured to control the emission of tessellated geometry data such that tessellated geometry data for a tessellation instance is not emitted until tessellated geometry data has been emitted for all previous tessellation instances.
20. A computer-readable storage medium having computer-readable code stored thereon, the computer-readable code being configured to cause execution of the method according to claim 1 or 2 when the code is executed.
Citation Information
Patent Citations
Accelerated Compute Tessellation by Compact Topological Data Structure
US20130169636A1
Load Balancing for Optimal Tessellation Performance
US20140152675A1
Method and apparatus for performing high throughput tessellation
US20170193697A1
Load-balanced tessellation distribution for parallel architectures
US20180075650A1