Graphics processor, index processing method, chip, server and electronic equipment

By decoupling resource dependencies in the graphics processor and using an index processing unit with storage unit capacity and triangle number thresholds as constraints, the index processing process is optimized, solving the problems of duplicate shading and pipeline inefficiency in triangle index processing and achieving more efficient graphics processing.

CN120598765AActive Publication Date: 2025-09-05MOORE THREADS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511093574.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-05
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

In the prior art, graphics processors have repeated shading operations when processing triangle indices, and pipeline performance is limited by the allocation of vertex data cache resources, resulting in low efficiency.

Method used

By decoupling the resource dependencies during batch splitting, the index processing unit is used to divide the batches based on the available capacity of the storage unit and/or the triangle number threshold. This optimizes the index processing process and allows the index processing of the next batch to start before the shading process of the previous batch is completed.

Benefits of technology

It improves the pipeline operation efficiency of the graphics processor, reduces the dependence on vertex data cache resources, and achieves more efficient index processing and shading operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598765A_ABST
    Figure CN120598765A_ABST
Patent Text Reader

Abstract

The invention discloses a graphics processor, an index processing method, a chip, a server and electronic equipment, and belongs to the field of processors. The graphics processor comprises an index processing unit; the index processing unit is used for executing index processing operation on each original index according to an arrangement sequence in the triangular index list; the index processing operation comprises the steps of distributing an update index for the original index under the condition that the update index corresponding to the original index is not stored in the storage unit, storing a mapping relation between the original index and the update index in the storage unit, and storing the mapping relation between the original index and the update index in the storage unit under the condition that the update index corresponding to the original index is stored in the storage unit. Querying from the storage unit to obtain an update index; and under the condition that the number of the mapping relationships currently stored in the storage unit reaches the available capacity of the storage unit, outputting the distributed or queried update indexes as update indexes of the same batch. According to the method, the resource dependency relationship during batch splitting is decoupled, and the operation efficiency of the assembly line is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of processors, and in particular to a graphics processor, an index processing method, a chip, a server, and an electronic device. Background Art

[0002] The Graphics Processing Unit (GPU) includes a geometry processing pipeline. The front-end hardware of the geometry processing pipeline will process the input vertex index stream. In the case where there are shared vertices among the primitives, the front-end hardware will copy the indices of the shared vertices and generate a triangle index list. The triangle index list includes the vertex index of each vertex of each triangle in multiple triangles. The front-end hardware then inputs the triangle index list into the geometry processing pipeline, and the shader of the geometry processing pipeline colors each vertex of each triangle included in the triangle index list. At this time, some vertices are repeatedly colored.

[0003] In the related technology, in order to optimize the above-mentioned repeated shading operations, after the triangle index list is expanded, the front-end hardware will also generate a unique vertex list and a mapping list. The unique vertex list stores the mapping relationship between the updated index of each vertex and the original index, and the mapping list stores a new triangle index list. The new triangle index list contains the updated index of each vertex of each triangle. Therefore, the unique vertex list and the mapping list can be used to restore the original index of each triangle vertex in the original triangle index list, and the shader only needs to color the vertices corresponding to the vertex index in the unique vertex list.

[0004] In the related technology, there is also a batch splitting operation. The expanded triangle index list will be split into different batches (batches) to execute the above optimization process in batches. The number of vertices corresponding to each batch depends on the size of the allocated downstream vertex data cache resources (used to store vertex data). In the related technology, only after the vertex index of one batch is processed and colored, the resources of the vertex data buffer area will be released, and then the vertex data cache resources corresponding to the next batch can be allocated. That is, the related technology sacrifices pipeline performance in exchange for the removal of duplicate vertex indexes as much as possible. Summary of the Invention

[0005] The present application provides a graphics processor, an index processing method, a chip, a server, and an electronic device. The present application decouples resource dependencies during batch splitting and improves the operating efficiency of the pipeline.

[0006] According to one aspect of the present application, a graphics processor is provided, the graphics processor including an index processing unit; An index processing unit, configured to obtain a triangle index list corresponding to a plurality of triangles, the triangle index list including an original index corresponding to each vertex of each triangle in the plurality of triangles; an index processing unit configured to perform an index processing operation on each original index according to the order of arrangement in the triangle index list; the index processing operation including assigning an update index to the original index if the storage unit does not store the update index corresponding to the original index, storing a mapping relationship between the original index and the update index in the storage unit, and obtaining the update index from the storage unit based on the original index if the storage unit already stores the update index corresponding to the original index; The index processing unit is used to output the updated indexes that have been allocated or queried as updated indexes of the same batch when the number of mapping relationships currently stored in the storage unit reaches the available capacity of the storage unit.

[0007] According to one aspect of the present application, a graphics processor is provided, the graphics processor including an index processing unit; An index processing unit, configured to obtain a triangle index list corresponding to a plurality of triangles, the triangle index list including an original index corresponding to each vertex of each triangle in the plurality of triangles; an index processing unit configured to perform an index processing operation on each original index according to the order of arrangement in the triangle index list; the index processing operation including assigning an update index to the original index if the storage unit does not store the update index corresponding to the original index, storing a mapping relationship between the original index and the update index in the storage unit, and obtaining the update index from the storage unit based on the original index if the storage unit already stores the update index corresponding to the original index; The index processing unit is configured to output the allocated or queried update indexes as update indexes of the same batch when the number of triangles corresponding to the mapping relationship currently stored in the storage unit reaches a triangle number threshold.

[0008] According to one aspect of the present application, an index processing method is provided, the method comprising: Get a triangle index list, where the triangle index list corresponds to multiple triangles and includes an original index corresponding to each vertex of each triangle in the multiple triangles; Performing an index processing operation on each original index in the order of arrangement in the triangle index list; the index processing operation includes assigning an update index to the original index if the storage unit does not store the update index corresponding to the original index, storing a mapping relationship between the original index and the update index in the storage unit, and querying the storage unit for the update index according to the original index if the storage unit already stores the update index corresponding to the original index; When the number of mapping relationships currently stored in the storage unit reaches the available capacity of the storage unit, the update indexes that have been allocated or queried are output as update indexes of the same batch.

[0009] According to one aspect of the present application, an index processing method is provided, which includes the following steps.

[0010] Get a triangle index list, where the triangle index list corresponds to multiple triangles and includes an original index corresponding to each vertex of each triangle in the multiple triangles; Performing an index processing operation on each original index in the order of arrangement in the triangle index list; the index processing operation includes assigning an update index to the original index if the storage unit does not store the update index corresponding to the original index, storing a mapping relationship between the original index and the update index in the storage unit, and querying the storage unit for the update index according to the original index if the storage unit already stores the update index corresponding to the original index; When the number of triangles corresponding to the mapping relationship currently stored in the storage unit reaches a triangle number threshold, the update indexes that have been allocated or queried are output as update indexes of the same batch.

[0011] According to one aspect of the present application, a chip is provided, which includes the above-mentioned graphics processor.

[0012] According to one aspect of the present application, a server is provided, and the server includes the above-mentioned graphics processor.

[0013] According to one aspect of the present application, an electronic device is provided, and the electronic device includes the above-mentioned graphics processor.

[0014] The beneficial effects brought about by the technical solutions provided in the embodiments of the present application include at least the following.

[0015] In an embodiment of the present application, when the index processing unit splits the original index in the triangle index list and maps out the updated index, it will be limited by the available capacity of the storage unit and / or the triangle number threshold, rather than being limited by the size of the allocated downstream vertex data cache resources in the related art. That is, the embodiment of the present application decouples the resource dependency between the index processing process and the shading process (in the shading process, vertex data will be taken out from the vertex data cache resources).

[0016] Compared with related technologies, after splitting a batch of update indexes, this application can continue to split the update indexes of the next batch without waiting for the coloring process corresponding to the previous batch to be completed. Therefore, the index processing unit provided by this application enables the pipeline to run more efficiently, thereby improving the operating efficiency of the hardware.

[0017] Moreover, in the related art, batches are divided based on the size of the allocated vertex data cache resources, and the batches divided in the related art are often larger, while the embodiment of the present application divides batches based on the available capacity of the storage unit and / or the triangle number threshold. The batch divided in the present application is smaller, so the vertex data cache resources corresponding to a batch in the present application can be quickly used, released and put into use in the next batch. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 It is a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application.

[0020] Figure 2 This is a schematic diagram of the principle of an index processing method provided by an embodiment of the present application.

[0021] Figure 3 This is a schematic diagram of the principle of an index processing method provided in another embodiment of the present application.

[0022] Figure 4 This is a schematic diagram of an application scenario corresponding to the index processing unit provided in one embodiment of the present application.

[0023] Figure 5 This is a schematic diagram of an application scenario corresponding to an index processing unit provided in another embodiment of the present application.

[0024] Figure 6This is a flowchart of an index processing method provided by an embodiment of the present application.

[0025] Figure 7 This is a flowchart of an index processing method provided by another embodiment of the present application.

[0026] Figure 8 This is a structural block diagram of an electronic device provided by an embodiment of the present application.

[0027] Figure 9 This is a schematic diagram of the structure of a server provided in one embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0029] First, some terms in the embodiments of this application are introduced: Vertex index: In the graphics processor, the vertex index and vertex data of the triangle will enter the geometry processing pipeline, and then the vertex shading, tessellation shading, geometry shading and other steps will be executed in the geometry processing pipeline. The vertex index refers to the index of the primitive vertex, and the vertex index is used to index the vertex data. Optionally, primitives include points, lines, and triangles. Optionally, the vertex index is usually stored in the vertex index buffer, or generated by the hardware during the input assembly stage. Vertex data includes vertex coordinates, vertex attributes and other data. Optionally, vertex data is usually stored in a vertex data buffer area independent of the control instructions. According to the vertex index, the hardware can obtain the corresponding vertex data from the vertex data buffer for shading processing.

[0030] In the related art, the vertex index needs to be processed before entering the geometry processing pipeline. In the related art, the vertex index is processed by the front-end hardware of the geometry processing pipeline. In the primitives described by the vertex index stream input by the front-end hardware, there are multiple primitives that share vertices. For all input primitive types, the front-end hardware will usually first convert them into a triangle type. Specifically, for point and line primitive types, they will be converted into triangles with one or two identical vertex indices. For input primitives with shared vertices, the front-end hardware will first copy the vertex indices of the shared vertices to generate independent triangle vertex indices, that is, a triangle index list. If not optimized, these copied vertices will each trigger a shading calculation. For example, the input vertex index stream is: (0, 10, 2, 32, 24, 65, 26), then the converted triangle index list is: Triangle 1: (0, 10, 2); Triangle 2: (10, 2, 32); Triangle 3: (2, 32, 24); Triangle 4: (32, 24, 65); Triangle 5: (24, 65, 26); As can be seen, if no additional optimization is performed, the input is equivalent to 7 vertex indices, and the front-end hardware will expand it to 15 vertex indices. The vertex data corresponding to the 15 vertex indices will be obtained and shaded. Among them, there are multiple vertices that are shaded repeatedly.

[0031] The related technology also provides an optimized processing process, pre-processing the vertex index. Specifically, after the vertex index stream enters the front-end hardware and is expanded into a triangle index list, the front-end hardware will process the vertex index again, converting the triangle index list into a unique vertex list (Unique Vertices List) and a mapping list (Mapping of The List). The unique vertex list stores the mapping relationship between the deduplicated vertex index (called the original index) and the new vertex index (called the updated index), while the mapping list uses the new vertex index (updated index) in the unique vertex list to map out a new triangle index list. The following table further processes the above example.

[0032] List of unique vertices:

[0033] Mapping list: Triangle 1: (0, 1, 2); Triangle 2: (1, 2, 3); Triangle 3: (2, 3, 4); Triangle 4: (3, 4, 5); Triangle 5: (4, 5, 6); By using the updated index in the mapping list, the corresponding original index in the unique vertex list can be found, and the triangle index list composed of the original index can be restored. When shading, only the vertices corresponding to the vertex index in the unique vertex list need to be shaded. The front-end hardware will input the new triangle index list in the mapping list into the geometry processing pipeline for shading.

[0034] In the related art, a batch splitting operation is also performed. The expanded triangle index list is split into different batches (batches) to perform the above optimization process in batches. For example, the original indices of triangles 1, 2, and 3 are used as the same batch of vertex indices to generate a vertex index list and a mapping list. The original indices of triangles 1, 2, and 3 are mapped to obtain the same batch of updated indices. The original indices of triangles 4 and 5 are used as another batch of vertex indices to generate a vertex index list and a mapping list, which are then mapped to obtain another batch of updated indices. In the related art, the number of vertices corresponding to each batch depends on the size of the downstream vertex data cache resources allocated to the current batch. In the related art, only after the vertex indexes of a batch are processed and shaded, the vertex data cache resources occupied by the current batch in the vertex data buffer area are released, and the vertex data cache resources corresponding to the next batch can then be allocated. The next batch can then be split according to the allocated vertex data cache resources. In other words, the related art sacrifices pipeline performance in exchange for the removal of duplicate vertex indices as much as possible.

[0035] like Figure 1 As shown, Figure 1 FIG2 is a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application. The graphics processor includes an index processing unit 101 , front-end hardware 102 of a geometry processing pipeline, and a geometry processing pipeline 103 .

[0036] The front-end hardware 102 receives the input vertex index stream, performs primitive assembly on the vertex index stream, and obtains a triangle index list. Optionally, the vertex index stream includes multiple vertex indices, and the vertex index is used to index vertex data. Optionally, the triangle index list corresponds to multiple triangles, and the triangle index list includes the original index corresponding to each vertex of each triangle in the multiple triangles. The front-end hardware 102 sends the triangle index list to the index processing unit 101, and the index processing unit 101 processes the triangle index list and sends the new triangle index list obtained by processing to the geometry processing pipeline 103.

[0037] Illustratively, the vertex index stream received by the front-end hardware 102 includes (12, 34, 25, 76, 89, 123, 245, 346, 9, 10, 21), a total of 11 vertex indices. The front-end hardware 102 expands the vertex index stream based on the triangles described in the vertex index stream to obtain a triangle index list, which includes: Triangle 1 (12, 34, 25); Triangle 2 (34, 25, 76); Triangle 3 (24, 76, 89); Triangle 4 (76, 89, 123); Triangle 5 (89, 123, 245); Triangle 6 (123, 245, 346); Triangle 7 (245, 346, 9); Triangle 8 (346, 9, 10); Triangle 9 (9, 10, 21); It can be seen that 11 vertex indices are expanded to 27 original indices. The index processing unit 101 provided in this application processes the triangle index list.

[0038] Batch division is limited by the available capacity of the storage unit Figure 2 A schematic diagram illustrating the principles of an index processing method provided by an exemplary embodiment of the present application is shown. A graphics processor includes an index processing unit 101 and a storage unit 104. Index processing unit 101 is used to control storage unit 104, storing, querying, and updating indexes therein. Storage unit 104 is used to generate and store the unique vertex list described above. Index processing unit 101 and storage unit 104 are logically located before the shading step performed by the geometry processing pipeline.

[0039] The index processing unit 101 obtains a triangle index list. The triangle index list refers to an index list obtained by expanding the vertex index stream according to the triangles described in the vertex index stream.

[0040] The index processing unit 101 performs an index processing operation on each original index according to the order of arrangement in the triangle index list. The index processing operation includes: if the storage unit 104 does not store the update index corresponding to the original index, assigning the update index to the original index; storing the mapping relationship between the original index and the update index in the storage unit 104; and if the storage unit 104 already stores the update index corresponding to the original index, querying the storage unit 104 based on the original index to obtain the update index.

[0041] The updated index refers to the vertex index reallocated by the index processing unit 101 .

[0042] The order in the triangle index list refers to the order in which the original indices in the triangle index list were arranged.

[0043] Schematically, according to the triangle index list above, when the index processing operation is performed for the first time, the first original index 12 is taken out. At this time, the storage unit 104 does not record the updated index corresponding to the original index 12, then the index processing unit 101 assigns the updated index 0 to the original index 12, and stores the mapping relationship between the original index 12 and the updated index 0 in the storage unit 104.

[0044] Then the second original index 34 is taken out. At this time, the storage unit 104 does not record the updated index corresponding to the original index 34. Similarly, the index processing unit 101 assigns the updated index 1 to the original index 34 and stores the mapping relationship between the original index 34 and the updated index 1 in the storage unit 104.

[0045] Then the third original index 25 is taken out. At this time, the storage unit 104 does not record the updated index corresponding to the original index 25. Similarly, the index processing unit 101 assigns the updated index 2 to the original index 25 and stores the mapping relationship between the original index 25 and the updated index 2 in the storage unit 104.

[0046] Next, the fourth original index 34 is retrieved. At this time, the storage unit 104 has stored the updated index 1 corresponding to the original index 34 . The index processing unit 101 then searches the storage unit 104 to obtain the updated index 1 .

[0047] The same process is repeated until the number of mapping relationships stored in the storage unit 104 reaches the available capacity of the storage unit 104 , and the update indexes that have been allocated or obtained through query are output as update indexes of the same batch.

[0048] In one embodiment, the available capacity of the storage unit 104 is pre-set. Alternatively, the available capacity of the storage unit 104 is set by a software program rather than being constrained by the hardware capacity of the storage unit 104. Alternatively, the available capacity is set by a graphics processor driver.

[0049] The optional usable capacity value is related to the number of threads that can be launched in parallel for a single task. GPUs can launch multiple threads in parallel, and the usable capacity value is used to ensure that the same batch of updated indexes can be processed in parallel by multiple threads without exceeding the load capacity of the multiple threads.

[0050] The value of the usable capacity is optionally related to the size of the downstream vertex data buffer, which is used to store vertex data. When shading is performed, vertex data needs to be retrieved from the vertex data buffer based on the vertex index. The value of the usable capacity is used to ensure that the vertex data corresponding to the same batch of update indices does not exceed the hardware capacity of the vertex data buffer.

[0051] like Figure 2 As shown, the available capacity of the storage unit 104 is 7. When the storage unit 104 stores 7 pairs of mapping relationships, the updated indexes of the first batch output are as follows: Triangle 1 (0, 1, 2); Triangle 2 (1, 2, 3); Triangle 3 (2, 3, 4); Triangle 4 (3, 4, 5); Triangle 5 (4, 5, 6); Afterwards, the index processing unit 101 resets the storage unit 104 and performs mapping processing on the next batch of original indexes. Similar to the above processing process, the second batch of updated indexes obtained are as follows: Triangle 6 (0, 1, 2); Triangle 7 (1, 2, 3); Triangle 8 (2, 3, 4); Triangle 9 (3, 4, 5).

[0052] In one embodiment, after forming the update indexes of the same batch, the index processing unit 101 resets the storage unit to prepare for forming the update indexes of the next batch.

[0053] It can be understood that the embodiment of the present application will divide the triangle index list into batches according to the available capacity of the storage unit. The batch division in the related art depends on the size of the allocated downstream vertex data cache resources, but the present application does not perform batch division in this way, that is, the embodiment of the present application decouples the upstream and downstream resource dependencies and separates the index processing process. In the related art, the index processing of the same batch must be completed and the corresponding shading process must be executed before the index processing of the next batch can be carried out. In the present application, after the first batch of update indexes are generated, the shading process corresponding to the first batch will be executed, and the second batch of update indexes can be generated at this time without waiting for the shading process corresponding to the first batch to be executed. That is, the embodiment of the present application improves the operating efficiency of the pipeline.

[0054] Batch division is performed based on the triangle number threshold exist Figure 2 In the scheme shown, the batch division is performed based on the available capacity of the storage unit. Figure 3 In the example, the index processing unit 101 divides the batches into groups based on the triangle number threshold.

[0055] With the above Figure 2 Similar to the related introduction, Figure 3 In the example, the index processing unit 101 performs index processing on each original index according to the order of arrangement in the triangle index list. When the number of triangles corresponding to the mapping relationship currently stored in the storage unit 104 reaches the triangle number threshold, the index processing unit 101 outputs the updated index that has been allocated or queried as the updated index of the same batch.

[0056] In one embodiment, the triangle number threshold is pre-set. Alternatively, the triangle number threshold is set by a graphics processor driver.

[0057] Optionally, in the case where the graphics processor is used to perform tessellation shading, the value of the triangle number threshold is associated with the number of threads that can be started in parallel by the shell shader located downstream. When the graphics processor performs tessellation shading, the tessellation shading will expand or add more triangles. Tessellation shading is performed by the tessellation shader, which includes a shell shader, a tessellation stage, and a domain shader. Therefore, by determining the triangle number threshold through the number of threads that can be started in parallel by the shell shader, it is possible to avoid setting a larger triangle number threshold for the upstream index processing unit, avoid updating indices containing more triangles in a batch divided upstream, and ensure that the threads started in parallel by the downstream shell shader can load the related tasks of the expanded or increased triangles.

[0058] Optionally, when a graphics processor is used to perform geometry shading, the value of the triangle count threshold is associated with the size of the storage area used by the output of the downstream geometry shader. Geometry shading is performed by the geometry shader, and the triangle count threshold is used to ensure that the output of the geometry shader does not exceed the size of the storage area, ensuring that the storage area of ​​the downstream geometry shader output is large enough.

[0059] Indicative, such as Figure 3 As shown, Figure 3 The triangle number threshold is 6, and the update index of the first batch obtained by partitioning is as follows: Triangle 1 (0, 1, 2); Triangle 2 (1, 2, 3); Triangle 3 (2, 3, 4); Triangle 4 (3, 4, 5); Triangle 5 (4, 5, 6); Triangle 6 (5, 6, 7); It can be seen that when the number of triangles corresponding to the mapping relationship stored in the storage unit 104 reaches 6, the update indexes of the same batch will be split out.

[0060] Afterwards, the index processing unit 101 resets the storage unit 104 and performs mapping of the next batch of original indexes. Similar to the above process, the second batch of updated indexes obtained are as follows: Triangle 7 (0, 1, 2); Triangle 8 (1, 2, 3); Triangle 9 (2, 3, 4); In one embodiment, after forming the update indexes of the same batch, the index processing unit 101 resets the storage unit to prepare for forming the update indexes of the next batch.

[0061] It is understandable that the embodiment of the present application will divide the index batches of the triangle index list according to the triangle number threshold. The batch division in the related art is based on the size of the allocated downstream vertex data cache resources, but the present application does not divide the batches in this way, that is, the embodiment of the present application decouples the upstream and downstream resource dependencies and separates the index processing process. In the related art, the index processing of the same batch must be completed and the corresponding shading process must be executed before the index processing of the next batch can be carried out. In the present application, after the update index of the first batch is generated, the shading process corresponding to the first batch will be executed, and the update index of the second batch can be generated at this time without waiting for the shading process corresponding to the first batch to be executed. That is, the embodiment of the present application improves the operating efficiency of the pipeline.

[0062] Moreover, splitting by the triangle number threshold can ensure that the update indexes of all vertices of the same primitive are in the same batch, ensuring the orderliness, accuracy and reliability of the subsequent shading process.

[0063] It should be noted that the above descriptions use the available storage capacity and triangle count threshold as constraints for batch division. In an optional embodiment, both the available storage capacity and the triangle count threshold are used as constraints. When either constraint is met, the allocated or queried update index is output as the update index for the same batch.

[0064] For example, continuing with the example above, assuming the available capacity is 7 and the triangle number threshold is 6, the updated indexes of the first batch obtained by partitioning are as follows: Triangle 1 (0, 1, 2); Triangle 2 (1, 2, 3); Triangle 3 (2, 3, 4); Triangle 4 (3, 4, 5); Triangle 5 (4, 5, 6); The updated index for the second batch is as follows: Triangle 6 (0, 1, 2); Triangle 7 (1, 2, 3); Triangle 8 (2, 3, 4); Triangle 9 (3, 4, 5); At this time, the constraint that "the number of mappings currently stored in the storage unit reaches the available capacity of the storage unit" is prioritized. This means that if the number of mappings currently stored in the storage unit reaches the available capacity and the number of triangles corresponding to the mappings currently stored in the storage unit does not reach the triangle number threshold, the index processing unit will output the allocated or queried update indexes as update indexes for the same batch.

[0065] For example, assuming the available capacity is 9 and the triangle number threshold is 6, the update index of the first batch obtained by partitioning is as follows: Triangle 1 (0, 1, 2); Triangle 2 (1, 2, 3); Triangle 3 (2, 3, 4); Triangle 4 (3, 4, 5); Triangle 5 (4, 5, 6); Triangle 6 (5, 6, 7); The updated index for the second batch is as follows: Triangle 7 (0, 1, 2); Triangle 8 (1, 2, 3); Triangle 9 (2, 3, 4); At this time, the constraint that "the number of triangles corresponding to the mappings currently stored in the storage unit reaches the triangle number threshold" is prioritized. This ensures that, if the number of triangles corresponding to the mappings currently stored in the storage unit reaches the triangle number threshold and the number of mappings currently stored in the storage unit does not reach the available capacity of the storage unit, the index processing unit outputs the allocated or queried update index as the update index for the same batch.

[0066] Usage scenarios of the index processing unit 101 In one embodiment, the index processing unit 101 can be used to Figure 1 The implementation environment shown can also be used in other implementation environments. Figure 1 In the implementation environment shown, Figure 4 As shown, the index processing unit 101 is used to process the triangle index list between the primitive assembly stage 401 and the vertex shading stage 402. The primitive assembly stage 401 is used to assemble the vertex index stream into a triangle index list. In the vertex shading stage 402, the corresponding vertex data is obtained according to the updated index for vertex shading.

[0067] It is understandable that when the index processing unit 101 is applied between the primitive assembly stage 401 and the vertex shading stage 402, since the mapping relationship within the storage unit is restricted from being replaced (which can be derived from the update index being split into the same batch when the available capacity of the storage unit is reached), the update index, primitive identifier, tessellation information, and geometric primitive information within a batch can be indexed with each other as independent units. Therefore, after the vertex shading corresponding to a batch is completed, the subsequent tessellation shading (the processing operation performed by the tessellation shader is called tessellation shading) and geometric shading can be started immediately. At this time, the vertex shading, tessellation shading, and geometric shading of different batches can be performed alternately in batches. That is, the index processing unit provided in the embodiment of the present application supports decoupling the operation of vertex shading, tessellation shading, and geometric shading. For example, the batch that executes vertex shading first can execute tessellation shading later. It is only necessary to ensure that the vertices of the same batch are executed in the order of vertex shading, tessellation shading, and geometric shading.

[0068] In one embodiment, the index processing unit 101 can be used to Figure 1 other implementation environments besides .

[0069] like Figure 5 As shown, index processing unit 101 is used to process triangle index lists between tessellation stage 501 and domain shading stage 502. Tessellation stage 501 is used to perform tessellation based on tessellation parameters output by the shell shader, generating more vertices and primitives. Domain shading stage 502 is used to process the vertices generated by tessellation stage 501. Tessellation stage 501 and domain shading stage 502 are processing stages executed within the geometry processing pipeline.

[0070] It is understood that in the tessellation stage 501, multiple new primitives are generated based on the input primitives, and the newly generated primitives also have a large number of shared vertices. In the above embodiment, the index processing unit 101 is used between the tessellation stage 501 and the domain shading stage 502 to perform deduplication and batch operations on the vertex indices of the newly generated vertices in the tessellation stage 501.

[0071] Optionally, when the index processing unit 101 is used between the tessellation stage 501 and the domain shading stage 502, the specific form of the original index in the triangle index list depends on the output of the tessellation stage 501. If the output of the tessellation stage 501 is directly the index value of the vertex, the original index in the triangle index list is the index value of the vertex. If the output of the tessellation stage 501 is the primitive identifier (patch id) and the UV coordinate of the vertex, the primitive identifier (patch id) and the UV coordinate of the vertex are merged into the vertex index of the vertex, thereby forming a triangle index list, and the index processing unit 101 performs deduplication and batching operations.

[0072] Optionally, between the primitive assembly stage 401 and the vertex shading stage 402, and between the surface subdivision stage 501 and the domain shading stage 502, the two places can use the same index processing unit 101 in a time-sharing manner, or use their own index processing units 101, and this application does not impose any restrictions on this.

[0073] It is understandable that the related art does not take into account the existence of tessellation shaders and geometry shaders in the graphics pipeline. After the introduction of tessellation shaders and geometry shaders, since they both generate new vertices, and the number of vertices generated is determined in real time based on calculation results during the shading stage, in the related art, the front-end hardware of the geometry processor can only estimate resources based on the maximum value limit of generated vertices, allocate vertex data cache resources based on the maximum number of vertices, and then divide batches based on the maximum vertex data cache resources. In the above embodiment, an index processing unit is inserted between the tessellation stage and the domain shading stage, and batch division can be performed based on the actual number of vertices generated in the tessellation stage, resulting in more accurate batch division results.

[0074] The index processing unit 101 processes multiple original indexes at a time In one embodiment, the index processing unit 101 is further configured to perform an index processing operation on each of the at least two original indices according to the order of arrangement in the triangle index list. In this case, performing an index processing operation on multiple original indices at a time can improve the throughput of the index processing unit.

[0075] Optionally, the index processing unit 101 is further configured to perform index processing operations on each original index of at least one triangle at a time, in the order of arrangement in the triangle index list, using triangles as granularity. In this case, performing index processing operations on multiple original indexes at a time using triangles as granularity can ensure that the addition and subtraction of the mapping relationships stored in the storage unit 104 are always performed in units of triangles, thereby ensuring that all updated indices for the same triangle are output in the same batch.

[0076] However, in the related art, index processing operations can only be performed on one original index at a time, as shown in the following example.

[0077] Assume that the technology receives the following raw index during processing: Triangle 12: (12, 25, 67); Moreover, the storage unit has stored the mapping relationship between the original index 25 and the updated index 0, and stored it in the first position; while the storage unit has not stored the mapping relationship between the original index 12 and the original index 67.

[0078] Assuming that the current storage unit is full, the existing mapping relationships in the storage unit include:

[0079] If you query these three original indexes at the same time, the results are:

[0080] When the three original indices are replaced at the same time, the original index 12 will replace the original index 25, and the original index 67 will replace the original index 34, while the original index 25 is considered to be in the storage unit. Therefore, the final result is: Triangle 12: (0, 0, 1); When it is restored later, the corresponding vertex index incorrectly becomes: (12, 12, 67).

[0081] Therefore, related technologies can only perform search operations on a single original index in sequence.

[0082] This application uses the preset available capacity of storage units and / or triangle number thresholds to divide batches, prohibiting storage unit replacement. When handling the above situation, since the storage unit is full, the current batch will be divided first, the updated index of the current batch will be sent to downstream processing, and the contents of the storage unit will be reset. When processing triangle 12, the index processing operation will be performed using the reset storage unit. At this time, the search situation for triangle 12 will become:

[0083] Then the updated index corresponding to triangle 12 is: Triangle 12: (0, 1, 2); The corresponding list of unique vertices is:

[0084] That is, the present application can avoid the limitation in the related art that index processing operations can only be performed on one original index at a time. The present application can theoretically support any number of original indexes to perform index processing operations at the same time.

[0085] Optionally, the usable capacity of the storage unit is in units of triangles. When index processing operations are performed simultaneously on multiple original indices, and the multiple original indices can form at least one triangle, the usable capacity of the storage unit needs to be in units of the number of triangles formed by the original indices used in one search.

[0086] Figure 6 A flowchart of an index processing method provided by an exemplary embodiment of the present application is shown. The method is illustrated by taking the index processing unit 101 as an example. The method includes: Step 620, obtain a triangle index list; A triangle index list is an index list obtained by expanding a vertex index stream based on the triangles described in the vertex index stream. A triangle index list corresponds to multiple triangles and includes the original index corresponding to each vertex of each triangle in the multiple triangles. A vertex index stream includes multiple vertex indices, which are used to index vertex data.

[0087] Step 640 , performing index processing operations on each original index according to the order of arrangement in the triangle index list; The order in the triangle index list refers to the order in which the original indices in the triangle index list were arranged.

[0088] The index processing operation includes assigning an updated index to the original index when the storage unit does not store the updated index corresponding to the original index, and storing the mapping relationship between the original index and the updated index in the storage unit. When the storage unit already stores the updated index corresponding to the original index, the updated index is obtained by querying from the storage unit according to the original index.

[0089] Update index refers to the reallocated vertex index.

[0090] In one embodiment, the index processing operation is performed on each of the at least two original indices according to the order of arrangement in the triangle index list. In this case, the index processing operation is performed on multiple original indices at a time, which can improve the throughput of the index processing unit.

[0091] Optionally, the index processing operation is performed on each original index of at least one triangle at a time, in the order of the triangle index list. In this case, the index processing operation is performed on multiple original indices at a time, in a triangle-based granularity. This ensures that the addition and subtraction of the mapping relationship stored in the storage unit is always performed on a triangle-by-triangle basis, and thus ensures that all updated indices for the same triangle are output in the same batch.

[0092] Step 660: When the number of mapping relationships currently stored in the storage unit reaches the available capacity of the storage unit, the update indexes that have been allocated or queried are output as update indexes of the same batch.

[0093] The available capacity is optionally configured by the GPU driver; the value of the available capacity is also associated with the number of threads that can be launched concurrently for a single task. GPUs can launch multiple threads concurrently, and the available capacity is used to ensure that the same batch of updated indexes can be processed concurrently by multiple threads without exceeding the threads' capacity.

[0094] The value of the usable capacity is optionally related to the size of the downstream vertex data buffer, which is used to store vertex data. When shading is performed, vertex data needs to be retrieved from the vertex data buffer based on the vertex index. The value of the usable capacity is used to ensure that the vertex data corresponding to the same batch of update indices does not exceed the hardware capacity of the vertex data buffer.

[0095] In one embodiment, if the number of mappings currently stored in a storage unit reaches the available capacity of the storage unit, and the number of triangles corresponding to the mappings currently stored in the storage unit does not reach a triangle number threshold, the allocated or queried update index is output as the update index for the same batch. At this time, the triangle number threshold is also considered to determine whether the output limit for the same batch has been reached.

[0096] In one embodiment, the method further includes resetting the storage unit after forming the update indexes of the same batch to prepare for forming the update indexes of the next batch.

[0097] To sum up, the above embodiment divides the triangle index list into batches according to the available capacity of the storage unit. The batch division in the related art depends on the size of the allocated downstream vertex data cache resources, but the present application does not perform batch division in this way, that is, the embodiment of the present application decouples the upstream and downstream resource dependencies and separates the index processing process. In the related art, the index processing of the same batch must be completed and the corresponding shading process must be executed before the index processing of the next batch can be carried out. In the present application, after the first batch of update indexes is generated, the shading process corresponding to the first batch will be executed, and the second batch of update indexes can be generated at this time without waiting for the shading process corresponding to the first batch to be executed. That is, the embodiment of the present application improves the operating efficiency of the pipeline.

[0098] In one embodiment, Figure 6 The index processing method shown is applied to process the triangle index list between the primitive assembly stage and the vertex shading stage.

[0099] The primitive assembly stage is used to assemble the vertex index stream into a triangle index list. In the vertex shading stage, the corresponding vertex data will be obtained according to the updated index for vertex shading.

[0100] In one embodiment, Figure 6 The index processing method shown is applied to process the triangle index list between the tessellation stage and the domain shading stage.

[0101] The tessellation stage is used to perform tessellation according to the tessellation parameters output by the shell shader to generate more vertices and primitives. The domain shading stage is used to process the vertices generated by the tessellation stage.

[0102] It is understandable that during the tessellation stage, multiple new primitives are generated based on the input primitives, and these newly generated primitives also have a large number of shared vertices. In the above embodiment, an index processing unit is used between the tessellation stage and the domain shading stage to similarly perform deduplication and batch operations on the vertex indices of the newly generated vertices during the tessellation stage.

[0103] Figure 7 A flowchart of an index processing method provided by an exemplary embodiment of the present application is shown. The method is illustrated by taking the index processing unit 101 as an example. The method includes: Step 720, obtain triangle index list; A triangle index list is an index list obtained by expanding a vertex index stream based on the triangles described in the vertex index stream. A triangle index list corresponds to multiple triangles and includes the original index corresponding to each vertex of each triangle in the multiple triangles. A vertex index stream includes multiple vertex indices, which are used to index vertex data.

[0104] Step 740 , performing index processing operations on each original index according to the order of arrangement in the triangle index list; The order in the triangle index list refers to the order in which the original indices in the triangle index list were arranged.

[0105] The index processing operation includes assigning an updated index to the original index when the storage unit does not store the updated index corresponding to the original index, and storing the mapping relationship between the original index and the updated index in the storage unit. When the storage unit already stores the updated index corresponding to the original index, the updated index is obtained by querying from the storage unit according to the original index.

[0106] Update index refers to the reallocated vertex index.

[0107] In one embodiment, the index processing operation is performed on each of the at least two original indices according to the order of arrangement in the triangle index list. In this case, the index processing operation is performed on multiple original indices at a time, which can improve the throughput of the index processing unit.

[0108] Optionally, the index processing operation is performed on each original index of at least one triangle at a time, in the order of the triangle index list. In this case, the index processing operation is performed on multiple original indices at a time, in a triangle-based granularity. This ensures that the addition and subtraction of the mapping relationship stored in the storage unit is always performed on a triangle-by-triangle basis, and thus ensures that all updated indices for the same triangle are output in the same batch.

[0109] Step 760 : When the number of triangles corresponding to the mapping relationship currently stored in the storage unit reaches a triangle number threshold, the allocated or queried update index is output as an update index of the same batch.

[0110] In one embodiment, the triangle number threshold is set by a graphics processor driver; optionally, when the graphics processor is used to perform surface tessellation shading, the value of the triangle number threshold is associated with the number of threads that can be launched in parallel by a downstream shell shader.

[0111] When the GPU performs tessellation shading, the tessellation shading will expand or add more triangles. Tessellation shading is performed by the tessellation shader, which includes the shell shader, the tessellation stage, and the domain shader. Therefore, by determining the triangle count threshold based on the number of threads that can be launched in parallel by the shell shader, it is possible to avoid setting a large triangle count threshold for the upstream index processing unit, preventing the upstream batch from containing many triangle update indices. This ensures that the threads launched in parallel by the downstream shell shader can handle the tasks related to the expansion or increase of triangles.

[0112] Optionally, when a graphics processor is used to perform geometry shading, the value of the triangle count threshold is associated with the size of the storage area used by the output of the downstream geometry shader. Geometry shading is performed by the geometry shader, and the triangle count threshold is used to ensure that the output of the geometry shader does not exceed the size of the storage area, ensuring that the storage area of ​​the downstream geometry shader output is large enough.

[0113] In one embodiment, if the number of triangles corresponding to the mappings currently stored in the storage unit reaches a triangle threshold, and the number of mappings currently stored in the storage unit does not reach the available capacity of the storage unit, the allocated or queried update indexes are output as update indexes for the same batch. At this point, the available capacity of the storage unit is also considered to determine whether the output limit for the same batch has been reached.

[0114] In one embodiment, the method further includes resetting the storage unit after forming the update indexes of the same batch to prepare for forming the update indexes of the next batch.

[0115] In summary, the embodiment of the present application will divide the index batches of the triangle index list according to the triangle number threshold. The batch division in the related art is based on the size of the allocated downstream vertex data cache resources, but the present application does not divide the batches in this way, that is, the embodiment of the present application decouples the upstream and downstream resource dependencies and separates the index processing process. In the related art, the index processing of the same batch must be completed and the corresponding shading process must be executed before the index processing of the next batch can be carried out. In the present application, after the first batch of update indexes is generated, the shading process corresponding to the first batch will be executed, and the second batch of update indexes can be generated at this time without waiting for the shading process corresponding to the first batch to be executed. That is, the embodiment of the present application improves the operating efficiency of the pipeline.

[0116] Moreover, splitting by the triangle number threshold can ensure that the update indexes of all vertices of the same primitive are in the same batch, ensuring the orderliness, accuracy and reliability of the subsequent shading process.

[0117] In one embodiment, Figure 7 The index processing method shown is applied to process the triangle index list between the primitive assembly stage and the vertex shading stage.

[0118] The primitive assembly stage is used to assemble the vertex index stream into a triangle index list. In the vertex shading stage, the corresponding vertex data will be obtained according to the updated index for vertex shading.

[0119] In one embodiment, Figure 7 The index processing method shown is applied to process the triangle index list between the tessellation stage and the domain shading stage.

[0120] The tessellation stage is used to perform tessellation according to the tessellation parameters output by the shell shader to generate more vertices and primitives. The domain shading stage is used to process the vertices generated by the tessellation stage.

[0121] It is understandable that during the tessellation stage, multiple new primitives are generated based on the input primitives, and these newly generated primitives also have a large number of shared vertices. In the above embodiment, an index processing unit is used between the tessellation stage and the domain shading stage to similarly perform deduplication and batch operations on the vertex indices of the newly generated vertices during the tessellation stage.

[0122] Figure 8 FIG. 8 is a block diagram of an electronic device 800 according to an exemplary embodiment of the present application. Optionally, the electronic device 800 includes a graphics processor according to an embodiment of the present application.

[0123] Alternatively, the electronic device may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Electronic device 800 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names. Generally, electronic device 800 includes a processor 801 and a memory 802.

[0124] Processor 801 may include one or more processing cores, such as a quad-core processor or a penta-core processor. Processor 801 may be implemented in hardware using at least one of the following: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). Processor 801 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content displayed on the display screen. In some embodiments, processor 801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0125] The memory 802 may include one or more computer-readable storage media, which may be non-transitory, and may also include high-speed random access memory and non-volatile memory, such as one or more magnetic disk storage devices and flash memory storage devices.

[0126] In some embodiments, the electronic device 800 may further include: a peripheral device interface 803 and at least one peripheral device. Figure 8 The structure shown in the figure does not constitute a limitation on the electronic device 800, and the electronic device 800 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0127] The present application also provides a chip, which includes the graphics processor described in the above embodiment.

[0128] The present application also provides a graphics card, which includes a graphics processor as described in the above embodiment.

[0129] The present application also provides a server, which includes a graphics processor as described in the above embodiment.

[0130] Figure 9A schematic diagram of the structure of a server provided by an exemplary embodiment of the present application is shown. The server 900 includes multiple graphics processors 901. At least one of the multiple graphics processors 901 includes the graphics processor introduced in the above embodiment of the present application.

[0131] The present application also provides a computing cluster, which includes multiple servers, at least one of which includes a graphics processor as described in the above embodiment.

[0132] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0133] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0134] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A graphics processor, characterized in that: The graphics processor includes an index processing unit; The index processing unit is configured to obtain a triangle index list corresponding to a plurality of triangles, wherein the triangle index list includes an original index corresponding to each vertex of each triangle in the plurality of triangles; The index processing unit is configured to perform an index processing operation on each original index according to the arrangement order in the triangle index list; the index processing operation includes assigning the update index to the original index if the storage unit does not store the update index corresponding to the original index, and storing a mapping relationship between the original index and the update index in the storage unit; and obtaining the update index from the storage unit according to the original index if the storage unit already stores the update index corresponding to the original index. The index processing unit is configured to output the allocated or queried update indexes as update indexes of the same batch when the number of mapping relationships currently stored in the storage unit reaches the available capacity of the storage unit.

2. The graphics processor according to claim 1, wherein: The available capacity is set by a driver of the graphics processor; The value of the available capacity is associated with the number of threads that can be started in parallel for a single task.

3. The graphics processor according to claim 1, wherein: The available capacity is set by a driver of the graphics processor; The value of the available capacity is associated with the size of a vertex data buffer area located downstream, and the vertex data buffer area is used to store vertex data.

4. The graphics processor according to any one of claims 1 to 3, characterized in that The index processing unit is also used to output the allocated or queried update index as an update index of the same batch when the number of mapping relationships currently stored in the storage unit reaches the available capacity of the storage unit and the number of triangles corresponding to the mapping relationships currently stored in the storage unit does not reach a triangle number threshold.

5. The graphics processor according to any one of claims 1 to 3, wherein: The index processing unit is further configured to perform the index processing operation on each of the at least two original indexes at a time according to the arrangement order in the triangle index list.

6. The graphics processor according to claim 5, wherein: The index processing unit is further configured to perform the index processing operation on each original index of at least one triangle at a time according to the arrangement order in the triangle index list and with triangles as granularity.

7. The graphics processor according to any one of claims 1 to 3, characterized in that: The index processing unit is used to perform processing on the triangle index list between the primitive assembly stage and the vertex shading stage; or, The index processing unit is configured to process the triangle index list between the tessellation stage and the domain shading stage; or, The index processing unit is configured to process the triangle index list between the primitive assembly stage and the vertex shading stage, and to process the triangle index list between the tessellation stage and the domain shading stage.

8. The graphics processor according to any one of claims 1 to 3, characterized in that: The index processing unit is further configured to reset the storage unit after forming the update indexes of the same batch.

9. A graphics processor, characterized in that: The graphics processor includes an index processing unit; The index processing unit is configured to obtain a triangle index list corresponding to a plurality of triangles, wherein the triangle index list includes an original index corresponding to each vertex of each triangle in the plurality of triangles; The index processing unit is configured to perform an index processing operation on each original index according to the arrangement order in the triangle index list; the index processing operation includes assigning the update index to the original index if the storage unit does not store the update index corresponding to the original index, and storing a mapping relationship between the original index and the update index in the storage unit; and obtaining the update index from the storage unit according to the original index if the storage unit already stores the update index corresponding to the original index. The index processing unit is configured to output the allocated or queried update indexes as update indexes of the same batch when the number of triangles corresponding to the mapping relationship currently stored in the storage unit reaches a triangle number threshold.

10. The graphics processor according to claim 9, wherein: The triangle quantity threshold is set by a driver of the graphics processor; When the graphics processor is used to perform tessellation shading, the value of the triangle quantity threshold is associated with the number of threads that can be launched in parallel for a downstream shell shader.

11. The graphics processor according to claim 9, wherein: The triangle quantity threshold is set by a driver of the graphics processor; When the graphics processor is used to perform geometry shading, the value of the triangle quantity threshold is associated with the size of a storage area used by the output of a downstream geometry shader.

12. The graphics processor according to any one of claims 9 to 11, characterized in that: The index processing unit is also used to output the allocated or queried update index as an update index of the same batch when the number of triangles corresponding to the mapping relationship currently stored in the storage unit reaches the triangle number threshold and the number of mapping relationships currently stored in the storage unit does not reach the available capacity of the storage unit.

13. The graphics processor according to any one of claims 9 to 11, characterized in that: The index processing unit is further configured to perform the index processing operation on each of the at least two original indexes at a time according to the arrangement order in the triangle index list.

14. The graphics processor according to claim 13, wherein: The index processing unit is further configured to perform the index processing operation on each original index of at least one triangle at a time according to the arrangement order in the triangle index list and with triangles as granularity.

15. The graphics processor according to any one of claims 9 to 11, characterized in that: The index processing unit is used to perform processing on the triangle index list between the primitive assembly stage and the vertex shading stage; or, The index processing unit is configured to process the triangle index list between the tessellation stage and the domain shading stage; or, The index processing unit is configured to process the triangle index list between the primitive assembly stage and the vertex shading stage, and to process the triangle index list between the tessellation stage and the domain shading stage.

16. The graphics processor according to any one of claims 9 to 11, characterized in that: The index processing unit is further configured to reset the storage unit after forming the update indexes of the same batch.

17. An index processing method, characterized in that: The method comprises: Obtain a triangle index list, where the triangle index list corresponds to a plurality of triangles, and the triangle index list includes an original index corresponding to each vertex of each triangle in the plurality of triangles; performing an index processing operation on each of the original indexes according to the arrangement order in the triangle index list; the index processing operation includes assigning the update index to the original index if a storage unit does not store an update index corresponding to the original index, and storing a mapping relationship between the original index and the update index in the storage unit; and obtaining the update index from the storage unit according to the original index if the storage unit already stores the update index corresponding to the original index. When the number of mapping relationships currently stored in the storage unit reaches the available capacity of the storage unit, the update indexes that have been allocated or obtained through query are output as update indexes of the same batch.

18. An index processing method, characterized in that: The method comprises: Obtain a triangle index list, where the triangle index list corresponds to a plurality of triangles, and the triangle index list includes an original index corresponding to each vertex of each triangle in the plurality of triangles; performing an index processing operation on each of the original indexes according to the arrangement order in the triangle index list; the index processing operation includes assigning the update index to the original index if a storage unit does not store an update index corresponding to the original index, and storing a mapping relationship between the original index and the update index in the storage unit; and obtaining the update index from the storage unit according to the original index if the storage unit already stores the update index corresponding to the original index. When the number of triangles corresponding to the mapping relationship currently stored in the storage unit reaches a triangle number threshold, the update indexes that have been allocated or queried are output as update indexes of the same batch.

19. A chip, characterized in that: The chip includes the graphics processor according to any one of claims 1 to 16.

20. A server, characterized in that: The server includes the graphics processor according to any one of claims 1 to 16.

21. An electronic device, characterized in that: The electronic device comprises the graphics processor according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Multi-view image generation device and graphics processor

    CN117952816A

  • GPU primitive preprocessing device, system and method

    CN118887071A

  • Geometric processing system, graphics processor and computing device

    CN119850404A

  • Methods and apparatus for scalable primitive rate architecture for geometry processing

    US20220327654A1