Geometry processing method and apparatus, device, storage medium, and program product

By designing multiple graphics pipeline clusters in the graphics processor, multi-level splitting and parallel processing of geometric data flow is solved, and the problem of insufficient geometric processing performance in the existing technology is significantly improved.

WO2025103319A1PCT designated stage expired Publication Date: 2025-05-22MOORE THREADS TECH CO LTD

Patent Information

Application Number
PCT/CN2024/131611
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-11-12
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

The geometric processing pipeline design in the prior art is difficult to provide higher geometric processing performance, which has become a bottleneck in graphics processing.

Method used

Multi-level splitting and parallel processing of geometric data flow is achieved by designing at least two graphics pipeline clusters in a graphics processor, each cluster including an element distribution unit, a geometric processing pipeline and a merger arbitrator.

Benefits of technology

It greatly improves the parallelism of the geometric processing stage, improves the geometric processing performance of the graphics processor, and facilitates determining the order between the geometric output data output from the graphics pipeline cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131611_22052025_PF_FP_ABST
    Figure CN2024131611_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a geometry processing method and apparatus, a device, a storage medium, and a program product. The geometry processing method comprises: each primitive distribution unit acquires primitive cluster data in a geometry data stream, wherein the primitive cluster data comprises a first label used for determining the sequence of the primitive cluster data in the geometry data stream; the primitive distribution unit further divides the primitive cluster data into primitive group data, and distributes the primitive group data to corresponding geometry processing pipelines; the geometry processing pipelines process the primitive group data to obtain primitive group processing results, and output the primitive group processing results to a corresponding merging arbiter; and the merging arbiter performs merging processing on the primitive group processing results on the basis of a second label, so as to obtain geometry output data corresponding to the primitive cluster data.
Need to check novelty before this filing date? Find Prior Art

Description

Geometry processing method, device, equipment, storage medium and program product

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on the Chinese patent application with application number 202311533612.2, application date November 17, 2023, and invention name “Geometric processing method, device, equipment and storage medium”, and claims the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this application as a reference. Technical Field

[0003] The present application relates to the field of graphics processing technology, and in particular to a geometry processing method, apparatus, device, storage medium, and program product. Background Art

[0004] A graphics processing unit (GPU) is a specialized graphics rendering device used to process and display computerized graphics. GPUs are constructed with a highly parallel architecture that provides more efficient processing of a range of complex algorithms than a typical general-purpose central processing unit (CPU). For example, these complex algorithms may correspond to the representation of two-dimensional or three-dimensional computerized graphics. A graphics processor typically includes a front-end geometry processing pipeline and a back-end pixel processing pipeline. In relevant application scenarios, the performance requirements for the geometry processing pipeline are relatively high. Existing geometry processing pipeline designs struggle to provide higher geometry processing performance, becoming a bottleneck for the overall graphics processing.

[0005] Summary of the Invention

[0006] In view of this, embodiments of the present application provide at least one geometry processing method, apparatus, device, storage medium, and program product.

[0007] The technical solution of the embodiment of the present application is implemented as follows:

[0008] On the one hand, an embodiment of the present application provides a geometry processing method, which is applied to a graphics processor, wherein the graphics processor includes at least two graphics pipeline clusters, the graphics pipeline cluster includes a primitive distribution unit, at least two geometry processing pipelines and a merging arbitrator, and the method includes: obtaining primitive block data in a geometry data stream through the primitive distribution unit; the primitive block data includes a first tag for determining the order of the primitive block data in the geometry data stream; dividing the primitive block data into primitive group data through the primitive distribution unit, and distributing the primitive group data to the geometry processing pipeline; the primitive The group data includes a second tag for determining the order of the primitive group data in the primitive block data; the primitive group data is processed by the geometry processing pipeline to obtain a primitive group processing result, and the primitive group processing result is output to the merge arbitrator; the primitive group processing result is merged based on the second tag by the merge arbitrator to obtain geometric output data with a second order corresponding to the primitive block data; the geometric output data includes the first tag corresponding to the primitive block data, which is used to determine the first order between the geometric output data output by each of the graphics pipeline clusters.

[0009] In some embodiments, obtaining the primitive block data in the geometry data stream by the primitive distribution unit includes: reading the primitive block data that the graphics pipeline cluster needs to process from the geometry data stream by the primitive distribution unit.

[0010] In some embodiments, the reading of the primitive block data from the geometric data stream through the primitive distribution unit includes: reading the geometric data stream through the primitive distribution unit; based on a preset primitive block acquisition strategy, discarding the data in the geometric data stream that does not belong to the current graphics pipeline cluster processing, and obtaining the primitive block data that the graphics pipeline cluster needs to process.

[0011] In some embodiments, the reading of the primitive block data from the geometric data stream through the primitive distribution unit includes: determining the address segment of the primitive block data that the graphics pipeline cluster needs to process based on a preset primitive block acquisition strategy; and reading the primitive block data that the graphics pipeline cluster needs to process in the geometric data stream based on the address segment.

[0012] In some embodiments, the graphics processor further includes a global distribution unit; the method further includes: reading the geometric data stream through the global distribution unit, determining the primitive block data that each graphics pipeline cluster needs to process from the geometric data stream based on a preset primitive block acquisition strategy, and distributing it to each graphics pipeline cluster; accordingly, obtaining the primitive block data in the geometric data stream through the primitive distribution unit includes: receiving the primitive block data that the graphics pipeline cluster needs to process sent by the global distribution unit through the primitive distribution unit.

[0013] In some embodiments, the graphics processor further includes at least one pixel processing pipeline, and the method further includes: caching the geometric output data corresponding to the primitive block data to a cache unit corresponding to the graphics pipeline cluster based on the first tag through the graphics pipeline cluster; and reading the geometric output data corresponding to each of the graphics pipeline clusters in a first order from the cache units corresponding to the at least two graphics pipeline clusters based on the first tag through the pixel processing pipeline.

[0014] In some embodiments, the graphics pipeline cluster further includes a blocker, and the caching of the geometric output data corresponding to the primitive block data to the cache unit corresponding to the graphics pipeline cluster based on the first tag by the graphics pipeline cluster includes: distributing the geometric output data corresponding to each of the graphics pipeline clusters in sequence based on the first tag by the blocker, and caching them to the cache unit corresponding to the graphics pipeline cluster; wherein, the distributing the geometric output data corresponding to each of the graphics pipeline clusters and caching them to the cache unit corresponding to the graphics pipeline cluster includes: determining the tile corresponding to each primitive block data in the geometric output data in sequence based on the second order by the blocker; and writing each primitive block data in the geometric output data into the polygon list of the corresponding tile in the cache unit corresponding to the graphics pipeline cluster.

[0015] In some embodiments, when there is at least one target element data belonging to a target tile in the geometric output data, the polygon list of the target tile includes a first label of the geometric output data and each target element data in the geometric output data belonging to the target tile; wherein the order of each target element data in the polygon list is the same as the second order.

[0016] In some embodiments, the pixel processing pipeline reads the geometric output data corresponding to each of the graphics pipeline clusters in a first order from the cache units corresponding to the at least two graphics pipeline clusters based on the first label, including: in the cache units corresponding to the at least two graphics pipeline clusters, the pixel processing pipeline traverses the first label in each polygon list corresponding to the target tile, and takes out the metadata corresponding to the target first label from the polygon list corresponding to the target first label until the metadata does not exist in each polygon list; wherein the target first label is determined based on the order of the first labels in the headers of each polygon list.

[0017] On the other hand, an embodiment of the present application provides a graphics processor, wherein the graphics processor includes at least two graphics pipeline clusters, wherein the graphics pipeline cluster includes a primitive distribution unit, at least two geometry processing pipelines and a merging arbiter; wherein the primitive distribution unit is configured to obtain primitive block data in a geometry data stream; the primitive block data includes a first tag for determining the order of the primitive block data in the geometry data stream; the primitive distribution unit is further configured to divide the primitive block data into primitive group data and distribute the primitive group data to the geometry processing pipeline; the primitive group ... The second tag is used to determine the order of the primitive group data in the primitive block data; the geometry processing pipeline is configured to process the primitive group data, obtain a primitive group processing result, and output the primitive group processing result to the merge arbitrator; the merge arbitrator is configured to merge the primitive group processing result based on the second tag to obtain geometric output data with a second order corresponding to the primitive block data; the geometric output data contains the first tag corresponding to the primitive block data, which is used to determine the first order between the geometric output data output by each of the graphics pipeline clusters.

[0018] On the other hand, an embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0019] On the other hand, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements some or all of the steps in the above method when executed by a processor.

[0020] On the other hand, an embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements some or all of the steps in the above method.

[0021] In an embodiment of the present application, at least two graphics pipeline clusters obtain the primitive block data that they need to process in the geometry data stream, thereby achieving a first split of the geometry data stream and achieving a first level of parallel processing through at least two graphics pipeline clusters; at the same time, within the graphics pipeline cluster, the primitive block data is divided and distributed to at least two geometry processing pipelines by a primitive distribution unit, thereby achieving a second split of the geometry data stream and performing a second level of parallel processing through at least two geometry processing pipelines. Thus, the present application can significantly improve the degree of parallelism in the geometry processing stage and improve the geometry processing performance of the graphics processor through two levels of data division and parallel processing; in addition, since the primitive block data includes a first tag for determining the order of the primitive block data in the geometry data stream, it is convenient to determine the first order between the geometry output data output by each graphics pipeline cluster; since the primitive group data includes a second tag for determining the order of the primitive group data in the primitive block data, it is convenient to reorder the processed primitive group data to restore the order of the processed primitive group data in the primitive block data.

[0022] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0024] FIG1 is a schematic diagram of a first implementation flow of a geometric processing method provided in an embodiment of the present application;

[0025] FIG2 is a second schematic diagram of an implementation flow of a geometric processing method provided in an embodiment of the present application;

[0026] FIG3 is a third schematic diagram of an implementation flow of a geometric processing method provided in an embodiment of the present application;

[0027] FIG4 is a fourth schematic diagram of an implementation flow of a geometric processing method provided in an embodiment of the present application;

[0028] FIG5 is a schematic diagram of a system architecture including a single geometry processing pipeline corresponding to one or more pixel processing pipelines provided by an embodiment of the present application;

[0029] FIG6 is a schematic diagram of a system architecture using a parallel geometry processing pipeline corresponding to one or more pixel processing pipelines according to an embodiment of the present application;

[0030] FIG7 is a schematic diagram of task distribution and merging of a geometry processing pipeline provided by an embodiment of the present application;

[0031] FIG8 is a first schematic diagram of a multi-level parallel GPU geometry processing pipeline provided by an embodiment of the present application;

[0032] FIG9 is a schematic diagram of data segmentation and distribution based on a round-robin approach provided by an embodiment of the present application;

[0033] FIG10 is a schematic diagram of data segmentation and distribution based on load balancing within a graphics pipeline cluster and on a round-robin basis between graphics pipeline clusters according to an embodiment of the present application;

[0034] FIG11 is a second schematic diagram of another multi-level parallel GPU geometry processing pipeline provided by an embodiment of the present application;

[0035] FIG12 is a schematic diagram of the structure of a graphics processor provided in an embodiment of the present application;

[0036] FIG13 is a schematic diagram of a hardware entity of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0038] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. The terms "first / second / third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or sequence of "first / second / third" may be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.

[0040] (1) Tile Based Rendering (TBR) is a scheme that divides the screen into tiles (also called tiles) so that each tile can fit into the on-chip cache. For example, if the on-chip cache can store 512kB of data, the screen can be divided into tiles so that the pixel data contained in each tile is less than or equal to 512kB. In this way, the scene can be rendered by dividing the screen into tiles that can be rendered into the on-chip cache and rendering each tile of the scene individually into the on-chip cache, storing the rendered tiles from the on-chip cache into the frame buffer, and repeating the rendering and storage for each tile of the screen. Therefore, the screen can be rendered tile by tile to render each tile of the scene. It can be understood that the TBR scheme is a mode of delayed reproduction of graphics. Due to its low power consumption, it is widely used in mobile devices, but it also has certain applications in desktop and server-level graphics processors.

[0041] (2) The block splitter is the last module in the front end of TBR. It is used to complete the screen segmentation, record the graphic data covering the tile (Tile), and write the generated information such as tile information (Primitive List) and vertex information (Vertex Data) into the system memory. Among them, the Primitive List is a fixed-length array with a length of Tile. In this array, each element is a linked list, which stores the pointers of all triangles intersecting with the current Tile. The pointer points to the Vertex Data; the Vertex Data stores the vertex and vertex attribute data.

[0042] In modern GPU rendering, the GPU first reads vertex information from a software-configured vertex buffer. The output is processed by the front-end (vertex shader, tessellation, geometry shader, and various fixed-function geometry processing and tiling functions) or the geometry processing pipeline, and the back-end (pixel processing pipeline). The output of the geometry processing pipeline serves as the input to the pixel processing pipeline.

[0043] In practice, some applications or games require heavy geometry processing pipeline workloads, such as large numbers of primitives, tessellation enabled, and complex vertex and geometry shaders. These requirements demand significant geometry pipeline processing throughput. Consequently, conventional graphics processing pipelines that rely solely on geometry processing cannot meet the performance requirements of the entire graphics processing task, and the geometry processing pipeline becomes a performance bottleneck.

[0044] The present invention provides a geometry processing method that can be executed by a processor of a computer device. The computer device may include a server, laptop, tablet, desktop computer, smart TV, set-top box, mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device), or other device with data processing capabilities.

[0045] FIG1 is a schematic diagram of a first implementation flow of a geometric processing method provided in an embodiment of the present application. As shown in FIG1 , the method includes the following steps S101 to S104:

[0046] Step S101: Acquire primitive block data in a geometry data stream through the primitive distribution unit; the primitive block data includes a first tag for determining the order of the primitive block data in the geometry data stream.

[0047] In some embodiments, a current graphics processor includes at least two graphics pipeline clusters. Here, the graphics pipeline clusters include a primitive distribution unit, at least two geometry processing pipelines, and a merge arbiter. Physically, the units within the same graphics pipeline cluster are located within the same area of ​​the chip. In other words, different graphics pipeline clusters can be located in different locations within the chip.

[0048] In some embodiments, the at least two graphics pipeline clusters included in the graphics processor may employ the same cluster configuration or different cluster configurations. The cluster configuration is used to determine at least the number of geometry processing pipelines included in the graphics pipeline cluster. For example, when the same cluster configuration is employed, different graphics pipeline clusters all include the same number of geometry processing pipelines, which is at least two. When different cluster configurations are employed, different graphics pipeline clusters may include different numbers of geometry processing pipelines. The graphics processor may include at least one graphics pipeline cluster, which may include at least two geometry processing pipelines. In some cases, the graphics processor may also include a graphics pipeline cluster containing one geometry processing pipeline.

[0049] In the embodiment of the present application, the aforementioned geometry data stream is the raw input data of the graphics processor. With respect to this geometry data stream, since the graphics processor includes at least two parallel graphics pipeline clusters, the primitive distribution unit within each graphics pipeline cluster can obtain the geometry data that its own graphics pipeline cluster needs to process, namely, the primitive block data within the geometry data stream. Here, the primitive block data can be obtained by the primitive distribution unit within the graphics pipeline cluster actively reading from the geometry data stream or passively receiving it, which is not limited in the present embodiment.

[0050] It is understandable that the primitive block data obtained by the primitive distribution units in different graphics pipeline clusters are different / non-overlapping, and the primitive block data obtained by the primitive distribution units in each graphics pipeline cluster can be combined to restore the original geometry data stream.

[0051] In some embodiments, the primitive block data includes a first tag for determining the order of the primitive block data in the geometric data stream. Exemplarily, the original geometric data stream may include primitive block data 1 to N arranged in sequence, wherein the first tag corresponding to primitive block data 1 can be set to "1",..., the first tag corresponding to primitive block data N can be set to "N". Thus, after the graphics pipeline cluster obtains the primitive block data that it needs to process, it can determine the position of the primitive block data that currently needs to be processed in the original geometric data stream based on the first tag in the current primitive block data. The above exemplary description of the first tag is only for the convenience of understanding the current implementation process, and does not limit the specific implementation method.

[0052] It should be noted that the primitive distribution unit here is described from the perspective of any one of the at least two graphics pipeline clusters in the graphics processor.

[0053] In the above step S101, at least two graphics pipeline clusters of the graphics processor obtain the primitive block data that they need to process from the geometric data stream, which is actually the first splitting of the geometric data stream, and performs the first level of parallel processing through at least two graphics pipeline clusters. The object of parallel processing is the primitive block data that each graphics pipeline cluster needs to process.

[0054] Step S102: Divide the primitive block data into primitive group data by the primitive distribution unit, and distribute the primitive group data to the geometry processing pipeline; the primitive group data includes a second tag for determining the order of the primitive group data in the primitive block data.

[0055] In some embodiments, after a primitive distribution unit within a graphics pipeline cluster receives primitive block data to be processed by the current graphics pipeline cluster, it further decomposes the primitive block data into primitive group data. The graphics pipeline cluster also includes at least two parallel geometry processing pipelines. The primitive distribution unit is connected to each geometry processing pipeline. After the primitive distribution unit divides the primitive block data into primitive group data, it distributes the resulting primitive group data to at least two subsequent geometry processing pipelines. It should be understood that the primitive group data received by different geometry processing pipelines does not overlap, thus avoiding resource waste.

[0056] In some embodiments, the primitive group data includes a second tag for determining the order of the primitive group data in the primitive block data. Exemplarily, the primitive block data obtained by the primitive distribution unit in the graphics pipeline cluster can be divided into primitive group data 1 to N arranged in sequence, wherein the second tag corresponding to primitive group data 1 can be set to "1",..., the second tag corresponding to primitive group data N can be set to "N". Thus, after the geometry processing pipeline processes the received primitive group data, it can determine the position of the processed primitive group data in the original primitive block data based on the second tag in the current primitive group data, thereby facilitating the subsequent restoration of the order of the primitive block data output by each geometry processing pipeline. The above exemplary description of the second tag is only for the purpose of facilitating the understanding of the current implementation process, and does not limit the specific implementation method.

[0057] In some embodiments, after dividing the primitive block data into primitive group data, the primitive distribution unit may sequentially send the obtained primitive group data to subsequent geometry processing pipelines in a round-robin manner. For example, after obtaining N primitive group data, if there are three geometry processing pipelines, the 1+3nth primitive group data may be distributed to the first geometry processing pipeline, the 2+3nth primitive group data may be distributed to the second geometry processing pipeline, and the 3rd (1+n)th primitive group data may be distributed to the third geometry processing pipeline, where n is 0, 1, 2, etc.

[0058] In other embodiments, after dividing the primitive block data into primitive group data, the primitive distribution unit may first obtain the load information of each subsequent geometry processing pipeline, and based on the load information corresponding to each geometry processing pipeline, distribute the primitive group data to the geometry processing pipeline with lower load, so as to achieve load balancing of at least two parallel geometry processing pipelines.

[0059] The above step S102 is actually a second splitting of the geometry data stream, and performing a second level of parallel processing through at least two geometry processing pipelines. The objects of the parallel processing are the primitive group data received by each geometry processing pipeline.

[0060] Step S103: Process the primitive group data through the geometry processing pipeline to obtain a primitive group processing result, and output the primitive group processing result to the merging arbitrator.

[0061] Here, at least two parallel geometry processing pipelines in each graphics pipeline cluster process the received primitive group data in parallel to obtain a processing result for the primitive group. It is understood that the primitive group processing result includes the processed primitive group data corresponding to each geometry processing pipeline. Here, the processed primitive group data also includes a second tag.

[0062] In the embodiment of the present application, the merging arbitrator is connected to each geometry processing pipeline and is configured to receive the primitive group processing result, that is, to receive the processed primitive group data output by each geometry processing pipeline.

[0063] Step S104: The processing results of the primitive groups are merged by the merge arbitrator based on the second label to obtain geometric output data with a second order corresponding to the primitive block data; the geometric output data includes the first label corresponding to the primitive block data, which is used to determine the first order between the geometric output data output by each of the graphics pipeline clusters.

[0064] In some embodiments, because at least two geometry processing pipelines process data in parallel, and considering differences in processing performance between the different geometry processing pipelines and differences in processing time between different primitive group data, when each geometry processing pipeline outputs processed primitive group data to the merge arbitrator, it is not possible to ensure that the output order is the same as the order of the primitive group data in the original primitive block data. Therefore, the merge arbitrator can reorder the processed primitive group data output by the at least two geometry processing pipelines based on the second tags included in each processed primitive group data to restore the order of the processed primitive group data in the primitive block data.

[0065] Here, the geometric output data having the second order is the processed primitive group data whose order has been restored.

[0066] Exemplarily, after the primitive block data to be processed by the current graphics pipeline cluster is divided into the 1st to 7th primitive group data, the 1st, 3rd, and 4th primitive group data are distributed to the first geometry processing pipeline for processing to obtain the processed 1st, 3rd, and 4th primitive group data; the 2nd, 5th, 6th, and 7th primitive group data are distributed to the second geometry processing pipeline for processing to obtain the processed 2nd, 5th, 6th, and 7th primitive group data. The merging arbitrator sequentially restores each processed primitive group data according to the second tag carried by each processed primitive group data, and the obtained geometric output data with the second order includes the processed 1st to 7th primitive group data. It can be understood that the above-mentioned geometric output data with the second order is actually the processed output data of the current graphics pipeline cluster for the primitive block data.

[0067] In the embodiment of the present application, the geometry output data with the second order corresponding to the primitive block data also includes the first tag corresponding to the primitive block data. Thus, for the geometry output data with the second order output by each of the at least two graphics pipeline clusters in the graphics processor, the first order between the geometry output data output by each graphics pipeline cluster can be determined based on the first tag.

[0068] In an embodiment of the present application, at least two graphics pipeline clusters obtain the primitive block data that they need to process in the geometry data stream, thereby achieving a first split of the geometry data stream and achieving a first level of parallel processing through at least two graphics pipeline clusters; at the same time, within the graphics pipeline cluster, the primitive block data is divided and distributed to at least two geometry processing pipelines by a primitive distribution unit, thereby achieving a second split of the geometry data stream and performing a second level of parallel processing through at least two geometry processing pipelines. Thus, the present application can significantly improve the degree of parallelism in the geometry processing stage and improve the geometry processing performance of the graphics processor through two levels of data division and parallel processing; in addition, since the primitive block data includes a first tag for determining the order of the primitive block data in the geometry data stream, it is convenient to determine the first order between the geometry output data output by each graphics pipeline cluster; since the primitive group data includes a second tag for determining the order of the primitive group data in the primitive block data, it is convenient to reorder the processed primitive group data to restore the order of the processed primitive group data in the primitive block data.

[0069] FIG2 is a second schematic diagram of a geometric processing method according to an embodiment of the present invention, which can be executed by a processor of a computer device. Based on FIG1 , S101 in FIG1 can be updated to S201 , which will be described in conjunction with the steps shown in FIG2 .

[0070] Step S201 : Reading primitive block data to be processed by the graphics pipeline cluster from the geometry data stream via the primitive distribution unit.

[0071] In the current embodiment, the primitive distribution unit can proactively obtain primitive block data required for processing by its current graphics pipeline cluster from the geometry data stream. For primitive distribution units located in different graphics pipeline clusters, these at least two primitive distribution units can each obtain primitive block data required for processing by their current graphics pipeline cluster from the geometry data stream using pre-agreed acquisition rules (i.e., subsequent primitive block acquisition strategies), and the primitive block data corresponding to each graphics pipeline cluster does not overlap.

[0072] In some embodiments, the primitive distribution unit may first obtain the entire geometry data stream and discard data not required for processing by the corresponding graphics pipeline cluster, thereby obtaining the primitive block data required for processing by the current graphics pipeline cluster. That is, reading the primitive block data required for processing by the graphics pipeline cluster from the geometry data stream by the primitive distribution unit may be accomplished through steps S2011 and S2012.

[0073] Step S2011: Read the geometric data stream through the primitive distribution unit.

[0074] In the embodiment of the present application, the geometric data stream includes multiple geometric data to be divided, and each geometric data to be divided is read by the primitive distribution unit of all graphics pipeline clusters in the form of an index. In the subsequent geometric processing process, the corresponding geometric data can be read through the index for subsequent processing.

[0075] Step S2012: Based on a preset primitive block acquisition strategy, discard the data in the geometric data stream that does not belong to the current graphics pipeline cluster for processing, and obtain the primitive block data that the graphics pipeline cluster needs to process.

[0076] In some embodiments, the primitive block acquisition strategy may be: based on the cluster identifier corresponding to the current graphics pipeline cluster and the size of the preset primitive block data, the data range of the primitive block data to be processed by the graphics pipeline cluster is determined in a round-robin manner, and the data acquired based on each data range is used as the primitive block data to be processed by the graphics pipeline cluster. During implementation, after obtaining the index of each piece of geometric data to be divided in the geometric data stream, it can be determined whether the index of the geometric data belongs to the data range corresponding to the current graphics pipeline cluster, and the geometric data to be divided that belongs to the data range of the primitive block data corresponding to the current graphics pipeline cluster is used as the primitive block data to be processed by the graphics pipeline cluster; and the geometric data to be divided that does not belong to the data range of the primitive block data corresponding to the current graphics pipeline cluster is discarded.

[0077] Exemplarily, when the geometry data stream includes 1 to 1000 (index) geometry data, there are 2 graphics pipeline clusters, and the preset size of the primitive block data is 200, the data range of the primitive block data corresponding to the first graphics pipeline cluster can be 1 to 200, 401 to 600, and 801 to 1000; the data range of the primitive block data corresponding to the second graphics pipeline cluster can be 201 to 400, 601 to 800. Therefore, for the first graphics pipeline cluster, after obtaining these 1000 geometric data, it can be determined that the 1st to 200th geometric data belong to the data range of the corresponding primitive block data, and the 1st to 200th geometric data are used as the first primitive block data that the first graphics pipeline cluster needs to process; for the 201st to 400th geometric data, since they do not belong to the data range of the corresponding primitive block data, the 201st to 400th geometric data are discarded, and so on, until the 801st to 100th geometric data are used as the third primitive block data that the first graphics pipeline cluster needs to process.

[0078] In the current embodiment, all geometric data streams are read through the primitive distribution unit, and data that does not belong to the current graphics pipeline cluster processing is discarded, thereby obtaining the primitive block data that the graphics pipeline cluster needs to process. Since this method filters from all the acquired geometric data, it can reduce the problem of data omission during the primitive block data distribution process.

[0079] In some embodiments, the primitive distribution unit may first determine the data range of primitive block data that the current graphics pipeline cluster needs to process, and then read only the primitive block data that the graphics pipeline cluster needs to process from the geometry data stream. That is, reading the primitive block data that the graphics pipeline cluster needs to process from the geometry data stream by the primitive distribution unit may also be accomplished through steps S2013 and S2014.

[0080] Step S2013: Based on a preset primitive block acquisition strategy, determine the address segment of the primitive block data that the graphics pipeline cluster needs to process.

[0081] In some embodiments, the primitive block acquisition strategy may be: based on the cluster identifier corresponding to the current graphics pipeline cluster and the size of the preset primitive block data, the address segments of the primitive block data that the graphics pipeline cluster needs to process are determined in a round-robin manner, and the data acquired based on each address segment is used as the primitive block data that the graphics pipeline cluster needs to process.

[0082] Step S2014: Reading primitive block data that needs to be processed by the graphics pipeline cluster from the geometry data stream based on the address segment.

[0083] Here, after obtaining the address segment of each primitive block data that needs to be processed by the graphics pipeline cluster, the corresponding primitive block data can be read from the memory based on the address segment of the primitive block data.

[0084] For example, when there are two graphics pipeline clusters, for the first graphics pipeline cluster, based on the starting address of the geometry data stream and the unit address offset corresponding to the preset primitive block data size, the address segments of the primitive block data corresponding to the first graphics pipeline cluster can be determined to include (starting address, starting and ending addresses + unit address offset - 1), (starting and ending addresses + 2 × unit address offsets, starting and ending addresses + 3 × unit address offset - 1), and so on; the address segments of the primitive block data corresponding to the second graphics pipeline cluster can be determined to include (starting and ending addresses + unit address offsets, starting and ending addresses + 2 × unit address offset - 1), (starting and ending addresses + 3 × unit address offsets, starting and ending addresses + 4 × unit address offset - 1), and so on. Subsequently, each graphics pipeline cluster reads the primitive block data to be processed from the memory based on the address segments of the corresponding primitive block data.

[0085] In the current embodiment, the primitive distribution unit does not need to read the entire geometric data stream. By calculating the address segment of the primitive block data that the graphics pipeline cluster needs to process, it can read the primitive block data that the graphics pipeline cluster needs to process, thereby reducing the need for memory access and also reducing the workload of the primitive distribution unit itself.

[0086] Figure 3 is a third schematic flow diagram of an implementation process of a geometry processing method provided in an embodiment of the present application. This method can be executed by a processor of a computer device. Based on Figure 1 , the graphics processor further includes a global distribution unit; the method in Figure 1 also includes step S301. Accordingly, step S101 can be updated to step S302. This will be explained in conjunction with the steps shown in Figure 3 .

[0087] Step S301 : The global distribution unit reads the geometry data stream, determines primitive block data to be processed by each graphics pipeline cluster based on a preset primitive block acquisition strategy, and distributes the primitive block data to each graphics pipeline cluster.

[0088] In an embodiment of the present application, the global distribution unit is connected to each graphics pipeline cluster and is configured to read the geometric data stream and determine the primitive block data that each graphics pipeline cluster needs to process from the geometric data stream based on a preset primitive block acquisition strategy; at the same time, the global distribution unit is also configured to distribute the primitive block data that each graphics pipeline cluster needs to process to the corresponding graphics pipeline cluster.

[0089] In the embodiment of the present application, the geometric data stream includes a plurality of geometric data to be divided, and each geometric data to be divided is read by the global distribution unit in the form of an index. In the subsequent geometric processing process, the corresponding geometric data can be read through the index for subsequent processing.

[0090] In some embodiments, the primitive block acquisition strategy may include determining the data range / address segment of the primitive block data to be processed by each graphics pipeline cluster in a round-robin manner based on the cluster identifier corresponding to each current graphics pipeline cluster and the preset size of the primitive block data. Subsequently, the data range / address segment of the primitive block data to be processed by each graphics pipeline cluster is sent to the corresponding graphics pipeline cluster. For a specific implementation of determining the data range / address segment of the primitive block data to be processed by each graphics pipeline cluster in a round-robin manner, please refer to the embodiment of FIG. 2 and will not be further described here.

[0091] In other embodiments, the primitive block acquisition strategy may be: based on the cluster identifier corresponding to each current graphics pipeline cluster and the preset size of the primitive block data, the geometric data stream is first divided, and the data range / address segment of each primitive block data at the divided point is determined. During the distribution of the primitive block data, the load information of each graphics pipeline cluster is obtained, and the primitive block data is distributed to the graphics pipeline cluster with the smallest load to achieve load balancing between graphics pipeline clusters.

[0092] Step S302: Receive, via the primitive distribution unit, primitive block data that needs to be processed by the graphics pipeline cluster and is sent by the global distribution unit.

[0093] In some embodiments, the global distribution unit can send the data range of the primitive block data that each graphics pipeline cluster needs to process, or it can send the address segment of the primitive block data to the primitive distribution unit of each graphics pipeline cluster. It can be understood that compared to the solution provided in Figure 2, in which the primitive distribution unit needs to read the primitive block data that the graphics pipeline cluster to which it belongs needs to process from the original geometry data stream, the primitive distribution units in each cluster in the current embodiment do not need to pre-agreed on the acquisition rules (i.e., primitive block acquisition strategy). Instead, the acquisition rules are stored in the global distribution unit to achieve the distribution of primitive block data from the geometry data stream to each graphics pipeline cluster.

[0094] In the current embodiment, each primitive distribution unit does not require a pre-agreed primitive block acquisition strategy, and the distribution of primitive block data from the geometric data stream to each graphics pipeline cluster is achieved through a global distribution unit. In this way, in the process of changing the primitive block acquisition strategy, only the global distribution unit needs to be configured, and there is no need to configure the primitive distribution unit in each graphics pipeline cluster, thereby improving the flexibility of the system.

[0095] FIG4 is a fourth schematic flow diagram of an implementation flow of a geometry processing method provided in an embodiment of the present application. The method can be executed by a processor of a computer device. The graphics processor also includes at least one pixel processing pipeline. Based on FIG1 , the method can also include steps S401 and S402, which will be described in conjunction with the steps shown in FIG4 .

[0096] Step S401 : Cache the geometric output data corresponding to the primitive block data to a cache unit corresponding to the graphics pipeline cluster based on the first tag through the graphics pipeline cluster.

[0097] In the embodiment of the present application, after completing geometric processing on the primitive block data, the graphics pipeline cluster generates geometric output data corresponding to the primitive block data in the second order. Accordingly, the geometric output data corresponding to the primitive block data in the second order also includes the first tag.

[0098] In some embodiments, different graphics pipeline clusters correspond to different cache units, and the cache units of different graphics pipeline clusters are independent of each other. The independent cache units can be independent address spaces or address segments of the global memory, or can be memories allocated to each graphics pipeline cluster.

[0099] In the process of caching the geometric output data with the second order corresponding to the primitive block data, the graphics pipeline cluster will store the geometric output data corresponding to the primitive block data in the cache unit corresponding to the current graphics pipeline cluster based on the first tag of the primitive block data.

[0100] For example, an original geometry data stream includes sequentially arranged primitive block data 1 to N. The first tag corresponding to primitive block data 1 can be set to "1," ..., and the first tag corresponding to primitive block data N can be set to "N." For the first graphics pipeline cluster, if the n+1th primitive block data is distributed sequentially to the first graphics pipeline cluster, where n is an integer greater than or equal to 0, then after the first graphics pipeline cluster processes each primitive block data sequentially and obtains the corresponding geometry output data, the geometry output data corresponding to each of the n+1 primitive blocks will be cached sequentially in the order of the first tags to the cache unit corresponding to the graphics pipeline cluster.

[0101] Step S402 : Reading, through the pixel processing pipeline based on the first tag, geometric output data corresponding to each of the graphics pipeline clusters from cache units corresponding to the at least two graphics pipeline clusters in a first order.

[0102] In some embodiments, different graphics pipeline clusters correspond to different cache units. That is, after all primitive block data corresponding to the original geometry data stream is processed, all processed geometry output data is dispersed into various cache units (of course, some cache units may not have processed geometry output data because the graphics pipeline cluster corresponding to the cache unit does not have primitive block data to process). Therefore, the pixel processing pipeline can restore the relative order of the geometry output data in each cache unit based on the first tag corresponding to each geometry output data, and sequentially read the geometry output data corresponding to each graphics pipeline cluster in the first order.

[0103] For example, taking the original geometric data stream including primitive block data 1 to 5 arranged in sequence as an example, if there are two graphics pipeline clusters, among which the 1st, 2nd and 5th primitive block data are distributed to the first graphics pipeline cluster; the 3rd and 4th primitive block data are distributed to the second graphics pipeline cluster; accordingly, the cache unit corresponding to the first graphics pipeline cluster stores the geometric output data corresponding to the 1st, 2nd and 5th primitive block data respectively, and the cache unit corresponding to the second graphics pipeline cluster stores the geometric output data corresponding to the 3rd and 4th primitive block data respectively. At this time, the pixel processing pipeline can read these 5 geometric output data in sequence based on the first labels corresponding to each geometric output data, according to the relative order of the primitive block data corresponding to the geometric output data, that is, the first order.

[0104] In the current embodiment, since the primitive block data includes a first tag for determining the order of the primitive block data in the geometric data stream, and at the same time, the first tag is passed through to the pixel processing pipeline as the primitive block data is processed by the graphics pipeline cluster to generate corresponding geometric output data, the pixel processing pipeline can determine the first order between the geometric output data output by each graphics pipeline cluster based on the first tag.

[0105] The above embodiments improve the parallelism of the geometry processing process. Accordingly, the parallelism of the back-end pixel processing process can also be improved through the TBR architecture. Therefore, in some embodiments, the graphics pipeline cluster further includes a block splitter. The block splitter is located between the merge arbiter and the cache unit. The aforementioned caching of the geometry output data corresponding to the primitive block data in the cache unit corresponding to the graphics pipeline cluster based on the first tag by the graphics pipeline cluster can be achieved through step S4011.

[0106] Step S4011 : Distribute the geometric output data corresponding to each of the graphics pipeline clusters in sequence based on the first tag through the block splitter, and cache the data in the cache unit corresponding to the graphics pipeline cluster.

[0107] The distributing of the geometric output data corresponding to each of the graphics pipeline clusters and caching it in the cache unit corresponding to the graphics pipeline cluster includes: determining, by the block splitter based on the second order, the tiles corresponding to each primitive in the geometric output data; and writing each primitive in the geometric output data into the polygon list of the corresponding tile in the cache unit corresponding to the graphics pipeline cluster.

[0108] In some embodiments, since the first tag can represent the prior order between the geometric output data, the blocker needs to process the distribution process of each geometric output data in sequence based on the first tag corresponding to each geometric output data. Specifically, within the geometric output data, the blocker can sequentially determine the tile corresponding to each primitive element in the geometric output data based on the second order. Therefore, from a holistic perspective, through the first tag and the second tag, the blocker can sequentially determine the tile to which each primitive element belongs according to the order in which it appears in the original geometric data stream, and store them sequentially in the polygon list corresponding to the tile.

[0109] In the current embodiment, the tile corresponding to each of the primitives in the geometric output data is determined in sequence by the block splitter based on the first label; each of the primitives in the geometric output data is written into the polygon list of the corresponding tile. In this way, the relative order between at least two primitives belonging to one geometric output data in the polygon list is the same as their order in the geometric data stream.

[0110] In some embodiments, each polygon list ends with a terminator; when the pixel processing pipeline extracts primitive block data from the polygon list, the terminator is used to indicate whether the primitive block data in the polygon list is completely extracted.

[0111] In some embodiments, when there is at least one target feature element in the geometry output data belonging to a target tile, the polygon list of the target tile includes a first tag of the geometry output data and each target feature element in the geometry output data belonging to the target tile.

[0112] The order of each target graphic element in the polygon list is the same as the second order.

[0113] In some embodiments, the first tag of the target pixel group data is located before the target pixel data. Thus, the pixel processing pipeline can first read the first tag corresponding to the pixel data and then determine whether to read the pixel data corresponding to the first tag. In other embodiments, the first tag of the target pixel group data can also be located after each target pixel data.

[0114] In some embodiments, the above-mentioned reading of the geometric output data corresponding to each of the graphics pipeline clusters in sequence according to the first order from the cache units corresponding to the at least two graphics pipeline clusters through the pixel processing pipeline based on the first tag can be achieved through step S4021.

[0115] Step S4021: In the cache units corresponding to the at least two graphics pipeline clusters, the first label in each polygon list corresponding to the target tile is traversed through the pixel processing pipeline, and the metadata corresponding to the target first label is taken out from the polygon list corresponding to the target first label until the metadata no longer exists in each polygon list.

[0116] In some embodiments, all first tags in all polygon lists may be obtained, and the target first tag may be determined based on the order of all first tags. In the current embodiment, the storage location of the first tag in the polygon list may be adaptively adjusted based on the actual scenario.

[0117] In other embodiments, when the first tag of the target primitive group data precedes the target primitive data, only the first tag at the header of each polygon list may be obtained, and the target first tag may be determined based on the order of precedence. In other words, when the first tag of the target primitive group data precedes the target primitive data, the pixel processing pipeline traverses the first tag at the header of each polygon list corresponding to the target tile, and the target first tag is determined based on the order of precedence of the first tags at the header of each polygon list.

[0118] In some embodiments, the above-mentioned step S4021 can be implemented by the following process: obtaining the first tag at the head of each polygon list corresponding to the tile through the pixel processing pipeline. In the case that there is at least one first tag in each polygon list, the target first tag is determined based on the order of the obtained first tags, and the target first tag and the metadata corresponding to the target first tag are sequentially taken out from the polygon list where the target first tag is located, and the process of obtaining the first tag at the head of each polygon list corresponding to the tile through the pixel processing pipeline is returned. In the case that there is no first tag and metadata in each polygon list, it indicates that the distribution process of each metadata has been completed.

[0119] In the current embodiment, by setting the first label of the primitive group corresponding to each primitive data before the primitive data in the polygon list, the order of the primitive data can be clarified in the process of merging the polygon list corresponding to the current block in the pixel processing pipeline, thereby effectively restoring the original input order.

[0120] The following describes the application of the geometric processing method provided in the embodiment of the present application in actual scenarios, involving a geometric processing method under the TBR architecture.

[0121] In modern GPU rendering, the GPU first reads vertex information from a software-configured vertex buffer. This information is then passed through the front-end (vertex shader, tessellation, geometry shader, and various fixed-function geometry processing and blocking functions)—the geometry processing pipeline—and the back-end (the pixel processing pipeline) to produce the output. The output of the geometry processing pipeline serves as the input for the pixel processing pipeline. To achieve parallel rendering, GPU architectures generally employ multiple pixel processing pipelines for tiled rendering. This involves dividing the entire screen coordinates into multiple tiles. After processing the vertex coordinates, the front-end module outputs the results to a specific data structure. Each tile has its own dedicated data structure, representing the primitive information covering the current tile. Each tile can be rendered independently in the fragment shader, while the geometry processing pipeline processes the entire screen's primitive information. Therefore, please refer to Figure 5 , which illustrates a system architecture diagram including a single geometry processing pipeline corresponding to one or more pixel processing pipelines.

[0122] As shown in Figure 5, the graphics processor pipeline includes a geometry processing pipeline 110. The input data of the geometry processing pipeline 110 is a geometry data stream. After completing the processing of the geometry data stream, the processed geometry data stream is sent to the blocker 120. The blocker 120 divides the processed geometry data stream into processed geometry data of different blocks according to a preset blocking strategy, and caches them in the cache unit 130; then, the processed geometry data of the corresponding blocks are read from the cache unit 130 through at least one parallel pixel processing pipeline 140 to complete the pixel processing process and obtain the final output data.

[0123] After research, it was found that the processing power of a single geometry processing pipeline is insufficient. Therefore, based on the graphics processing pipeline in Figure 5, this application provides a solution that includes a parallel geometry processing pipeline. Please refer to Figure 6, which shows a schematic diagram of a system architecture that uses a parallel geometry processing pipeline corresponding to one or more pixel processing pipelines.

[0124] As shown in Figure 6 , compared to the pipeline shown in Figure 5 , the geometry processing pipeline 110 and tiler 120 in Figure 5 are updated to form a graphics pipeline cluster 20. The graphics pipeline cluster 20 includes a primitive distribution unit 210, at least one geometry processing pipeline 220, a merge arbiter 230, and a tiler 240. The primitive distribution unit 210 is configured to split the input geometry data stream into multiple parts and send each part to the at least one geometry processing pipeline 220. After processing by the at least one geometry processing pipeline 220 is completed, the output data of the at least one geometry processing pipeline 220 is merged in the original order by the merge arbiter 230 and then processed by the tiler 240. The blocker 240 divides the processed geometric data stream into processed geometric data of different blocks according to a preset blocking strategy, and caches them in the cache unit 130; then, the processed geometric data of the corresponding blocks are read from the cache unit 130 through at least one parallel pixel processing pipeline 140 to complete the pixel processing process and obtain the final output data.

[0125] It's understandable that the aforementioned graphics pipeline cluster is both a logical concept and a concept in chip physical design. Physically, the geometry processing pipelines and units within the same graphics pipeline cluster are physically located in the same area on the chip die (a single wafer region encompassing a complete functional unit or a group of related functional units). Logically, the primitive distribution unit and merge arbiter within the graphics pipeline cluster merge the inputs and outputs of multiple geometry processing pipelines, allowing at least one geometry processing pipeline to share the same input and output interfaces as an existing geometry processing pipeline, allowing it to directly replace a single existing geometry processing pipeline to increase throughput.

[0126] The current design of the parallel geometry processing pipeline is to split the input geometry data stream into multiple primitive groups (PG) according to smaller granularity (such as a group of hundreds of triangles), and then send each primitive group to a geometry processing pipeline. After processing, the output results of each geometry processing pipeline are merged and sent out in the original input order through the merging arbitrator. In some embodiments, the merging arbitrator restores the original input order by attaching a label number to each primitive group through the primitive distribution unit. This label number is passed through the geometry processing pipeline to the merging arbitrator, so that the merging arbitrator can identify the mutual order of the processed primitive group data output by each geometry processing pipeline. It can be understood that the above-mentioned mechanism of restoring the order through the label number is only for exemplary purposes, and the present application can also restore the order between each primitive group through other mechanisms.

[0127] In some embodiments, after the input geometry data stream is split into multiple primitive groups, various mechanisms can be used to distribute them to the geometry processing pipelines. For example, round robin sequential distribution can be used. Alternatively, dynamic load balancing can be used for distribution. Specifically, when a primitive distribution unit generates a primitive group, it selects the one with the least outstanding tasks among the n geometry processing pipelines, thereby balancing the processing load across the multiple processing pipelines.

[0128] Please refer to Figure 7, which is a schematic diagram of task distribution and merging in a geometry processing pipeline provided by an embodiment of the present application. The primitive distribution unit 310 is configured to split the input primitive data into multiple primitive groups, such as primitive groups 1 through 3n in Figure 7 . The divided primitive groups are then distributed to different geometry processing pipelines. For example, primitive group 1, primitive group n+1, primitive group 2n+1, and so on are distributed to geometry processing pipeline 321; primitive group 2, primitive group n+2, primitive group 2n+2, and so on are distributed to geometry processing pipeline 322; and finally, primitive group n, primitive group 2n, primitive group 3n, and so on are distributed to geometry processing pipeline 32n. Each geometry processing pipeline sends the processed data to a merge arbiter 330, which merges and arbitrates the processed data, restoring the primitive groups to their original order, and then sends it to a blocker 340. It is understandable that the above process needs to ensure that the output order is completely consistent with the input order, which is commonly known as the GPU API Order preservation requirement.

[0129] In actual applications, some applications or games will have very heavy geometry processing pipeline workloads, such as a large number of primitives, surface tessellation enabled, complex vertex shaders, and geometry shaders, all of which require a large processing throughput of the geometry pipeline. The scalability of the geometry processing pipeline in related technical solutions is limited. Even if there is a design with multiple geometry processing pipelines, it is limited by the throughput of the two splitting and merging units, the primitive distribution unit and the merge arbitrator. It is difficult to continue to expand the geometry processing pipeline. In addition, when the number of geometry processing pipelines increases, they are usually distributed in multiple graphics pipeline clusters (due to the needs of architectural design or chip physical design), and are physically separated by a large distance. Connecting the input and output of the geometry processing pipelines distributed in multiple clusters with a distributor and a blocker to ensure large bandwidth while handling the task imbalance between multiple pipelines is also difficult (the buffer size required for merge arbitration increases dramatically).

[0130] Based on the above reasons, the embodiment of the present application proposes a method / device that can realize parallel processing of geometry processing work by multiple graphics pipeline clusters, thereby improving the geometry processing performance of the entire GPU. There are multiple parallel geometry processing pipelines (Geometry Processing Pipe, GPP) in each graphics pipeline cluster, ensuring that each graphics pipeline cluster has good throughput and computing power. At the same time, the parallelization of multiple graphics pipeline clusters further improves the overall geometry processing throughput, solving the problem of insufficient throughput improvement of parallel geometry processing pipelines within one level. This also achieves the effect of configurable geometry throughput and computing power at both levels of the entire GPU segmentation.

[0131] Please refer to FIG8 , which is a schematic diagram of a multi-level parallel GPU geometry processing pipeline according to an embodiment of the present application. Among them, the first layer is the Graphics Pipelines Cluster (GPC), also called GPU core or GPU by some manufacturers. The Graphics Pipelines Cluster layer includes n graphics pipeline clusters, namely Graphics Pipeline Cluster 41 to Graphics Pipeline Cluster 4n. The first Graphics Pipeline Cluster includes a primitive distribution unit 411, at least two geometry processing pipelines 412, a merge arbitrator 413 and a blocker 414, ..., the nth Graphics Pipeline Cluster includes a primitive distribution unit 4n1, at least two geometry processing pipelines 4n2, a merge arbitrator 4n3 and a blocker 4n4. At the same time, the output data of each Graphics Pipeline Cluster can be stored in a corresponding cache unit, such as the first Graphics Pipeline Cluster corresponds to cache unit 415, and the nth Graphics Pipeline Cluster corresponds to cache unit 4n5. The second layer is the Geometry Processing Pipeline (GPP), which can include multiple pixel processing pipelines (clusters) 46. Each GPC can have multiple GPPs. There is a primitive distribution unit in the front stage of all GPCs. It is responsible for grabbing vertex information from memory and splitting the data stream using a fixed algorithm (such as round-robin) for two-layer data partitioning and distribution.

[0132] In order to facilitate the understanding of the above scheme, the following will be explained by taking the architecture including m graphics pipeline clusters and the graphics pipeline cluster including n geometry processing pipelines as an example. The input primitive data stream is first evenly divided into primitive blocks (PC) of fixed size according to a certain rule, and each primitive block is divided into a graphics pipeline cluster in a round-robin order. For the reading process, there are two reading methods. One is that each graphics pipeline cluster reads all the input geometric data streams, and then discards the data that does not belong to the graphics pipeline cluster to be processed; the other is that each graphics pipeline cluster calculates the range of data that needs to be processed by the graphics pipeline cluster, and only reads the input geometric data stream within this range.

[0133] Within the graphics pipeline cluster, a primitive block is further divided into n primitive groups, each assigned to a geometry processing pipeline. After processing, the blocks are merged and restored to their original order by the merge arbiter before being sent to the block splitter. The distribution and merge arbitration mechanisms here are identical to those in the prior art (possibly using round-robin or dynamic load balancing), and will not be further elaborated here.

[0134] Finally, each blocker outputs a polygon list to a cache space unique to each graphics pipeline cluster. The subsequent pixel processing pipeline (also known as the rasterization processing pipeline or fragment processing pipeline), whether in the form of a single channel, multiple channels, or a multi-channel cluster, will merge the polygon lists output by each graphics pipeline cluster, restore the order, and perform rasterization processing.

[0135] Please refer to FIG9 , which is a schematic diagram of data segmentation and distribution based on a round-robin approach provided in an embodiment of the present application.

[0136] There are m graphics pipeline clusters, from graphics pipeline cluster 51 to graphics pipeline cluster 5m. The primitive distribution unit in each graphics pipeline cluster reads the primitive blocks that its own graphics pipeline cluster needs to process from the input geometry data stream. As shown in Figure 9, the current division / reading of primitive blocks uses a round-robin order. That is, for graphics pipeline cluster 51, the primitive blocks it needs to process are the 1st primitive block, the m+1th primitive block, and so on; for graphics pipeline cluster 5m, the primitive blocks it needs to process are the mth primitive block, the 2mth primitive block, and so on. At this time, the primitive distribution unit in a graphics pipeline cluster is configured to obtain the primitive blocks that its own graphics pipeline cluster needs to process from the geometry data stream.

[0137] Inside the graphics pipeline cluster, taking the graphics pipeline cluster 51 as an example, the primitive distribution unit 511 inside the graphics pipeline cluster 51 will divide the obtained 1st primitive block, m+1th primitive block and 2m+1th primitive block, etc., respectively, and then distribute them to the internal n geometry processing pipelines (geometry processing pipeline 5121 to geometry processing pipeline 512n). For example, for the 1st primitive block, the primitive distribution unit 511 divides the 1st primitive block into n primitive groups, distributes primitive group 1 of the 1st primitive block to the geometry processing pipeline 5121, distributes primitive group 2 of the 1st primitive block to the geometry processing pipeline 5122, and distributes primitive group n of the 1st primitive block to the geometry processing pipeline 512n.

[0138] The merging arbiter 513 is configured to rearrange the primitive groups output by the geometry processing pipelines in the order in which they were input, and output them to the block splitter. For example, for the 1st and m+1th primitive blocks, the geometry processing pipelines 5121 to 512n send the processed primitive groups 1 to n of the 1st primitive block and the processed primitive groups 1 to n of the m+1th primitive block to the merging arbiter 513, respectively. The merging arbiter 513 rearranges these 2n primitive groups in order to obtain: primitive group 1 of the 1st primitive block, primitive group 2 of the 1st primitive block, ..., primitive group n of the 1st primitive block, primitive group 1 of the m+1th primitive block, primitive group 2 of the m+1th primitive block, ..., primitive group n of the m+1th primitive block, and further merges them into the 1st primitive block and the m+1th primitive block.

[0139] The blocker 514 is configured to distribute the multiple data to be distributed in each primitive group based on the above order, and generate a polygon list for each tile in the cache unit 515. The polygon list includes the distributed data corresponding to the tile. Afterwards, the pixel processing pipeline can obtain the distributed data corresponding to each tile from each polygon list corresponding to the tile for rasterization of the tile. It can be understood that different cache units are set for different blockers (actually for different graphics pipeline clusters), and the cache units corresponding to different blockers (graphics pipeline clusters) are independent of each other. The independent cache units here can be independent address spaces or address segments of the global memory, or can be allocated to each cluster.

[0140] Correspondingly, the graphics pipeline cluster 5m also generates a polygon list for each tile in the cache unit 5m5, wherein the polygon list includes the distributed data corresponding to the tile.

[0141] Please refer to Figure 10, which is a schematic diagram of data partitioning and distribution based on load balancing within a graphics pipeline cluster and between graphics pipeline clusters in a round-robin manner, provided by an embodiment of the present application. It can be seen that, unlike Figure 9, the input data of the geometry processing pipeline within the graphics pipeline cluster 51 is no longer distributed in a round-robin manner, but is distributed based on the load of the geometry processing pipeline. The primitive distribution unit 511 will distribute the primitive groups obtained after the division based on the loads corresponding to the geometry processing pipelines 5121 to 512n respectively. Due to the different loads, the primitive groups corresponding to the primitive blocks will be distributed to the subsequent geometry processing pipelines with lower loads according to the load level (rather than in round-robin distribution), such as the first one. 1 primitive block is distributed to geometry processing pipeline 5122, primitive group 3 of the first primitive block is distributed to geometry processing pipeline 5121, and primitive group 5 of the first primitive block is distributed to geometry processing pipeline 512n. Other primitive groups of the first primitive block, such as primitive group 2 of the first primitive block, are distributed to other geometry processing pipelines not shown in FIG10 (which may be geometry processing pipeline 5123, not shown in FIG10). Similarly, other primitive groups are also distributed to subsequent geometry processing pipelines with lower loads based on load, which will not be described in detail here. It should be understood that FIG10 is an example based on load conditions. Therefore, it appears that there is no regular pattern in the input data of each geometry processing pipeline.

[0142] In the above embodiment, the primitive distribution units in each graphics pipeline cluster independently read the input geometric data stream and select the portion to be processed.

[0143] In other embodiments, a two-stage primitive distribution unit can be used. The first-stage primitive distribution unit is responsible for reading the input geometry data stream, splitting it into primitive blocks, and distributing it to each graphics pipeline cluster. The second-stage primitive distribution unit within the graphics pipeline cluster further splits the input primitive blocks into primitive groups and distributes them to the geometry processing pipeline. Subsequent merging and other operations are consistent with the above-mentioned embodiments.

[0144] Please refer to Figure 11, which shows another schematic diagram of a multi-level parallel GPU geometry processing pipeline. It can be seen that compared to Figure 8, the original primitive distribution unit 411 to primitive distribution unit 4n1 have been updated to the current primitive level 1 distribution unit 71, primitive level 2 distribution unit 721 to primitive level 2 distribution unit 72n. The primitive level 1 distribution unit 71 is configured to read the entire input geometry data stream, then determine the data that each graphics pipeline cluster needs to process, and distribute the data to each graphics pipeline cluster; the primitive level 2 distribution unit is configured to receive the data that the graphics pipeline cluster needs to process, and further divide it into n primitive groups, which are respectively allocated to each geometry processing pipeline within the graphics pipeline cluster. At this time, the graphics pipeline cluster layer includes n graphics pipeline clusters, namely graphics pipeline cluster 41 to graphics pipeline cluster 4n, wherein the first graphics pipeline cluster includes a primitive secondary distribution unit 721, at least two geometry processing pipelines 412, a merge arbitrator 413 and a blocker 414, ..., the nth graphics pipeline cluster includes a primitive secondary distribution unit 72n, at least two geometry processing pipelines 4n2, a merge arbitrator 4n3 and a blocker 4n4. At the same time, the output data of each graphics pipeline cluster can be stored in a corresponding cache unit, such as the first graphics pipeline cluster corresponds to cache unit 415, and the nth graphics pipeline cluster corresponds to cache unit 4n5; the geometry processing pipeline layer remains unchanged and includes multiple pixel processing pipelines (clusters) 46.

[0145] Through the above-mentioned embodiment, a two-stage parallel processing geometry pipeline is realized, which can not only flexibly configure the number of graphics pipeline clusters, but also improve the processing efficiency of the geometry pipeline. Compared with the parallel method of the geometry processing pipeline in the related art, the parallelism of the geometry processing pipeline can be further improved. At the same time, a two-stage parallel geometry processing pipeline is adopted. Compared with the one-stage parallel through the merge arbitration method, the blocker and the merge arbitrator will not need to process too many outputs of the geometry processing pipeline and will not become a performance bottleneck (too many pipeline output merges will cause pipeline disconnection and serialization effects); in addition, compared with the one-stage parallel through the front end of the pixel processing pipeline to merge multiple polygon lists, the pixel processor stage does not need to merge too many copies of the geometry processing pipeline output data and will not become a performance bottleneck. The embodiment of the present application can enhance the configurability of the geometry processing capability of the entire system and facilitate the configuration of clusters together with the pixel processing pipeline.

[0146] Based on the foregoing embodiments, an embodiment of the present application provides a geometry processing device, which includes the various units included and the various modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0147] FIG12 is a schematic diagram of the composition structure of a graphics processor provided by an embodiment of the present application. As shown in FIG12 , a graphics processor 1200 includes at least two graphics pipeline clusters 1210. The graphics pipeline clusters 1210 include a primitive distribution unit 1211, at least two geometry processing pipelines 1212, and a merging arbiter 1213.

[0148] The primitive distribution unit 1211 is configured to obtain primitive block data in the geometry data stream; the primitive block data includes a first tag for determining the order of the primitive block data in the geometry data stream;

[0149] The primitive distribution unit 1211 is further configured to divide the primitive block data into primitive group data and distribute the primitive group data to the geometry processing pipeline; the primitive group data includes a second tag for determining the order of the primitive group data in the primitive block data;

[0150] The geometry processing pipeline 1212 is configured to process the primitive group data to obtain a primitive group processing result, and output the primitive group processing result to the merging arbitrator;

[0151] The merge arbiter 1213 is configured to merge the processing results of the primitive group based on the second label to obtain geometric output data with a second order corresponding to the primitive block data; the geometric output data contains the first label corresponding to the primitive block data, and is configured to determine the first order between the geometric output data output by each of the graphics pipeline clusters.

[0152] In some embodiments, the primitive distribution unit is further configured to read primitive block data that the graphics pipeline cluster needs to process from the geometry data stream.

[0153] In some embodiments, the primitive distribution unit is further configured to read the geometric data stream; based on a preset primitive block acquisition strategy, discard the data in the geometric data stream that does not belong to the current graphics pipeline cluster processing, and obtain the primitive block data that the graphics pipeline cluster needs to process.

[0154] In some embodiments, the primitive distribution unit is further configured to determine the address segment of the primitive block data that the graphics pipeline cluster needs to process based on a preset primitive block acquisition strategy; and read the primitive block data that the graphics pipeline cluster needs to process in the geometry data stream based on the address segment.

[0155] In some embodiments, the graphics processor further includes a global distribution unit; the global distribution unit is configured to read the geometric data stream, determine the primitive block data that each of the graphics pipeline clusters needs to process from the geometric data stream based on a preset primitive block acquisition strategy, and distribute the primitive block data to each of the graphics pipeline clusters; accordingly, the primitive distribution unit is further configured to receive the primitive block data that the graphics pipeline cluster needs to process sent by the global distribution unit.

[0156] In some embodiments, the graphics processor further includes at least one pixel processing pipeline, and the graphics pipeline cluster is further configured to cache the geometric output data corresponding to the primitive block data to a cache unit corresponding to the graphics pipeline cluster based on the first tag; the pixel processing pipeline is configured to read the geometric output data corresponding to each of the graphics pipeline clusters in a first order from the cache units corresponding to the at least two graphics pipeline clusters based on the first tag.

[0157] In some embodiments, the graphics pipeline cluster further includes a blocker, which is configured to distribute the geometric output data corresponding to each of the graphics pipeline clusters in sequence based on the first label, and cache the data to the cache unit corresponding to the graphics pipeline cluster; wherein the blocker is further configured to determine the tile corresponding to each primitive in the geometric output data in sequence based on the second order; and write each primitive in the geometric output data into the polygon list of the corresponding tile in the cache unit corresponding to the graphics pipeline cluster.

[0158] In some embodiments, when there is at least one target element data belonging to a target tile in the geometric output data, the polygon list of the target tile includes a first label of the geometric output data and each target element data in the geometric output data belonging to the target tile; wherein the order of each target element data in the polygon list is the same as the second order.

[0159] In some embodiments, the pixel processing pipeline is further configured to traverse the first tag of the header of each polygon list corresponding to the target tile in the cache units corresponding to the at least two graphics pipeline clusters respectively, and take out the metadata corresponding to the target first tag from the polygon list corresponding to the target first tag until the metadata does not exist in each polygon list; wherein, the target first tag is determined based on the order of the first tags in the headers of each polygon list.

[0160] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or units included in the device provided in the embodiments of the present application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0161] It should be noted that, in the embodiment of the present application, if the above-mentioned geometric processing method is implemented in the form of a software functional unit and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.

[0162] An embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0163] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.

[0164] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.

[0165] An embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK), etc.

[0166] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.

[0167] Figure 13 is a schematic diagram of the hardware entity of a computer device provided in an embodiment of the present application. As shown in Figure 13, the hardware entity of the computer device 1300 includes: a processor 1301 and a memory 1302, wherein the memory 1302 stores a computer program that can be run on the processor 1301, and when the processor 1301 executes the program, the steps in the method of any of the above embodiments are implemented.

[0168] The memory 1302 stores computer programs that can be run on the processor. The memory 1302 is configured to store instructions and applications executable by the processor 1301. It can also cache data to be processed or processed by the processor 1301 and each unit in the computer device 1300 (for example, image data, audio data, voice communication data and video communication data). It can be implemented through flash memory (FLASH) or random access memory (RAM).

[0169] When the processor 1301 executes the program, the steps of any of the above-mentioned geometry processing methods are implemented. The processor 1301 generally controls the overall operation of the computer device 1300.

[0170] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the geometric processing method of any of the above embodiments.

[0171] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0172] The processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device that implements the functions of the processor may also be other electronic devices, which are not specifically limited in the embodiments of the present application.

[0173] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0174] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0175] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0176] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0177] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0178] In addition, the functional units in the embodiments of the present application can all be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units. It can be understood by those skilled in the art that all or part of the steps of the above-mentioned method embodiments can be completed by hardware related to program instructions, and the above-mentioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiments; and the above-mentioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), magnetic disks or optical disks.

[0179] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0180] The above is only an implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application. Industrial Applicability

[0181] The present application discloses a geometry processing method, apparatus, device, storage medium, and program product, wherein the geometry processing method includes: obtaining primitive block data in a geometry data stream through a primitive distribution unit; the primitive block data includes a first tag for determining the order of the primitive block data in the geometry data stream; the primitive distribution unit further divides the primitive block data into primitive group data, and distributes the primitive group data to a geometry processing pipeline; the geometry processing pipeline processes the primitive group data to obtain a primitive group processing result, and outputs the primitive group processing result to a merging arbitrator; the merging arbitrator merges the primitive group processing results based on the second tag to obtain geometry output data corresponding to the primitive block data. The above method can improve the geometry processing performance of a graphics processor.

Claims

1. A geometry processing method, applied to a graphics processor, wherein the graphics processor comprises at least two graphics pipeline clusters, wherein the graphics pipeline clusters comprise a primitive distribution unit, at least two geometry processing pipelines and a merging arbitrator, wherein the geometry processing method comprises: Acquire the primitive block data in the geometric data stream through the primitive distribution unit; The primitive block data includes a first tag for determining the order of the primitive block data in the geometry data stream; The primitive block data is divided into primitive group data by the primitive distribution unit, and the primitive group data is distributed to the geometry processing pipeline; the primitive group data includes a second tag for determining the order of the primitive group data in the primitive block data; Processing the primitive group data through the geometry processing pipeline to obtain a primitive group processing result, and outputting the primitive group processing result to the merging arbitrator; By means of the merging arbitrator, the processing results of the primitive groups are merged based on the second label to obtain geometric output data with a second order corresponding to the primitive block data; The geometry output data includes the first tag corresponding to the primitive block data, which is used to determine a first order between the geometry output data output by each of the graphics pipeline clusters.

2. The geometric processing method according to claim 1, wherein: The obtaining of primitive block data in the geometric data stream by the primitive distribution unit includes: The primitive block data to be processed by the graphics pipeline cluster is read from the geometry data stream through the primitive distribution unit.

3. The geometric processing method according to claim 2, wherein: The step of reading the primitive block data to be processed by the graphics pipeline cluster from the geometry data stream by the primitive distribution unit includes: Reading the geometric data stream through the primitive distribution unit; Based on a preset primitive block acquisition strategy, data in the geometric data stream that does not belong to the current graphics pipeline cluster for processing is discarded to obtain primitive block data that the graphics pipeline cluster needs to process.

4. The geometric processing method according to claim 2, wherein: The step of reading the primitive block data to be processed by the graphics pipeline cluster from the geometry data stream by the primitive distribution unit includes: Based on a preset primitive block acquisition strategy, determining an address segment of primitive block data that needs to be processed by the graphics pipeline cluster; The primitive block data required to be processed by the graphics pipeline cluster is read from the geometry data stream based on the address segment.

5. The geometric processing method according to claim 1, wherein: The graphics processor also includes a global distribution unit; The geometry processing method further comprises: reading the geometry data stream through the global distribution unit, determining the primitive block data to be processed by each of the graphics pipeline clusters from the geometry data stream based on a preset primitive block acquisition strategy, and distributing the primitive block data to each of the graphics pipeline clusters; Correspondingly, the obtaining of primitive block data in the geometry data stream through the primitive distribution unit includes: receiving, through the primitive distribution unit, primitive block data that needs to be processed by the graphics pipeline cluster and is sent by the global distribution unit.

6. The geometric processing method according to any one of claims 1 to 5, wherein: The graphics processor further includes at least one pixel processing pipeline, and the geometry processing method further includes: caching the geometric output data corresponding to the primitive block data to a cache unit corresponding to the graphics pipeline cluster based on the first tag through the graphics pipeline cluster; The pixel processing pipeline sequentially reads the geometric output data corresponding to each of the graphics pipeline clusters from the cache units corresponding to the at least two graphics pipeline clusters respectively based on the first tag in a first order.

7. The geometric processing method according to claim 6, wherein: The graphics pipeline cluster further includes a block splitter, and caching the geometric output data corresponding to the primitive block data to a cache unit corresponding to the graphics pipeline cluster based on the first tag through the graphics pipeline cluster includes: Distributing the geometric output data corresponding to each of the graphics pipeline clusters in sequence based on the first tag by the block divider, and caching the data to a cache unit corresponding to the graphics pipeline cluster; Among them, the geometric output data corresponding to each of the graphics pipeline clusters is distributed and cached in the cache unit corresponding to the graphics pipeline cluster, including: determining the blocks corresponding to each primitive in the geometric output data in sequence based on the second order by the block divider; and writing each primitive in the geometric output data into the polygon list of the corresponding block in the cache unit corresponding to the graphics pipeline cluster.

8. The geometric processing method according to claim 7, wherein: In the case where there is at least one target primitive element belonging to a target tile in the geometry output data, the polygon list of the target tile comprises a first tag of the geometry output data and each target primitive element belonging to the target tile in the geometry output data; The order of each of the target graphic elements in the polygon list is the same as the second order.

9. The geometric processing method according to claim 7, wherein: The step of sequentially reading, by the pixel processing pipeline and based on the first tag, from cache units corresponding to the at least two graphics pipeline clusters in accordance with a first order, geometric output data corresponding to each of the graphics pipeline clusters comprises: In the cache units respectively corresponding to the at least two graphics pipeline clusters, the first label in each polygon list corresponding to the target tile is traversed through the pixel processing pipeline, and the primitive data corresponding to the target first label is taken out from the polygon list corresponding to the target first label until the primitive data does not exist in each polygon list; The target first label is determined based on the sequence of the first labels in the headers of the polygon lists.

10. A graphics processor, the graphics processor comprising at least two graphics pipeline clusters, the graphics pipeline clusters comprising a primitive distribution unit, at least two geometry processing pipelines and a merge arbitrator; wherein, The primitive distribution unit is configured to obtain primitive block data in a geometry data stream; the primitive block data includes a first tag for determining an order of the primitive block data in the geometry data stream; The primitive distribution unit is further configured to divide the primitive block data into primitive group data and distribute the primitive group data to the geometry processing pipeline; the primitive group data includes a second tag for determining the order of the primitive group data in the primitive block data; The geometry processing pipeline is configured to process the primitive group data to obtain a primitive group processing result, and output the primitive group processing result to the merging arbitrator; The merging arbitrator is configured to merge the processing results of the primitive groups based on the second tag to obtain geometric output data with a second order corresponding to the primitive block data; The geometry output data includes the first tag corresponding to the primitive block data, which is used to determine a first order between the geometry output data output by each of the graphics pipeline clusters.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 9 when executing the program.

12. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program or instructions, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Graph processing method and system

    CN115908102A

  • Graphics processor, operation method and machine readable storage medium

    CN116188241A

  • Control flow stitching for multi-core three-dimensional graphics rendering

    CN116894901A

  • Graphics processor and method, multi-core graphics processing system, electronic device and equipment

    CN117058288A

  • Geometry processing method and device, equipment and storage medium

    CN117252751A

Cited By

  • Chip element task processing method and device, equipment, storage medium and program product

    CN120823089A