Task allocation method and device

By reasonably allocating vertex and primitive processing tasks in the graphics processor system and utilizing mapping relationships and sorting methods, the problem of unbalanced task distribution in the graphics processor system is solved, and load balancing and efficient graphics display are achieved.

CN120704812APending Publication Date: 2025-09-26HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410361936.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-09-26

Smart Images

  • Figure CN120704812A_ABST
    Figure CN120704812A_ABST
Patent Text Reader

Abstract

The invention provides a task allocation method and device. The method comprises the following steps: receiving a vertex processing task; and allocating a primitive processing task for the target tile to the target graphics processor. The vertex processing task instructs the graphics processor Ni to perform vertex coloring processing on a target vertex to obtain a vertex position of the target vertex after the vertex coloring processing, and the target vertex is determined according to a vertex associated with the drawing call instruction Ci; the vertex processing task is distributed to the graphics processor Ni according to the vertex load of each graphics processor of the graphics processor system and the number of the target vertexes. The primitive processing task for the target tile indicates the target graphics processor to carry out fragmentation processing on the target primitive of which the bounding box is located in the coverage area of the target tile, the target tile is matched with the target graphics processor, and the vertexes forming the target primitive are subjected to vertex coloring processing. According to the task allocation method, the parallel computing resources of a plurality of graphics processors of the graphics processor system can be efficiently utilized to complete bin division.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the fields of graphics processing technology and graphics processor technology, and more particularly to a task allocation method and device. Background Art

[0002] A graphics processing unit (GPU), also known as a display core, visual processor, or display chip, is a microprocessor used to perform image and graphics-related calculations on personal computers, workstations, game consoles, and some mobile devices (such as tablets and smartphones).

[0003] With the increasing demand for GPU computing power and the increasing use of multiple GPUs (also known as multi-cluster GPUs) in cloud scenarios, GPUs face the challenge of properly allocating tasks to achieve linear performance gains. How to properly allocate processing tasks during the binning phase to multiple GPUs in a GPU system, effectively utilizing the parallel computing resources of the GPUs, has become a pressing technical challenge. Summary of the Invention

[0004] The present application provides a task allocation method, a task allocation device, an electronic device, a chip module, a graphics processor system, a computer-readable storage medium, and a computer program product.

[0005] In a first aspect, the present application relates to a task allocation method, which is applied to any graphics processor N of a graphics processor system. i The task allocation method includes: receiving a vertex processing task; and allocating a primitive processing task for a target tile to a target graphics processor.

[0006] Vertex processing task instructs GPU N i Perform vertex shading on the target vertex to obtain the vertex position of the target vertex after vertex shading. The target vertex is based on the draw call instruction C i The vertex processing task is assigned to the graphics processor N according to the vertex load of each graphics processor in the graphics processor system and the number of target vertices. i of.

[0007] The primitive processing task for the target tile instructs the target GPU to perform tile processing on the target primitive whose bounding box is within the coverage area of ​​the target tile. The target tile is matched with the target GPU, and the vertices constituting the target primitive are subjected to vertex shading processing.

[0008] According to the task allocation method of the present application, since the bounding box of the target primitive is located within the coverage of the target tile, and the vertices constituting the target primitive have undergone the first stage of vertex shading processing, for example, after the vertex has undergone the first stage of vertex shading processing, the position of the vertex and the position of the target primitive composed of the vertex may change, the bounding box of the target primitive can fuzzily represent the position of the target primitive, and the bounding box of the target primitive is located within the coverage of the target tile, and the target primitive matches the target graphics processor, so the task allocation method of the present application can fuzzily allocate the target primitive to the corresponding target tile according to the position of the target primitive composed of the vertices obtained by the first stage of vertex shading processing (the position of the target primitive is represented by the bounding box of the target primitive), and then allocate it to the target graphics processor corresponding to the target tile. Since the target primitive that has undergone vertex shading processing has been allocated to the corresponding target tile and the target graphics processor corresponding to the target tile according to the fuzzy position represented by the bounding box, the target graphics processor can accurately perform fragmentation processing after receiving the primitive processing task for the target tile.

[0009] In a possible example, a target tile and a target graphics processor have a mapping relationship, and the mapping relationship is determined by a row number and a column number of the target tile relative to a frame image.

[0010] Exemplarily, the mapping relationship between the target tile and the target GPU can be represented by a function.

[0011] For example, some embodiments sequentially allocate multiple tiles to multiple GPUs of a graphics processor in a "Z" shape. For a cylindrical graphic shape, for example, this embodiment causes the tiles to be centrally allocated to a specific GPU, which results in an unbalanced tile load between the various GPUs of the graphics processor system. According to the task allocation method of the present application, a mapping relationship is established between the target tile and the target GPU, and the mapping relationship is determined by the row and column numbers of the target tile relative to the frame image. This can disrupt the regular allocation of tiles, allowing the various GPUs of the graphics processor system to be more randomly allocated to tiles, and making the tile loads of the various GPUs of the graphics processor system more balanced.

[0012] In one possible example, the communication method further includes: receiving primitive processing tasks for X allocated tiles; and sorting the allocated primitives corresponding to all the allocated tiles so that an order in which the allocated primitives corresponding to all the allocated tiles are processed in tiles is consistent with an order in which draw call instructions are received by the graphics processor system.

[0013] The primitive processing task for the X allocated tiles is assigned by the GPU of the GPU system to GPU N iThe primitive processing task for X allocated tiles instructs the graphics processor N i Each allocated primitive whose bounding box is within the coverage area of ​​the allocated tile is sliced, and any allocated tile and graphics processor N i Matching, the vertices that make up the assigned primitive are processed by vertex shading.

[0014] The drawing call instructions sent by the CPU have a timing relationship (i.e., a sequence). If the timing relationship changes, it may cause errors in the graphics display. Since the task allocation method of the present application uses the various graphics processors of the graphics processor system to perform the second stage of primitive processing tasks on the tiles corresponding to the primitives composed of vertices processed by the first stage vertex shading in parallel, it involves allocation among multiple graphics processors. For any graphics processor N i For example, the graphics processor N may receive primitive processing tasks for the assigned tiles from itself and / or from other graphics processors. The allocation of primitive processing tasks for the assigned tiles between multiple graphics processors may disrupt the graphics processor N. i The order of primitive processing tasks received for the X allocated tiles can be made consistent with the order of draw call instructions received by the graphics processor system by sorting the allocated primitives corresponding to the X allocated tiles according to the task allocation method of the present application, thereby ensuring that graphics display is correct.

[0015] In a possible example, the vertex processing task further instructs to perform culling processing on primitives composed of vertices processed by vertex shading.

[0016] According to the task allocation method of the present application, by performing vertex shading on the vertices, the vertex positions after vertex shading can be obtained, and then the graphics primitives composed of the vertices after vertex shading can be obtained. For example, in order to reduce the graphics primitive processing tasks in the second stage and improve the processing efficiency of the binning stage, the graphics primitives composed of the vertices that have been vertex shading can be culled.

[0017] In a possible example, the graphics processor system includes a master graphics processor and at least one slave graphics processor, wherein the graphics processor N i In the case of a main graphics processor, the communication method further includes: receiving a drawing call instruction C i ; According to the drawing call instruction C i Vertex processing tasks are assigned to graphics processors of the graphics processor system based on the associated number of vertices and the vertex load of each graphics processor of the graphics processor system.

[0018] Vertex processing task instructions for drawing call instruction C i The associated vertices are vertex shaded.

[0019] The operation performed by the graphics processor of the graphics processor system is instructed by the draw call instruction. In order to uniformly allocate the vertices associated with the draw call instruction, the task allocation method of the present application receives the draw call instruction C from the main graphics processor. i And according to the drawing call instruction C i The number of associated vertices and the vertex load of each GPU in the GPU system can be used to uniformly distribute vertex processing tasks to the GPU system at the granularity of draw call instructions, so that the vertex loads of each GPU in the GPU system after being assigned vertex processing tasks tend to be consistent, that is, to achieve vertex load balancing among the GPUs in the GPU system. Each GPU can then perform subsequent operations such as vertex shading in parallel more synchronously.

[0020] In one possible example, according to the draw call instruction C i The number of associated vertices and the vertex load of each graphics processor of the graphics processor system, and the distribution of vertex processing tasks to the graphics processor of the graphics processor system include: according to the number of accumulated vertices in the current batch, the drawing call instruction C i The number of associated vertices and the batch vertex number threshold determine the vertices associated with the current batch. The cumulative number of vertices in the current batch is obtained by accumulating the number of vertices associated with the historical draw call instructions of the current batch. The reception time of the historical draw call instructions is at the draw call instruction C i Before the receiving moment; according to the vertex load of each graphics processor of the graphics processor system and the number of vertices associated with the current batch, a vertex processing task is assigned to the graphics processor of the graphics processor system, and the vertex processing task indicates vertex shading processing of the vertices associated with the current batch.

[0021] According to the number of accumulated vertices in the current batch, the drawing call instruction C i The number of associated vertices and the threshold number of vertices in the batch, determining the vertices associated with the current batch can also be understood as: drawing call instructions C i The associated vertices are assigned to the current batch, so that the cumulative number of vertices in the current batch and the assigned draw call instructions C i The sum of the number of associated vertices approaches the vertex number threshold, which makes the difference in the number of associated vertices in each batch smaller.

[0022] The graphics processor to which the current batch is assigned is determined based on the vertex load of each graphics processor in the graphics processor system. When vertex tasks are assigned on a batch basis, for example, vertices associated with the current batch can be assigned to graphics processors with smaller vertex loads in the graphics processor system, thereby making the vertex load of each graphics processor in the graphics processor system more balanced.

[0023] In one possible example, based on the cumulative number of vertices in the current batch, the draw call instruction C i The number of associated vertices and the vertex number threshold of the batch, determining the vertices associated with the current batch includes: the number of accumulated vertices in the current batch and the draw call instruction C i When the sum of the number of associated vertices is less than or equal to the vertex number threshold, the cumulative vertices of the current batch and the draw call instruction C i The associated vertex updates the cumulative vertex of the current batch; the cumulative number of vertices in the current batch is equal to the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is less than or equal to the vertex number threshold, the draw call instruction C i The associated vertices are assigned to the downstream batch after the current batch; the cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is greater than the vertex number threshold, the draw call instruction C is sent according to the vertex number threshold. i The associated vertices are assigned to the current batch and at least one downstream batch after the current batch.

[0024] In summary, the task allocation method of the present application can balance the vertex load of each graphics processor of the graphics processor system through vertex processing task allocation. Each graphics processor can perform vertex shading processing on the target vertices with a consistent number of allocated tasks in parallel. Through primitive processing task allocation, the target primitives composed of vertices after vertex shading processing can be fuzzy allocated to the target graphics processor corresponding to the target tile. The target graphics processor can accurately perform tile processing after receiving the primitive processing task for the target tile, thereby making full use of the parallel computing resources of each graphics processor of the graphics processor system to perform balanced and efficient binning.

[0025] In a second aspect, the present application relates to a task allocation device, which is applied to any graphics processor N of a graphics processor system. i The task allocation device includes: a vertex processing task receiving module and a primitive processing task allocation module.

[0026] Vertex processing task receiving module, used to receive vertex processing tasks, vertex processing tasks indicate graphics processor N i Perform vertex shading on the target vertex to obtain the vertex position of the target vertex after vertex shading. The target vertex is based on the draw call instruction C i The vertex processing task is assigned to the graphics processor N according to the vertex load of each graphics processor in the graphics processor system and the number of target vertices. i of.

[0027] The primitive processing task allocation module is used to allocate primitive processing tasks for target tiles to the target graphics processor. The primitive processing tasks for the target tiles instruct the target graphics processor to perform tile processing on the target primitives whose bounding boxes are located within the coverage area of ​​the target tiles. The target tiles are matched with the target graphics processors, and the vertices that make up the target primitives are subjected to vertex shading processing.

[0028] According to the task allocation method of the present application, when vertices are allocated at batch granularity, for example, the number of vertices associated with some draw call instructions is small, so a current batch can accumulate vertices associated with multiple draw call instructions. For example, the number of vertices associated with some draw call instructions is large, so the vertices associated with the draw call instructions can be allocated to the current batch and the downstream batch. Therefore, according to the task allocation method of the present application, according to the cumulative number of vertices in the current batch and the number of draw call instructions C, the vertices associated with the draw call instructions C can be allocated to the current batch and the downstream batch. i The relative size relationship between the sum of the number of associated vertices and the vertex number threshold and the draw call instruction C i The relative size relationship between the number of associated vertices and the vertex number threshold can accurately and efficiently allocate vertices with the same vertex number threshold to the current batch. In addition, for the above 2) and 3) cases, you can also try to make a draw call instruction C i Assigning related vertices to the same batch can more reasonably and efficiently assign the same number of vertices as the vertex threshold to the current batch.

[0029] In a possible example, a target tile and a target graphics processor have a mapping relationship, and the mapping relationship is determined by a row number and a column number of the target tile relative to a frame image.

[0030] In a possible example, the communication device further includes: a graphics primitive processing task receiving module and a sorting module.

[0031] The primitive processing task receiving module is used to receive primitive processing tasks for X allocated tiles, where the primitive processing tasks for X allocated tiles are allocated by the graphics processor of the graphics processor system to the graphics processor N. iThe primitive processing task for X allocated tiles instructs the graphics processor N i Each allocated primitive whose bounding box is within the coverage area of ​​the allocated tile is sliced, and any allocated tile and graphics processor N i Matching, the vertices that make up the assigned primitive are processed by vertex shading.

[0032] The sorting module is used to sort the allocated primitives corresponding to all the allocated tiles so that the order of slicing the allocated primitives corresponding to all the allocated tiles is consistent with the order of draw call instructions received by the graphics processor system.

[0033] In a possible example, the vertex processing task further instructs to perform culling processing on primitives composed of vertices processed by vertex shading.

[0034] In a possible example, the graphics processor system includes a master graphics processor and at least one slave graphics processor, wherein the graphics processor N i In the case of a master graphics processor, the communication device further includes: a drawing call instruction receiving module and a vertex processing task allocating module.

[0035] The drawing call instruction receiving module is used to receive the drawing call instruction C i .

[0036] Vertex processing task allocation module is used to allocate the vertex processing task according to the drawing call instruction C i The number of associated vertices and the vertex load of each graphics processor of the graphics processor system are used to assign vertex processing tasks to the graphics processors of the graphics processor system. The vertex processing tasks indicate the response to the draw call instruction C i The associated vertices are vertex shaded.

[0037] In a possible example, the vertex processing task allocation module includes: a vertex determination submodule and a vertex processing task allocation submodule.

[0038] Vertex determination submodule is used to determine the number of accumulated vertices in the current batch and the drawing call instruction C i The number of associated vertices and the batch vertex number threshold determine the vertices associated with the current batch. The cumulative number of vertices in the current batch is obtained by accumulating the number of vertices associated with the historical draw call instructions of the current batch. The reception time of the historical draw call instructions is at the draw call instruction C i Before the reception time.

[0039] The vertex processing task allocation submodule allocates vertex processing tasks to the graphics processors of the graphics processor system according to the vertex load of each graphics processor of the graphics processor system and the number of vertices associated with the current batch, and the vertex processing tasks indicate vertex shading processing for the vertices associated with the current batch.

[0040] In one possible example, the vertex determination submodule is used to: i When the sum of the number of associated vertices is less than or equal to the vertex number threshold, the cumulative vertices of the current batch and the draw call instruction C i The associated vertex updates the cumulative vertex of the current batch; the cumulative number of vertices in the current batch is equal to the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is less than or equal to the vertex number threshold, the draw call instruction C i The associated vertices are assigned to the downstream batch after the current batch; the cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is greater than the vertex number threshold, the draw call instruction C is sent according to the vertex number threshold. i The associated vertices are assigned to the current batch and at least one downstream batch after the current batch.

[0041] In a third aspect, the present application relates to a chip module, which includes a transceiver component and a chip, and the chip can be used to execute the task allocation method of the first aspect.

[0042] In a fourth aspect, the present application relates to an electronic device comprising a processor and an interface circuit, wherein the interface circuit is used to receive signals from electronic devices other than the electronic device and transmit them to the processor or to send signals from the processor to electronic devices other than the electronic device, and the processor is used to implement the task allocation method of the first aspect through logic circuits or execution code commands.

[0043] Exemplarily, the processor is a graphics processor.

[0044] Exemplarily, the electronic device is a chip.

[0045] In a fifth aspect, the present application relates to a graphics processor system, comprising a master graphics processor and at least one slave graphics processor.

[0046] In a sixth aspect, the present application relates to a computer-readable storage medium storing computer instructions, wherein when the computer instructions are executed, the computer executes the task allocation method of the first aspect. In some exemplary embodiments, the computer-readable storage medium is a non-transitory storage medium.

[0047] In a seventh aspect, the present application relates to a computer program product, comprising a computer program stored on a readable storage medium, which enables a computer to implement the task allocation method of the first aspect when executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The following is an introduction to the drawings used in the embodiments of this application.

[0049] Figure 1 A schematic diagram schematically illustrates a task allocation method according to an embodiment of the present disclosure;

[0050] Figure 2 A schematic diagram schematically illustrates determining vertices associated with a current batch according to a task allocation method according to an embodiment of the present disclosure;

[0051] Figure 3 Schematically shows a block diagram of a task allocation device according to an embodiment of the present disclosure;

[0052] Figure 4A An electronic device according to an embodiment is schematically shown;

[0053] Figure 4B A schematic block diagram of an electronic device according to another embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0055] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0056] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first target object" and "second target object" are used to distinguish different objects, rather than to describe a specific order of objects.

[0057] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0058] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0059] The following describes in detail the background of the task allocation method according to the embodiment of the present disclosure.

[0060] A graphics processing unit (GPU), also known as a display core, visual processor, or display chip, is a microprocessor used to perform image and graphics-related calculations on personal computers, workstations, game consoles, and some mobile devices (such as tablets and smartphones).

[0061] In the field of graphics rendering technology, tile-based rendering (TBR) pipelines are widely used. Specifically, the GPU contains an ultra-fast on-chip memory, which functions similarly to cache memory, but has a relatively small capacity. TBR accelerates rendering by migrating the frequent interactions between immediate mode rendering (IMR) and system memory to this high-speed on-chip buffer, bypassing the often slow and obstructive interactions with system memory. Because the on-chip memory capacity is small and cannot accommodate the data of the frame buffer of an entire screen, in order to fully utilize the high speed advantage of on-chip memory, the frame buffer can be divided into non-overlapping tiles, so that each tile can be stored in static random access memory (SRAM) closer to the GPU. The GPU can access the frame buffer on the SRAM tile by tile, and after accessing the entire frame, the entire frame is transferred back to the dynamic random access memory (DRAM), thus avoiding slow interaction with system memory and reducing bandwidth.

[0062] In order to implement TBR-based graphics rendering, vertex shading processing must be completed first to determine the primitive position. Based on the primitive position, primitives located in different tiles can be filtered, and then subsequent frame shading processing can be performed tile by tile. The above stage is called binning.

[0063] With the increasing demand for GPU computing power and the increasing use of multiple GPUs (also known as multi-cluster GPUs) in cloud scenarios, GPUs face the challenge of properly allocating tasks to achieve linear performance gains. How to properly allocate processing tasks during the binning phase to multiple GPUs in a GPU system, effectively utilizing the parallel computing resources of the GPUs, has become a pressing technical challenge.

[0064] The task processing method will be described below in conjunction with a system architecture (ie, a graphics processor system) to which the task allocation method according to an embodiment of the present disclosure is applied.

[0065] Figure 1 The following schematically illustrates a task allocation method according to an embodiment of the present disclosure. Figure 1The figure also schematically shows a specific example in which the graphics processor system includes one master graphics processor (master GPU) and n slave graphics processors from GPU1 to GPUn.

[0066] Combine Figure 1 The binning stage includes a first stage and a second stage. The first stage is used to perform vertex shading processing on the vertices that make up the primitives, and the second stage is used to perform fragmentation processing on the primitives. According to the task allocation method of the embodiment of the present disclosure, the tasks of the first stage (referred to as vertex processing tasks in the embodiment of the present disclosure) and the tasks of the second stage (referred to as primitive processing tasks in the embodiment of the present disclosure) are allocated to the graphics processor of the graphics processor system.

[0067] Combine Figure 1 , the binning stage of the task allocation method of the embodiment of the present disclosure is started by the central processing unit sending a draw call instruction to the main graphics processor. The draw call instruction is an instruction of the Application Programming Interface (API) for graphics drawing sent by the engine to the GPU. The draw call instruction can, for example, instruct the GPU to draw the specified primitives and the drawing method. Vertices can constitute primitives, and primitives include, for example, points, line segments, and triangles. For example, the primitives can be represented by the index of the vertices that constitute the primitives. For example, in the binning stage, the draw call instruction instructs the primitives associated with the draw call instruction to perform vertex shading processing, which can also be understood as the draw call instruction instructing the associated vertices that constitute the primitives to perform vertex shading processing.

[0068] The first stage of task allocation in the embodiment of the present disclosure may be performed by a main graphics processor of the graphics processor system.

[0069] The main graphics processor can receive all the drawing call instructions sent by the CPU. The following will be any drawing call instruction C i Take this as an example to illustrate.

[0070] exist Figure 1 In the example, the main graphics processor receives the draw call instruction C i , and the vertex processing task distribution module set in the main graphics processor executes the drawing call instruction C i The vertex processing task is assigned to the graphics processor of the graphics processor system based on the number of associated vertices and the vertex load of each graphics processor of the graphics processor system. The graphics processor that assigns the vertex processing task is the master graphics processor, and the graphics processor to which the vertex processing task is assigned can be the master graphics processor or the slave graphics processor.

[0071] Vertex processing task instructions for drawing call instruction C iThe associated vertices are vertex shaded.

[0072] The vertex load of a GPU can be characterized by the number of vertices that are assigned to the GPU but have not yet been processed.

[0073] For example, the drawing call instruction C i The associated vertices are assigned to the GPU with the smallest vertex load (ie, the smallest number of vertices) in the GPU system.

[0074] The operation executed by the graphics processor of the graphics processor system is instructed by the draw call instruction. In order to uniformly allocate the vertices associated with the draw call instruction, the task allocation method of the embodiment of the present disclosure receives the draw call instruction C from the main graphics processor. i And according to the drawing call instruction C i The number of associated vertices and the vertex load of each GPU in the GPU system can be used to uniformly distribute vertex processing tasks to the GPU system at the granularity of draw call instructions, so that the vertex loads of each GPU in the GPU system after being assigned vertex processing tasks tend to be consistent, that is, to achieve vertex load balancing among the GPUs in the GPU system. Each GPU can then perform subsequent operations such as vertex shading in parallel more synchronously.

[0075] According to another embodiment of the present disclosure, a task allocation method can be implemented by, for example, using the following embodiment according to the draw call instruction C i A specific example of assigning vertex processing tasks to the graphics processors of the graphics processor system based on the number of associated vertices and the vertex load of each graphics processor of the graphics processor system: according to the number of accumulated vertices in the current batch, the draw call instruction C i The number of associated vertices and a threshold value of the number of vertices in the batch are used to determine vertices associated with the current batch. A vertex processing task is assigned to a graphics processor of the graphics processor system based on a vertex load of each graphics processor of the graphics processor system and the number of vertices associated with the current batch. The vertex processing task indicates that vertex shading processing is performed on the vertices associated with the current batch.

[0076] The vertex count threshold for each batch can be the same. This is explained below using the same vertex count threshold for each batch as an example. Of course, the vertex count threshold for each batch can also be different. In this case, the vertex count thresholds for each batch can be similar.

[0077] According to the number of accumulated vertices in the current batch, the drawing call instruction C iThe number of associated vertices and the threshold number of vertices in the batch, determining the vertices associated with the current batch can also be understood as: drawing call instructions C i The associated vertices are assigned to the current batch, so that the cumulative number of vertices in the current batch and the assigned draw call instructions C i The sum of the number of associated vertices approaches the vertex number threshold, which makes the difference in the number of associated vertices in each batch smaller.

[0078] In addition, the graphics processors to which the current batch is assigned are determined based on the vertex load of each graphics processor in the graphics processor system. When vertex tasks are assigned based on batches, the vertex load of each graphics processor can also be understood as the batch load. For example, the vertices associated with the current batch can be assigned to graphics processors with smaller vertex loads (i.e., smaller batch loads) in the graphics processor system, thereby making the vertex load of each graphics processor in the graphics processor system more balanced.

[0079] According to the task allocation method of the embodiment of the present disclosure, for example, the following embodiments can be used to implement the task allocation based on the number of accumulated vertices in the current batch, the draw call instruction C i The number of associated vertices and the vertex number threshold for the batch determine the specific examples of vertices associated with the current batch:

[0080] 1) The cumulative number of vertices in the current batch and the draw call instruction C i When the sum of the number of associated vertices is less than or equal to the vertex number threshold, the cumulative vertices of the current batch and the draw call instruction C i The associated vertices update the accumulated vertices of the current batch.

[0081] For example, the number of accumulated vertices in the current batch is 100, and the draw call instruction C i The number of associated vertices is 100, and the vertex threshold is 500. The current batch of accumulated vertices and draw call instructions C i The associated vertex updates the cumulative vertex of the current batch, so that the number of cumulative vertices in the updated current batch is 200, which includes the draw call instruction C i Associated vertices.

[0082] 2) The cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is less than or equal to the vertex number threshold, the draw call instruction C i The associated vertices are assigned to the downstream batch after the current batch.

[0083] For example, the current batchi The cumulative number of vertices is 400, and the draw call instruction C i The number of associated vertices is 200, the vertex threshold is 500, and the current batch is i The cumulative number of vertices is 400 and the draw call instruction C i The sum of the associated vertices (200) is 600, which is greater than the vertex threshold of 500 and the draw call instruction C i The number of associated vertices 200 is less than the vertex number threshold of 500, and the draw call instruction C i The associated vertices are assigned to the current batch i Subsequent downstream batches i+1 , that is, the current batch i The number of associated vertices remains at 400, and the draw call instruction C i Associated vertices are assigned to downstream batches i+1 It is understandable that the current draw call instruction C i Associated vertices are assigned to downstream batches i+1 , if the next draw call instruction C i The next draw call instruction C i+1 When the associated vertices are assigned, batch i+1 For the current batch.

[0084] 3) The cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is greater than the vertex number threshold, the draw call instruction C is sent according to the vertex number threshold. i The associated vertices are assigned to the current batch and at least one downstream batch after the current batch.

[0085] For example, the current batch i The cumulative number of vertices is 450, and the draw call instruction C i The number of associated vertices is 600, the vertex threshold is 500, and the current batch is i The cumulative number of vertices is 450 and the draw call instruction C i The sum of the associated vertices (600) is 1050, which is greater than the vertex threshold of 500 and the draw call instruction C i The number of associated vertices 600 is greater than the vertex number threshold 500, and the draw call instruction C i The associated vertices are assigned to the current batch iAnd at least one downstream batch after the current batch. That is, the drawing call instruction C i Associated vertices to the current batch i Allocate 50 so that the current batch i The number of associated vertices is 500, and the draw call instruction C i Associated vertices to the downstream batch i+1 Allocate 500 so that the downstream batch i+1 The number of associated vertices is 500, and the draw call instruction C i If there are 50 associated vertices left, continue to batch them downstream. i+2 Allocate 50.

[0086] It should be noted that in the above example, the cumulative number of vertices in the current batch and the downstream batch is iteratively accumulated by the number of vertices associated with the assigned draw call instructions. Each batch can be marked with a unique identifier.

[0087] According to the task allocation method of the embodiment of the present disclosure, when vertices are allocated at a batch granularity, for example, the number of vertices associated with some draw call instructions is small, so a current batch can accumulate vertices associated with multiple draw call instructions. For example, the number of vertices associated with some draw call instructions is large, so the vertices associated with the draw call instructions can be allocated to the current batch and the downstream batch. Therefore, according to the task allocation method of the embodiment of the present disclosure, based on the cumulative number of vertices in the current batch and the number of draw call instructions C, the vertices associated with the draw call instructions C can be allocated to the current batch and the downstream batch. i The relative size relationship between the sum of the number of associated vertices and the vertex number threshold and the draw call instruction C i The relative size relationship between the number of associated vertices and the vertex number threshold can accurately and efficiently allocate vertices with the same vertex number threshold to the current batch. In addition, for the above 2) and 3) cases, you can also try to make a draw call instruction C i Assigning related vertices to the same batch can more reasonably and efficiently assign the same number of vertices as the vertex threshold to the current batch.

[0088] Figure 2 The following schematic diagram shows a method for determining vertices associated with a current batch according to an embodiment of the present disclosure. For details, please refer to the description of the above embodiment and will not be repeated here.

[0089] In summary, the above embodiments can realize the task allocation of the first stage of the warehouse division phase.

[0090] Combine Figure 1 Schematic diagram, after the first stage of task allocation, any graphics processor N of the graphics processor system can be assignedi Assign vertex processing tasks. Graphics processor N to which vertex processing tasks are assigned i Can be a master graphics processor or a slave graphics processor.

[0091] Vertex processing task instructs GPU N i Perform vertex shading processing on the target vertex to obtain a vertex position of the target vertex after the vertex shading processing.

[0092] In combination with the above embodiment, the vertex processing task is allocated to the graphics processor N according to the vertex load of each graphics processor of the graphics processor system and the number of target vertices. i The target vertex is based on the draw call instruction C i The associated vertex is determined. For example, the target vertex can be the draw call instruction C i The associated vertex and the target vertex may also be the associated vertices of the current batch determined by the above embodiment.

[0093] For example, a vertex shader can be used to perform vertex shading. Vertex shading can be used for coordinate conversion, so that each vertex associated with a vertex processing task obtains the vertex position after vertex shading through vertex shading. Since the actual drawing call instruction instructs to perform vertex shading on the vertices that make up the primitives in the bin phase, the vertex processing task instructs to perform vertex shading on the target vertex, and the target vertex with a changed vertex position can be obtained. It can also be understood that after vertex shading, the primitive composed of the target vertex completes the position change. At this point, the first stage of vertex shading in the bin phase is completed.

[0094] After the first phase, the second phase of the binning phase includes sharding, which assigns each tile to a corresponding GPU (a matching relationship between the tile and the GPU). Hereinafter, these matching tiles and GPUs are referred to as target tiles and target GPUs, respectively. It should be noted that a target tile only matches one target GPU, but a target GPU can match multiple target tiles.

[0095] Graphics Processor N i With tile T i For example, the two can be understood as target graphics processors N with matching relationship. i and the target tile T i . Still with graphics processor N i For example, in the first stage, the graphics processor N i The primitives composed of the target vertices processed are not necessarily located in tile T iTherefore, the task allocation method of the embodiment of the present disclosure is also composed of Figure 1 The primitive processing task assignment module assigns primitive processing tasks for target tiles to the target GPU before the second stage of tile processing. The primitive processing tasks for the target tiles instruct the target GPU to perform tile processing on target primitives whose bounding boxes are within the coverage area of ​​the target tiles.

[0096] The primitive processing task for the target tile instructs the target graphics processor to perform slicing processing on the target primitive whose bounding box is within the coverage area of ​​the target tile, which can be understood as: the primitive processing task for the target tile instructs the target graphics processor to perform slicing processing on the target primitive, and the bounding box of the target primitive is within the coverage area of ​​the target tile. Here, the bounding box of the target primitive is within the coverage area of ​​the target tile, which can be understood as at least a portion of the bounding box of the target primitive is within the coverage area of ​​the target tile, for example, the entire bounding box of the target tile is within the coverage area of ​​the target tile or a portion of the bounding box of the target tile is within the coverage area of ​​the target tile.

[0097] The primitive processing task for the target tile can be understood as a processing task for the target tile, and can also be understood as a processing task for the target primitive located within the coverage area of ​​the target tile. The processing task can, for example, be related operations for performing tiling processing on the target primitive. The related operations for tiling processing can include: primitive establishment (taking the primitive as a triangle as an example, i.e., triangle setup), raster processing, depth testing, and tiling. For example, tiling processing can be performed using a tiler.

[0098] According to the task allocation method of the embodiment of the present disclosure, since the bounding box of the target primitive is located within the coverage of the target tile, and the vertices constituting the target primitive have undergone the first stage of vertex shading processing, for example, after the vertex has undergone the first stage of vertex shading processing, the position of the vertex and the position of the target primitive composed of the vertices may change, the bounding box of the target primitive can fuzzily represent the position of the target primitive, and the bounding box of the target primitive is located within the coverage of the target tile, and the target primitive matches the target graphics processor, so the task allocation method of the embodiment of the present disclosure can fuzzily allocate the target primitive to the corresponding target tile based on the position of the target primitive composed of the vertices obtained by the first stage of vertex shading processing (the position of the target primitive is represented by the bounding box of the target primitive), and then allocate it to the target graphics processor corresponding to the target tile. Since the target primitive that has undergone vertex shading processing has been allocated to the corresponding target tile and the target graphics processor corresponding to the target tile based on the fuzzy position represented by the bounding box, the target graphics processor can accurately perform fragmentation processing after receiving the primitive processing task for the target tile.

[0099] In summary, the task allocation method of the embodiment of the present invention can balance the vertex load of each graphics processor of the graphics processor system through vertex processing task allocation. Each graphics processor can perform vertex shading processing on the target vertices with a consistent number of allocated targets in parallel. Through primitive processing task allocation, the target primitives composed of vertices after vertex shading processing can be fuzzy allocated to the target graphics processor corresponding to the target tile. The target graphics processor can accurately perform tile processing after receiving the primitive processing task for the target tile, thereby making full use of the parallel computing resources of each graphics processor of the graphics processor system to perform balanced and efficient binning.

[0100] According to a task allocation method according to yet another embodiment of the present disclosure, a target tile and a target graphics processor have a mapping relationship, and the mapping relationship is determined by a row number and a column number of the target tile relative to a frame image.

[0101] Exemplarily, the mapping relationship between the target tile and the target GPU can be represented by a function.

[0102] In TBR-based graphics rendering, each GPU performs subsequent frame shading operations at a tile granularity. Therefore, how to allocate tiles to balance the tile loads of the various GPUs in the GPU system is also an aspect of improving the efficiency of the GPU system.

[0103] For example, some implementations sequentially allocate multiple tiles to the various GPUs of a graphics processor in a "Z" shape. For a cylindrical graphic shape, for example, this implementation causes the tiles to be centrally allocated to a specific GPU, resulting in an unbalanced tile load across the various GPUs in the graphics processor system. According to the task allocation method of an embodiment of the present disclosure, a mapping relationship is established between the target tile and the target GPU, and this mapping relationship is determined by the row and column numbers of the target tile relative to the frame image. This can disrupt the regular allocation of tiles, resulting in a more balanced tile load across the various GPUs in the graphics processor system.

[0104] Combine Figure 1 Schematic diagram, through the above embodiment of "distributing the primitive processing task for the target tile to the target graphics processor" will be any graphics processor N i Allocate X primitive processing tasks for the assigned tile. X can be an integer greater than or equal to 1.

[0105] According to a task allocation method according to another embodiment of the present disclosure, for any graphics processor N i The method may further include receiving primitive processing tasks for X allocated tiles, and sorting the allocated primitives corresponding to all the allocated tiles so that the order in which the allocated primitives corresponding to all the allocated tiles are processed in slices is consistent with the order in which the draw call instructions are received by the graphics processor system.

[0106] For example, when the target vertices associated with the vertex processing tasks in the first stage are of batch granularity, the allocated primitives corresponding to all the allocated tiles can be sorted according to the batch order corresponding to each allocated primitive, so that the order in which the allocated primitives corresponding to all the allocated tiles are sliced ​​and processed is consistent with the order in which the drawing call instructions are received by the graphics processor system.

[0107] The primitive processing task for the X allocated tiles is assigned by the GPU of the GPU system to GPU N i The primitive processing task for X allocated tiles instructs the graphics processor N i Each tile whose bounding box is within the coverage area of ​​the assigned tile is fragmented. There is also the following situation: the bounding box of a tile is within the coverage area of ​​multiple tiles, and the multiple tiles here are also matched with different GPUs, so this tile will be assigned to multiple GPUs.

[0108] The vertices that make up the assigned primitives are processed by vertex shading.

[0109] The graphics processor N is instructed to process the primitives for the assigned tiles. iThe fragmentation processing of the allocated primitives whose bounding boxes are located within the coverage area of ​​the allocated tiles can be understood as: the primitive processing task for the allocated tiles instructs the graphics processor N i The allocated primitive is sliced, and the bounding box of the allocated primitive is located within the coverage area of ​​the allocated tile. Here, the bounding box of the allocated primitive is located within the coverage area of ​​the allocated tile, which can be understood as at least a portion of the bounding box of the allocated primitive is located within the coverage area of ​​the allocated tile, for example, the entire bounding box of the allocated tile is located within the coverage area of ​​the allocated tile, or a portion of the bounding box of the allocated tile is located within the coverage area of ​​the allocated tile.

[0110] The drawing call instructions sent by the CPU have a timing relationship (i.e., a sequence). If the timing relationship changes, graphics display errors may occur. Since the task allocation method of the embodiment of the present disclosure uses the various graphics processors of the graphics processor system to perform the second stage of primitive processing tasks on the tiles corresponding to the primitives composed of vertices processed by the first stage vertex shading in parallel, it involves allocation among multiple graphics processors. For any graphics processor N i For example, the graphics processor N may receive primitive processing tasks for the assigned tiles from itself and / or from other graphics processors. The allocation of primitive processing tasks for the assigned tiles between multiple graphics processors may disrupt the graphics processor N. i Regarding the order of primitive processing tasks received for the X allocated tiles, according to the task allocation method of an embodiment of the present disclosure, by sorting the allocated primitives corresponding to the X allocated tiles, the order in which the X allocated tiles are processed in tiles can be made consistent with the order in which draw call instructions are received by the graphics processor system, thereby ensuring correct graphics display.

[0111] According to a task allocation method according to yet another embodiment of the present disclosure, the vertex processing task further instructs to perform culling (Cull) on primitives composed of vertices that have undergone vertex shading processing.

[0112] By culling the primitives composed of vertices processed by vertex shading, a culled primitive can be obtained. The target tile is matched with the target GPU, and the bounding box of the culled primitive is within the coverage of the target tile.

[0113] According to the task allocation method of the embodiment of the present disclosure, by performing vertex shading on the vertices, the vertex positions after vertex shading can be obtained, and then the graphics primitives composed of the vertices after vertex shading can be obtained. For example, in order to reduce the graphics primitive processing tasks in the second stage and improve the processing efficiency of the binning stage, the graphics primitives composed of the vertices after vertex shading can be culled.

[0114] The embodiment of the present disclosure also provides a task allocation device applied to a graphics processor N of a graphics processor system. i .

[0115] Figure 3 The block diagram of the task allocation device according to an embodiment of the present disclosure is schematically shown.

[0116] like Figure 3 As shown, the task allocation device 300 according to an embodiment of the present disclosure includes: a vertex processing task receiving module 310 and a primitive processing task allocation module 320.

[0117] Vertex processing task receiving module 310 is used to receive vertex processing tasks, and the vertex processing tasks indicate the graphics processor N i Perform vertex shading on the target vertex to obtain the vertex position of the target vertex after vertex shading. The target vertex is based on the draw call instruction C i The vertex processing task is assigned to the graphics processor N according to the vertex load of each graphics processor in the graphics processor system and the number of target vertices. i of.

[0118] The primitive processing task assignment module 320 is used to assign primitive processing tasks for target tiles to the target graphics processor. The primitive processing tasks for the target tiles instruct the target graphics processor to perform tile processing on the target primitives whose bounding boxes are within the coverage area of ​​the target tiles. The target tiles are matched with the target graphics processors, and the vertices that make up the target primitives undergo vertex shading processing.

[0119] For example, vertex shading processing may be performed by a graphics processor N i The vertex shader is executed, and the fragmentation processing can be performed by the tiler of the graphics processor, for example.

[0120] According to a task allocation device according to another embodiment of the present disclosure, a target tile and a target graphics processor have a mapping relationship, and the mapping relationship is determined by a row number and a column number of the target tile relative to a frame image.

[0121] According to another embodiment of the present disclosure, the task allocation device further includes: a primitive processing task receiving module and a sorting module.

[0122] The primitive processing task receiving module is used to receive primitive processing tasks for X allocated tiles, where the primitive processing tasks for X allocated tiles are allocated by the graphics processor of the graphics processor system to the graphics processor N. i The primitive processing task for X allocated tiles instructs the graphics processor N iEach allocated primitive whose bounding box is within the coverage area of ​​the allocated tile is sliced, and any allocated tile and graphics processor N i Matching, the vertices that make up the assigned primitive are processed by vertex shading.

[0123] The sorting module is used to sort the allocated primitives corresponding to all the allocated tiles so that the order of slicing the allocated primitives corresponding to all the allocated tiles is consistent with the order of draw call instructions received by the graphics processor system.

[0124] Exemplarily, the primitive processing task receiving module may also be used to sort the allocated primitives corresponding to all the allocated tiles, that is, the primitive processing task receiving module also has the function of a sorting module.

[0125] According to a task allocation device according to yet another embodiment of the present disclosure, the vertex processing task further instructs to perform culling processing on primitives composed of vertices that have undergone vertex shading processing.

[0126] For example, a vertex shader may be used to perform culling processing on primitives composed of vertices that have undergone vertex shading processing.

[0127] According to another embodiment of the present disclosure, a task allocation device is provided. The graphics processor system includes a master graphics processor and at least one slave graphics processor. The graphics processor N i In the case of a main graphics processor, the device further includes: a drawing call instruction receiving module and a vertex processing task allocating module.

[0128] The drawing call instruction receiving module is used to receive the drawing call instruction C i .

[0129] Vertex processing task allocation module is used to allocate the vertex processing task according to the drawing call instruction C i The number of associated vertices and the vertex load of each graphics processor of the graphics processor system are used to assign vertex processing tasks to the graphics processors of the graphics processor system. The vertex processing tasks indicate the response to the draw call instruction C i The associated vertices are vertex shaded.

[0130] For example, the structures of the master graphics processor and the slave graphics processors may be the same, that is, the master graphics processor and each slave graphics processor include a vertex processing task allocation module, but the vertex processing task allocation module of the slave graphics processor may be disabled.

[0131] According to a task allocation device according to yet another embodiment of the present disclosure, a vertex processing task allocation module includes: a vertex determination submodule and a vertex processing task allocation submodule.

[0132] Vertex determination submodule is used to determine the number of accumulated vertices in the current batch and the drawing call instruction C i The number of associated vertices and the batch vertex number threshold determine the vertices associated with the current batch. The cumulative number of vertices in the current batch is obtained by accumulating the number of vertices associated with the historical draw call instructions of the current batch. The reception time of the historical draw call instructions is at the draw call instruction C i Before the reception time.

[0133] The vertex processing task allocation submodule allocates vertex processing tasks to the graphics processors of the graphics processor system according to the vertex load of each graphics processor of the graphics processor system and the number of vertices associated with the current batch, and the vertex processing tasks indicate vertex shading processing for the vertices associated with the current batch.

[0134] According to a task allocation device according to another embodiment of the present disclosure, the vertex determination submodule is configured to:

[0135] The cumulative number of vertices in the current batch and the draw call instruction C i When the sum of the number of associated vertices is less than or equal to the vertex number threshold, the cumulative vertices of the current batch and the draw call instruction C i The associated vertices update the accumulated vertices of the current batch.

[0136] The cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is less than or equal to the vertex number threshold, the draw call instruction C i The associated vertices are assigned to the downstream batch after the current batch.

[0137] The cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is greater than the vertex number threshold, the draw call instruction C is sent according to the vertex number threshold. i The associated vertices are assigned to the current batch and at least one downstream batch after the current batch.

[0138] It should be understood that Figure 3 The embodiments of the apparatus portion of the present disclosure shown are identical or similar to the embodiments of the method portion of the present disclosure, and are not described in detail herein.

[0139] According to an embodiment of the present disclosure, the present disclosure also provides a chip module, a graphics processor system, an electronic device, a computer-readable storage medium, and a computer program product.

[0140] The chip module according to the embodiment of the present disclosure includes a transceiver component and a chip, and the chip can be used to execute the task allocation method of any of the above embodiments.

[0141] The graphics processor system according to an embodiment of the present disclosure includes a master graphics processor and at least one slave graphics processor. The operations performed by the master graphics processor and the slave graphics processor have been described in detail in the above embodiments and will not be repeated here.

[0142] An embodiment of the present disclosure also provides an electronic device.

[0143] Figure 4A An electronic device of an embodiment is schematically shown, which includes a processor and an interface circuit. The interface circuit is used to receive signals from other electronic devices outside the electronic device and transmit them to the processor or send signals from the processor to other electronic devices outside the electronic device. The processor is used to implement a task allocation method such as any of the above embodiments through logic circuits or execution code commands.

[0144] Exemplarily, the processor is a graphics processor.

[0145] Exemplarily, the electronic device is a chip.

[0146] Figure 4B A schematic block diagram of an electronic device 400 according to another embodiment of the present disclosure is shown.

[0147] Electronic devices include various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also include various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0148] like Figure 4B As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0149] Multiple components in electronic device 400 are connected to I / O interface 405, including: an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless communication transceiver, etc. The communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0150] The computing unit 401 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above, such as the task allocation method. For example, in some embodiments, the aforementioned method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the task allocation method described above can be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the task allocation method in any other appropriate manner (eg, by means of firmware).

[0151] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and commands from a storage system, at least one input device, and at least one output device, and transmit data and commands to the storage system, the at least one input device, and the at least one output device.

[0152] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0153] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with a command execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, a flash memory, or any suitable combination of the foregoing.

[0154] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0155] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0156] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0157] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0158] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A task allocation method, characterized in that: Any graphics processor N used in the graphics processor system i , the task allocation method includes: Receive a vertex processing task, the vertex processing task instructs the graphics processor N i Perform vertex shading processing on the target vertex to obtain the vertex position of the target vertex after the vertex shading processing, and the target vertex is based on the draw call instruction C i The vertex processing task is determined by the associated vertices, and the vertex processing task is allocated to the graphics processor N according to the vertex load of each graphics processor of the graphics processor system and the number of the target vertices. i of; A primitive processing task for a target tile is assigned to a target graphics processor, wherein the primitive processing task for the target tile instructs the target graphics processor to perform tile processing on a target primitive whose bounding box is within a coverage area of ​​the target tile, the target tile is matched with the target graphics processor, and vertices constituting the target primitive undergo vertex shading processing.

2. The method according to claim 1, characterized in that The target tile and the target graphics processor have a mapping relationship, and the mapping relationship is determined by the row number and column number of the target tile relative to the frame image.

3. The method according to claim 1, characterized in that Also includes: Receive primitive processing tasks for X allocated tiles, wherein the primitive processing tasks for the X allocated tiles are allocated by the graphics processor of the graphics processor system to the graphics processor N. i The primitive processing task for the X allocated tiles indicates the graphics processor N i The slicing process is performed on each allocated primitive whose bounding box is within the coverage area of ​​the allocated tile. Any of the allocated tiles and the graphics processor N i Matching, the vertices constituting the allocated primitives undergo the vertex shading process; The allocated primitives corresponding to all the allocated tiles are sorted so that an order in which the allocated primitives corresponding to all the allocated tiles are subjected to the tile processing is consistent with an order in which the draw call instructions are received by the graphics processor system.

4. The method according to claim 1, wherein The vertex processing task further instructs to perform culling processing on the primitives composed of the vertices processed by the vertex shading process.

5. The method according to claim 1, wherein The graphics processor system includes a master graphics processor and at least one slave graphics processor. i In the case of the master graphics processor, the method further includes: Receive draw call instruction C i ; According to the drawing call instruction C i The number of associated vertices and the vertex load of each of the graphics processors of the graphics processor system are used to assign the vertex processing task to the graphics processor of the graphics processor system, wherein the vertex processing task indicates the processing of the draw call instruction C. i The associated vertex is subjected to the vertex shading process.

6. The method according to claim 5, characterized in that According to the drawing call instruction C i The number of associated vertices and the vertex load of each graphics processor of the graphics processor system, and allocating the vertex processing task to the graphics processor of the graphics processor system includes: According to the number of accumulated vertices in the current batch, the drawing call instruction C i The number of associated vertices and the vertex number threshold of the batch are used to determine the vertices associated with the current batch, wherein the cumulative number of vertices in the current batch is obtained by accumulating the number of vertices associated with the historical draw call instructions of the current batch, and the reception time of the historical draw call instructions is within the time period of the draw call instruction C. i Before the moment of receipt; The vertex processing task is allocated to the graphics processor of the graphics processor system according to the vertex load of each graphics processor of the graphics processor system and the number of vertices associated with the current batch, wherein the vertex processing task instructs to perform the vertex shading process on the vertices associated with the current batch.

7. The method according to claim 6, characterized in that The cumulative number of vertices in the current batch, the drawing call instruction C i The number of associated vertices and the threshold number of vertices in the batch, determining the vertices associated with the current batch includes: The cumulative number of vertices in the current batch and the draw call instruction C i When the sum of the number of associated vertices is less than or equal to the vertex number threshold, the current batch of accumulated vertices and the draw call instruction C i The associated vertices update the cumulative vertices of the current batch; The cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i If the number of associated vertices is less than or equal to the vertex number threshold, the draw call instruction C i The associated vertices are assigned to a downstream batch following the current batch; The cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is greater than the vertex number threshold, the draw call instruction C is set according to the vertex number threshold. i The associated vertices are allocated to the current batch and at least one of the downstream batches after the current batch.

8. A task allocation device, characterized in that: Any graphics processor N used in the graphics processor system i , the task allocation device includes: A vertex processing task receiving module is used to receive a vertex processing task, wherein the vertex processing task indicates that the graphics processor N i Perform vertex shading processing on the target vertex to obtain the vertex position of the target vertex after the vertex shading processing, and the target vertex is based on the draw call instruction C i The vertex processing task is determined by the associated vertices, and the vertex processing task is allocated to the graphics processor N according to the vertex load of each graphics processor of the graphics processor system and the number of the target vertices. i of; A primitive processing task assignment module is configured to assign a primitive processing task for a target tile to a target graphics processor, wherein the primitive processing task for the target tile instructs the target graphics processor to perform tile processing on a target primitive whose bounding box is within the coverage area of ​​the target tile, the target tile is matched with the target graphics processor, and the vertices constituting the target primitive undergo vertex shading processing.

9. The device according to claim 8, characterized in that The target tile and the target graphics processor have a mapping relationship, and the mapping relationship is determined by the row number and column number of the target tile relative to the frame image.

10. The device according to claim 8, characterized in that Also includes: A primitive processing task receiving module is configured to receive primitive processing tasks for X allocated tiles, wherein the primitive processing tasks for the X allocated tiles are allocated by the graphics processor of the graphics processor system to the graphics processor N. i The primitive processing task for the X allocated tiles indicates the graphics processor N i The slicing process is performed on each allocated primitive whose bounding box is within the coverage area of ​​the allocated tile. Any of the allocated tiles and the graphics processor N i Matching, the vertices constituting the allocated primitives undergo the vertex shading process; A sorting module is configured to sort the allocated primitives corresponding to all the allocated tiles so that an order in which the allocated primitives corresponding to all the allocated tiles are subjected to sharding processing is consistent with an order in which the draw call instructions are received by the graphics processor system.

11. The device according to claim 8, characterized in that The vertex processing task further instructs to perform culling processing on the primitives composed of the vertices processed by the vertex shading process.

12. The device according to claim 8, characterized in that The graphics processor system includes a master graphics processor and at least one slave graphics processor. i In the case of the main graphics processor, the device further includes: The drawing call instruction receiving module is used to receive the drawing call instruction C i ; Vertex processing task allocation module, used for according to the drawing call instruction C i The number of associated vertices and the vertex load of each of the graphics processors of the graphics processor system are used to assign the vertex processing task to the graphics processor of the graphics processor system, wherein the vertex processing task indicates the processing of the draw call instruction C. i The associated vertex is subjected to the vertex shading process.

13. The device according to claim 12, characterized in that The vertex processing task allocation module includes: The vertex determination submodule is used to determine the number of accumulated vertices in the current batch, the drawing call instruction C i The number of associated vertices and the vertex number threshold of the batch are used to determine the vertices associated with the current batch, wherein the cumulative number of vertices in the current batch is obtained by accumulating the number of vertices associated with the historical draw call instructions of the current batch, and the reception time of the historical draw call instructions is within the time period of the draw call instruction C. i Before the moment of receipt; A vertex processing task allocation submodule allocates the vertex processing tasks to the graphics processors of the graphics processor system according to the vertex load of each graphics processor of the graphics processor system and the number of vertices associated with the current batch, wherein the vertex processing tasks instruct to perform vertex shading processing on the vertices associated with the current batch.

14. The device according to claim 13, characterized in that The vertex determination submodule is used for: The cumulative number of vertices in the current batch and the draw call instruction C i When the sum of the number of associated vertices is less than or equal to the vertex number threshold, the current batch of accumulated vertices and the draw call instruction C i The associated vertices update the cumulative vertices of the current batch; The cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i If the number of associated vertices is less than or equal to the vertex number threshold, the draw call instruction C i The associated vertices are assigned to a downstream batch following the current batch; The cumulative number of vertices in the current batch and the draw call instruction C i The sum of the number of associated vertices is greater than the vertex number threshold and the draw call instruction C i When the number of associated vertices is greater than the vertex number threshold, the draw call instruction C is set according to the vertex number threshold. i The associated vertices are allocated to the current batch and at least one of the downstream batches after the current batch.

15. An electronic device, characterized in that: The device comprises a processor and an interface circuit, wherein the interface circuit is used to receive signals from electronic devices other than the electronic device and transmit them to the processor or send signals from the processor to electronic devices other than the electronic device, and the processor implements the method according to any one of claims 1 to 7 through a logic circuit or executing code instructions.

16. The electronic device according to claim 15, characterized in that The processor is a graphics processor.

17. The electronic device according to claim 15, characterized in that The electronic device is a chip.

18. A chip module, characterized in that: The method comprises a transceiver component and a chip, wherein the chip is used to execute the method according to any one of claims 1 to 7.

19. A graphics processor system, characterized in that: The system comprises a master graphics processor and at least one slave graphics processor, wherein the master graphics processor and the slave graphics processor are used to implement the method according to any one of claims 1 to 4, and the master graphics processor is further used to implement the method according to any one of claims 5 to 7.

20. A computer-readable storage medium storing computer instructions, characterized in that: include: Computer instructions, wherein when the computer instructions are executed, the computer is caused to perform the method according to any one of claims 1 to 7.