Method and device for rendering task allocation, electronic device, and storage medium

By determining the task load level of computing unit clusters within the processing core of the graphics processing device and allocating tasks accordingly, the problem of uneven distribution of rendering tasks is solved, achieving load balancing of computing resources and improving system performance.

CN120523610BActive Publication Date: 2025-11-21MOORE THREADS TECHNOLOGY (SHANGHAI) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511029301.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-21
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

In existing technologies, the allocation of rendering tasks in graphics processing devices suffers from uneven distribution of computing resources, which affects the utilization of computing resources and system performance.

Method used

By establishing rendering tasks in the processing core of the graphics processing device, determining the task load value and task load level of the computing unit cluster, and selecting the target computing unit cluster from multiple computing unit clusters based on the task load level for task allocation, the load balancing of computing resources is achieved.

Benefits of technology

It improves the utilization of computing resources and system performance, and enhances the processing efficiency of rendering tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523610B_ABST
    Figure CN120523610B_ABST
Patent Text Reader

Abstract

The present disclosure provides a rendering task allocation method and device, an electronic device and a storage medium. The method is applied to a processing core of a graphics processing device. The processing core includes a plurality of computing unit clusters. The method includes: establishing a first rendering task of a first tile in a rendering object according to primitive set information of the first tile to be rendered; determining a task load value and a task load level of each of the plurality of computing unit clusters according to a set task allocation mode; determining a target computing unit cluster from the plurality of computing unit clusters according to the task load levels of the plurality of computing unit clusters; and sending the first rendering task to a task queue of the target computing unit cluster, so that the target computing unit cluster executes the first rendering task. According to the embodiments of the present disclosure, the load balancing of the computing resources can be improved, and the utilization rate of the computing resources can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of graphics processing, and particularly relates to a rendering task allocation method and device, electronic equipment and computer readable storage medium. BACKGROUND

[0002] Graphics processing technology can refer to a technology of processing two-dimensional or three-dimensional vision, image, etc. by using various hardware and software devices to achieve a specific rendering effect. A graphics processing device can refer to a device used for performing graphics processing related operations, such as a graphics processing unit (GPU) or other similar devices. Generally, a graphics processing device can include multiple processing cores, each processing core including multiple clusters of computing units, and each cluster of computing units including multiple computing units, so as to achieve higher graphics processing efficiency through parallel operation. Therefore, how to allocate graphics processing tasks to maximize the computing power of the graphics processing device has become a problem that attracts much attention. SUMMARY

[0003] The present disclosure provides a rendering task allocation method and device, electronic equipment and computer readable storage medium.

[0004] In a first aspect, the present disclosure provides a rendering task allocation method applied to a processing core of a graphics processing device, the processing core including multiple clusters of computing units. The method includes: establishing a first rendering task of a first tile to be rendered in a rendering object according to primitive set information of the first tile, the first rendering task including the primitive set information of the first tile and hardware related information required for executing the first rendering task; determining task load values and task load levels of the multiple clusters of computing units respectively according to a set task allocation mode; determining a target cluster of computing units from the multiple clusters of computing units according to the task load levels of the multiple clusters of computing units; and sending the first rendering task to a task queue of the target cluster of computing units, so as to enable the target cluster of computing units to execute the first rendering task of the first tile.

[0005] In a second aspect, the present disclosure provides a rendering task allocation apparatus applied to a processing core of a graphics processing device, the processing core comprising a plurality of compute unit clusters, the apparatus comprising: a task establishing module configured to establish a first rendering task of a first tile to be rendered in a rendering object according to primitive set information of the first tile, the first rendering task comprising primitive information of the first tile and hardware related information required for executing the first rendering task; and a task allocation module configured to determine task load values and task load levels of the plurality of compute unit clusters according to a set allocation manner, determine a target compute unit cluster from the plurality of compute unit clusters according to the task load levels of the plurality of compute unit clusters, and send the first rendering task to a task queue of the target compute unit cluster so as to enable the target compute unit cluster to execute the first rendering task of the first tile.

[0006] In a third aspect, the present disclosure provides an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the rendering task allocation method described above.

[0007] In a fourth aspect, the present disclosure provides a computer readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the rendering task allocation method described above.

[0008] The embodiments provided by the present disclosure can establish a rendering task for a tile in a rendering object, determine task load values and task load levels of a plurality of compute unit clusters in a processing core, and then determine a target compute unit cluster for executing the rendering task from the plurality of compute unit clusters based on the task load levels. The task allocation based on the task load levels of the compute unit clusters can improve load balancing of computing resources, improve utilization of computing resources, and improve system performance.

[0009] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, which together with the embodiments of the present disclosure serve to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent from the detailed description of the specific embodiments described below, taken in conjunction with the accompanying drawings, in which:

[0011] Figure 1 FIG. 1 is a schematic diagram of a rendering architecture of a related art;

[0012] Figure 2 FIG. 2 is a schematic diagram of a processing core of a graphics processing device according to an embodiment of the present disclosure;

[0013] Figure 3 FIG. 3 is a flowchart of a rendering task allocation method according to an embodiment of the present disclosure;

[0014] Figure 4 FIG. 4 is a schematic diagram of an architecture of a control module according to an embodiment of the present disclosure;

[0015] Figure 5 FIG. 5 is a schematic diagram of a task allocation process according to an embodiment of the present disclosure;

[0016] Figure 6 FIG. 6 is a schematic diagram of a task load value determination process according to an embodiment of the present disclosure;

[0017] Figure 7 FIG. 7 is a schematic diagram of a task load state according to an embodiment of the present disclosure;

[0018] Figure 8 FIG. 8 is a schematic diagram of a rendering task allocation state according to an embodiment of the present disclosure;

[0019] Figure 9 FIG. 9 is a block diagram of a rendering task allocation apparatus according to an embodiment of the present disclosure;

[0020] Figure 10 FIG. 10 is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to help understanding, which should be considered only as exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0022] In the case of no conflict, each embodiment of the present disclosure and each feature in the embodiments can be combined with each other.

[0023] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. "Coupled" or "connected" or similar terms are not restricted to physical or mechanical connections or associations, but can also include electrical connections, whether direct or indirect.

[0025] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.

[0026] Before embodiments of the present disclosure are described in detail, some related concepts are first explained for the sake of clarity.

[0027] A render target can refer to an object to be rendered, which can be a part of a virtual scene that needs to be rendered, or can also be a picture or a video frame, etc. that needs to be rendered.

[0028] A primitive can refer to a basic geometric shape that constitutes a figure, such as a point, a line, a triangle, etc. For example, for a figure drawn by an application, it can be represented in a computer by splicing a large number of basic geometries to represent the figure.

[0029] A rendering pipeline, also known as a graphics pipeline or a graphics rendering pipeline, is a key concept in computer graphics. It is a stage in a graphics processing unit (GPU) that is responsible for processing and converting graphics data for rendering. The main task of the rendering pipeline is to convert input geometric primitives (such as points, lines, triangles, etc.) into pixels visible on the screen. The goal of the rendering pipeline is to process graphics in an efficient way and generate the final image. Through parallel processing and specialized hardware support, the GPU can quickly perform these calculations to achieve real-time graphics rendering.

[0030] In related technologies, the rendering process can be implemented using a rendering pipeline. In a standard rendering pipeline, the primitives included in the rendering object can be rasterized, and then the valid pixels are sent to the pixel shader for coloring.

[0031] To improve processing performance and reduce power consumption, a technique was proposed that divides the rendering object into multiple tiles and performs tile-based rendering (TBR) on the rendering object, where each tile can be used to render a small portion of the scene. Building upon TBR, tile-based deferred rendering (TBDR) was further developed. TBDR further delays fragment shading processing on top of TBR and uses hardware-level features to perform hidden surface removal (HSR), thus solving the overdraw problem.

[0032] Figure 1 This is a schematic diagram of a rendering architecture based on related technologies, illustrating the TBDR rendering pipeline. For example... Figure 1 As shown, the rendering pipeline may include vertex processing11; clip, project / cull12; tiling13; rasterization14; early depth test (also known as early visibility test)15; texture and shade16; alpha test17; late depth test (also known as late visibility test)18; and alpha blend19.

[0033] The rendering pipeline can be divided into front-end and back-end processing. In front-end processing, the graphics processing device can read geometry data from system memory, and obtain the primitives of the rendering object through vertex processing 11 and clipping, projection and culling 12 stages; in the tileization 13 stage, the rendering object is segmented, and the graphic data covering each tile, including the primitive list, vertex data, etc., are recorded and written to system memory.

[0034] For a tile, the rendering pipeline loads all the primitives (such as triangle primitives) contained in the tile to the fragment stage for processing, all the primitives covering the tile can be directly read from the corresponding primitive list, and after all the primitives of the tile are processed, the next tile can be processed.

[0035] In the back-end processing, the primitives of each tile are rasterized 14 to obtain corresponding fragments, and the fragments are compared by early depth test 15 to remove the occluded fragments. The fragments passing the early depth test 15 are processed in the texture and shading 16 stage based on the texture data in the system memory, and then are subjected to alpha test 17. The fragments passing the alpha test 17 are subjected to late depth test 18 based on the on-chip depth buffer, and the fragments passing the late depth test 18 are subjected to alpha blending 19 with the data in the on-chip color buffer to obtain the rendering result, such as the pixel data to be output. The on-chip depth buffer interacts with the depth buffer in the system memory, and the on-chip color buffer interacts with the frame buffer in the system memory.

[0036] It should be understood that the above is only a schematic of the TBDR rendering pipeline, and those skilled in the art can set the specific structure of various rendering pipelines according to actual conditions, and the present disclosure does not limit this.

[0037] In the TBR and TBDR rendering pipelines, tiled-based rendering can be performed on the rendering objects. In the processing, according to the complexity and importance of each tile, a tile-level scheduler in a graphics processing unit (GPU) can be used to schedule GPU resources, so as to allocate more computing resources to the tiles that need more details. This method can implement tile-level resource allocation at the GPU hardware level, so as to make greater use of the computing power of the GPU and help reduce power consumption, and therefore it is often used in scenarios where the computing resources are limited, such as the GPU of a mobile device.

[0038] For tiled-based task allocation, related technologies usually include static analysis-based task allocation and dynamic feedback-based task allocation.

[0039] Among them, the static analysis-based task allocation scheme determines the complexity and importance of each tile through static analysis of the scene in advance, and allocates computing resources according to these information. This scheme can allocate tasks according to the geometric complexity of the scene, the number of textures, lighting requirements and other factors.

[0040] In the task allocation scheme based on dynamic feedback, the task allocation strategy is dynamically adjusted according to the real-time performance feedback of each tile. For example, if the computation time of a certain tile exceeds a predetermined threshold, more resources can be allocated to the tile to speed up the rendering process. This scheme can be dynamically optimized according to the actual rendering situation.

[0041] In the task allocation scheme based on static analysis, the entire screen is divided into fixed-size tiles, and the tiles are evenly distributed to each processing core in a predefined order (such as scan line order or space-filling curve order). The advantage is that the scheduling is simple and the implementation overhead is low; but the disadvantage is that it cannot adapt to the case where the load of each tile (such as the number of fragments and the operation complexity) varies greatly in the actual scene, and it is easy to cause some processing cores to be overloaded while other processing cores are idle, thereby affecting the overall rendering performance and energy consumption.

[0042] In the task allocation scheme based on dynamic feedback, load balancing can be achieved based on dynamic priority queue or dynamic task division, by estimating the load of each tile and then adjusting the priority of the task in real time, or by the processing core with lower load actively grabbing the unallocated or partially completed task. This approach can counteract the unevenness of tile rendering tasks to some extent, but it often requires a more complex scheduler, additional hardware resources, and higher synchronization (such as atomic operation) overhead, thereby increasing system complexity and resource occupation.

[0043] As can be seen, although these task allocation schemes can help improve the utilization of computing resources and reduce power consumption, they still have the problem of uneven allocation of computing tasks, which affects the utilization of computing resources and system performance.

[0044] According to the rendering task allocation method of the embodiments of the present disclosure, a task allocation method for balancing the load of tile rendering tasks within a processing core can be proposed based on the TBR and TBDR rendering architectures, the rendering task of a tile can be established, the task load value and task load level of multiple computing unit clusters in the processing core can be determined, and then based on the task load level, a target computing unit cluster for executing the rendering task is determined from the multiple computing unit clusters, so that the task allocation is based on the task load level of the computing unit cluster, the load balancing of computing resources can be achieved, the utilization of computing resources can be improved, and the system performance can be improved.

[0045] The rendering task allocation method according to the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server, and the terminal device can be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor invoking computer readable program instructions stored in a memory. Alternatively, the method can be executed by a server.

[0046] The rendering task allocation method according to the embodiments of the present disclosure can be applied to a processing core of a graphics processing device (e.g., a GPU) in an electronic device. The graphics processing device can include a plurality of processing cores, each processing core including a plurality of compute unit clusters, and each compute unit cluster including a plurality of compute units for performing various computing functions.

[0047] Figure 2 A schematic diagram of a processing core of a graphics processing device according to an embodiment of the present disclosure is provided. As shown in Figure 2 , the processing core (core) can include a control module 21, a compute unit cluster 22, a compute unit cluster 23, and an internal storage unit 24. The control module 21 is configured to receive information such as a control stream transmitted by a rendering pipeline front end, and perform task allocation and control. The compute unit cluster 22 and the compute unit cluster 23 each include a plurality of compute units (compute units) configured to perform rendering tasks, such as rasterizing primitives in a tile. The internal storage unit 24 is configured to store various data in rendering processing.

[0048] It should be understood that the processing core including two compute unit clusters is only a schematic diagram in Figure 2 , and the present disclosure does not limit the number of compute unit clusters included in the processing core and the number of compute units included in each compute unit cluster.

[0049] Figure 3 A flowchart of a rendering task allocation method according to an embodiment of the present disclosure is provided. Referring to Figure 3 , the method includes:

[0050] In step S31, a first rendering task of a first tile in a rendering object is established according to primitive set information of the first tile to be rendered, and the first rendering task includes the primitive set information of the first tile and hardware-related information required for executing the first rendering task.

[0051] In step S32, according to the set task allocation mode, the task load values and the task load levels of the plurality of computing unit clusters are determined respectively.

[0052] In step S33, according to the task load levels of the plurality of computing unit clusters, a target computing unit cluster is determined from the plurality of computing unit clusters.

[0053] In step S34, the first rendering task is sent to the task queue of the target computing unit cluster, so that the target computing unit cluster executes the first rendering task of the first tile.

[0054] For example, after the rendering object is tiled in the rendering pipeline, a plurality of tiles of the rendering object can be obtained, and information related to the primitive set of each tile is determined, including the number of primitives in the primitive set, the primitive address and the primitive offset, and the size information, the coordinates in the screen, the index, and the like of the tile. The size of the tile can be 16x16, 32x32, 64x64, and the like. The specific size of the tile is not limited in the present disclosure.

[0055] In some possible implementation manners, the control stream of the front end of the rendering pipeline can transmit the information to the processing core of the graphics processing device. When the control module of the processing core receives the control stream information of a tile, the first size of the tile can be obtained, and compared with the processing size of the processing core. If the first size is greater than the processing size, the tile needs to be split to obtain a sub-tile of the tile as the first tile actually processed by the processing core. If the first size is less than or equal to the processing size, the tile does not need to be split, and the tile is directly taken as the first tile actually processed by the processing core.

[0056] The processing size of the processing core can be determined according to the hardware configuration of the processing core, for example, the maximum size that can be processed by the computing unit cluster of the processing core in a single clock cycle. The second size of the first tile is less than or equal to the processing size and is also less than the first size of the tile. The second size of the first tile is, for example, 8x8, 16x16, and the like. The determination manner of the processing size of the processing core, the processing size of the processing core, and the specific values of the second size of the tile are not limited in the present disclosure.

[0057] In some possible implementation manners, after the first tile is determined, the primitive set information of the first tile can be determined, including the number of primitive sets (there can be one or more primitive sets) of the first tile, the number of primitives contained in each primitive set, the primitive address and the primitive offset in the primitive set, and the like. The primitive can be, for example, a triangle, and the primitive address in the primitive set can be calculated from the memory address of the rendering object and the relative coordinates of the first tile and the like.

[0058] In some possible implementation manners, in step S31, according to the primitive set information of the first tile and the hardware related information required for performing the rendering task and obtained from the system memory, the first rendering task of the first tile can be packaged and established.

[0059] The hardware related information required for performing the rendering task can include hardware state information and register information. The hardware state information can include, for example, fragment state words (used for storing configuration parameters and intermediate results related to the fragment processing stage), event status (used for recording the occurrence state or completion flag of a specific event in the rendering process), and the like. The register information includes various parameters and configuration information required for hardware operation. It should be understood that the hardware related information can be set by the person skilled in the art according to the actual situation, and the specific content of the hardware related information is not limited in the present disclosure.

[0060] According to the embodiments of the present disclosure, the software is allowed to configure parameters through the hardware interface, and different task allocation manners are supported for task scheduling in different scenarios, so that the task allocation can match the scene computing power demand, and the average processing efficiency is improved.

[0061] In some possible implementation manners, the task allocation manner can include an upstream allocation manner and a total allocation manner. The upstream allocation manner refers to allocating related information generated in the task allocation by the upstream (i.e., the control module) to the downstream (i.e., the computing unit cluster) to determine the task load and perform task allocation. The total allocation manner refers to combining the related information generated in the task allocation by the upstream to the downstream and the task state information fed back by the downstream to determine the task load and perform task allocation. The task allocation manner can also include other allocation manners. It should be understood that the person skilled in the art can set the task allocation manner according to the actual situation, and the specific category and quantity of the task allocation manner are not limited in the present disclosure.

[0062] In some possible implementation manners, according to the set task allocation manner, the task load values of the plurality of computing unit clusters can be determined respectively in step S32, and the task load level of each computing unit cluster can be determined according to the task load values. Further, in step S33, the target computing unit cluster can be determined from the plurality of computing unit clusters according to the task load levels of the plurality of computing unit clusters.

[0063] The task load levels can be categorized into three levels: Level 1, Level 2, and Level 3. Level 1 corresponds to a low task load, and tasks can be preferentially allocated to Level 1 computing unit clusters. Level 2 corresponds to a medium task load, and tasks can be allocated to Level 2 computing unit clusters in turn. Level 3 corresponds to a high task load, and tasks can be excluded from allocation to Level 3 computing unit clusters. This approach achieves balanced task distribution and improves the utilization of computing resources.

[0064] In some possible implementations, in step S34, the first rendering task can be sent to the task queue of the target computing unit cluster so that the target computing unit cluster can execute the first rendering task of the first tile, thereby completing the rendering task allocation process.

[0065] According to embodiments of this disclosure, a rendering task for a tile can be established; the task load value and task load level of multiple computing unit clusters in the processing core can be determined; and then, based on the task load level, a target computing unit cluster for executing the rendering task can be determined from the multiple computing unit clusters, thereby allocating tasks based on the task load level of the computing unit clusters, which can achieve load balancing of computing resources, improve computing resource utilization, and improve system performance.

[0066] The rendering task allocation method according to embodiments of this disclosure will now be described in detail.

[0067] Figure 4 This is a schematic diagram of the architecture of a control module provided in an embodiment of this disclosure. Figure 4 As shown, the control module may include a task creation module 41, a task allocation module 42, and a post-processing module 43. The control logic of this module is set before rasterization processing and can also be called a Raster Parameter Processing (RPP) module. This control module is used to create and allocate tile rendering tasks and can also be called a rendering task allocation module or device. Specifically, the task creation module 41 receives control flow information, determines the primitive set of the tiles, and creates rendering tasks. The primitives are mainly triangles, and the task creation module 41 can also be called a triangle setup module. The task allocation module 42 determines the task allocation method, also known as the allocation algorithm, to implement distributed algorithm control. The task allocation module 42 is also used to implement in-core pipe dispatcher task scheduling.

[0068] In an example, the post-processing module 43 is configured to implement tile post processing, maintain and manage state information in the process of tile rendering processing. For example, after the tiles are assigned to the cluster of computing units, the post-processing module 43 can update the information of the rendering state, receive special events, report abnormal conditions, etc., such as maintaining the current processing tile index, tile processing state, error processing, etc.

[0069] In some possible implementation manners, in a case where the rendering manner of the rendering object is tile-based rendering (TBR) or tile-based deferred rendering (TBDR), as shown in Figure 1 , the geometry data of the rendering object is processed via the vertex processing 11 and clipping, projection and culling 12 to obtain a plurality of primitives of the rendering object; and then the rendering object is divided into a plurality of tiles in the tiling 13, each of the plurality of tiles is covered by at least one primitive of the rendering object.

[0070] In some possible implementation manners, the task establishment module 41 can input an interface to obtain the control stream information of the front end of the graphics rendering pipeline from the previous module (for example, the tiling module) or in other manners (as shown in Figure 4 ). For any tile (hereinafter referred to as a second tile) of the plurality of tiles, the control stream information includes the number of primitives in the primitive set of the second tile, the primitive address and the primitive offset of the primitive set, and the first size of the second tile. Wherein, the primitives in the primitive set are mainly triangles.

[0071] In some possible implementation manners, the rendering task allocation method according to the embodiments of the present disclosure can further include: in a case where the control stream information of a second tile to be rendered in the rendering object is received, determining a first size of the second tile; the control stream information includes the number of primitives in the primitive set of the second tile, the primitive address and the primitive offset of the primitive set, and the first size; in a case where the processing size of the processing core is smaller than the first size of the second tile, splitting the second tile to obtain a plurality of first tiles of the second tile; the second size of the first tile is smaller than or equal to the processing size of the processing core; according to the control stream information of the second tile, determining the primitive set information of a plurality of the first tiles respectively.

[0072] In other words, when the task creation module 41 receives the control flow information of the second tile, it can obtain the first size of the second tile and compare it with the size that the processing core can process. If the first size is larger than the processing size, the tile needs to be split to obtain the sub-tile of the tile, which is used as the tile that the processing core actually processes. This sub-tile is referred to as the first tile. If the first size is less than or equal to the processing size, the tile does not need to be split and the second tile is directly used as the first tile that the processing core actually processes.

[0073] The processing size of the processing core can be determined according to the hardware configuration of the processing core, such as the maximum size that the computing unit cluster or counting unit of the processing core can process in a single clock cycle; the second size of the first block is less than or equal to the processing size and also less than the first size of the block. The second size of the first block is, for example, 8×8, 16×16, etc. This disclosure does not limit the method of determining the processing size of the processing core, the specific values ​​of the processing size of the processing core and the second size of the block.

[0074] In some possible implementations, after determining the first tile, the primitive set information of the first tile can be determined, including the number of primitive sets of the first tile (there may be one or more primitive sets), the number of primitives contained in each primitive set, the primitive addresses in the primitive sets, and the primitive offsets, etc. Here, a primitive can be, for example, a triangle, and the primitive addresses in the primitive sets can be calculated from the memory address of the rendering object and the relative coordinates of the first tile.

[0075] In this way, the tile size can be matched with the processing power of the processing core, thereby improving the efficiency of rendering.

[0076] In some possible implementations, in step S31, based on the primitive set information of the first block, the task establishment module 41 can determine the primitive set information from, for example... Figure 4 The system memory shown retrieves the hardware-related information required to execute the rendering task, packages the primitive set information and hardware-related information of the first tile, and establishes the first rendering task for the first tile.

[0077] The hardware-related information includes hardware status information and register information. Hardware status information includes, for example, fragment state words (used to store configuration parameters and intermediate results related to the fragment processing stage) and event status (used to record the occurrence status or completion flags of specific events during rendering). Register information includes various parameters and configurations required for hardware operation. It should be understood that those skilled in the art can set the hardware-related information according to actual circumstances, and this disclosure does not limit the specific content of the hardware-related information.

[0078] In some possible implementation manners, the task allocation module 42 can determine the task allocation manner, for example, select to implement a certain task allocation manner or combine two or more task allocation manners according to different scenes to which the rendering objects belong, or can also be specified by a user. Among them, the user can specify a certain task allocation manner through a configuration register, write an API (Application Program Interface), and the like, and can control the parameters of the executed task allocation manner, which is transmitted to the task allocation module 42 through the software control in the configuration register, and the task allocation module 42 determines the task allocation manner according to the input configuration information. In this way, the flexibility of the task allocation manner configuration can be improved. Figure 4

[0079] In some possible implementation manners, the task allocation manner can include an upstream allocation manner and a total allocation manner. The upstream allocation manner refers to determining a task load and performing task allocation by allocating related information generated in the task allocation from an upstream (that is, a control module) to a downstream (that is, a computing unit cluster). The total allocation manner refers to determining a task load and performing task allocation by combining related information generated in the task allocation from the upstream to the downstream and feedback information of a task execution state of the downstream. The task allocation manner can also include other allocation manners, and it should be understood that a person skilled in the art can set the task allocation manner according to actual conditions, and the disclosure does not limit the specific categories and quantities of the task allocation manner.

[0080] By supporting different task allocation manners for task scheduling in different scenes, the task allocation can be matched with the scene computing power demand, and the overall processing efficiency is improved.

[0081] In some possible implementation manners, after the task allocation manner is determined, in step S32, the task load values of the plurality of computing unit clusters can be respectively determined according to the set task allocation manner, and the task load level of each computing unit cluster can be determined according to the task load values.

[0082] In some possible implementation manners, a queue task counter can be set for each computing unit cluster in the task allocation module 42, so as to perform scheduling according to the count value of the queue task counter. The queue task counter can also be referred to as a water level counter, which is used to indicate the number of allocated tasks in a task queue (FIFO) between the upstream and the downstream.

[0083] For any computing unit cluster, when the task allocation module 42 sends a rendering task of a tile to the task queue of the computing unit cluster, the count value of the queue task counter is incremented by 1, indicating that a rendering task is added to the task queue.​

[0084] In some possible implementation manners, the method for rendering task allocation according to the embodiments of the present disclosure further includes: in the case that a credit feedback signal sent by any computing unit cluster is received, reducing the first count value of the queue task counter of the computing unit cluster by 1; the credit feedback signal is used to indicate that the computing unit cluster acquires one rendering task from the task queue of the computing unit cluster.

[0085] That is to say, for any computing unit cluster, the computing unit cluster acquires one rendering task of a tile from the task queue and performs rasterization processing each time, and then sends one credit feedback signal to the task allocation module 42. When the task allocation module 42 receives the credit feedback signal, the count value of the queue task counter of the computing unit cluster is reduced by 1, indicating that one rendering task is reduced in the task queue.

[0086] The count value of the queue task counter represents the task load level of the computing unit cluster, which can improve the accuracy of task scheduling, thereby improving the balance of task allocation and improving the utilization rate of computing resources.

[0087] In some possible implementation manners, in the case that the task allocation manner is the upstream allocation manner, the step of determining the task load value and the task load level of the plurality of computing unit clusters in step S32 can include: for any computing unit cluster, in the case that the task allocation manner is the upstream allocation manner, determining the task load value of the computing unit cluster according to the first count value of the queue task counter of the computing unit cluster; the queue task counter is arranged in the control module of the processing core; and determining the task load level of the computing unit cluster according to the task load value of the computing unit cluster and the preset upper limit value and lower limit value of the task.

[0088] For example, the upper limit value (limit_high) and the lower limit value (limit_low) of the task queue of the computing unit cluster can be preset, which are respectively used to represent the critical values of the number of tasks in the task queue, and the upper limit value and the lower limit value are respectively set to 5 and 10, and the present disclosure does not limit the specific values of the upper limit value and the lower limit value.

[0089] In some possible implementation manners, if the task allocation manner is the upstream allocation manner, the first count value of the queue task counter of the computing unit cluster can be directly used as the task load value of the computing unit cluster, and then the task load value is compared with the upper limit value and the lower limit value of the task to obtain the task load level of the computing unit cluster, which can be denoted as FreeEntryCounter.

[0090] In this way, the task load level of the computing unit cluster can be determined directly according to the number of tasks in the upstream task queue, and the scheduling is simpler, thereby reducing the resource overhead in task scheduling.

[0091] In some possible implementation manners, the step of determining the task load level of the computing unit cluster according to the task load value of the computing unit cluster and the preset upper limit value and lower limit value of the task includes: determining the task load level of the computing unit cluster as a first level in a case where the task load value of the computing unit cluster is less than the lower limit value of the task; determining the task load level of the computing unit cluster as a second level in a case where the task load value of the computing unit cluster is greater than or equal to the lower limit value of the task and less than or equal to the upper limit value of the task; and determining the task load level of the computing unit cluster as a third level in a case where the task load value of the computing unit cluster is greater than the upper limit value of the task, where the task load of the third level is higher than that of the second level, and the task load of the second level is higher than that of the first level.

[0092] That is, the task load level can include a first level, a second level and a third level, the first level corresponding to low task load of the computing unit cluster, the second level corresponding to medium task load of the computing unit cluster, and the third level corresponding to high task load of the computing unit cluster. That is, the task load of the third level is higher than that of the second level, and the task load of the second level is higher than that of the first level.

[0093] For any computing unit cluster, if the task load value of the computing unit cluster is less than the lower limit value of the task, the task load level of the computing unit cluster is determined as the first level, the task load of the computing unit cluster is low, and the computing unit cluster is preferentially allocated with tasks; if the task load value of the computing unit cluster is greater than or equal to the lower limit value of the task and less than or equal to the upper limit value of the task, the task load level of the computing unit cluster is determined as the second level, the task load of the computing unit cluster is medium, and tasks can be alternately allocated to each computing unit cluster of the second level; and if the task load value of the computing unit cluster is greater than the upper limit value of the task, the task load level of the computing unit cluster is determined as the third level, the task load of the computing unit cluster is high, and no task is allocated to the computing unit cluster in the current round of task allocation.

[0094] In this way, the task load level of the computing unit cluster can be determined for scheduling tasks, thereby achieving load balancing and improving the utilization rate of computing resources.

[0095] In this way, the task load level of each computing unit cluster can be determined through the above steps, and then the target computing unit cluster is determined from the plurality of computing unit clusters according to the task load levels of the plurality of computing unit clusters in step S33.

[0096] In some possible implementation manners, step S33 can include: in a case where there is a first computing unit cluster with a task load level of a first level in the plurality of computing unit clusters, determining, as the target computing unit cluster, a computing unit cluster with a lowest task load value in the first computing unit cluster; in a case where there is no first computing unit cluster with a task load level of a first level in the plurality of computing unit clusters, and there is a second computing unit cluster with a task load level of a second level, polling to determine the target computing unit cluster in the second computing unit cluster; and in a case where the task load levels of the plurality of computing unit clusters are all of a third level, skipping the current round of task allocation, and determining the task load levels of the plurality of computing unit clusters again in a next round of task allocation, and determining the target computing unit cluster according to the task load levels.

[0097] For example, the task allocation module can preferentially allocate tasks to computing unit clusters with low task loads. If there is a first computing unit cluster with a task load level of a first level in the plurality of computing unit clusters, a target computing unit cluster is selected in the first computing unit cluster. In this case, if the first computing unit cluster is one, the first computing unit cluster is directly taken as the target computing unit cluster; or if the first computing unit cluster is multiple, a computing unit cluster with a lowest task load value in the first computing unit cluster is taken as the target computing unit cluster.

[0098] In some possible implementation manners, if there is no first computing unit cluster with a task load level of a first level in the plurality of computing unit clusters, it is further determined whether there is a second computing unit cluster with a task load level of a second level; and if there is a second computing unit cluster with a task load level of a second level, the target computing unit cluster is polled and determined in the second computing unit cluster. In this case, if the second computing unit cluster is one, the second computing unit cluster is directly taken as the target computing unit cluster; or if the second computing unit cluster is multiple, different computing unit clusters are selected as the target computing unit cluster in each round of task allocation by polling.

[0099] For example, as shown in FIG. 22, if the task load levels of the computing unit clusters 22 and 23 are both of a second level, the computing unit cluster 23 is taken as the target computing unit cluster in the current round of task allocation, and the computing unit cluster 22 is taken as the target computing unit cluster in a next round of task allocation, so that the tasks are allocated in turns. Figure 2

[0100] In some possible implementation manners, if the task load levels of the plurality of computing unit clusters are all of a third level, the current round of task allocation is skipped, the step of determining the task load levels of the plurality of computing unit clusters is performed again in a next round of task allocation, and the target computing unit cluster is determined according to the task load levels.​

[0101] Figure 5 This is a schematic diagram illustrating a task allocation process provided in an embodiment of this disclosure. Figure 5 As shown, at the start of task allocation, the validity of the next task (the next task after the previous task, i.e., the current task) is first verified; if invalid, the process returns; if valid, the process proceeds to the next step, determining whether there is a first-level computing unit cluster, i.e., a computing unit cluster whose task load value is less than the task's lower limit; if so, the current task is allocated to the computing unit cluster with the lowest task load value; if not, the process continues to determine whether there is a second-level computing unit cluster, i.e., a computing unit cluster whose task load value is greater than or equal to the task's lower limit and less than or equal to the task's upper limit.

[0102] In the example, if a computing unit cluster with a task load level of the second level exists, the target computing unit cluster is determined by polling in the second computing unit cluster and the current task is assigned; if it does not exist, the task assignment in this round is skipped and the process returns to the step of determining whether a computing unit cluster with the first level exists, and the determination is made again in the next round of task assignment.

[0103] This method enables a balanced distribution of tasks and improves the utilization of computing resources.

[0104] In some application scenarios with high requirements for rendering efficiency, it may not be sufficient to consider only the upstream load factor in task allocation. The task load of the computing unit cluster also includes the tasks that are being executed. These tasks may include heavy raw data rendering tasks, which will lead to computing latency.

[0105] In some possible implementations, task allocation methods also include overall allocation methods, which combine relevant information generated during the upstream-to-downstream task allocation with the downstream feedback on the status of tasks currently being executed to determine the task load and allocate tasks.

[0106] In the case that the task allocation mode is the global allocation mode, the step of determining the task load value and the task load level of the plurality of computing unit clusters in step S32 can include: for any computing unit cluster, in the case that the task allocation mode is the global allocation mode, obtaining a first count value of a queue task counter of the computing unit cluster and a second count value of an inflight task counter sent by the computing unit cluster; wherein the queue task counter is arranged in the control module of the processing core, and the inflight task counter is arranged in the computing unit cluster; determining the task load value of the computing unit cluster according to the first count value and a queue load weight, and the second count value and an inflight load weight; and determining the task load level of the computing unit cluster according to the task load value of the computing unit cluster and a preset upper limit value and a lower limit value of the task.

[0107] For example, the inflight task counter can be arranged in each computing unit cluster to count the number of tasks being executed by the computing unit cluster; when the computing unit cluster obtains a rendering task of a tile from the task queue and performs rasterization processing, the count value of the inflight task counter is increased by 1, indicating that the number of rendering tasks in execution is increased by 1; after the computing unit cluster completes the rasterization processing of a task and sends the task to the downstream, for example, initiates a call of a pixel shader, the count value of the inflight task counter is decreased by 1, indicating that the number of rendering tasks in execution is decreased by 1.

[0108] In some possible implementation manners, the computing unit cluster can send the second count value of the inflight task counter to the task allocation module at a certain period, so that the task allocation module schedules tasks based on the second count value.

[0109] For any computing unit cluster, if the task allocation mode is the global allocation mode, the task allocation module can obtain the first count value of the queue task counter of the computing unit cluster and the received second count value of the inflight task counter; and determine the task load value of the computing unit cluster according to the first count value and a preset queue load weight, and the second count value and a preset inflight load weight, and the task load value is expressed as:

[0110] total_workload_size = (fifo_task_counter)×(fifo_workload_weight) +(inflight_task_counter)×(inflight_workload_weight) (1)

[0111] In formula (1), total_workload_size represents the task load value, which can also be called the overall task load level; fifo_task_counter represents the first count value; fifo_workload_weight represents the queue load weight; inflight_task_counter represents the second count value; and inflight_workload_weight represents the execution load weight.

[0112] The sum of the queue load weight and the execution load weight is 1. For example, the queue load weight is set to 0.4 and the execution load weight is set to 0.6. It should be understood that those skilled in the art can set the specific values ​​of the queue load weight and the execution load weight according to the actual situation, and this disclosure does not impose any restrictions on this.

[0113] Figure 6 This is a schematic diagram illustrating a task load value determination process provided in an embodiment of this disclosure. Figure 6 As shown, the processing core comprises N computational unit clusters, numbered 0, 1, ..., N-1, where N is an integer greater than 1. Each computational unit cluster has a task queue FIFO, denoted as task queue 0, task queue 1, ..., task queue N-1; each computational unit within a cluster contains multiple tasks in execution, performing rasterization processing on the tiles, and the count value of the in-process task counter (... Figure 6 The execution task count values ​​(0, 1, ..., N-1) are fed back to the control module so that the control module can calculate the task load value by combining the count value of the queue task counter of the corresponding task queue and realize task allocation.

[0114] The tasks completed by the computing unit cluster are sent to the backend for further processing. Backend processing includes, for example... Figure 1 Early depth testing and texture and shadow shading processing, etc.

[0115] In some possible implementations, after determining the task load value, the task load level of the computing unit cluster can be determined based on the task load value of the computing unit cluster and preset task upper and lower limits. The task upper and lower limits here may be the same as or different from the task upper and lower limits described above, and those skilled in the art can set them according to actual circumstances.

[0116] The process of determining the task load level based on the task load value, the upper limit value, and the lower limit value, as well as the process of determining the target computing unit cluster based on the task load levels of multiple computing unit clusters in step S33, can be found in the previous descriptions and will not be repeated here.

[0117] Figure 7A schematic diagram of a task load state is provided for an embodiment of the present disclosure. As shown in Figure 7 The processing core includes a computing unit cluster 0 and a computing unit cluster 1, the computing unit cluster 0 includes a computing unit 0 and a computing unit 1, and the computing unit cluster 1 includes a computing unit 2 and a computing unit 3. There are multiple task queues from the control module to the computing unit cluster 0 and the computing unit cluster 1, the shaded part 71 in the task queue represents the task load (FIFO workload) of the task to be processed in the task queue, and the shaded part 72 in the computing unit represents the task load (inflight workload) of the task in execution. It can be seen that the task load (FIFO workload) of the task to be processed and the task load (inflight workload) of the task in execution can better reflect the actual task load state of the computing unit cluster.

[0118] By combining the relevant information generated in the task allocation from upstream to downstream and the task state information fed back by downstream, the fine degree of task scheduling can be improved, the balance of task allocation can be further improved, and the utilization rate of computing resources can be further improved.

[0119] After the target computing unit cluster to be allocated is determined, the first rendering task can be sent to the task queue of the target computing unit cluster in step S34, so that the target computing unit cluster executes the first rendering task of the first tile, thereby completing the rendering task allocation process.

[0120] In some possible implementation manners, the target computing unit cluster can perform rasterization and other processing on each primitive of the first tile by the computing units therein, and the present disclosure does not limit the specific rendering processing content and processing manner.

[0121] In some possible implementation manners, the rendering task allocation method according to the embodiment of the present disclosure further includes: in the case of sending the first rendering task to the target computing unit cluster, adding 1 to the first count value of the queue task counter of the target computing unit cluster.

[0122] That is, after the task allocation module sends the rendering task of the first tile to the task queue of the target computing unit cluster, the first count value of the queue task counter of the target computing unit cluster can be added by 1, indicating that a rendering task is added to the task queue. In this way, the accuracy of task queue counting can be improved.

[0123] Figure 8 A schematic diagram of a rendering task allocation state is provided for an embodiment of the present disclosure. As shown in Figure 8 The rendering object includes 6x6=36 tiles, and the processing core for processing multiple tiles of the rendering object includes four computing unit clusters 0, 1, 2 and 3.

[0124] In an example, the information of each tile in the rendering object is sent to the control module of the processing core in the control flow information of the front end of the rendering pipeline; the task establishment module in the control module packs and establishes the rendering task of the tile according to the primitive set information and the hardware related information of the tile.

[0125] In an example, the task distribution module in the control module determines the task distribution mode, respectively determines the task load values of the computing unit clusters 0, 1, 2 and 3, and determines the task load level of each computing unit cluster according to the task load values; the target computing unit cluster is determined according to the task load level and the rendering task of the tile is distributed. The rendering task of each tile is independently judged.

[0126] As shown in Figure 8 , the numbers 0, 1, 2 and 3 in the tile represent the number of the computing unit cluster that executes the rendering task of the tile. As can be seen from the figure, the distribution of tasks of each tile in the computing unit cluster is relatively balanced.

[0127] According to the rendering task distribution method of the embodiment of the present disclosure, on the basis of the rendering architecture such as TBR and TBDR, the scheduling control logic of the control module of the processing core is configured before the rasterization stage, the task distribution mode suitable for the application scenario is selected, and the flexibility and adaptability of the scheduling control are improved. And a new linking mechanism is established between the control module of the processing core and the downstream computing unit cluster that executes rasterization, the current task load can be fed back from the downstream to the control module of the upstream, the upstream and downstream load coordination based on the feedback mechanism is realized, so that the control module of the upstream can determine the task load of the downstream in real time, and the task scheduling is performed based on the task load of the upstream and downstream, thereby realizing the load balancing of the computing resources, improving the utilization rate of the computing resources, improving the overall performance of the tile processing in the graphics rendering pipeline, being easy to implement on the GPU hardware, and improving the rendering effect.

[0128] According to the rendering task distribution method of the embodiment of the present disclosure, the tile can also be split based on the processing capacity of the processing core, the rendering task is established in units of tiles or sub-tiles, so that the size of the tile matches the processing capacity of the processing core, the granularity of the task distribution is smaller, the efficiency of the task distribution is improved, and the computing resources are concentrated in the part with more primitives (especially triangles), thereby further improving the utilization rate of the computing resources.

[0129] According to the rendering task allocation method provided in the embodiments of the present disclosure, the tasks and loads can be balanced, the problem of unbalanced computing resources in processing cores can be solved, the idle time of the computing unit can be reduced, the computing resources in the GPU can be more fully utilized, the parallel computing capability of the GPU can be maximized, and the utilization rate of the computing resources can be improved. Moreover, by balancing the allocation of computing resources, the excessive load of some computing units or the serial dependence of computing tasks can be reduced, so that the time of the entire computing process can be shortened. Parallelization of tasks and load balancing can help improve the efficiency of computing and reduce the computing time. Moreover, balanced allocation of computing resources can reduce resource contention and conflict, optimize the parallelism of the computing process, improve the throughput and response speed of the system, and improve the performance of the system; by balancing the allocation of computing resources, the excessive load of some computing units can be reduced, so that the energy consumption of the system can be reduced, the computing resources can be more effectively utilized, and waste of resources and unnecessary energy consumption can be avoided, which helps improve the energy efficiency of the system.

[0130] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without deviating from the principle logic. Limited by the length of the present disclosure, the present disclosure will not be described again. Those skilled in the art can understand that in the above-mentioned method of the specific implementation, the specific execution order of each step should be determined according to its function and possible internal logic.

[0131] In addition, the present disclosure also provides a rendering task allocation device, an electronic device, and a computer readable storage medium, which can be used to implement any one of the rendering task allocation methods provided by the present disclosure. The corresponding technical solutions and descriptions are described in the method part and are not described again.

[0132] Figure 9 A block diagram of a rendering task allocation device provided by an embodiment of the present disclosure is provided.

[0133] With reference to Figure 9 The present disclosure provides a rendering task allocation device applied to a processing core of a graphics processing device, wherein the processing core comprises a plurality of computing unit clusters. The device comprises: a task establishment module 91 configured to establish a first rendering task of a first tile in a rendering object according to primitive set information of the first tile, wherein the first rendering task comprises primitive information of the first tile and hardware-related information required for executing the first rendering task; and a task allocation module 92 configured to determine task load values and task load levels of the plurality of computing unit clusters according to a set allocation mode, determine a target computing unit cluster from the plurality of computing unit clusters according to the task load levels of the plurality of computing unit clusters, and send the first rendering task to a task queue of the target computing unit cluster, so that the target computing unit cluster executes the first rendering task of the first tile.

[0134] In some possible implementation manners, the task allocation module 92 is configured to: for any computing unit cluster, when the task allocation mode is the upstream allocation mode, determine a task load value of the computing unit cluster according to a first count value of a queue task counter of the computing unit cluster; the queue task counter is arranged in a control module of the processing core; and determine a task load level of the computing unit cluster according to the task load value of the computing unit cluster and preset upper and lower task limit values.

[0135] In some possible implementation manners, the task allocation module 92 is configured to: for any computing unit cluster, when the task allocation mode is the overall allocation mode, obtain a first count value of a queue task counter of the computing unit cluster and a second count value of an executing task counter sent by the computing unit cluster; the queue task counter is arranged in a control module of the processing core, and the executing task counter is arranged in the computing unit cluster; determine a task load value of the computing unit cluster according to the first count value and a queue load weight, and the second count value and an executing load weight; and determine a task load level of the computing unit cluster according to the task load value of the computing unit cluster and preset upper and lower task limit values.

[0136] In some possible implementation manners, the task allocation module 92 is configured to: when the task load value of the computing unit cluster is less than the lower task limit value, determine that the task load level of the computing unit cluster is a first level; when the task load value of the computing unit cluster is greater than or equal to the lower task limit value and less than or equal to the upper task limit value, determine that the task load level of the computing unit cluster is a second level; and when the task load value of the computing unit cluster is greater than the upper task limit value, determine that the task load level of the computing unit cluster is a third level, where the third level has a higher task load than the second level, and the second level has a higher task load than the first level.

[0137] In some possible implementation manners, the task load levels include a first level, a second level and a third level, the task load of the third level is higher than that of the second level, and the task load of the second level is higher than that of the first level. The task allocation module 92 is configured to: in a case where there is a first computing unit cluster with the task load level of the first level in the plurality of computing unit clusters, determine a computing unit cluster with the lowest task load value in the first computing unit cluster as the target computing unit cluster; in a case where there is no first computing unit cluster with the task load level of the first level in the plurality of computing unit clusters, and there is a second computing unit cluster with the task load level of the second level, poll the second computing unit cluster to determine the target computing unit cluster; and in a case where the task load levels of the plurality of computing unit clusters are all the third level, skip the current round of task allocation, determine the task load levels of the plurality of computing unit clusters again in the next round of task allocation, and determine the target computing unit cluster according to the task load levels.

[0138] In some possible implementation manners, the apparatus further includes a first counting module configured to, in a case where the first rendering task is sent to the target computing unit cluster, add 1 to a first counting value of a queue task counter of the target computing unit cluster.

[0139] In some possible implementation manners, the apparatus further includes a second counting module configured to, in a case where a credit feedback signal sent by any computing unit cluster is received, subtract 1 from a first counting value of a queue task counter of the computing unit cluster; and the credit feedback signal is used to indicate that the computing unit cluster obtains one rendering task from a task queue of the computing unit cluster.

[0140] In some possible implementation manners, the task establishment module 91 is further configured to: in a case where control flow information of a second tile to be rendered in the rendering object is received, determine a first size of the second tile; the control flow information includes a number of primitives in a primitive set of the second tile, a primitive address, a primitive offset and the first size; in a case where a processing size of the processing core is smaller than the first size of the second tile, split the second tile to obtain a plurality of first tiles of the second tile; a second size of the first tile is smaller than or equal to the processing size of the processing core; and determine primitive set information of the plurality of first tiles according to the control flow information of the second tile.

[0141] In some possible implementation manners, the rendering manner of the rendering object includes tile-based rendering (TBR) or tile-based deferred rendering (TBDR), wherein the rendering object is divided into a plurality of tiles, and the rendering object includes a plurality of primitives, each of the plurality of tiles is covered by at least one primitive of the rendering object, and the second tile is any one of the plurality of tiles.

[0142] Figure 10 A block diagram of an electronic device is provided for the embodiments of the present disclosure.

[0143] With reference to Figure 10 The embodiments of the present disclosure provide an electronic device, which includes at least one processor 701, at least one memory 702, and one or more I / O interfaces 703 connected between the processor 701 and the memory 702; wherein the memory 702 stores one or more computer programs executable by the at least one processor 701, and the one or more computer programs are executed by the at least one processor 701 to enable the at least one processor 701 to perform the rendering task allocation method described above.

[0144] The embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the rendering task allocation method described above. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0145] The embodiments of the present disclosure also provide a computer program product including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the rendering task allocation method described above.

[0146] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the functions of the modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations. In the hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer readable storage medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media).

[0147] As those skilled in the art will appreciate, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable program instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read only memory (CD-ROM), digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, as those skilled in the art will appreciate, communication media typically embodies computer readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" means a signal that has one or more of its characteristics changed or set in a manner so as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as wireless networks, cellular telephone networks, code division multiple access (CDMA) networks, and other terrestrial and satellite radio frequency communication networks. Thus the computer readable program instructions and / or other program modules can be embodied in a computer readable storage medium, which can be any device or article that is enab!ed to store and / or carry computer readable program instructions and / or data structures. The computer readable storage medium can also be distributed over networked computer systems so that the computer readable program instructions and / or other program modules are stored and executed in a distributed fashion.

[0148] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0149] Computer readable program instructions for carrying out operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or any combination of one or more of the above in any combination, written in any combination of one or more programming languages, including object oriented programming languages such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0150] The computer program product described herein can be embodied in a specific manner by hardware, software, or a combination thereof. In an optional embodiment, the computer program product is embodied as a computer storage medium. In another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK), and the like.

[0151] The various aspects of the present disclosure are described herein with reference to flow diagrams and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer readable program instructions.

[0152] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, to cause a series of operational elements or steps to be performed on the computer to produce a computer implemented process; such that the instructions, which execute via one or more computer program product, implement a computer implemented process for performing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0153] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational elements or steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer implemented process such that the instructions that execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0154] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational elements or steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer implemented process such that the instructions that execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0155] Example embodiments have been disclosed and although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that features, characteristics or elements described with respect to a particular embodiment can be used, alone or in combination, with other embodiments unless specifically recited otherwise. Accordingly, various modifications, alterations, and improvements will readily occur to those skilled in the art with the foregoing description. Accordingly, the present disclosure is not intended to be limited by the method, system, and apparatus disclosed herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for allocating rendering tasks, characterized in that, A processing core applied to a graphics processing device, the processing core comprising multiple clusters of computing units, the method comprising: Based on the primitive set information of the first tile to be rendered in the rendering object, a first rendering task for the first tile is established. The first rendering task includes the primitive set information of the first tile and the hardware-related information required to execute the first rendering task. Based on the set task allocation method, the task load values ​​and task load levels of multiple computing unit clusters are determined respectively; the task allocation method includes any one of the upstream allocation method based on queued tasks and the overall allocation method based on queued tasks and tasks in execution. Based on the task load levels of the multiple computing unit clusters, a target computing unit cluster is determined from the multiple computing unit clusters; The first rendering task is sent to the task queue of the target computing unit cluster so that the target computing unit cluster executes the first rendering task of the first tile.

2. The method according to claim 1, characterized in that, The step of determining the task load values ​​and task load levels of the multiple computing unit clusters according to the set allocation method includes: For any computing unit cluster, when the task allocation method is upstream allocation, the task load value of the computing unit cluster is determined according to the first count value of the queue task counter for the computing unit cluster; the queue task counter is set in the control module of the processing core. The task load level of the computing unit cluster is determined based on the task load value of the computing unit cluster and the preset upper and lower task limits.

3. The method according to claim 1, characterized in that, The step of determining the task load values ​​and task load levels of the multiple computing unit clusters according to the set allocation method includes: For any computing unit cluster, when the task allocation method is the overall allocation method, the first count value of the queue task counter for the computing unit cluster and the second count value of the executing task counter sent by the computing unit cluster are obtained; wherein, the queue task counter is set in the control module of the processing core, and the executing task counter is set in the computing unit cluster; The task load value of the computing unit cluster is determined based on the first count value and the queue load weight, the second count value and the execution load weight; The task load level of the computing unit cluster is determined based on the task load value of the computing unit cluster and the preset upper and lower task limits.

4. The method according to claim 2 or 3, characterized in that, The step of determining the task load level of the computing unit cluster based on the task load value of the computing unit cluster and preset task upper and lower limits includes: If the task load value of the computing unit cluster is less than the task lower limit value, the task load level of the computing unit cluster is determined to be the first level; If the task load value of the computing unit cluster is greater than or equal to the lower limit of the task and less than or equal to the upper limit of the task, the task load level of the computing unit cluster is determined to be the second level. If the task load value of the computing unit cluster is greater than the task upper limit value, the task load level of the computing unit cluster is determined to be level three. The workload of the third level is higher than that of the second level, and the workload of the second level is higher than that of the first level.

5. The method according to claim 1, characterized in that, The task load levels include a first level, a second level, and a third level. The task load of the third level is higher than that of the second level, and the task load of the second level is higher than that of the first level. The step of determining the target computing unit cluster from the plurality of computing unit clusters based on the task load levels of the plurality of computing unit clusters includes: If there is a first computing unit cluster with a task load level of first level among the multiple computing unit clusters, the computing unit cluster with the lowest task load value among the first computing unit clusters shall be determined as the target computing unit cluster. If there is no first computing unit cluster with a task load level of the first level among the multiple computing unit clusters, but there is a second computing unit cluster with a task load level of the second level, the target computing unit cluster is determined by polling among the second computing unit clusters. If the task load level of multiple computing unit clusters is all at level three, skip the current round of task allocation, determine the task load level of multiple computing unit clusters again in the next round of task allocation, and determine the target computing unit cluster based on the task load level.

6. The method according to claim 2 or 3, characterized in that, The method further includes: When the first rendering task is sent to the target computing unit cluster, the first count value of the queue task counter of the target computing unit cluster is incremented by 1.

7. The method according to claim 2 or 3, characterized in that, The method further includes: Upon receiving a credit feedback signal from any computing unit cluster, the first count value of the queue task counter of the computing unit cluster is decremented by 1. The credit feedback signal is used to instruct the computing unit cluster to obtain a rendering task from the task queue of the computing unit cluster.

8. The method according to claim 1, characterized in that, The method further includes: Upon receiving control flow information of the second tile to be rendered in the rendering object, a first size of the second tile is determined; the control flow information includes the number of primitives in the primitive set of the second tile, primitive addresses, primitive offsets, and the first size; If the processing size of the processing core is smaller than the first size of the second block, the second block is split into multiple first blocks; the second size of the first block is smaller than or equal to the processing size of the processing core. Based on the control flow information of the second block, the primitive set information of multiple first blocks is determined respectively.

9. The method according to claim 8, characterized in that, The rendering method for the rendered object includes tile-based rendering (TBR) or tile-based deferred rendering (TBDR). The rendering object is divided into multiple tiles, and the rendering object includes multiple primitives. Each tile in the multiple tiles is covered by at least one primitive of the rendering object, and the second tile is any one of the multiple tiles.

10. A rendering task allocation device, characterized in that, A processing core for a graphics processing device, the processing core comprising multiple clusters of computing units, the device comprising: The task creation module is used to: create a first rendering task for the first tile based on the primitive set information of the first tile to be rendered in the rendering object. The first rendering task includes the primitive information of the first tile and the hardware-related information required to execute the first rendering task. The task allocation module is configured to: determine the task load values ​​and task load levels of multiple computing unit clusters according to a set allocation method; the task allocation method includes any one of an upstream allocation method based on queued tasks and an overall allocation method based on queued tasks and tasks in execution; determine a target computing unit cluster from the multiple computing unit clusters according to the task load levels of the multiple computing unit clusters; and send the first rendering task to the task queue of the target computing unit cluster so that the target computing unit cluster executes the first rendering task of the first tile.

11. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the rendering task allocation method as described in any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the rendering task allocation method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Rendering control method and device and rendering system

    CN116775298A