Core computing processor and device and method for dynamically distributing workload

By using the method of dynamically allocating work blocks with load credit values ​​in multi-core GPU systems, the problem of execution time difference caused by fixed mode allocation is solved, and the overall computing efficiency of the GPU is improved.

CN120144281APending Publication Date: 2025-06-13BEIJING FENGHUA CHUANGZHI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510170828.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The multi-core GPU system distributes workloads in a fixed mode during the computing processing stage, resulting in a large difference in the total execution time between GPU cores, affecting the overall computing efficiency.

Method used

By calculating the load credit value of the processor, the work block is dynamically allocated, ensuring that the difference in workload between each GPU core is reduced, thereby achieving dynamic allocation of workloads.

Benefits of technology

It effectively reduces the difference in workload in each working group, ensures the overall computing efficiency of the GPU, and solves the problem of total execution time delay caused by fixed mode allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144281A_ABST
    Figure CN120144281A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of high-performance computing, and provides a core computing processor and a device and method for dynamically distributing workloads. In the invention, a primary core computing processor is connected with each secondary core computing processor; the secondary core computing processor determines a load credit value of the secondary core computing processor and sends the load credit value of the secondary core computing processor to the primary core computing processor; and the main core computing processor determines a load credit value of the main core computing processor, allocates a working block to the secondary core computing processor or the main core computing processor meeting a preset condition according to the load credit values from other core calculators and the load credit value of the main core computing processor, and further allocates the working block to the secondary core computing processor or the main core computing processor according to the offset of the current working group and the offset of the current working block. Selectively processing the current work group in the primary core computing processor or processing the current work group in the secondary core computing processor; therefore, the dynamic distribution of workload is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of high-performance computing, and particularly to a core computing processor, a device and a method for dynamic workload allocation. Background Art

[0002] A Graphics Processing Unit (GPU) is a hardware specifically responsible for graphics computing and processing. With the development of high-performance parallel computing, GPUs are widely used to accelerate general computing applications. For example, currently popular artificial intelligence applications such as Large Language Models (LLMs).

[0003] The workload of a GPU refers to the amount of computing resources used by the GPU, that is, the load of the GPU when performing a computing task. The level of GPU load is an important indicator to measure the usage efficiency of the GPU. When the GPU load is high, it means that the GPU is processing large and complex computing tasks, and at the same time, it also means that the GPU is making full use of computer resources, thus improving the computing efficiency. However, when the GPU load is too high, it may also cause the computer to overheat, freeze or even crash. Therefore, it is necessary to reasonably control the GPU load when performing computing tasks.

[0004] In the prior art, in the entire computing and processing stage, a multi-core GPU system often allocates workloads to each GPU core in a fixed mode. For example, a constant number of executions is predefined through a software driver, and workgroups are allocated among GPU cores according to this constant number of executions. When the workloads in each workgroup are equivalent, this fixed mode can ensure GPU performance, and there will be no significant difference in the total execution time among GPU cores. However, when there are differences in the workloads among workgroups, the execution cycles of the workloads among GPU cores are very likely to vary greatly. Therefore, there is a significant difference in the total execution time among GPU cores. Furthermore, due to the difference in the completion time of execution, the later-completed GPU core will cause an additional delay in the total execution time, affecting the overall computing efficiency of the GPU.

[0005] In view of this, overcoming the defects of the prior art is an urgent problem to be solved in this technical field. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a core computing processor, a device and a method for dynamic workload allocation. The purpose is to dynamically allocate work blocks for each GPU core according to the load credit value, realize dynamic workload allocation, and ensure the overall computing efficiency of the GPU by minimizing the workload difference in each workgroup, so as to solve the problem of additional delay in the total execution time caused by the fixed-mode workload allocation in the entire computing and processing stage of the multi-core GPU system.

[0007] The present invention adopts the following technical solutions:

[0008] In the first aspect, the present invention provides a core computing processor, including a computing and processing module;

[0009] The computing and processing module is used to determine the load credit value of the core computing processor where the computing and processing module is located; and is also used to allocate work blocks for the core computing processors that meet the preset conditions according to the load credit values from other core calculators and its own load credit value; wherein, one work block includes multiple workgroups.

[0010] It is also used to selectively process the current workgroup by the core computing processor where it is located or by other core computing processors according to the current workgroup offset and the current work block offset.

[0011] Further, the computing and processing module includes a work block allocation unit and a work block buffer unit;

[0012] The work block allocation unit is used to determine the difference between the maximum number of work blocks in the work block buffer unit and the current number of work blocks, and obtain the buffer space credit value of the core computing processor;

[0013] The work block allocation unit is also used to determine the workload corresponding to the work blocks of each core computing processor, and obtain the corresponding work credit value;

[0014] The work block allocation unit is also used to obtain the load credit value according to the buffer space credit value and the work credit value corresponding to the work blocks of the core computing processor;

[0015] The work block allocation unit is also used to determine the maximum load credit value among the multiple load credit values from other core calculators and its own load credit value, and allocate work blocks for the core computing processor corresponding to the maximum load credit value.

[0016] Further, it further includes a shader execution unit; the computing and processing module further includes a workgroup allocation unit;

[0017] The working block buffer unit is used to store the working block offset of the working block allocated to the core computing processor;

[0018] The workgroup allocation unit is used to obtain the current workgroup offset of the current workgroup; when the current workgroup offset is within the range of the workgroups included in the current working block stored in the working block buffer unit, it is determined that the core computing processor processes the current workgroup; when the current workgroup offset is not within the range of the workgroups included in the current working block, the other core computing processors process the current workgroup;

[0019] The workgroup allocation unit is also used to allocate the workgroups to be processed by the core computing processor to multiple shader execution units;

[0020] The shader execution unit is used to process the workgroups allocated by the workgroup allocation unit.

[0021] In a second aspect, the present invention also provides a device for dynamically allocating workloads, including a main core computing processor and multiple secondary core computing processors;

[0022] The main core computing processor is connected to each secondary core computing processor;

[0023] The secondary core computing processor is used to determine its own load credit value and send the own load credit value to the main core computing processor;

[0024] The main core computing processor is used to determine its own load credit value; it is also used to allocate working blocks to the secondary core computing processors or the main core computing processor that meet the preset conditions according to the load credit values from other core calculators and its own load credit value; it is also used to selectively process the current workgroup in the main core computing processor or in the secondary core computing processor according to the current workgroup offset and the current working block offset.

[0025] The main core computing processor is also used to determine the number of workgroups included in the working block according to the workload of the workgroup; among them, when the workload is greater than the first threshold, the number of workgroups included in the working block is less than the number of workgroups included in the working block when the workload is less than the second threshold, and the first threshold is greater than or equal to the second threshold;

[0026] The main core computing processor is also used to determine the processing order of the working blocks according to the relevance of the workloads of the working blocks; when there is no relevance between the corresponding workloads, the working blocks with workloads greater than the third threshold are preferentially processed.

[0027] Further, the work block allocation unit of the main core computing processor is configured to receive the load credit values sent by the work block allocation units of each sub-core computing processor, determine its own load credit value, determine the maximum load credit value among the multiple load credit values from each sub-core calculator and its own load credit value, and allocate work blocks to the main core computing processor or sub-core computing processor corresponding to the maximum load credit value;

[0028] The work block buffer unit of the main core computing processor and / or the work block buffer units of each sub-core computing processor are configured to receive the offsets of the allocated work blocks.

[0029] In a third aspect, the present invention further provides a method for dynamic workload allocation, including:

[0030] Allocating work blocks to the main core computing processor or sub-core computing processor that meets the preset conditions according to the load credit value of the main core computing processor and the load credit values of each sub-core computing processor;

[0031] Processing the current work group in the main core computing processor or in the sub-core computing processor selectively according to the current work group offset and the current work block offset.

[0032] Further, the step of allocating work blocks to the main core computing processor or sub-core computing processor that meets the preset conditions according to the load credit value of the main core computing processor and the load credit values of each sub-core computing processor includes:

[0033] Determining the difference between the maximum number of work blocks of the main core computing processor and the current number of work blocks to obtain the buffer space credit value of the main core computing processor; determining the difference between the maximum number of work blocks of each sub-core computing processor and the current number of work blocks to obtain the buffer space credit value of the corresponding sub-core computing processor;

[0034] Determining the workload corresponding to the work blocks of the main core computing processor to obtain the work credit value of the main core computing processor; determining the workloads corresponding to the work blocks of each sub-core computing processor to obtain the work credit values of the corresponding sub-core computing processors;

[0035] Obtaining the load credit value of the main core computing processor according to the buffer space credit value and work credit value of the main core computing processor; obtaining the load credit values of the corresponding sub-core computing processors according to the buffer space credit values and work credit values of each sub-core computing processor;

[0036] Determine the maximum load credit value among multiple load credit values from each sub-core calculator and its own load credit value; allocate a work block to the main core computing processor or sub-core computing processor corresponding to the maximum load credit value.

[0037] Further, the step of allocating a work block to the main core computing processor or sub-core computing processor that meets the preset conditions according to the load credit value of the main core computing processor and the load credit values of each sub-core computing processor further includes:

[0038] The main core computing processor determines the number of workgroups included in the work block according to the workload of the workgroup; wherein, when the workload is greater than the first threshold, the number of workgroups included in the work block is less than the number of workgroups included in the work block when the workload is less than the second threshold, and the first threshold is greater than or equal to the second threshold;

[0039] The main core computing processor determines the processing order of the work block according to the relevance of the workloads of the work block; when there is no relevance between the corresponding workloads, the work block with a workload greater than the third threshold is preferentially processed;

[0040] When the number of instructions included in the shader is less than the fourth threshold and the workloads in each workgroup are equal, use a fixed allocation mode to allocate workgroups to the core computing processors;

[0041] When the number of instructions included in the shader is greater than the fifth threshold and the workloads in each workgroup are different, use a dynamic allocation mode to allocate workgroups to the core computing processors; wherein, the fifth threshold is greater than or equal to the fourth threshold.

[0042] Further, the step of using a dynamic allocation mode to allocate workgroups to the core computing processors when the number of instructions included in the shader is greater than the fifth threshold and the workloads in each workgroup are different includes:

[0043] In the dynamic allocation mode, use a fixed allocation mode to allocate initialization work blocks to each core computing processor, and each core computing processor starts processing the work block immediately at the start of processing.

[0044] Further, it further includes:

[0045] When the current computing task is paused and another computing task needs to be started, allocate all work block entry fields to the corresponding core computing processors for processing; wherein, the work block entry field represents the current work block offset of the corresponding core computing processor when the current computing task is paused.

[0046] Determine the work block corresponding to the last work block entry field among all the work block entry fields, and starting from the determined work block, allocate the subsequent work blocks to the main core computing processor and / or the secondary core computing processor for processing.

[0047] In a fourth aspect, the present invention further provides a non-volatile computer storage medium storing computer-executable instructions that, when executed by one or more processors, are used to complete the method for dynamically allocating workloads described in the first aspect.

[0048] In a fifth aspect, there is provided a chip including a processor and an interface for calling and running a computer program stored in a memory from the memory and executing the method for dynamically allocating workloads as described in the first aspect.

[0049] In a sixth aspect, there is provided a computer program product containing instructions that, when run on a computer or a processor, cause the computer or the processor to execute the method for dynamically allocating workloads as described in any one of the first aspect to the fourth aspect.

[0050] In a seventh aspect, there is provided a system for dynamically allocating workloads, including the device for dynamically allocating workloads as described in the second aspect, and using the method for dynamically allocating workloads as described in the third aspect to complete the interaction of the device for dynamically allocating workloads as described in the second aspect.

[0051] Different from the prior art, the present invention has at least the following beneficial effects:

[0052] The present invention provides a device for dynamically allocating workloads, wherein a main core computing processor is connected to each secondary core computing processor; the secondary core computing processor determines its own load credit value and sends its own load credit value to the main core computing processor; the main core computing processor determines its own load credit value, and based on the load credit values from other core calculators and its own load credit value, allocates work blocks to the secondary core computing processors or the main core computing processor that meet the preset conditions, and then, according to the current workgroup offset and the current work block offset, selectively processes the current workgroup in the main core computing processor or in the secondary core computing processor to achieve dynamic workload allocation, by minimizing the difference in the amount of work in each workgroup as much as possible, ensuring the overall computing efficiency of the GPU, and solving the problem of additional delay in the total execution time caused by the fixed-mode workload allocation in the entire computing and processing stage of the multi-core GPU system. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0054] Figure 1 is a schematic diagram of a computing load pipeline of a multi-core GPU in the prior art provided by an embodiment of the present invention;

[0055] Figure 2 is a schematic diagram of the workgroup distribution in a multi-core GPU in the prior art provided by an embodiment of the present invention;

[0056] Figure 3 is a schematic diagram of a core computing processor provided by an embodiment of the present invention;

[0057] Figure 4 is a specific example of a device for dynamically allocating workloads provided by an embodiment of the present invention;

[0058] Figure 5 is a schematic flowchart of a method for dynamically allocating workloads provided by an embodiment of the present invention;

[0059] Figure 6 is a schematic flowchart of step 10 provided by an embodiment of the present invention;

[0060] Figure 7 is another schematic flowchart of step 10 provided by an embodiment of the present invention;

[0061] Figure 8 is a schematic flowchart of workgroup allocation provided by an embodiment of the present invention;

[0062] Figure 9 is a schematic diagram of a context state memory provided by an embodiment of the present invention;

[0063] Figure 10 is a schematic flowchart of another method for dynamically allocating workloads provided by an embodiment of the present invention;

[0064] Figure 11 is a schematic diagram of the architecture of a core computing processor provided by an embodiment of the present invention. Detailed Embodiments

[0065] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0066] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0067] Unless the context otherwise requires, throughout the specification and claims, the term "comprising" is interpreted in an open, inclusive sense, i.e., "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "examples", "specific examples" or "some examples", etc. are intended to indicate that the specific features, structures, materials or characteristics related to the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representations of the above terms are not necessarily directed to the same embodiment or example. In addition, the specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any appropriate manner, that is, although they may be carried in the above-mentioned embodiments or examples due to reasons such as the order of appearance and position, etc., but it is not limited that they can be carried by one embodiment or example in a combined manner.

[0068] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present disclosure.

[0069] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "a plurality" is two or more. In addition, for example, in the description, for the same type of nouns, the method of adding "A" and "B" at the end is used to describe them as two independent individuals. In this case, the features defined with "A" and "B" are only used for the purpose of distinguishing similar individuals and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features.

[0070] In describing some embodiments, expressions such as "coupled", "coupling", and "connected" and their derivatives may be used. For example, in describing some embodiments, the term "connected" may be used to indicate that two or more components have direct physical or electrical contact with each other. Another example is that in describing some embodiments, the term "coupled" may be used to indicate that two or more components have direct physical or electrical contact. However, the term "connected" or "coupled" may also mean that two or more components do not have direct contact with each other, but still cooperate or interact with each other, such as "optical path coupling", "wireless connection", etc. The embodiments disclosed herein are not necessarily limited to the content of the present invention.

[0071] In the description of the present invention, the expression "A and / or B" (where A and B are used to formally represent specific feature contents) will be involved. The corresponding expression includes the following three combinations: only A, only B, and the combination of A and B.

[0072] As used in the present invention, "about", "substantially", or "approximately" includes the stated value and the average value within an acceptable deviation range of the specific value, where the acceptable deviation range is determined by a person of ordinary skill in the art considering the measurement being discussed and the errors associated with the measurement of a particular quantity (i.e., the limitations of the measurement system).

[0073] To execute the computing tasks of an application, the corresponding workload is sent to the GPU hardware through the computing control flow. The GPU executes the computing tasks according to its own state, data, and the code of the shader and generates the final result. The computing control flow may include multiple compute kicks, each compute kick may include multiple work groups, and each work group may include multiple work items executed in the GPU hardware.

[0074] Such as Figure 1As shown, in a typical multi-core GPU architecture, the workload is arranged into workgroups by a Compute Processing Model (CPM). The CPM selects one Shader Execution Unit (SEU) from multiple SEUs of the GPU core where it is located for the corresponding workgroup, and then the workgroup is assigned to the selected downstream SEU to execute computational tasks. The distribution of workgroups is as follows: (1) The CPMs from each GPU core read the same computational control flow; (2) During the entire computational processing stage, the workload of the GPU core is in a fixed mode; (3) The workgroups allocated between GPU cores are determined by a predefined constant number of executions controlled by a software driver; (4) The offset of the workgroups assigned to each GPU core can be calculated based on the execution count and the offset of the GPU core. Here, the execution count refers to the number of workgroups processed by the GPU core each time.

[0075] When the execution count is 2, a specific example of the distribution of workgroups during the calculation process is as Figure 2 shown. The above typical multi-core GPU architecture solution has the lowest requirement for communication between GPU cores, so it is easier to implement in GPU hardware. When the workloads in the workgroups are equivalent, this solution should work well, so there is not much difference in the total execution time between GPU cores. However, when the workloads of the workgroups are not similar and the difference cannot be ignored, the execution cycles of the workloads between GPU cores may vary greatly, so the total execution times of multiple GPU cores may vary too much, and thus the GPU core that finishes execution later will bring an additional total execution time delay.

[0076] To solve the above problems, as Figure 3 shown, an embodiment of the present invention provides a core computing processor, including a computational processing module;

[0077] The computational processing module is used to determine the load credit value of the core computing processor where the computational processing module is located; and is also used to allocate work blocks to the core computing processors that meet the preset conditions according to the load credit values from other core calculators and its own load credit value. One work block includes multiple workgroups.

[0078] It is also used to selectively process the current workgroup by the core computing processor where it is located or by other core computing processors according to the current workgroup offset and the current work block offset.

[0079] Among them, after improving the CPM, the calculation and processing module implemented by the present invention can be obtained. For the specific units included in the calculation and processing module, please refer to the following introduction. The core calculation processor is essentially the GPU kernel.

[0080] The preset conditions are selected by those skilled in the art according to the specific usage scenarios.

[0081] The embodiment of the present invention provides a concept of a work block, as Figure 2 shown. The definition of a work block is the basic unit for the workload distribution of the multi-core CPM. A work block contains several workgroups, and the number of workgroups can be the same as the execution count; the execution count is predefined by the software driver. The number of workgroups included in a work block can be determined according to the workload size of the workgroups. When the workload is larger, the number of workgroups included in the work block is less than that when the workload is smaller. It should be noted that the concept of a work block can be the same as the concept of the execution count in the prior art; in an alternative embodiment, the work block can also be set to a parameter different from the execution count. The workgroups in the embodiment of the present invention have the same concept as the workgroups in the prior art. The workgroups in the work block are processed by the CPM of the GPU kernel in the multi-core GPU system. The offset of the work block in the entire calculation load (workgroup_X; workgroup_Y; workgroup_Z) can be the coordinates of the first workgroup in the work block.

[0082] As Figure 2 shown, each work block includes 2 workgroups, that is, the core calculation processor (equivalent to the GPU kernel in the prior art) processes 2 workgroups each time. The GPU kernel offset refers to the serial number of the core calculation processor. For example, core calculation processor 0, core calculation processor 1, etc. The workgroup offset refers to the serial number of the workgroup where a core calculation processor starts to process. For example Figure 2 in which core calculation processor 0 starts to process workgroup 0, so the workgroup offset is 0; core calculation processor 1 starts to process workgroup 2, so the workgroup offset is 2; for any core calculation processor, the workgroup offset refers to the following formula:

[0083] The workgroup offset of core calculation processor n = the number of workgroups in the work block × n

[0084] The work block offset refers to the serial number of the work block. According to Figure 2 the fixed allocation mode shown (that is, the fixed mode used to allocate the GPU kernel in the prior art), for any core calculation processor, the work block offset refers to the following formula:

[0085] The work block offset of core calculation processor n = n

[0086] Workgroup offset of the core computing processor n

[0087] = Workblock offset of the core computing processor n × Number of workgroups in the workblock

[0088] For each core computing processor, the current workgroup offset is the coordinate of the workgroup allocated to this core computing processor this time; the current workblock offset is the coordinate of the first workgroup in the workblock already stored by this core computing processor. Each core computing processor parses the computing control flow to obtain the current workgroup offset, compares the current workgroup offset with the current workblock offset in its own workblock buffer unit, and then determines whether it is suitable for processing itself.

[0089] In an alternative embodiment, as Figure 3 shown, the computing processing module includes a workblock allocation unit and a workblock buffer unit. The workblock allocation unit is used to determine the difference between the maximum number of workblocks of the workblock buffer unit and the current number of workblocks, so as to obtain the buffer space credit value of the core computing processor; the workblock allocation unit is also used to determine the workload corresponding to the workblocks of each core computing processor, so as to obtain the corresponding work credit value; the workblock allocation unit is also used to obtain the load credit value according to the buffer space credit value and the work credit value corresponding to the workblocks of the core computing processor; the workblock allocation unit is also used to determine the maximum load credit value among the multiple load credit values from other core calculators and its own load credit value, and allocate workblocks to the core computing processor corresponding to the maximum load credit value.

[0090] As Figure 3 shown, the core computing processor further includes a shader execution unit; the computing processing module further includes a workgroup allocation unit.

[0091] The workblock buffer unit is used to store the workblock offset of the workblock allocated to the core computing processor.

[0092] The workgroup allocation unit is used to obtain the current workgroup offset of the current workgroup; when the current workgroup offset is within the range of the workgroups included in the current workblock stored in the workblock buffer unit, it is determined that the core computing processor processes the current workgroup; when the current workgroup offset is not within the range of the workgroups included in the current workblock, the current workgroup is processed by other core computing processors.

[0093] The workgroup allocation unit is also used to allocate the workgroups that the core computing processor needs to process to multiple shader execution units.

[0094] The shader execution unit is used to process the workgroups assigned by the workgroup allocation unit.

[0095] In the multi-core GPU system according to the embodiment of the present invention, the computing processing module of the corresponding core computing processor processes the computing control flow of the entire computing processing module from beginning to end. The core computing processor processes the workgroups in the computing control flow in a fixed allocation mode, which is controlled by an execution count parameter. When a workgroup is processed by a certain core computing processor within the fixed allocation mode, the workgroup will be sent from the computing processing module of the core computing processor to a more downstream shader execution unit (i.e., a SEU provided by the embodiment of the present invention) pipeline for processing, otherwise the workgroup will be skipped and not processed. Whether the corresponding core computing processor skips the last workgroup or not, the computing processing modules of all core computing processors know the information of the last workgroup and the termination of the computing control flow.

[0096] The computing processing module in each core computing processor parses the computing control flow for each computing kick. Therefore, when parsing the computing control flow, each core computing processor can know the information of the last workgroup in its own computing processing module, even if the core computing processor skips the workgroup.

[0097] Based on the above core computing processor architecture, the embodiment of the present invention provides a device for dynamic workload allocation, including a main core computing processor and multiple secondary core computing processors.

[0098] The main core computing processor is connected to each secondary core computing processor.

[0099] The secondary core computing processor is used to determine its own load credit value and send the own load credit value to the main core computing processor.

[0100] The main core computing processor is used to determine its own load credit value; is also used to allocate work blocks to the secondary core computing processors or the main core computing processor that meet the preset conditions according to the load credit values from other core calculators and its own load credit value; is also used to selectively process the current workgroup in the main core computing processor or in the secondary core computing processor according to the current workgroup offset and the current work block offset.

[0101] Wherein, the number of secondary core computing processors is selected by those skilled in the art according to specific usage scenarios and is not limited herein. For example Figure 4The following is a specific example of a device for dynamically allocating workloads provided by an embodiment of the present invention. In this specific example, there are 4 core computing processors, including one main core computing processor and 3 secondary core computing processors. It should be noted that in the embodiments of the present invention, the design of the computing processing modules in each core computing processor is symmetric, and each core computing processor has the functions of both the main core computing processor and the secondary core computing processor. Therefore, any core computing processor can be set as the main core computing processor or the secondary core computing processor; each core computing processor can implement the functions of both the main core computing processor and the secondary core computing processor. After setting the primary and secondary, the only difference among the core computing processors lies in their operating modes.

[0102] In the embodiments of the present invention, the computing workload in the workgroup is allocated by the computing processing module in the main core computing processor to each core computing processor (including the main core computing processor and the secondary core computing processors), thereby reducing the workload pressure on the main core computing processor and making the workloads of the workgroups in each core computing processor relatively equal, so as to achieve the purpose of minimizing the difference in the total execution time of each core computing processor as much as possible. As Figure 4 shown, the computing processing module of the main core computing processor allocates work blocks to the computing processing module of the secondary core computing processor through an inter-core connection; the computing processing module of the secondary core computing processor indicates the completion status of the workgroup processing to the computing processing module of the main core computing processor through the inter-core connection. According to the workload status of the core computing processor, taking the work block as the basic unit, the current workgroup parsed from the computing control flow is dynamically allocated to the next best available core computing processor (including the main core computing processor and the secondary core computing processors). In an optional embodiment, the number of workgroups allocated to each core computing processor each time can be configured through software settings. The number of workgroups included in a work block can be determined according to the workload size of the workgroup; when the workload is larger, the number of workgroups included in a work block is less than when the workload is smaller.

[0103] The present invention provides a device for dynamically allocating workloads. Among them, a main core computing processor is connected to each secondary core computing processor; the secondary core computing processor determines its own load credit value and sends its own load credit value to the main core computing processor; the main core computing processor determines its own load credit value, and based on the load credit values from other core calculators and its own load credit value, allocates work blocks to the secondary core computing processors or the main core computing processor that meet the preset conditions, and then, according to the current workgroup offset and the current work block offset, selectively processes the current workgroup in the main core computing processor or in the secondary core computing processor to achieve dynamic workload allocation. The workgroups that are kicked in the calculation will be preferentially allocated to the core computing processors with less workload and stronger computing processing capabilities. By minimizing the difference in workload among each workgroup as much as possible, the overall computing efficiency of the GPU is guaranteed, and the problem of additional delay in the total execution time caused by the fixed allocation mode for allocating workloads in the multi-core GPU system during the entire computing processing stage is solved.

[0104] The work block allocation unit of the main core computing processor is used to receive the load credit values sent by the work block allocation units of each secondary core computing processor, and determine its own load credit value, determine the maximum load credit value among the multiple load credit values from each secondary core calculator and its own load credit value, and allocate work blocks to the main core computing processor or secondary core computing processor corresponding to the maximum load credit value.

[0105] As Figure 4 shown, in the multi-core GPU architecture of the embodiment of the present invention, inter-core communication is required between each core computing processor to facilitate the dynamic allocation of work blocks. The computing processing module of each core computing processor can calculate the load credit value of the core computing processor where it is located; among them, the role of the load credit value is equivalent to the weight coefficient for dynamically allocating work blocks. The load credit value of the computing processing module of each secondary core computing processor is sent to the computing processing module of the main core computing processor through an inter-core connection. The computing processing module of the main core computing processor also calculates its own load credit value. During the process of executing the computing task, the computing processing module of the main core computing processor will dynamically allocate work blocks to each core computing processor that meets the preset conditions according to the load credit scores of itself and all other secondary core computing processors according to the workload allocation strategy, so that each core computing processor can be utilized evenly.

[0106] The working block buffer unit of the main core computing processor and / or the working block buffer unit of each sub-core computing processor are used to receive the offset of the allocated working block; the main core computing processor is further used to determine the number of workgroups included in the working block according to the workload of the workgroup; wherein, when the workload is greater than the first threshold, the number of workgroups included in the working block is less than that when the workload is less than the second threshold, and the first threshold is greater than or equal to the second threshold; the main core computing processor is further used to determine the processing order of the working block according to the correlation of the workloads of the working blocks; when there is no correlation between the corresponding workloads, the working block with a workload greater than the third threshold is preferentially processed. Wherein, the first threshold, the second threshold and the third threshold are selected by those skilled in the art according to the specific usage scenario and are not limited herein.

[0107] In order to minimize the workgroup allocation delay between the computing processing module of the main core computing processor and the computing processing module of the main core computing processor or the sub-core computing processor as much as possible, an embodiment of the present invention adds a working block buffer unit to the computing processing module of each core computing processor to store the working block offset of the working block allocated to the corresponding core computing processor.

[0108] It should be noted that only the working block offset needs to be stored in the working block buffer unit, so the hardware cost of the working block buffer unit is very low. The maximum depth of the working block buffer unit can be predefined by those skilled in the art through a software driver in combination with the specific usage scenario, and the depth of the working block buffer unit can be determined through experiments to reduce the delay of dynamic working block allocation, so as to achieve the best performance on a multi-core GPU system.

[0109] Based on the above multi-core GPU architecture, as Figure 5 shown, an embodiment of the present invention provides a method for dynamic workload allocation, including:

[0110] Step 10: Allocate working blocks to the main core computing processor or the sub-core computing processor that meets the preset conditions according to the load credit value of the main core computing processor and the load credit values of each sub-core computing processor.

[0111] Wherein, the preset conditions are selected by those skilled in the art according to the specific usage scenario.

[0112] Step 20: Selectively process the current workgroup in the main core computing processor or process the current workgroup in the sub-core computing processor according to the current workgroup offset and the current working block offset.

[0113] To reduce the latency of allocating work blocks from the computing processing module of the main core computing processor to each core computing processor at the start of a compute kick, embodiments of the present invention are statically started at the start of the compute kick. That is, each core computing processor, whether it is the main core computing processor or the secondary core computing processor, will automatically start processing work groups according to the fixed allocation mode as shown in Figure 2 and will not receive the allocation from the main core computing processor through the inter-core connection. In this way, all core computing processors can immediately start processing work groups at the start of the compute kick. Among them, the number of work groups allocated in the fixed allocation mode can be a constant predefined by the software driver or the same as the maximum number of work blocks in the work block buffer unit.

[0114] To illustrate the process of dynamic allocation of the work block allocation unit in the main core computing processor, as shown in Figure 6 , step 10 includes:

[0115] Step 101a: Determine the difference between the maximum number of work blocks and the current number of work blocks of the main core computing processor to obtain the buffer space credit value of the main core computing processor; determine the difference between the maximum number of work blocks and the current number of work blocks of each secondary core computing processor to obtain the corresponding buffer space credit value of the secondary core computing processor.

[0116] Embodiments of the present invention use the difference between the maximum number of work blocks and the current number of work blocks to measure the load credit value (i.e., buffer space credit value) of the buffer space in the core computing processor. The calculation methods of the buffer space credit values in the main core computing processor and the secondary core computing processor are the same, and both are calculated according to the state of the work block buffer unit; it can be determined according to how many work blocks are in the work block buffer unit or how many empty spaces are in the work block buffer unit. In an alternative embodiment, if the maximum number of work blocks in the work block buffer unit of a certain core computing processor n is WB max , and if the current number of work blocks in the work block buffer unit of this core computing processor n is WB n , then the buffer space credit value CRS n of this core computing processor n is:

[0117] CRS n = WB max - WB n

[0118] When there are more blank spaces in the work block buffer of the core computing processor, the corresponding buffer space credit value is higher.

[0119] When the shader corresponding to the work block contains a large number of instructions, the workload of the work block increases, and the processing time in the GPU will be extended. Therefore, the work credit value CRWn of the main core computing processor can be obtained according to the number of instructions (i.e., workload) contained in the shader corresponding to the work block of the main core computing processor:

[0120]

[0121] where NIn is the sum of the number of instructions contained in the shaders corresponding to all work blocks in the work block buffer unit of the main core computing processor. M is a preset constant, such as 50 or 100; to avoid the denominator being 0 when NIn is 0, NIn + 1 is used as the denominator of CRWn. When the sum of the number of instructions increases, the value of CRWn will decrease. When the sum of the number of instructions contained in the shaders of all work groups in the work block buffer of the core computing processor is small, the work credit value CRWn is high.

[0122] Finally, the load credit value CRn of the main core computing processor can be determined according to the buffer space credit value CRS n and the work credit value CRWn of the main core computing processor:

[0123] CRn = Ws * CRS n + Ww * CRWn

[0124] where Ws is the weighting factor corresponding to the buffer space credit value CRS n and Ww is the weighting factor corresponding to the work credit value CRW n The load credit value CR n of the main core computing processor will increase with the increase of CRS n and CRWn.

[0125] Step 102a: Determine the workload corresponding to the work block of the main core computing processor to obtain the work credit value of the main core computing processor; determine the workload corresponding to the work block of each sub-core computing processor to obtain the corresponding work credit value of each sub-core computing processor.

[0126] Step 103a: Obtain the load credit value of the main core computing processor according to the buffer space credit value and work credit value of the main core computing processor; obtain the load credit value of the corresponding sub-core computing processor according to the buffer space credit value and work credit value of each sub-core computing processor.

[0127] Step 104a: Determine the maximum load credit value among multiple load credit values from each sub-core calculator and its own load credit value; allocate a work block to the main core computing processor or sub-core computing processor corresponding to the maximum load credit value.

[0128] The dynamic allocation strategy of the embodiments of the present invention is: for a multi-core GPU system with m core computing processors, each time the next work block to be allocated is allocated to the core computing processor with the largest load credit value.

[0129] Max(CR 1 , CR 2 ,..., CR m )

[0130] where CR n is the load credit value of core computing processor n. During the process of allocating a work block once, when the maximum load credit value comes from multiple core computing processors, the next work block can be allocated to the core computing processor with the largest load credit value in a round-robin order.

[0131] After the main core computing processor allocates a work block to the main core computing processor or sub-core computing processor each time, the workgroup allocation unit of the corresponding core computing processor needs to determine whether it can handle it. In the main core computing processor and sub-core computing processors, it is determined in parallel whether the core computing processor itself needs to process the current workgroup. Each work block will only be allocated to one core computing processor. For a core computing processor, if the current workgroup offset is within the range of the workgroups of the work block in the work block buffer unit, then the current workgroup is processed in this core computing processor. Otherwise, this core computing processor will skip the current workgroup after parsing the compute control flow data. The embodiments of the present invention provide a specific example for implementing this judgment logic as follows:

[0132] If(Workgroup_Offset_Current >= Work_Block_Offset &&

[0133] Workgroup_Offset_Current < Work_Block_Offset + Work_Block_Size)

[0134] {

[0135] Process the workgroup

[0136] }

[0137] Else

[0138] {

[0139] Skip the workgroup

[0140] }

[0141] To specifically illustrate the process of allocating workgroups, as Figure 7 and Figure 8 shown, step 10 further includes:

[0142] Step 101b: The main core computing processor determines the number of workgroups included in the work block according to the workload of the workgroup; wherein, when the workload is greater than the first threshold, the number of workgroups included in the work block is less than when the workload is less than the second threshold, and the first threshold is greater than or equal to the second threshold.

[0143] The main core computing processor of the embodiment of the present invention determines the number of workgroups included in the work block according to the size of the workload, and the principle for determining the number of workgroups is: when the workload is larger, the number of workgroups included in the work block is less than when the workload is smaller.

[0144] Step 102b: The main core computing processor determines the processing order of the work block according to the correlation of the workloads of the work block; when there is no correlation between the corresponding workloads, the work block with a workload greater than the third threshold is preferentially processed.

[0145] The main core computing processor processes in the GPU hardware in the order in which the workloads are submitted. In the embodiment of the present invention, the software driver can mark the dependency relationship between the workloads before submitting the workloads to the GPU hardware. However, if there is no correlation between two workloads (i.e., there is no dependency relationship between them), then their processing order has no impact on the final result.

[0146] To avoid processing the workgroup with a heavier workload at the end, resulting in a delay in the total processing time, when there is no correlation between the workloads of the work block, the work block with a heavier load can be preferentially processed. That is, in the dynamic allocation mode, the main core computing processor will preferentially submit the workgroups from the shader with more instructions and no dependency relationship, and then submit the workgroups from the shader with fewer instructions and no dependency relationship.

[0147] Step 103b: When the number of instructions included in the shader is less than the fourth threshold and the workloads in each workgroup are equivalent, the fixed allocation mode is used to allocate workgroups to the core computing processor.

[0148] Among them, the fourth threshold is selected by those skilled in the art according to the specific usage scenario and is not limited herein. In the embodiments of the present invention, according to the distribution state of the workload, a dynamic allocation mode or a fixed allocation mode is used to allocate workgroups to the core computing processors; the distribution state of the workload is the number of instructions in the shader and the workload in each workgroup.

[0149] Since in the fixed allocation mode, the workgroups are allocated to each main core computing processor or secondary core computing processor from the main core computing processor in a strict round-robin order, and a fixed number of workgroups are allocated to each main core computing processor or secondary core computing processor. The fixed allocation mode is the simplest way in hardware implementation, and there is no need to connect between the core computing processors; although this allocation mode may not make the GPU performance reach the optimal, it still has a good effect in some cases.

[0150] The following describes the workload applicable to the fixed allocation mode: In the case of using a simple shader and multiple workgroups to execute a computing task, the requirements for computing power and storage resources are not high, so the processing time of the workgroups in the core computing processor will not be very long. At the same time, due to the dynamic limitations of computing power and resources, the possibility of blocking processing threads and thus causing processing time delay is also small. Therefore, in this case, each workgroup will often complete processing within a similar time, and the performance will not be reduced due to the delay of some workgroups.

[0151] To simplify the processing flow, the fixed allocation mode is adopted when the workload in the workgroup is light and evenly distributed. If the workgroup contains a complex shader, due to the dynamic limitations of GPU computing power and resources, the execution time of each workgroup on the GPU may be different. However, when the workload in the workgroup is light, the execution time of each workgroup on each core computing processor is often very similar. When there are simple shaders with large X, Y, and Z dimensions in the core computing processor, the workload in the workgroup is light and evenly distributed; in this case, each workgroup will often complete the calculation within a similar time, so the impact of using different workgroup allocation modes on GPU performance is very small.

[0152] It should be noted that the size of the work block or the number of workgroups in the work block can be used as a configuration parameter set by the software driver. In the case of using a simple shader and multiple workgroups to execute a computing task, the size of the work block can be set to a larger value. On the other hand, in the case of using a complex shader to execute a computing task, since some workgroups may complete relatively late, in order to shorten the delay in this part of the overall processing time, the size of the work block can be set to a smaller value.

[0153] Step 104b: When the number of instructions included in the shader is greater than the fifth threshold and the workloads in each workgroup are different, use the dynamic allocation mode to allocate workgroups to the core computing processor; wherein, the fifth threshold is greater than or equal to the fourth threshold.

[0154] The fifth threshold is selected by those skilled in the art according to the specific usage scenario and is not limited herein.

[0155] In the case of using a complex shader to execute a computing task, the requirements for computing power and storage resources are relatively high. Therefore, the processing time of the workgroups in the core computing processor will be relatively long. At the same time, due to the dynamic limitation of computing power and resources, the possibility of processing thread blocking and thus processing time delay is also greater. In this case, each workgroup may not be able to complete processing within a similar time, and the delay of some workgroups will result in performance loss.

[0156] When there are more instructions in the shader and the workloads in each workgroup are different, the dynamic allocation mode can be used to allocate workgroups to the core computing processor to balance the workload and processing time on each core computing processor. In the dynamic allocation mode, the fixed allocation mode is used to allocate the initialization work blocks to each core computing processor, and each core computing processor starts processing the work blocks immediately at the beginning of processing to avoid the delay required for each core computing processor to wait for the main core computing processor in the dynamic allocation mode to allocate work blocks according to the load credit value at the beginning of processing.

[0157] In the application scenario of the GPU, when the host pauses the current computing application running on the GPU and starts another computing application, context switching is required. During the context switching process, the context information of the current computing application running on the GPU (such as GPU status and workgroup processing status, etc.) will be stored by the computing processing module in the GPU into the external memory of the context state memory, then the current computing application is terminated, and a new computing application is started in the GPU. When the paused computing application resumes running, the computing processing module in the GPU reads the context information stored in the context state memory from the external memory as the context and loads it, and then resumes normal computing processing in the GPU.

[0158] The embodiment of the present invention provides a context switching scheme:

[0159] When storing the context, the state data and the information of the current workgroup processing state are stored by the computing processing module in the external memory allocated by the software driver. The data structure shown in the following table is an example of the context state structure that can be used for computing context switching in the computing processing module. The computing processing module will fill the context information for the current computing application in the stored context. The software driver should allocate memory space for the context state structure of the context of each computing application.

[0160] In a multi-core GPU system, each core computing processor has independent context information. The context state structure shown in the following table is stored by core computing processor, and its fixed memory stride starts from the base address of the context state memory of the computing processing module. The current workgroup processed by the core computing processor during context switching will be set in the current workgroup_X, current workgroup_Y, and current workgroup_Z fields in the corresponding context state memory. The total workgroup settings in the X, Y, and Z dimensions are in the workgroup_X, workgroup_Y, and workgroup_Z fields. The software driver should be responsible for allocating sufficient memory for all context state memories.

[0161] In an optional embodiment, the embodiment of the present invention provides a specific instance of implementing the context state structure; the "current workgroup_X", "current workgroup_Y", and "current workgroup_Z" fields in the context state structure shown in the following table should be set by the computing processing module of each core computing processor of the GPU, and are used to set the workgroup offset of the current workgroup that has been processed during context switching. The other fields of the context state structure should be set to the same value in all computing processing modules.

[0162] Field Number of bits occupied Meaning represented Workgroup_Z 32 Number of workgroups in the Z dimension Workgroup_Y 32 Number of workgroups in the Y dimension Workgroup_X 32 Number of workgroups in the X dimension Current workgroup_Z 32 Z coordinate in the current workgroup Current workgroup_Y 32 Y coordinate in the current workgroup Current workgroup_X 32 X coordinate in the current workgroup Needs_to_be_restored 1 Whether there is remaining computational work

[0163] Since it may not be possible to process the workgroup in the work block even if it is allocated to the work block buffer unit; therefore, the work block offset stored in the work block buffer unit should be stored in the context state memory.

[0164] As shown in the following table, in an optional embodiment, additional information on the workgroup processing state in each core computing processor of the embodiment of the present invention can be stored in the context state structure of the computing processing module.

[0165] The field "Next Working Block" is added to the context state structure shown in the following table, which is used to indicate that there is an additional working block entry field set at the end of the context state memory; this field can only be set to 1 in the context state memory by the computing processing module of the main core computing processor. This additional working block entry field is used to represent the working block offset of the first working block, which has not been assigned to any computing processing module during context storage and from which the normal computing process will resume during context recovery.

[0166] The "Number of Working Blocks" field is added to the context state structure shown in the following table, which is used to indicate the number of working blocks stored in the context state memory. The total number of working block entry fields and the total number of working block mask words are equal to the value of the number of working blocks; the value of the number of working blocks should be less than or equal to the maximum working block size, which is a constant predefined by the software driver.

[0167] When the number of working blocks field is set to 0, it means that there are no working blocks in the working block buffer unit of the context state memory. In this case, all working blocks have been sent to the downstream shader execution unit for computing processing by the computing processing module during context storage, and no working blocks are stored in the context state memory.

[0168]

[0169] As shown in the following table, working block entry fields can be added to the context state structure as 3 32-bit words, which contain the working block offsets in the working block buffer unit during context storage by the core computing processor. The offsets of the first workgroup in the working block are stored in the working block entry field in the form of working block offset_X, working block offset_Y, and working block offset_Z.

[0170]

[0171]

[0172] As shown in the following table, the working block mask field shown in the following table can be added to the context state structure. This field has 32 bits and is used to indicate whether each workgroup in the working block needs to be processed during context recovery. The working block mask is located after the working block entry field of the working block. If there are more than 32 workgroups in a working block, multiple working block mask fields may need to be set.

[0173] Due to the addition of extra data in the context state structure, the software driver should reserve extra memory space for the context state memory; as Figure 9The following shows a specific example of a context state memory provided by an embodiment of the present invention.

[0174] For each working block stored in the context state memory, the working block entry field is three 32-bit fields. The size of the working block mask field is: (working block size + 31 / 32) fields. The total number of fields of data in the context state memory of a working block is 32 bits, specifically as follows:

[0175] The total size of a working block = 3 + ((working block size + 31 / 32))

[0176] The working block in the core computing processor also requires data in other additional context state memories, including the working block entry field, the working block mask field, and the working block entry field of the last working block in the first working block not assigned to any core computing processor. These fields altogether occupy 32 bits, specifically as follows:

[0177] The total size of the working block = 3 + the total size of a working block × the number of working blocks

[0178] The maximum working block size occupied by the data in the other additional context state memory is:

[0179] The maximum working block size = 3 + the total size of a working block * the maximum working block buffer unit size

[0180] Wherein, the maximum working block buffer unit size is a predefined constant, and this predefined constant represents the maximum number of working blocks in the working block buffer unit. In specific implementation, the software driver reserves context state storage spaces with additional field numbers (i.e., the maximum working block size) for each core computing processor.

[0181] To illustrate the process of loading during context restoration, as Figure 10 shown, it further includes:

[0182] In step 301, when the current computing task pauses and another computing task needs to be started, all the working block entry fields are assigned to the corresponding core computing processors for processing; wherein, the working block entry field represents the current working block offset of the corresponding core computing processor when the current computing task pauses.

[0183] In step 302, determine the working block corresponding to the last working block entry field among all the working block entry fields, and starting from the determined working block, assign the subsequent working blocks to the main core computing processor and / or the secondary core computing processor for processing.

[0184] When the context is loaded, in each context state memory, the work block offsets assigned to the work block buffer unit before context switching are read back by the computing processing modules of the main core computing processor and the secondary core computing processor. According to the work block offsets and work block mask fields stored in the context state memory during context storage, the computing processing modules of the main core computing processor and the secondary core computing processor will process the work groups in the stored work blocks.

[0185] After all the work groups in the work block entry field are assigned to the computing processing module for computing processing, on the premise that the next work block field is set to 1, the work blocks starting from the last work block entry field will be assigned to each core computing processor by the computing processing module of the main core computing processor as usual.

[0186] The embodiment of the present invention also provides precautions when the solution of the embodiment of the present invention is implemented in hardware, which are as follows:

[0187] In an alternative embodiment, additional control parameters are added in the GPU hardware, as shown in the following table:

[0188]

[0189]

[0190] As shown in the above table, the workload type field is a control parameter for the workload allocation type of the embodiment of the present invention; when the value of this field is 0, it represents the existing computing workload allocation mode, that is, the GPU kernel processes work groups in a fixed allocation mode according to the "execution count" field; when the value of this field is 1, the dynamic allocation mode of computing workload is enabled.

[0191] The "number of executions" field is used to control the number of work groups assigned to the core computing processor each time when allocating workloads. In both the fixed allocation mode and the dynamic allocation mode, the value of this field can be the number of work groups assigned to the core computing processor each time.

[0192] The value of the "maximum work block buffer unit size" field is the maximum number of work blocks that the work block buffer unit can store; the number of work groups in each work block is the value of the "statistical count" field. The optimal size of the value of the "maximum work block buffer unit size" field can be determined through experiments to reduce the dynamic latency of work block allocation, thereby improving the performance of the multi-core GPU system.

[0193] In another alternative embodiment, as Figure 11 shown, is a schematic diagram of the architecture of the core computing processor of the embodiment of the present invention. The core computing processor of this embodiment includes one or more processors 21 and a memory 22. Among them, Figure 11Take a processor 21 as an example.

[0194] The processor 21 and the memory 22 can be connected by a bus or other means. Figure 11 Take the connection by bus as an example.

[0195] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the method for dynamic workload allocation in this embodiment. The processor 21 executes the method for dynamic workload allocation by running the non-volatile software programs and instructions stored in the memory 22.

[0196] The memory 22 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 may optionally include a memory remotely located relative to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0197] The program instructions / modules are stored in the memory 22 and, when executed by the one or more processors 21, execute the method for dynamic workload allocation in the above embodiments. For example, execute the Figure 5 , Figure 6 , Figure 7 and Figure 10 each step shown.

[0198] The embodiment of the present invention also provides a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, for example Figure 11 a processor 21, it enables the above one or more processors to execute the method for dynamic workload allocation in the specific embodiments of the present invention. For example, execute the Figure 5 , Figure 6 , Figure 7 and Figure 10 each step shown; it can also implement Figure 11 each module and unit described; or execute the method for dynamic workload allocation in the specific embodiments of the present invention. For example, execute the Figure 5 , Figure 6 , Figure 7 and Figure 10 each step shown; it can also implement Figure 11 each module and unit described.

[0199] It should be noted that, for the information interaction, execution process, etc. between the modules and units in the above-mentioned device and system, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.

[0200] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0201] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A core computing processor, characterized in that: including a computing and processing module; The computing processing module is used to determine the load credit value of the core computing processor where the computing processing module is located; and is also used to allocate work blocks to the core computing processor that meets the preset conditions according to the load credit values ​​from other core computers and its own load credit value; wherein a work block includes multiple work groups; It is also used to selectively process the current work group by its own core computing processor or by other core computing processors according to the current work group offset and the current work block offset.

2. The core computing processor according to claim 1, characterized in that: The computing processing module includes a work block allocation unit and a work block buffer unit; The work block allocation unit is used to determine the difference between the maximum number of work blocks and the current number of work blocks of the work block buffer unit to obtain a buffer space credit value of the core computing processor; The work block allocation unit is further used to determine the workload corresponding to the work block of each core computing processor and obtain the corresponding work credit value; The work block allocation unit is further used to obtain the load credit value according to the buffer space credit value and the work credit value corresponding to the work block of the core computing processor; The work block allocation unit is further used to determine a maximum load credit value from multiple load credit values ​​from other core computing processors and its own load credit value, and allocate work blocks to the core computing processor corresponding to the maximum load credit value.

3. The core computing processor according to claim 2, characterized in that: It also includes a shader execution unit; the computing processing module also includes a work group allocation unit; The work block buffer unit is used to store the work block offsets of the work blocks allocated to the core computing processor; The work group allocation unit is used to obtain a current work group offset of a current work group; When the current work group offset is within the range of the work groups included in the current work block stored in the work block buffer unit, it is determined that the current work group is processed by the core computing processor; When the current work group offset is not within the range of the work groups included in the current work block, the current work group is processed by other core computing processors; The workgroup allocation unit is further used to allocate the workgroups to be processed by the core computing processor to the multiple shader execution units; The shader execution unit is used to process the work group allocated by the work group allocation unit.

4. A device for dynamically allocating workloads, characterized in that: It includes a main core computing processor and a plurality of sub-core computing processors; The main core computing processor is connected to each secondary core computing processor; The secondary core computing processor is used to determine its own load credit value and send its own load credit value to the primary core computing processor; The main core computing processor is used to determine its own load credit value; and is also used to allocate work blocks to the secondary core computing processors or the main core computing processor that meet the preset conditions according to the load credit values ​​from other core computing processors and its own load credit value; It is also used to selectively process the current working group in the main core computing processor or in the secondary core computing processor according to the current working group offset and the current working block offset.

5. The device for dynamically allocating workload according to claim 4, characterized in that: The work block allocation unit of the main core computing processor is used to receive the load credit value sent by the work block allocation unit of each sub-core computing processor, and determine its own load credit value, determine the maximum load credit value among the multiple load credit values ​​from each sub-core computing processor and its own load credit value, and allocate a work block to the main core computing processor or the sub-core computing processor corresponding to the maximum load credit value; The work block buffer unit of the main core computing processor and / or the work block buffer units of each secondary core computing processor are used to receive the offset of the allocated work block; The main core computing processor is further used to determine the number of work groups included in the work block according to the workload of the work group; wherein when the workload is greater than a first threshold, the number of work groups included in the work block is less than the number of work groups included in the work block when the workload is less than a second threshold, and the first threshold is greater than or equal to the second threshold; The main core computing processor is also used to determine the processing order of the work blocks according to the correlation of the workloads of the work blocks; when there is no correlation between the corresponding workloads, the work blocks with workloads greater than a third threshold are processed preferentially.

6. A method for dynamic workload allocation, characterized in that: include: Allocate work blocks to the main core computing processor or the sub-core computing processor that meets preset conditions according to the load credit value of the main core computing processor and the load credit values ​​of each sub-core computing processor; According to the current work group offset and the current work block offset, the current work group is selectively processed in the main core computing processor or in the secondary core computing processor.

7. The method for dynamic workload allocation according to claim 6, characterized in that: The allocating work blocks to the main core computing processor or the secondary core computing processor that meets the preset conditions according to the load credit value of the main core computing processor and the load credit values ​​of each secondary core computing processor includes: Determine the difference between the maximum number of working blocks of the main core computing processor and the current number of working blocks to obtain a buffer space credit value of the main core computing processor; determine the difference between the maximum number of working blocks of each secondary core computing processor and the current number of working blocks to obtain a buffer space credit value of the corresponding secondary core computing processor; Determine the workload corresponding to the work block of the main core computing processor to obtain the work credit value of the main core computing processor; determine the workload corresponding to the work block of each sub-core computing processor to obtain the corresponding work credit value of each sub-core computing processor; According to the buffer space credit value and the working credit value of the main core computing processor, a load credit value of the main core computing processor is obtained; according to the buffer space credit value and the working credit value of each secondary core computing processor, a load credit value of the corresponding secondary core computing processor is obtained; Determine a maximum load credit value from a plurality of load credit values ​​of each sub-core calculator and its own load credit value; and allocate a work block to a main core computing processor or a sub-core computing processor corresponding to the maximum load credit value.

8. The method for dynamic workload allocation according to claim 6, characterized in that: The allocating work blocks to the main core computing processor or the secondary core computing processor that meets the preset conditions according to the load credit value of the main core computing processor and the load credit values ​​of each secondary core computing processor includes: The main core computing processor determines the number of working groups included in the working block according to the workload of the working group; wherein when the workload is greater than a first threshold, the number of working groups included in the working block is less than the number of working groups included in the working block when the workload is less than a second threshold, and the first threshold is greater than or equal to the second threshold; The main core computing processor determines the processing order of the work blocks according to the correlation of the workloads of the work blocks; when there is no correlation between the corresponding workloads, the work blocks with a workload greater than a third threshold are processed preferentially; When the instructions included in the shader are less than a fourth threshold and the workloads in the various working groups are similar, the working groups are allocated to the core computing processors using a fixed allocation mode; When the instructions included in the shader are greater than a fifth threshold and the workloads in the various working groups are different, a dynamic allocation mode is used to allocate working groups to the core computing processors; wherein the fifth threshold is greater than or equal to the fourth threshold.

9. The method for dynamic workload allocation according to claim 8, characterized in that: When the instructions included in the shader are greater than the fifth threshold and the workloads in the various work groups are different, using the dynamic allocation mode to allocate work groups to the core computing processors includes: In the dynamic allocation mode, the fixed allocation mode is used to allocate the initialization work block to each core computing processor, and each core computing processor starts processing the work block immediately when the processing starts.

10. The method for dynamic workload allocation according to claim 6, characterized in that: Also includes: When the current computing task is suspended and another computing task needs to be started, all work block entry fields are assigned to the corresponding core computing processor for processing; wherein the work block entry field indicates the current work block offset of the corresponding core computing processor when the current computing task is suspended; Determine the work block corresponding to the last work block entry field among all the work block entry fields, and, starting from the determined work block, assign subsequent work blocks to the main core computing processor and / or the secondary core computing processor for processing.