Calculation processing unit, workload processing device and method
By introducing a computing processing unit into the GPU to process the computation trigger signals from multiple applications, decompose them into multiple working groups and allocate them to the shader processing cluster, the problem of only serial processing in the existing GPU architecture is solved, and parallel processing of multiple applications in the GPU is realized, and GPU performance is improved.
Patent Information
- Application Number
- CN202510171378.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-13
AI Technical Summary
The existing GPU architecture can only process input data from one computing control flow, resulting in all computing tasks performed in the GPU coming from the computing context of the same application, and multiple applications can only be executed serially, resulting in low SEU utilization and GPU performance affected.
The computing processing unit processes the computation trigger signals from multiple applications, determines multiple working groups for each independent computing context, and assigns the corresponding working groups to the shader processing cluster processing, so as to realize the parallel processing of multiple applications in the GPU.
It solves the problem of low SEU utilization and impact of GPU performance due to only serial processing of applications, and realizes parallel processing of multiple applications in the GPU, improving the overall performance of the GPU.
Smart Images

Figure CN120144282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high-performance computing technology, and particularly to a computing processing unit, a workload processing device, and a method. Background Art
[0002] With the development of high-performance parallel computing, Graphics Processing Unit (GPU) is widely used to accelerate general computing, such as in the current popular Large Language Model (LLM).
[0003] The GPU completes general computing operations through a Compute Processing Module (CPM), such as arithmetic logic unit (ALU) operations in floating-point and integer formats. The computing workload is driven by instructions in a Compute Shader through application programming interface (API) platforms such as Open Computing Language (OpenCL) and Vulkan. The Compute Shader is executed in the Shader Execution Unit (SEU) of the GPU hardware.
[0004] In the existing GPU architecture, the CPM can only process input data from one computing control flow. Therefore, all computing tasks executed in the GPU come from the computing context of the same application, for example, from the same OpenCL program. Multiple applications can only be executed serially one by one in the GPU. Since the computing workload of processing a small application can be completed with only a small number of SEUs, the computing performance of the GPU will not be significantly reduced in the scenario of processing small applications. Especially if small applications are serially processed on a GPU with more SEUs, the SEU utilization rate will be very low, resulting in a significant reduction in the computing performance of the GPU.
[0005] In view of this, overcoming the defects of the existing technology is an urgent problem to be solved in this technical field. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a computing processing unit, a workload processing device and a method. The purpose is to provide a GPU architecture. By the computing processing unit processing the computing trigger signals from multiple application programs, respectively determining multiple workgroups corresponding to each independent computing context therein, and allocating the corresponding workgroups to the shader processing clusters for processing, realizing parallel processing of multiple application programs in the GPU, and solving the problems of too low utilization rate of SEUs and affected GPU performance caused by only serially processing application programs on a large number of SEUs in the GPU.
[0007] The present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a computing processing unit, including a control module and multiple independent processing modules;
[0009] The output end of the control module is respectively connected to the input ends of the multiple independent processing modules;
[0010] The control module is used to obtain computing trigger signals, and respectively send the independent computing contexts corresponding to the multiple computing trigger signals to the corresponding independent processing modules;
[0011] The independent processing module is used to process the received independent computing context.
[0012] Further, it further includes an arbitration module; the independent processing module includes a direct memory access module and a computing workgroup generator;
[0013] The output end of the computing processing unit is respectively connected to the input ends of the multiple direct memory access modules, and the output end of the direct memory access module is connected to the input end of the computing workgroup generator in the independent processing module where it is located;
[0014] The direct memory access module is used to receive the independent computing context allocated by the control module, and obtain the independent computing control flow corresponding to the independent computing context from the memory through the arbitration module;
[0015] The computing workgroup generator is used to obtain the computing kernels in the independent computing control flow, and decompose the computing kernels into multiple workgroups.
[0016] In a second aspect, the present invention further provides a workload processing device based on multiple computing contexts, including a computing processing unit as described in the first aspect and at least one shader processing cluster; the output end of the computing processing unit is connected to the input end of the at least one shader processing cluster;
[0017] The computing processing unit is used to obtain a computing trigger signal, decompose the computing kernels of the independent computing contexts corresponding to the multiple computing trigger signals into multiple workgroups; and is also used to send the multiple workgroups to the shader processing cluster;
[0018] The shader processing cluster is used to process the multiple workgroups.
[0019] Further, it further includes a configuration register and a secondary cache unit; the shader processing cluster includes a task scheduler, a shader execution unit, and a primary cache unit;
[0020] The output end of the configuration register is connected to the input end of the control module; the output ends of the independent processing modules are respectively connected to the input ends of multiple task schedulers; the output ends of the task schedulers are respectively connected to the input ends of multiple shader execution units; the primary cache unit is respectively connected to the multiple shader execution units; the task scheduler is also connected to the primary cache unit; the secondary cache unit is respectively connected to multiple primary cache units;
[0021] The configuration register is used to control the computing trigger signal based on the status control information and send the computing trigger signal to the control module;
[0022] The task scheduler is used to generate multiple computing subtasks according to the received workgroups, allocate the multiple computing subtasks to at least one shader execution unit; and is also used to obtain computing associated data from the primary cache unit to facilitate generating corresponding computing subtasks;
[0023] The shader execution unit is used to process the corresponding computing subtasks; and is also used to obtain the required computing associated data from the primary cache unit when the computing subtask requires computing associated data;
[0024] The primary cache unit is also used to obtain the corresponding computing associated data from the secondary cache unit when the required computing associated data does not exist in the primary cache unit.
[0025] In a third aspect, the present invention further provides a workload processing method based on multiple computing contexts, including:
[0026] Obtain independent computing contexts corresponding to multiple computing trigger signals, and allocate corresponding independent processing modules to the independent computing contexts;
[0027] Process the independent computing control flow corresponding to the independent computing context in the independent processing module to obtain multiple workgroups;
[0028] Process the multiple workgroups to complete multiple computing subtasks corresponding to the computing trigger signals.
[0029] Further, obtaining independent computing contexts corresponding to multiple computing trigger signals and allocating corresponding independent processing modules to the independent computing contexts includes:
[0030] When the computing trigger signal corresponds to multiple computing contexts, check the dependency relationships between the multiple computing contexts to determine whether there are multiple independent computing contexts;
[0031] When there are multiple independent computing contexts, determine whether there are available independent processing modules;
[0032] If there are available independent processing modules, allocate the corresponding independent processing modules to the independent computing contexts;
[0033] If there are no available independent processing modules, wait until an available independent processing module appears, and allocate the appeared available independent processing module to the independent computing context.
[0034] Further, processing the independent computing control flow corresponding to the independent computing context in the independent processing module to obtain multiple workgroups includes:
[0035] Obtain the maximum number of shader units corresponding to the independent computing context from the configuration register;
[0036] According to the current available quantity and the maximum number of shader units, allocate corresponding shader processing clusters to each independent processing module; wherein, the shader processing cluster includes multiple shader execution units; the number of shader execution units allocated to the independent processing module is less than or equal to its corresponding maximum number of shader units;
[0037] According to the memory base address in its own configuration register, obtain the corresponding independent computing control flow for the independent computing context from the memory;
[0038] Obtain the computing kernels in the independent computing control flow, and decompose the computing kernels into multiple workgroups to facilitate sending the corresponding workgroups to the allocated shader processing clusters.
[0039] Further, the multiple workgroups corresponding to the computing trigger signal carry the current trigger ID of the computing trigger signal;
[0040] The allocating corresponding shader processing clusters to each independent processing module according to the current available quantity and the maximum number of shader units includes:
[0041] Determine the minimum value between the maximum number of shader units and the current available number, and use this minimum value as the total number of shaders for the shader execution units in the shader processing clusters assigned to each individual processing module;
[0042] According to the current trigger ID, determine the parallel trigger IDs of multiple workgroups that are processed in parallel with the multiple workgroups, and obtain the parallel trigger signals corresponding to the parallel trigger IDs;
[0043] Compare the computational workload corresponding to the computational trigger signal with the computational workload corresponding to the parallel trigger signal, and determine the trigger information with a greater computational workload and / or the trigger information with a higher computational priority among them;
[0044] Estimate the computational workload of the independent computational context, and dynamically allocate shader execution units to the independent processing modules of the independent computational context based on the workload level of the estimated computational workload;
[0045] Based on the total number of shaders, according to the current trigger ID or the parallel trigger ID, preferentially allocate more shader execution units to the determined trigger information.
[0046] Further, the estimating the computational workload of the independent computational context and dynamically allocating shader execution units to the independent processing modules of the independent computational context based on the workload level of the estimated computational workload includes:
[0047] Based on the number of shader instructions included in the independent computational context, and the product of the size of the main computational kernel of the independent computational context in the X direction, the size of the main computational kernel in the Y direction, and the size of the main computational kernel in the Z direction, obtain the workload estimation value of the computational workload; where the main computational kernel is: one of the computational kernels included in the independent computational context; or, based on the number of shader instructions, and the product of the size of the computational kernel included in the independent computational context in the X direction, the size of the computational kernel in the Y direction, and the size of the computational kernel in the Z direction, obtain the kernel workload value of the computational kernel; determine the sum of the kernel workload values of all the computational kernels included in the independent computational context as the workload estimation value of the computational workload;
[0048] Map the workload estimation value to the workload level range to obtain the workload level of the independent computational context;
[0049] Determine the accumulation of the workload levels of all independent computing contexts as a first intermediate value; obtain a second intermediate value based on the ratio of the workload level of the independent computing context to be allocated to the first intermediate value; obtain the number of shader execution units allocated to the independent computing context to be allocated based on the product of the second intermediate value and the current available quantity, so as to allocate more shader execution units to the independent computing context with a high workload level.
[0050] Further, the processing of the multiple workgroups includes:
[0051] Define the computing context ID of the independent computing context in the configuration register; wherein, one configuration register corresponds to one independent computing context;
[0052] Determine the independent processing module for processing the independent computing context corresponding to the computing trigger signal, and determine the computing context ID of the independent computing context;
[0053] Establish a mapping relationship between the computing context ID and the current trigger ID of the computing trigger signal, and store the mapping relationship in the configuration register corresponding to the independent computing context, so as to obtain the computing context ID according to the mapping relationship, and obtain the corresponding computing association data based on the computing context ID;
[0054] Generate a plurality of computing subtasks according to the received workgroups;
[0055] Process the multiple computing subtasks, obtain computing association data according to the mapping relationship, and complete the multiple computing subtasks based on the computing association data.
[0056] In a fourth aspect, the present invention further provides a non-volatile computer storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors to complete the workload processing method based on multiple computing contexts described in the first aspect.
[0057] In a fifth aspect, a chip is provided, including: a processor and an interface, for calling and running a computer program stored in a memory from the memory, and executing the workload processing method based on multiple computing contexts as described in the first aspect.
[0058] In a sixth aspect, a computer program product including instructions is provided. When the instructions run on a computer or a processor, the computer or the processor is caused to execute the workload processing method based on multiple computing contexts as described in the first aspect to the fourth aspect and any one of them.
[0059] In a seventh aspect, a workload processing method and system based on multiple computing contexts are provided, including a workload processing device based on multiple computing contexts as in the second aspect, and using the workload processing method based on multiple computing contexts as described in the first aspect to complete the interaction of the workload processing device based on multiple computing contexts in the second aspect.
[0060] Different from the prior art, the present invention has at least the following beneficial effects:
[0061] The present invention obtains independent computing contexts corresponding to multiple computing trigger signals and assigns corresponding independent processing modules to them; processes the independent computing control flows corresponding to the independent computing contexts in the independent processing modules to obtain multiple working groups; and processes the multiple working groups to complete multiple computing subtasks corresponding to the computing trigger signals. Among them, the computing processing unit processes the computing trigger signals from multiple application programs, respectively determines multiple working groups corresponding to each independent computing context therein, and assigns the corresponding working groups to the shader processing cluster for processing, realizing the parallel processing of multiple application programs in the GPU, and solving the problems of too low utilization rate of SEUs and affected GPU performance caused by the serial processing of application programs only on a large number of SEUs in the GPU. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0063] Figure 1 is a schematic diagram of a GPU workload pipeline in the prior art provided by an embodiment of the present invention;
[0064] Figure 2 is a schematic diagram of the architecture of a computing processing unit with 4 independent processing modules provided by an embodiment of the present invention;
[0065] Figure 3 is a schematic diagram of the architecture of a workload processing device based on multiple computing contexts provided by an embodiment of the present invention;
[0066] Figure 4 is a schematic diagram of a specific example of a workload processing device based on multiple computing contexts provided by an embodiment of the present invention;
[0067] Figure 5 is a schematic diagram of a specific example of a workload processing device based on multiple computing contexts with 4 independent processing modules provided by an embodiment of the present invention;
[0068] Figure 6 It is a schematic diagram of a workload processing method based on multiple computing contexts provided by an embodiment of the present invention;
[0069] Figure 7 It is a schematic flowchart of step 10 provided by an embodiment of the present invention;
[0070] Figure 8 It is a schematic diagram of the allocation process of a control module provided by an embodiment of the present invention;
[0071] Figure 9 It is a schematic flowchart of step 20 provided by an embodiment of the present invention;
[0072] Figure 10 It is a schematic flowchart of step 202 provided by an embodiment of the present invention;
[0073] Figure 11 It is a schematic flowchart of step 2024 provided by an embodiment of the present invention;
[0074] Figure 12 It is a schematic flowchart of step 30 provided by an embodiment of the present invention;
[0075] Figure 13 It is a schematic flowchart of step 301 provided by an embodiment of the present invention;
[0076] Figure 14 It is a schematic diagram of the mapping relationship of an independent computing context provided by an embodiment of the present invention;
[0077] Figure 15 It is a schematic diagram of the architecture of another workload processing device based on multiple computing contexts provided by an embodiment of the present invention. Detailed implementation manners
[0078] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0079] Unless otherwise required by the calculation context, throughout the specification and claims, the term "comprising" is interpreted in an open - inclusive sense, that is, "including, but not limited to". In the description of the specification, terms such as "one embodiment", "some embodiments", "exemplary embodiments", "examples", "specific examples", or "some examples" are intended to indicate that specific features, structures, materials, or characteristics related to the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representations of the above - mentioned terms do not necessarily refer to the same embodiment or example. In addition, the specific features, structures, materials, or characteristics may be included in any one or more embodiments or examples in any appropriate manner. That is, although they may be carried in the embodiments or examples of the above - mentioned terms due to reasons such as the order of appearance and position, etc., there is no limitation that they can be carried by one embodiment or example in a combined manner.
[0080] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present disclosure.
[0081] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise stated, the meaning of "a plurality" is two or more. In addition, for example, in the description, for the same type of nouns, the method of adding "A" and "B" at the end is used to describe them as two independent individuals. In this case, the features defined with "A" and "B" are only used for the purpose of distinguishing similar individuals in the description and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features.
[0082] In the description of some embodiments, the expressions "coupled", "coupling" and "connected" and their derivatives may be used. For example, in the description of some embodiments, the term "connected" may be used to indicate that two or more components have direct physical or electrical contact with each other. Another example is that in the description of some embodiments, the term "coupled" may be used to indicate that two or more components have direct physical or electrical contact. However, the term "connected" or "coupled" may also mean that two or more components do not have direct contact with each other, but still cooperate or interact with each other, such as "optical path coupling", "wireless connection", etc. The embodiments disclosed herein are not necessarily limited to the content of the present invention.
[0083] In the description of the present invention, the expression "A and / or B" (where A and B are used to formally represent specific feature contents) will be involved, and the corresponding expression includes the following three combinations: only A, only B, and the combination of A and B.
[0084] As used in the present invention, "about", "substantially" or "approximately" includes the stated value and the average value within an acceptable deviation range of the specific value, where the acceptable deviation range is determined by those of ordinary skill in the art considering the measurement being discussed and the errors associated with the measurement of a particular quantity (i.e., the limitations of the measurement system).
[0085] As Figure 1 shown, the computational workload of the application is sent to the GPU hardware through the Compute Control Stream (i.e., the CPM in the GPU as Figure 1 shown). The Compute Control Stream includes status, compute shader code, and input data so that the GPU can generate the final result by performing computational operations in the compute shader. The Compute Control Stream may contain multiple Compute kicks, each Compute kick may contain multiple workgroups, each workgroup may contain multiple work items or instances, and the compute shader is executed in the SEU of the GPU hardware.
[0086] In a typical GPU architecture, the CPM schedules the workload executed by the compute shader into workgroups, and then assigns the workgroups to the Task Scheduler (abbreviated as TS) in each SPC. The TS schedules the computational workload of the workgroup into computational tasks, and then sends the computational tasks to multiple downstream SEU pipelines, where the instances in the computational tasks are processed by the compute shader. For workload allocation, the CPM assigns an SPC to the workgroup and an SEU to the computational task.
[0087] However, in the current GPU architecture, the CPM can only process input data from one computational control flow. Therefore, all computational tasks executed on the GPU come from the computational context of the same application, and multiple applications can only be executed serially one by one in the GPU system. If multiple applications can be processed in parallel on the GPU, it will be beneficial to improve GPU performance.
[0088] To solve the above problems, an embodiment of the present invention provides a computational processing unit, including a control module and multiple independent processing modules.
[0089] The output end of the control module is respectively connected to the input ends of the multiple independent processing modules.
[0090] The control module is used to obtain computational trigger signals, and respectively send the independent computational contexts corresponding to the multiple computational trigger signals to the corresponding independent processing modules.
[0091] The independent processing module is used to process the received independent computational context.
[0092] Among them, the computational processing unit is an improved CPM provided by an embodiment of the present invention; the computational trigger signal refers to the information contained in the computational kick; each independent computational context has a corresponding computational trigger signal; the independent computational contexts corresponding to multiple computational trigger signals refer to: among the computational contexts parsed from the computational trigger signal, the computational contexts that do not have a dependency relationship with each other; the computational context generally refers to the environment and state information required when executing the computational trigger signal on the GPU, which ensures that the computational trigger signal is executed in the correct environment and state.
[0093] In an embodiment of the present invention, the computational processing unit processes computational kernels from the computational control flows of multiple applications, and then processes them in parallel in the downstream GPU pipeline. The computational processing unit can accept multiple independent computational trigger signals with independent configuration register settings from multiple applications. Some functions in the computational processing unit are merged into an independent processing module, and this independent processing module works in an independent computational context. A computational processing unit may include multiple independent processing modules. As Figure 2 Shown is a specific example of a computational processing unit provided by an embodiment of the present invention, which includes four independent processing modules.
[0094] When the control module receives a computational trigger signal in a new independent computational context, the control module will allocate an available independent processing module for the independent computational context to which the computational trigger signal belongs; multiple independent processing modules are parallel, and each independent processing module processes the independent computational context received by itself.
[0095] As Figure 2As shown in the figure, the computing and processing unit of the embodiment of the present invention further includes an arbitration module; the independent processing module includes a direct memory access module and a compute workgroup generator.
[0096] The output end of the computing and processing unit is respectively connected to the input ends of a plurality of direct memory access modules, and the output end of the direct memory access module is connected to the input end of the compute workgroup generator in the independent processing module where it is located.
[0097] The direct memory access module is used to receive the independent computing context allocated by the control module, and obtain the independent computing control flow corresponding to the independent computing context from the memory through the arbitration module.
[0098] The compute workgroup generator is used to obtain the compute kernels in the independent computing control flow, and decompose the compute kernels into multiple workgroups.
[0099] After receiving the computing trigger signal, the present invention starts the computing process in the computing and processing unit. The control module allocates an independent processing module for each independent computing context, and the workload of the independent computing context is processed in the independent processing module, that is, the computing workload of the independent processing module is limited to one computing context corresponding to one independent computing control flow.
[0100] For each independent processing module, the direct memory access (DMA) module obtains the corresponding computing control flow from the memory through interaction with the control module and the arbitration module; for example, the input end of the direct memory access module of the independent processing module 0 receives the independent computing context 0 allocated by the control module, and obtains the independent computing control flow 0 corresponding to the independent computing context 0 from the memory through the arbitration module. The compute workgroup generator (CWG) receives the computing control flow from the direct memory access module, obtains the compute kernels from it, and decomposes them into multiple workgroups.
[0101] The arbitration module (Arbiter) receives the memory operations of the direct memory access modules and compute workgroup generators from each independent processing module, and arbitrates them into a memory request or response interface. It should be noted that the memory access addresses of different independent computing contexts should not overlap.
[0102] Based on the above computing and processing unit, as Figure 3 shown, the present invention further provides a workload processing device based on multiple computing contexts, including a computing and processing unit as described in the embodiment of the present invention and at least one shader processing cluster; the output end of the computing and processing unit is connected to the input end of the at least one shader processing cluster.
[0103] The computing processing unit is used to obtain a computing trigger signal, decompose the computing kernels of the independent computing contexts corresponding to the multiple computing trigger signals into multiple workgroups; and is also used to send the multiple workgroups to the shader processing cluster.
[0104] The shader processing cluster is used to process the multiple workgroups.
[0105] To improve computing performance and scalability, multiple shader execution units (i.e., SEUs) can be pipelined and combined into a shader processing cluster (abbreviated as SPC). Configuring the GPU system according to different performance goals, a GPU architecture can include multiple shader processing clusters. Among them, one SPC group can be configured with one or more shader processing clusters to process workgroups from independent computing control flows. Therefore, the shader processing clusters and shader execution units in the SPC group can process multiple computing contexts in parallel. Each SPC group corresponds to an independent processing module, and the computing workloads from different independent computing contexts will be sent from the corresponding independent processing modules to the corresponding SPC groups for parallel processing.
[0106] As Figure 4 shown, the workload processing device based on multiple computing contexts in the embodiment of the present invention further includes a configuration register and a secondary cache unit; the shader processing cluster includes a task scheduler, shader execution units, and a primary cache unit.
[0107] The output end of the configuration register is connected to the input end of the control module; the output ends of the independent processing modules are respectively connected to the input ends of multiple task schedulers; the output ends of the task schedulers are respectively connected to the input ends of multiple shader execution units; the primary cache unit is respectively connected to the multiple shader execution units; the task scheduler is also connected to the primary cache unit; the secondary cache unit is respectively connected to multiple primary cache units.
[0108] The configuration register is used to control the computing trigger signal based on the status control information and send the computing trigger signal to the control module.
[0109] The control module is used to obtain the corresponding independent computing context according to the computing trigger signal and send the obtained multiple independent computing contexts to the corresponding independent processing modules. The independent processing module is used to determine the multiple workgroups corresponding to the independent computing context and allocate the multiple workgroups to the corresponding task schedulers. In addition, the control module is also used to determine and allocate the number of shader execution units for the independent computing context; this process will be described below.
[0110] The task scheduler is used to generate a plurality of computing subtasks according to the received work groups, and allocate the plurality of computing subtasks to at least one shader execution unit; it is also used to obtain computing correlation data from the first-level cache unit to facilitate the generation of corresponding computing subtasks.
[0111] The workload processing device based on multiple computing contexts according to an embodiment of the present invention is provided with two levels of caches; there is a corresponding first-level cache unit in each shader processing cluster. When the task scheduler generates computing subtasks, it performs data interaction with the first-level cache unit in the shader processing cluster where it is located to obtain the cached computing correlation data from the first-level cache unit.
[0112] The shader execution unit is used to process corresponding computing subtasks; it is also used to obtain the required computing correlation data from the first-level cache unit when the computing subtasks require computing correlation data.
[0113] When the shader execution unit processes computing subtasks, it performs data interaction with the first-level cache unit in the shader processing cluster where it is located to obtain the cached computing correlation data from the first-level cache unit.
[0114] The first-level cache unit is further used to obtain the corresponding computing correlation data from the second-level cache unit when the required computing correlation data does not exist in the first-level cache unit.
[0115] Among them, there is a second-level cache unit in the workload processing device based on multiple computing contexts according to an embodiment of the present invention, and all shader processing clusters share the second-level cache unit. When the required computing correlation data does not exist in the first-level cache unit of any shader processing cluster, according to the computing context ID of the independent computing context processed by the shader processing cluster, query the required computing correlation data in the second-level cache unit, and after obtaining the required computing correlation data from the second-level cache unit, return it to the task scheduler or the corresponding shader execution unit.
[0116] The configuration register is a hardware internal control register used to control various functions of the GPU architecture. Some state and control parameters required for processing workloads in an independent computing context can be defined in the configuration register. The state and control parameters of different independent computing contexts can be defined using multiple configuration register groups.
[0117] To enable the workloads of computing processing units from multiple independent computing contexts to be processed in parallel in the GPU, different versions of configuration registers should be maintained in the configuration register groups of each independent computing context in the GPU system; the computing workloads of different independent computing contexts processed in the GPU are controlled by their respective configuration registers; for each independent computing context processed in parallel in the GPU system, the software driver should map a separate memory space to the configuration registers for it.
[0118] The direct memory access module is responsible for fetching the independent computing control flow from memory. It can fetch the independent computing control flow of the independent computing context from memory according to the memory base address defined in its corresponding configuration register group to obtain the corresponding input data. The memory base address and size of the independent computing control flow are defined in the corresponding configuration register group. Each independent computing control flow uses a different configuration register group, so the independent processing module of the GPU system can separately control the workload processing of the independent computing context.
[0119] As Figure 5 shown, in an alternative embodiment, a corresponding trigger ID can be set for each computing trigger signal as the index of the corresponding independent processing module; for example, as Figure 5 shown, the corresponding independent processing modules have indexes with trigger IDs equal to 1, 2, 3, or 4 respectively; the trigger ID is also passed to the task scheduler in the SPC group connected to the independent processing module.
[0120] To distinguish data operations of different independent computing contexts in the secondary cache unit and avoid overlapping virtual memory addresses, a corresponding computing context ID can also be set for each independent computing context. The task scheduler and shader execution unit in the SPC group use the trigger ID index to read the computing context ID of the currently processed independent computing context from the corresponding configuration register.
[0121] In the embodiment of the present invention, the secondary cache unit will process multiple context IDs instead of only one computing context ID; the SPC group determines the computing context ID of the corresponding independent computing context based on the trigger ID of different computing trigger signals and sends it to the secondary cache unit when accessing memory to access the corresponding virtual memory address.
[0122] In the GPU architecture, the maximum number of independent processing modules is fixed. The number of currently used independent processing modules can be configured by the software driver by setting the configuration registers. The number of shader processing clusters in the SPC group corresponding to the independent processing module will be dynamically configured in the hardware of the computing processing unit according to the workload status of the shader processing clusters. The computing workloads of multiple applications can be dynamically allocated to each shader execution unit in the shader processing clusters.
[0123] In the embodiments of the present invention, the workload processed in the shader processing cluster is limited to the workload of the same independent computing control flow from the same independent computing context (for example, having the same computing context ID). In this case, the task scheduler that allocates computing subtasks to the downstream shader execution units will be consistent with the GPU architecture that processes independent computing contexts in the computing processing unit.
[0124] It should be noted that the embodiments of the present invention can effectively improve the overall performance of GPU computing applications, especially when multiple independent small applications need to be processed in the GPU system.
[0125] Based on the above GPU architecture, as Figure 6 shown, the present invention also provides a workload processing method based on multiple computing contexts, including:
[0126] Step 10: Obtain the independent computing contexts corresponding to multiple computing trigger signals, and allocate corresponding independent processing modules to the independent computing contexts.
[0127] Among them, through the control module of the computing processing unit, an independent processing module corresponding to each independent computing context is allocated.
[0128] Step 20: Process the independent computing control flow corresponding to the independent computing context in the independent processing module to obtain multiple workgroups.
[0129] In an optional embodiment, a compute kernel is obtained from the independent compute control flow, the compute kernel is unpacked into workgroups, and then the workgroups are unpacked into instances. The IDs and instances of the workgroups are submitted to the task scheduler.
[0130] Step 30: Process the multiple workgroups to complete multiple computing subtasks corresponding to the computing trigger signals.
[0131] The task scheduler packs the instances into tasks, allocates computing resources, and then generates multiple computing subtasks; the corresponding computing subtasks are allocated to the shader execution units scheduled by the task scheduler for execution, and the computing results of the corresponding computing subtasks are obtained and written into the external memory.
[0132] The present invention obtains independent computing contexts corresponding to multiple computing trigger signals, and allocates corresponding independent processing modules thereto; processes the independent computing control flows corresponding to the independent computing contexts in the independent processing modules to obtain multiple workgroups; and processes the multiple workgroups to complete multiple computing subtasks corresponding to the computing trigger signals. Among them, the computing trigger signals from multiple application programs are processed by a computing processing unit, multiple workgroups corresponding to each independent computing context are respectively determined, and the corresponding workgroups are allocated to a shader processing cluster for processing, so as to realize parallel processing of multiple application programs in a GPU, and solve the problems of too low utilization rate of single-event upsets (SEUs) and affected GPU performance caused by the fact that application programs can only be serially processed on a large number of SEUs in the GPU.
[0133] To illustrate the process of allocating independent processing modules, as Figure 7 and Figure 8 shown, step 10 includes:
[0134] Step 101: When the multiple computing trigger signals correspond to corresponding multiple computing contexts, check the dependency relationship between the multiple computing contexts, and determine whether there are multiple independent computing contexts.
[0135] Among them, there may be a dependency relationship between multiple computing contexts belonging to the same application program. For example, only after a certain computing context is executed, the computing context depending on it can be normally executed; when there is no dependency relationship between one computing context and other computing contexts, then this computing context is an independent computing context. The dependency relationship between multiple computing contexts can be checked by a software driver. When there is more than one independent computing context, the multiple independent computing contexts can be simultaneously processed in parallel in the GPU according to the method of the embodiment of the present invention.
[0136] It should be noted that if there is no dependency relationship between computing contexts, some context switches on the computing contexts can be eliminated by parallel processing of independent computing contexts. In this case, the performance of workload processing will be further improved without being affected by context storage and context loading.
[0137] Step 102: When there are multiple independent computing contexts, determine whether there are available independent processing modules.
[0138] Among them, when there are multiple independent computing contexts, the software driver can choose to set them to be processed in parallel in the GPU.
[0139] Step 103: If there are available independent processing modules, allocate the corresponding independent processing modules to the independent computing contexts.
[0140] Step 104: If there is no available independent processing module, wait until an available independent processing module appears, and assign the appeared available independent processing module to the independent computing context.
[0141] In the embodiment of the present invention, the control module extends the function of the processing configuration register in the existing CPM. The control module can process the configuration register groups corresponding to different independent computing contexts. As Figure 8 shown, when the control module receives a new computing trigger signal, it first checks whether there is an available independent processing module. When there is an available independent processing module, it assigns the available independent processing module to the independent computing context belonging to the computing trigger signal; otherwise, it waits until an available independent processing module appears.
[0142] To illustrate the processing process in the independent processing module, as Figure 9 shown, the step 20 includes:
[0143] Step 201: Obtain the maximum number of shader units corresponding to the independent computing context from the configuration register.
[0144] Among them, the maximum number of shader units is the maximum number of shader execution units in the SPC group corresponding to the independent processing module that processes a certain independent computing context.
[0145] Step 202: According to the current available quantity and the maximum number of shader units, assign corresponding shader processing clusters to each independent processing module; wherein, the shader processing cluster includes a plurality of shader execution units; the number of shader execution units assigned to the independent processing module is less than or equal to its corresponding maximum number of shader units.
[0146] Among them, the current available quantity refers to: the number of currently available shader execution units in all the shader processing clusters.
[0147] It should be noted that when the current available quantity is greater than the maximum number of shader units corresponding to the independent computing context, corresponding shader processing clusters can still be assigned to each independent processing module.
[0148] Step 203: Obtain the corresponding independent computing control flow for the independent computing context from the memory according to the memory base address in its own configuration register.
[0149] Step 204: Obtain the computing kernels in the independent computing control flow, and decompose the computing kernels into multiple workgroups so as to send the corresponding workgroups to the assigned shader processing clusters.
[0150] In an alternative embodiment, the "CONFIG_CPM_START" field can be set in the configuration register, and the content contained in this field is shown in the following table:
[0151]
[0152]
[0153] Among them, when an event that triggers the start of the computing processing unit is set through the configuration register, the computing processing unit will start processing the corresponding computing task; for example, writing 1 in the "START_CPM" field in the "CONFIG_CPM_START" field of the configuration register will trigger an event to start the computing processing unit, and then an independent computing module will be started to process the computing task of this independent computing context.
[0154] The "VALID_CPM_CONTEXTS" field in the "CONFIG_CPM_START" field is used to define the number of independent computing contexts that can be processed in parallel in the GPU; the default value of this field is 0, and a value of 0 for this field means that the GPU can only process one computing context at a time.
[0155] The "MAX_SPCS" field in the "CONFIG_CPM_START" field is used to define the maximum number of shader units in the SPC group corresponding to the independent computing module that the computing processing unit can allocate. When the number of independent computing contexts processed by the GPU is equal to the maximum number of independent computing contexts supported by the GPU defined in the "VALID_CPM_CONTEXTS" field in the "CONFIG_CPM_START" field, the computing trigger signals from different independent computing contexts are not started. If the control module receives a running instruction for the computing processing unit but there is no available independent processing module, this running instruction will be stopped until an available independent processing module appears. When the control module receives a computing trigger signal in a new independent computing context and allocates an independent processing module, the control module will check the number of available shader execution units; the control module allocates the number of shader execution units in the SPC group to the independent processing module corresponding to this computing trigger signal. The control unit allocates the number of shader execution units to the SPC groups of each independent processing module based on the currently available shader execution units and the value of the "MAX_SPCS" field. It should be noted that when the value of the "VALID_CPM_CONTEXTS" field is set to 0, the "MAX_SPCS" field is ignored. In this case, all shader execution units in the GPU are in the only SPC group of the first independent processing module, and the GPU only processes one independent computing context.
[0156] In actual implementation, when a software driver wishes to parallel - process multiple independent computing contexts in a GPU, the software driver should set the value of the "VALID_CPM_CONTEXTS" field to a value greater than 0; meanwhile, the software driver should set the value of "MAX_SPCS" to a value less than the total number of shader execution units in the GPU.
[0157] The control module sends the calculation trigger signal and the definition of the corresponding calculation control flow in the configuration register corresponding to the independent computing context to the allocated independent processing module. Meanwhile, the shader execution units allocated to the independent processing module in the SPC group are also indicated to the independent processing module through the control module using an SPC mask; the SPC mask will be described in detail below.
[0158] The software driver sets configuration registers for the calculation trigger signals of each independent computing context in the memory mapping of a separate configuration register, such as calculation trigger signals, the base address and size of the calculation control flow, etc. All settings of the calculation trigger signals from independent computing contexts are completely independent to ensure that they can be parallel - processed in the GPU.
[0159] To illustrate the process of allocating corresponding shader processing clusters to each independent processing module according to the current available quantity and the maximum number of shader units, as Figure 10 shown, step 202 includes:
[0160] Step 2021: Determine the minimum value between the maximum number of shader units and the current available quantity, and use the minimum value as the total number of shaders of the shader execution units in the shader processing cluster allocated to each independent processing module.
[0161] In an alternative embodiment, the strategy for allocating shader execution units to the SPC group can be to allocate the maximum number of available shader execution units within the limit of the maximum number of shader units in the SPC group; the shader execution units allocated to the SPC group can be expressed as the following formula:
[0162] Group_SPCs = minimum(Available_SPCs, MAX_SPCS)
[0163] where Group_SPCs is the number of shader execution units allocated to the SPC group, and Available_SPCs is the number of currently available shader execution units.
[0164] Step 2022: According to the current trigger ID, determine the parallel trigger IDs of multiple workgroups that are parallel - processed with the multiple workgroups, and obtain the parallel trigger signals corresponding to the parallel trigger IDs.
[0165] Among them, the multiple working groups corresponding to the calculation trigger signal carry the current trigger ID of the calculation trigger signal.
[0166] In an alternative embodiment, the cumulative total of "MAX_SPCS" of the independent calculation contexts corresponding to the number of valid calculation contexts in all "VALID_CPM_CONTEXTS" fields should be equal to the total number of shader execution units in the GPU system; otherwise, the calculation resources of the shader execution units in the GPU may not be fully utilized.
[0167] For example, if there are a total of 10 shader execution units in the GPU system and the number of valid calculation contexts in the "VALID_CPM_CONTEXTS" field is set to 4, then the software driver can set the values of the "MAX_SPCS" fields of each independent calculation context to {4, 2, 2, 2} or {4, 3, 2, 1}, but cannot set it to {2, 2, 2, 2}, because in the case of setting it to {2, 2, 2, 2}, 2 shader execution units will not be used to process the calculation workload.
[0168] Step 2023: Compare the calculation workload corresponding to the calculation trigger signal with the calculation workload corresponding to the parallel trigger signal, and determine the trigger information with a greater calculation workload and / or the trigger information with a higher calculation priority among them.
[0169] Step 2024: Estimate the calculation workload of the independent calculation context, and dynamically allocate shader execution units to the independent processing module of the independent calculation context based on the workload level of the estimated calculation workload.
[0170] When allocating shader execution units, the control module in the embodiment of the present invention considers the workload in the calculation environment, that is, by estimating the calculation workload, more shader execution units are allocated to the independent calculation context with a greater calculation workload to improve the calculation efficiency, and the corresponding allocation strategy will be described below.
[0171] Step 2025: Based on the total number of shaders, preferentially allocate more shader execution units to the determined trigger information according to the current trigger ID or the parallel trigger ID.
[0172] The software driver should set the value of the "MAX_SPCS" field in the "CONFIG_CPM_START" field of the configuration register according to the workload and priority of the compute trigger signal, and compare it with the value of the "MAX_SPCS" field of the compute trigger signals of other independent compute contexts to be processed in parallel in the GPU. Independent compute contexts with large workloads and / or high priorities should be allocated more shader execution units for faster processing in the GPU to improve GPU performance. When multiple independent compute contexts with similar workloads and priorities are set to be processed in parallel in the GPU, the software driver should evenly distribute the total number of shader execution units in the GPU to each independent compute context and accordingly set the value of the "MAX_SPCS" field in the "CONFIG_CPM_START" field of the configuration register for the independent compute context.
[0173] For example, in a GPU system with 16 shader execution units and 4 independent processing modules in the compute processing unit. When there are 10 independent small applications, the software driver can set the value of the "VALID_CPM_CONTEXTS" field to 4 to process 4 independent compute contexts in parallel in the GPU. The value of the "MAX_SPCS" field can be set to 4, so that the 16 shader execution units can be evenly divided into 4 SPC groups to process 4 independent compute contexts in parallel at a time. Another example, if there are two very small applications and one large application, the software driver can set the value of the "VALID_CPM_CONTEXTS" field to 2 to process 2 independent compute contexts in parallel in the GPU. For the small applications, the value of the "MAX_SPCS" field can be set to 1; while for the large application, the value of the "MAX_SPCS" field can be set to 15. Therefore, the 16 shader execution units can be divided into 2 SPC groups to process 2 independent compute contexts in parallel in the GPU at a time. Most of the SPUs in the GPU can be allocated to the SPC group for processing the large application with a large compute workload. The 2 small applications can be successively processed with 1 shader execution unit in the parallel SPC group. Since the processing cycles of the 2 small applications can be masked by the processing cycle of the large application, the overall performance of the GPU can be effectively improved.
[0174] To dynamically allocate more shader execution units to independent compute contexts with a larger compute workload, specifically, as Figure 11 shown, step 2024 includes:
[0175] An estimated workload value of the computing workload is obtained based on the number of shader instructions included in the independent computing context and the product of the size of the main computing kernel of the independent computing context in the X direction, the size of the main computing kernel in the Y direction, and the size of the main computing kernel in the Z direction; wherein, the main computing kernel is: one of the computing kernels included in the independent computing context.
[0176] Alternatively, an estimated kernel workload value of the computing kernel is obtained based on the number of shader instructions and the product of the size of the computing kernel included in the independent computing context in the X direction, the size of the computing kernel in the Y direction, and the size of the computing kernel in the Z direction; the sum of the estimated kernel workload values of all the computing kernels included in the independent computing context is determined as the estimated workload value of the computing workload.
[0177] The size of the computing workload of an application can be estimated by the number of shader instructions included in the independent computing context corresponding to the application and the size of the computing kernel. The size of the computing workload in an independent computing context of an application refers to: the number of operations required for the GPU to process the computing kernel in the independent computing context. When dynamically allocating shader execution units, the estimation of the computing workload does not need to be very accurate, as shown in the following formula:
[0178] W = numInst * DimX * DimY * DimZ
[0179] where, W is the estimated workload value corresponding to the independent computing context, numInst is the number of shader instructions included in the independent computing context
[0180] text, DimX is the size of the computing kernel included in the independent computing context in the X direction, DimY is the size of the computing kernel included in the independent computing context in the Y direction, and DimZ is the size of the computing kernel included in the independent computing context in the Z direction.
[0181] The estimated workload value W corresponding to the independent computing context can be represented by the cumulative value of the estimated workload values of all the computing kernels included in the independent computing context, or can be represented by the estimated workload value of the main computing kernel included in the independent computing context. The estimated workload value corresponding to the independent computing context can be calculated by the software driver according to the application and provided as an input parameter to the control module in the GPU.
[0182] The estimated workload value is mapped to a workload level range to obtain the workload level of the independent computing context.
[0183] To facilitate the simplification of GPU hardware implementation, the workload estimation value corresponding to an independent computing context can be divided into a fixed number of workload levels N, which is used to dynamically allocate the shader execution units included in an independent processing module according to the size of the workload level of the independent computing context. For a GPU, the workload level N is a preset parameter, such as N = 4 or N = 8.
[0184] For an independent computing context i, the corresponding workload level Li can be determined according to the following formula:
[0185] Li = (Wi + m - 1) / m
[0186] Where Wi is the workload estimation value corresponding to the independent computing context i, and m is a preset constant. By the constant m, the workload estimation values Wi corresponding to all independent computing contexts are mapped to the workload level range of (1, N).
[0187] For an independent computing context, the corresponding workload level can be determined by the software driver according to the workload estimation value calculated by the application program and provided to the control module in the GPU as an input parameter. The workload level corresponding to the independent computing context can also be determined by the control module in the GPU according to the workload estimation value calculated by the application program provided by the software driver.
[0188] The control module in the GPU can dynamically allocate the number of shader execution units included in each independent processing module according to the workload level corresponding to the independent computing context. The sum of the workload levels of all independent computing contexts is determined as the first intermediate value; based on the ratio of the workload level of the independent computing context to be allocated to the first intermediate value, a second intermediate value is obtained; based on the product of the second intermediate value and the current available quantity, the number of shader execution units allocated to the independent computing context to be allocated is obtained, so as to allocate more shader execution units to the independent computing context with a higher workload level.
[0189] For example, when allocating the available shader execution units with the quantity of validSPCs to numContexts independent computing contexts according to the workload level, the number of shader execution units numSPCi allocated to the independent processing module of the independent computing context i with the workload level Li can be determined as:
[0190]
[0191] Therefore, the independent processing module corresponding to the independent computing context with a high workload level will be allocated more shader execution units than the independent processing module corresponding to the independent computing context with a low workload level. In the embodiment of the present invention, the available shader execution units are allocated to each independent computing context according to the workload level, avoiding the problem of work processing blockage of the independent processing module in the GPU caused by uneven allocation of shader execution units, thereby improving the overall processing performance of the GPU system.
[0192] In an alternative embodiment, it is also possible to combine the maximum value of the number of shader execution units to dynamically allocate the number of shader execution units included in each independent processing module according to the workload level corresponding to the independent computing context.
[0193] It should be noted that the number of shader execution units numSPCi allocated to the independent processing module of the independent computing context i calculated above should not exceed the maximum value of the shader execution units in the "MAX_SPCS" field of CONFIG_CPM_START in the configuration register corresponding to the independent computing context.
[0194] The control module in the embodiment of the present invention allocates more shader execution units to the independent computing context with a greater computing workload. Due to the stronger computing power from more shader execution units, the computing context with a high workload will be processed faster; and since the allocation is based on the workload level corresponding to the independent computing context, the possibility of too many shader execution units being allocated to the computing context with a high workload is greatly reduced, and the processing of the computing contexts parallelly processed in the GPU, whether their workloads are high or low, will be more balanced.
[0195] In an alternative embodiment, the control module allocates shader execution units for each independent processing module through the SPC mask field, and each independent processing module only works on the corresponding in-use shader execution units in the SPC mask.
[0196] The embodiment of the present invention newly adds configuration registers CONFIG_CPM_SPC_ENABLEO, CONFIG_CPM_SPC_ENABLE1, CONFIG_CPM_SPC_ENABLE2, and CONFIG_CPM_SPC_ENABLE3 for allocating shader execution units, as shown in the following table.
[0197]
[0198] After the control module allocates shader execution units to independent processing modules for independent computing contexts, it writes the value of the SPC mask into the configuration register. The SPC mask field in the configuration register can be read in the secondary cache unit to filter out the memory accesses of the shader execution units used for barrier, flush, and invalidate operations in the independent computing context.
[0199] To illustrate the process of the shader processing cluster handling workgroups, as Figure 12 shown, step 30 includes:
[0200] Step 301: Define the mapping relationship corresponding to the multiple workgroups in the configuration register.
[0201] Step 302: Generate multiple computing subtasks according to the received workgroups.
[0202] Step 303: Process the multiple computing subtasks, obtain computing associated data according to the mapping relationship, and complete the multiple computing subtasks based on the computing associated data.
[0203] As Figure 5 shown, during the process of the shader execution unit completing the corresponding computing subtasks, it may be necessary to quickly obtain the corresponding computing associated data in addition to the input data through the cache; in the embodiment of the present invention, a secondary cache is designed, and the primary cache unit is set in each shader processing cluster, and the secondary cache unit is independent of the shader processing cluster, and all shader processing clusters share the secondary cache unit.
[0204] To illustrate the process of defining the mapping relationship, as Figure 13 shown, step 301 includes:
[0205] Step 3011: Define the computing context ID of the independent computing context in the configuration register; wherein, one configuration register corresponds to one independent computing context.
[0206] The configuration registers of different independent computing contexts can be accessed using the computing context ID (the index of the computing context processed in the GPU).
[0207] Step 3012: Determine the independent processing module of the independent computing context corresponding to the computing trigger signal, and determine the computing context ID of the independent computing context.
[0208] Step 3013: Establish the mapping relationship between the computing context ID and the current trigger ID of the computing trigger signal, and store the mapping relationship in the configuration register corresponding to the independent computing context, so as to obtain the computing context ID according to the mapping relationship and obtain the corresponding computing associated data based on the computing context ID.
[0209] As Figure 14 shown, the secondary cache unit needs to store the compute context IDs of multiple independent compute contexts, so as to facilitate the parallel processing of memory accesses to virtual memory addresses from multiple independent compute contexts in the GPU, in order to avoid the overlap of virtual memory addresses from different independent compute contexts; in order to distinguish memory accesses from different independent compute contexts in the secondary cache unit, embodiments of the present invention map the compute context IDs of the independent compute contexts processed in parallel in the GPU to the corresponding trigger IDs in the independent processing modules through configuration registers. For example, the compute context ID of the independent compute context processed in independent processing module 0 (this trigger ID = 0) can be defined in the configuration register "CONFIG_CONTEXT_MAPPING0", as Figure 5 and the following table shows:
[0210]
[0211] The "context_ID" field is the compute context ID of the compute context processed in the independent processing module when mapped to trigger ID = 0. This compute context ID can be used by the compute processing unit to access memory when accessing data (including operations of the secondary cache unit).
[0212] In order to support the parallel processing of memory accesses from multiple independent compute contexts in the GPU, additional configuration registers can be added for the compute context IDs of multiple independent compute contexts. For a GPU system such as Figure 5 shown, three configuration registers are added on the basis of the original one configuration register to store the corresponding mapping relationships. After the addition, it is shown in the following table.
[0213]
[0214]
[0215] By reading the compute context ID from the mapping relationship stored in the configuration register, the compute context ID for memory access is set in the task scheduler and the shader execution unit. The compute context ID read back from the task scheduler and the shader execution unit will be sent to the secondary cache unit together with the virtual memory address of the memory to be accessed.
[0216] To handle barrier and Cache Flush / Invalidate (CFI) operations in the secondary cache unit, the secondary cache unit needs to obtain information about the shader execution units allocated by the compute processing unit for the SPC groups of independent compute contexts. Among them, a barrier is a synchronization mechanism used to ensure that all relevant threads or processors have reached a certain execution point before performing certain operations, thus avoiding data races or data inconsistencies; the CFI operation refers to the process of clearing or invalidating the data in the cache to ensure that when the data in the memory is modified, the data in the cache will not continue to be misused, thereby ensuring data consistency and accuracy.
[0217] For a GPU architecture such as Figure 5 shown, the SPC mask in the four configuration registers in the above text is used to indicate the shader execution units allocated for the independent compute contexts in the independent processing modules; after allocating the shader execution units for the independent compute contexts to the corresponding independent processing modules, the control module sets the value of the SPC mask in the configuration register.
[0218] The SPC mask in the configuration register will be read in the secondary cache unit to filter out each shader execution unit involved in the shielding and CFI operations in the specified independent compute context.
[0219] The operation signals of the global shielding or CFI operations are broadcast to the shader execution units active in the independent compute contexts in the GPU system; in an optional embodiment, the shader execution units enabled for processing the workload can be defined by setting the "SEU_ENABLE_MASK" field in the "CONFIG_CPM_SEU_ENABLE" field in the configuration register. The SPC mask in the "CONFIG_CPM_SPC_ENABLE" field in the configuration register can be extended to an SEU bit mask according to the configured number of shader execution units in each shader processing cluster of the GPU system. For example, when there are 4 shader execution units in each shader processing cluster in the GPU system, each bit in the SPC mask is repeated 4 times to generate an SEU bit mask for the shader execution units occupied by the independent compute context. Performing an AND operation on the "SEU_ENABLE_MASK" field and the SEU bit mask generates the combined SEU mask for the shader execution units occupied by this independent compute context.
[0220] The operation signals for global masking or CFI operations are broadcast only to multiple shader execution units set in the combined SEU mask of the independent compute context that issues the masking or CFI operation. In the secondary cache unit, the operation signal acts only on the memory interfaces of multiple shader execution units set in the combined SEU mask of the independent compute context that issues the masking or CFI operation. The operation signal is not broadcast to the shader processing clusters and shader execution units occupied by other independent compute contexts, nor is it executed.
[0221] In an alternative embodiment, a page translation table for virtual-to-physical address is set for each independent compute context processed in parallel in the secondary cache unit. The software driver should use the compute context ID to map each independent compute context to its own page translation table. In the prior art, there is often only one compute context ID for an independent compute context in the secondary cache unit. However, in the embodiments of the present invention, there may be multiple compute context IDs in the secondary cache unit. Therefore, the compute context ID of the circuit compute context processed in the secondary cache unit should be extended from one to the maximum number of compute contexts processed in the GPU.
[0222] It should be noted that the solution of the embodiments of the present invention is more applicable to the scenario of multiple independent small computing processing units. Independent compute contexts can be processed in parallel in the GPU through multiple independent compute control flows. In the present invention, the processing performance of mutually dependent compute contexts is not affected because they must be processed sequentially. For an independent compute context with a large computing workload, the improvement in GPU performance obtained by using the workload processing method based on multiple compute contexts in the embodiments of the present invention may not be significant because an independent compute context with a large workload requires more GPU computing resources, and the corresponding independent processing module should be allocated a larger number of shader execution units. Therefore, for an independent compute context with a large computing workload, compared with a GPU system that uses all the shader execution units in the GPU to process only one independent compute context, the improvement in GPU performance obtained by using the workload processing method based on multiple compute contexts in the embodiments of the present invention may not be significant.
[0223] In another alternative embodiment, as Figure 15 shown, is a schematic diagram of the architecture of the workload processing device based on multiple compute contexts according to the embodiments of the present invention. The workload processing device based on multiple compute contexts of this embodiment includes one or more processors 21 and a memory 22. Among them, Figure 15 one processor 21 is taken as an example.
[0224] The processor 21 and the memory 22 can be connected through a bus or other means, Figure 15 and taking the connection through the bus as an example.
[0225] The memory 22 serves as a non-volatile computer-readable storage medium and can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the workload processing method based on multiple computing contexts in this embodiment. The processor 21 executes the workload processing method based on multiple computing contexts by running the non-volatile software programs and instructions stored in the memory 22.
[0226] The memory 22 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 optionally includes a memory remotely disposed relative to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0227] The program instructions / modules are stored in the memory 22 and, when executed by the one or more processors 21, execute the workload processing method based on multiple computing contexts in the above embodiments. For example, execute the Figures 6 to 7 and Figures 9 to 13 steps shown respectively.
[0228] The embodiment of the present invention also provides a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, such as Figure 13 a processor 21, it enables the above one or more processors to execute the workload processing method based on multiple computing contexts in the specific embodiments of the present invention. For example, execute the Figures 6 to 7 and Figures 9 to 13 steps shown respectively; it can also implement Figure 13 the various modules and units described above; or execute the workload processing method based on multiple computing contexts in the specific embodiments of the present invention. For example, execute the Figures 6 to 7 and Figures 9 to 13 steps shown respectively; it can also implement Figure 13 the various modules and units described above.
[0229] It should be noted that for the information interaction, execution process, etc. between the modules and units in the above device and system, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.
[0230] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disks or optical discs, etc.
[0231] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A computing processing unit, characterized in that: It includes a control module and multiple independent processing modules; The output end of the control module is connected to the input ends of the plurality of independent processing modules respectively; The control module is used to obtain a calculation trigger signal, and send the independent calculation contexts corresponding to the multiple calculation trigger signals to the corresponding independent processing modules respectively; The independent processing module is used to process the received independent computing context.
2. The computing processing unit according to claim 1, characterized in that: Also includes an arbitration module; the independent processing module includes a direct memory access module and a computing workgroup generator; The output end of the computing processing unit is connected to the input end of a plurality of direct memory access modules respectively, and the output end of the direct memory access module is connected to the input end of the computing work group generator in the independent processing module where the module is located; The direct memory access module is used to receive the independent computing context allocated by the control module, and obtain the independent computing control flow corresponding to the independent computing context from the memory through the arbitration module; The computing workgroup generator is used to obtain the computing kernel in the independent computing control flow and decompose the computing kernel into multiple working groups.
3. A workload processing device based on multiple computing contexts, characterized in that: The method comprises a computing processing unit as claimed in claim 1 or 2 and at least one shader processing cluster; the output end of the computing processing unit is connected to the input end of the at least one shader processing cluster; The computing processing unit is used to obtain a computing trigger signal, and decompose the computing kernels of the independent computing contexts corresponding to the multiple computing trigger signals into multiple working groups; Also used for sending the plurality of working groups to the shader processing cluster; The shader processing cluster is used to process the plurality of work groups.
4. The workload processing device based on multiple computing contexts according to claim 3, characterized in that: It also includes a configuration register and a secondary cache unit; the shader processing cluster includes a task scheduler, a shader execution unit and a primary cache unit; The output end of the configuration register is connected to the input end of the control module; the output end of the independent processing module is respectively connected to the input end of multiple task schedulers; the output end of the task scheduler is respectively connected to the input end of multiple shader execution units; the first-level cache unit is respectively connected to the multiple shader execution units; the task scheduler is also connected to the first-level cache unit; the second-level cache unit is respectively connected to multiple first-level cache units; The configuration register is used to control the calculation trigger signal based on the state control information, and send the calculation trigger signal to the control module; The task scheduler is used to generate a plurality of computing subtasks according to the received work group, and assign the plurality of computing subtasks to at least one shader execution unit; Also used for acquiring calculation-related data from the first-level cache unit to generate corresponding calculation subtasks; The shader execution unit is used to process the corresponding calculation subtask; and is also used to obtain the required calculation-related data from the first-level cache unit when the calculation subtask needs to calculate the associated data; The first-level cache unit is further configured to obtain corresponding calculation-related data from the second-level cache unit when the required calculation-related data does not exist in the first-level cache unit.
5. A workload processing method based on multiple computing contexts, characterized in that: include: Acquire independent computing contexts corresponding to a plurality of computing trigger signals, and assign corresponding independent processing modules to the independent computing contexts; Processing the independent computing control flow corresponding to the independent computing context in the independent processing module to obtain a plurality of working groups; The multiple working groups are processed to complete the multiple computing subtasks corresponding to the computing trigger signal.
6. The workload processing method based on multiple computing contexts according to claim 5, characterized in that: The step of obtaining independent computing contexts corresponding to a plurality of computing trigger signals and allocating corresponding independent processing modules to the independent computing contexts comprises: When the computing trigger signal corresponds to multiple computing contexts, checking the dependency relationship between the multiple computing contexts to determine whether there are multiple independent computing contexts; When there are multiple independent computing contexts, determining whether there is an available independent processing module; If there is an available independent processing module, assigning the corresponding independent processing module to the independent computing context; If there is no available independent processing module, wait until an available independent processing module appears, and allocate the available independent processing module to the independent computing context.
7. The workload processing method based on multiple computing contexts according to claim 5, characterized in that: The independent computing control flow corresponding to the independent computing context is processed in the independent processing module to obtain multiple working groups, including: Get the maximum number of shader units corresponding to the independent computing context from the configuration register; According to the currently available number and the maximum number of shader units, a corresponding shader processing cluster is allocated to each independent processing module; wherein the shader processing cluster includes a plurality of shader execution units; the number of shader execution units allocated to the independent processing module is less than or equal to the maximum number of shader units corresponding to it; Obtaining a corresponding independent computing control flow for the independent computing context from the memory according to the memory base address in its own configuration register; A computing kernel in the independent computing control flow is obtained, and the computing kernel is decomposed into a plurality of working groups, so as to send corresponding working groups to the assigned shader processing cluster.
8. The workload processing method based on multiple computing contexts according to claim 7, characterized in that: The multiple working groups corresponding to the calculation trigger signal carry the current trigger ID of the calculation trigger signal; The allocating a corresponding shader processing cluster to each independent processing module according to the currently available number and the maximum number of shader units includes: Determine a minimum value between the maximum number of shader units and the currently available number, and use the minimum value as the total number of shaders of the shader execution units in the shader processing cluster allocated to each independent processing module; According to the current trigger ID, determining the parallel trigger IDs of the multiple working groups processed in parallel with the multiple working groups, and obtaining the parallel trigger signal corresponding to the parallel trigger ID; Comparing the computation workload corresponding to the computation trigger signal with the computation workload corresponding to the parallel trigger signal, and determining the trigger information with a greater computation workload and / or the trigger information with a higher computation priority; estimating a computational workload of the independent computational context, and dynamically allocating shader execution units to independent processing modules of the independent computational context based on a workload level of the estimated computational workload; Based on the total number of shaders, more shader execution units are preferentially allocated to the determined trigger information according to the current trigger ID or the parallel trigger ID.
9. The workload processing method based on multiple computing contexts according to claim 7, characterized in that: The estimating the computational workload of the independent computational context and dynamically allocating a shader execution unit to an independent processing module of the independent computational context based on the estimated workload level of the computational workload comprises: Based on the number of shader instructions contained in the independent computing context, and the product of the size of the main computing kernel of the independent computing context in the X direction, the size of the main computing kernel in the Y direction, and the size of the main computing kernel in the Z direction, the workload estimation value of the computing workload is obtained; wherein the main computing kernel is: one of the computing kernels contained in the independent computing context; or, based on the number of shader instructions, and the product of the size of the computing kernel in the X direction, the size of the computing kernel in the Y direction, and the size of the computing kernel in the Z direction, the kernel workload value of the computing kernel is obtained; the sum of the kernel workload values of all the computing kernels contained in the independent computing context is determined as the workload estimation value of the computing workload; Mapping the workload estimate value into a workload level range to obtain the workload level of the independent computing context; The accumulation of the workload levels of all independent computing contexts is determined as a first intermediate value; based on the ratio of the workload level of the independent computing context to be allocated to the first intermediate value, a second intermediate value is obtained; based on the product of the second intermediate value and the currently available number, the number of shader execution units allocated to the independent computing context to be allocated is obtained, so as to allocate more shader execution units to the independent computing context with a high workload level.
10. The workload processing method based on multiple computing contexts according to claim 7, characterized in that: The processing of the plurality of working groups comprises: Defining a computing context ID of the independent computing context in the configuration register; wherein one configuration register corresponds to one independent computing context; Determine an independent processing module for processing the independent computing context corresponding to the computing trigger signal, and determine a computing context ID of the independent computing context; Establishing a mapping relationship between the computing context ID and the current trigger ID of the computing trigger signal, and storing the mapping relationship in a configuration register corresponding to the independent computing context, so as to obtain the computing context ID according to the mapping relationship, and obtain corresponding computing-related data based on the computing context ID; Generate multiple computing subtasks according to the received work group; The plurality of computing subtasks are processed, and computing-related data are acquired according to the mapping relationship, so as to complete the plurality of computing subtasks based on the computing-related data.