Computing task scheduling method and device and parallel processing system
By breaking down large matrix operations into subsets on the GPU and accumulating intermediate results in the local memory of the computing core, the problems of memory bandwidth pressure and low computational efficiency are solved, achieving more efficient computing performance.
Patent Information
- Application Number
- CN202511900872.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-02-06
AI Technical Summary
Existing GPU computing strategies lead to increased memory bandwidth pressure and reduced computational efficiency when performing large matrix operations. This is mainly because intermediate results require frequent access to global memory, which introduces additional read/write latency and synchronization overhead.
Large matrix multiplication and addition operations are broken down into task subsets and these subtasks are scheduled to be executed on the same computing core. The local memory of the computing core is used to accumulate intermediate results, thus avoiding frequent access to global memory.
By using local memory, the pressure on video memory bandwidth is reduced, unnecessary synchronization and read/write latency are eliminated, and the overall computational efficiency of large matrix multiplication and addition operations is improved.
Smart Images

Figure CN121478451A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a computing task scheduling method, apparatus, and parallel processing system. Background Technology
[0002] Modern graphics processing units (GPUs) are widely used in artificial intelligence computing and general-purpose computing. These GPUs typically contain a large number of dedicated computing cores, such as tensor cores or matrix cores, to accelerate matrix operations.
[0003] In applications such as deep learning, the matrices that need to be processed are often huge. Therefore, it is necessary to break down large matrix operation tasks into multiple smaller computational subtasks, and then schedule these computational subtasks to numerous computing cores for parallel computation.
[0004] However, existing scheduling strategies have efficiency bottlenecks. For example, when calculating a 128×128 matrix multiplication, in order to calculate a 16×16 sub-block (e.g., D[0-15:0-15]) in the result matrix D, it may be necessary to execute multiple computational subtasks (e.g., an A[0-15:0-63]×B[0-63:0-15] task and an A[0-15:64-127]×B[64-127:0-15] task) and accumulate their results.
[0005] In existing technologies, logically related computational subtasks are assigned to arbitrarily different, currently idle computing cores. Each core performs parallel computation, writes intermediate results to GPU memory, waits for all cores to complete their computations, and then performs a final accumulation operation. However, this approach increases the pressure on GPU memory bandwidth and introduces additional read / write latency and synchronization overhead, thus reducing the overall computational efficiency of the GPU.
[0006] Therefore, how to effectively reduce the access pressure on global video memory and improve computational efficiency when performing large matrix operations in a parallel processing system is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] This disclosure provides a computing task scheduling method, apparatus, and parallel processing system; it can effectively reduce the access pressure on global video memory and improve computing efficiency when performing large matrix operations in a parallel processing system.
[0008] The technical solution disclosed herein is implemented as follows: In a first aspect, this disclosure provides a computational task scheduling method applied to a parallel processing system comprising multiple computational cores configured with local memory, the method comprising: Large matrix multiplication and addition operations are broken down into one or more task subsets, each of which includes at least one computational subtask. Schedule each subset of tasks to the corresponding computing core; Each computational subtask in the corresponding task subset is executed using the computational core; Specifically, during the execution of the current computational subtask by the computational core, the first computation result of the previous computational subtask is read from the local memory configured in the computational core, and the first computation result is added to the second computation result of the current computational subtask to obtain the third computation result, which is then stored in the local memory.
[0009] Secondly, this disclosure provides a computing task scheduling apparatus applied to a parallel processing system including multiple computing cores configured with local memory, the apparatus comprising: The splitting module is used to break down large matrix multiplication and addition operations into one or more task subsets, each task subset including at least one computational subtask. The scheduling module is used to schedule each subset of tasks to the corresponding computing core; The computing module is used to execute each computing subtask in the corresponding task subset using the computing core; Specifically, during the execution of the current computational subtask by the computational core, the first computation result of the previous computational subtask is read from the local memory configured in the computational core, and the first computation result is added to the second computation result of the current computational subtask to obtain the third computation result, which is then stored in the local memory.
[0010] Thirdly, this disclosure provides a parallel processing system, including: Global memory; Multiple computing cores, each configured with local memory; and The computing task scheduling device disclosed herein.
[0011] Fourthly, this disclosure provides a computing device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the computing task scheduling method described in this disclosure.
[0012] Fifthly, this disclosure provides a computer-readable storage medium storing at least one instruction, which is executed by a processor to implement the computing task scheduling method described in this disclosure.
[0013] This disclosure provides a computational task scheduling method, apparatus, and parallel processing system. Specifically, this disclosure is applied to a parallel processing system including multiple computational cores configured with local memory. First, a large matrix multiplication-addition operation is divided into one or more task subsets, each including at least one computational subtask. All computational subtasks in each task subset are scheduled to corresponding computational cores, and all computational subtasks in the corresponding task subset are executed sequentially by the same computational core. When executing the current computational subtask, the computational core directly reads the calculation result of the previous computational subtask from its configured local memory, accumulates it, and stores the result of the current computational subtask in the local memory. By using local memory to store and read intermediate results, the intermediate result read / write operations that originally had to be performed through the graphics card are moved to the local memory of the computational core, reducing the pressure on graphics memory bandwidth, eliminating unnecessary synchronization and read / write latency, thereby improving the overall computational efficiency of large matrix multiplication-addition operations. Attached Figure Description
[0014] Figure 1 This is a diagram of the parallel processing system architecture provided in this disclosure.
[0015] Figure 2 A flowchart of the computational task scheduling method provided in this disclosure.
[0016] Figure 3 This is a schematic diagram of all computational subtasks in the subset of tasks performed by the computational core provided in this disclosure.
[0017] Figure 4 This is a flowchart of the large matrix multiplication and addition operation breakdown provided in this disclosure.
[0018] Figure 5 The flowchart of the two-level scheduling method provided in this disclosure.
[0019] Figure 6 This is a schematic diagram of the computing task scheduling device provided in this disclosure.
[0020] Figure 7 This is a schematic diagram of the structure of the computing device provided in this disclosure.
[0021] Figure 8 This is a schematic diagram of a multiply-accumulate operation in the prior art provided in this disclosure that does not exceed the maximum processing capacity of the computing core.
[0022] Figure 9 This is a schematic diagram illustrating the multiply-accumulate operation performed by the first idle computing core when the processing capacity of the computing core exceeds the maximum processing capacity provided in the prior art of this disclosure.
[0023] Figure 10This is a schematic diagram illustrating the multiplication and addition operation performed by a second idle computing core when the processing capacity of the computing core exceeds the maximum processing capacity of the computing core, as provided in the prior art of this disclosure. Detailed Implementation
[0024] The technical solutions in this disclosure will now be clearly and completely described with reference to the accompanying drawings.
[0025] Figure 1 This is a diagram illustrating the architecture of a parallel processing system provided in this disclosure. The parallel processing system 10 can be a graphics processing unit (GPU), a computing array in a system-on-a-chip (SoC), or other type of processor hardware integrating multiple computing cores. This parallel processing system 10 serves as the execution environment for all subsequent method embodiments. Figure 1 As shown, the parallel processing system 10 may include: a global memory 101, a command processor 102, and one or more computing clusters 103.
[0026] In GPU scenarios, Global Memory 101 typically refers to Video Random Access Memory (VRAM). Global Memory 101 is a system-wide shared storage space with a large capacity (e.g., several GB to tens of GB), but its access speed is relatively slow, and its access bandwidth is a bottleneck resource shared by the entire system.
[0027] Command processor 102 is the top-level scheduling unit of parallel processing system 10. It is responsible for receiving high-level computational tasks from the central processing unit (CPU) or drivers, such as a complete large matrix multiplication and addition operation (e.g., D=A×B+C). Command processor 102 is used to break down large matrix multiplication and addition operations into multiple subtasks and assign the subtasks to computation cluster 103.
[0028] The parallel processing system 10 comprises multiple computing clusters 103. The computing cluster 103 is the basic object for the command processor 102 to allocate tasks. Each computing cluster 103 has a certain upper limit on its processing capacity; for example, a computing cluster 103 may be designed to handle matrix operations of a maximum size of 16×16×64.
[0029] Command processor 102 communicates with all computing clusters 103 via control bus 104. This control bus 104 can be a high-speed bus dedicated to transmitting control data (such as task instructions and thread number ranges), such as the AXI (Advanced Xtensible Interface) bus. Command processor 102 sends task instructions to designated computing clusters 103 via this control bus 104.
[0030] Figure 1 The internal structure of the computing cluster is shown in detail. Each computing cluster 103 may include: a dispatch unit 105 and multiple computing cores 106.
[0031] The scheduling unit 105 receives subtasks from the command processor 102 and is responsible for allocating the subtasks to the computing cores 106 under its jurisdiction. Specifically, the scheduling unit 105 can poll and allocate tasks to idle computing cores 106 based on the idle status bit of the computing cores 106 under its jurisdiction. In this disclosure, the scheduling unit 105 can be a fixed-function unit implemented through hardware logic (e.g., a state machine, ASIC) or an embedded microcontroller that implements its scheduling function by executing firmware instructions.
[0032] The computation core 106 is the smallest hardware unit that performs computations. The computation core 106 contains an arithmetic logic unit and high-speed registers for performing computations, and the computation core 106 is configured with local memory 107.
[0033] Based on the parallel processing system 10 described above, assuming the task assigned to computing core 106 is matrix block operations with a basic unit size of 16×16×64, then given two matrices of size 64×64 performing a multiplication-addition operation: D=A×B+C, the general scheduling strategy is: like Figure 8 As shown, the first idle computing core 106 calculates the first 0-15 rows of matrix A multiplied by the first 0-15 columns of matrix B, and then adds the corresponding elements of matrix C; The second idle computing core 106 calculates the first 0-15 rows of matrix A multiplied by the first 16-31 columns of matrix B, and then adds the corresponding elements of matrix C; The third idle computing core 106 calculates the first 0-15 rows of matrix A multiplied by the first 32-47 columns of matrix B, and then adds the corresponding elements of matrix C; The fourth idle computation core 106 calculates the first 0-15 rows of matrix A multiplied by the first 48-63 columns of matrix B, and then adds the corresponding elements of matrix C. The fifth idle computation core 106 calculates the first 16-31 rows of matrix A multiplied by the first 0-15 columns of matrix B, and then adds the corresponding elements of matrix C. The sixth idle computation core 106 calculates the first 16-31 rows of matrix A multiplied by the first 16-31 columns of matrix B, plus the corresponding elements of C; The seventh idle computation core 106 calculates the first 16-31 rows of matrix A multiplied by the first 32-47 columns of matrix B, plus the corresponding elements of C; The eighth idle computation core 106 calculates the first 16-31 rows of matrix A multiplied by the first 48-63 columns of matrix B, plus the corresponding elements of C; This process continues until the complete matrix D is obtained.
[0034] The basic unit of scheduling for the computing core 106 is a 16×16×64 matrix block operation, which is determined by the hardware design. The size of the register resources in the computing core 106 determines the maximum number of threads that can run in parallel (Workgroup / Threadblock), meaning that the maximum processing capacity of the computing core 106 is fixed.
[0035] When the matrix becomes 128×128 in size, if the scheduling strategy above is still applied, data dependencies between cores will occur, thus reducing computational parallelism. For example: like Figure 9 As shown, the first idle computing core 106 calculates A[0-15:0-63]×B[0-63:0-15]+C[0-15:0-15]; like Figure 10 As shown, the second idle computing core 106 calculates A[0-15:64-127]×B[64-127:0-15]+C[0-15:0-15]; The third idle computing core 106 calculates A[0-15:0-63]×B[0-63:16-31]+C[0-15:16-31]; The fourth idle computing core 106 calculates A[0-15:64-127]×B[64-127:16-31]+C[0-15:16-31]; This process continues until the complete matrix D is obtained.
[0036] The result of the first computation core needs to be added to the result of the second computation core to obtain the final result: D[0-15:0-15]. The common approach is to have each computation core perform the computation in parallel, write the result to video memory, wait for all cores to complete their computations, and then perform a final accumulation operation. However, this method increases the pressure on video memory bandwidth and reduces computational efficiency.
[0037] Figure 2 A flowchart illustrating the computational task scheduling method provided in this disclosure. This method is applied to a parallel processing system comprising multiple computational cores configured with local memory, and can be used by… Figure 1 The parallel processing system 10 shown may be executed by one or more controllers, for example, primarily by command processor 102, or by command processor 102 and scheduling unit 105 in cooperation.
[0038] like Figure 2 As shown, the method includes: Step S201: The large matrix multiplication and addition operation is split into one or more task subsets, each task subset including at least one computational subtask.
[0039] In this disclosure, the command processor 102 receives a large matrix multiplication and addition operation, such as a 128×128 dimension D=A×B+C operation, and then analyzes the computational dependencies of the task.
[0040] Large matrix multiplication and addition operations refer to computational operations whose scale (e.g., the M, N, K dimensions of the matrix) exceeds the processing capacity of a single computing core 106 (or even a computing cluster 103) in the parallel processing system 10, and therefore must be broken down into multiple computational subtasks. For example, a 16×16×128 matrix is broken down into 16×16×64 or 16×16×16 matrix blocks. Each sub-block in the resulting matrix D is obtained by multiplying at least one row from the source matrix A and at least one column from the source matrix B. When the multiplication of a row in source matrix A and a column in source matrix B exceeds the maximum processing capacity of a single computing core 106, it needs to be broken down into multiple computational subtasks. The results of these computational subtasks are added together to obtain the corresponding sub-block in the resulting matrix D. In this disclosure, it is assumed that these computational subtasks required to obtain the corresponding sub-block in the resulting matrix D are dependent on each other, and therefore these computational subtasks are grouped into the same task subset. One task subset corresponds to a single sub-block of the resulting matrix D.
[0041] Taking a large matrix multiplication and addition operation of 16×16×128 as an example, in order to calculate the result matrix sub-block of size D[0-15:0-15] in the result matrix D, the command processor recognizes that the result matrix sub-block requires the accumulation of the following two computational subtasks: Calculate subtask 1: A[0-15:0-63]×B[0-63:0-15]; Calculate subtask 2: A[0-15:64-127]×B[64-127:0-15].
[0042] Computational subtasks 1 and 2 together constitute a task subset of sub-block D[0-15:0-15] in the guiding result matrix D. The scheduler determines which computational subtasks should belong to a task subset by analyzing the computation graph, the instructions of the shader program, or based on the inherent mathematical properties of matrix operations (i.e., K-dimensional splitting). For example, when generating tasks for a general-purpose computing program (such as a CUDA or OpenCL kernel), the command processor 102 can resolve to a pattern of cyclic accumulation or K-dimensional splitting in the program, thereby marking all computational subtasks involved in the cycle or split as belonging to the same task subset.
[0043] Step S202: Schedule each subset of tasks to the corresponding computing core.
[0044] In this disclosure, all computational subtasks in the task subset are assigned to the corresponding computational core 106 for execution. If the assigned computational core 106 is in a "busy" state while executing computational subtask 1, then computational subtask 2 will not be assigned to other idle computational cores 106, but will be queued in the waiting queue of its superior (e.g., scheduling unit 105) until the assigned computational core 106 becomes idle.
[0045] This scheduling can be implemented in various ways. For example, the command processor 102 can assign the same affinity ID to all computational subtasks in the task subset, while the scheduling unit 105 is configured such that all computational subtasks with the same affinity ID must be executed on the same computational core 106.
[0046] Step S203: Execute each computational subtask in the corresponding task subset using the computational kernel; During the execution of the current computational subtask by the computing core 106, the first computation result of the previous computational subtask is read from the local memory 107 configured in the computing core 106, and the first computation result is added to the second computation result of the current computational subtask to obtain the third computation result, which is then stored in the local memory 107.
[0047] Figure 3 This is a schematic diagram illustrating all computational subtasks within the task subset executed by the computing core 106 provided in this disclosure. In this disclosure, computational tasks are executed using the private resources of the computing core 106, as shown below. Figure 3 The specific process is as follows: The computing core 106 executes the computing subtask 1: A[0-15:0-63]×B[0-63:0-15], and stores the first computing result Partial_D1 obtained therefrom in its configured local memory 107.
[0048] The computation core 106 executes computation subtask 2: A[0-15:64-127]×B[64-127:0-15], and obtains the second computation result Partial_D2. During accumulation, it directly reads the previously stored first computation result Partial_D1 from its local memory 107 as matrix C[0-15: 0-15] for accumulation, that is, calculates Partial_D2+Partial_D1, and stores the calculated third computation result Partial_D3 in its configured local memory 107.
[0049] In this way, both the write and read operations of the intermediate result Partial_D1 are completed in a closed loop within the local memory 107 of the computing core 106. This avoids accessing the global memory for intermediate accumulated results.
[0050] Therefore, in this disclosure, by identifying task subsets and binding all computational subtasks within the task subsets to the same computational core 106, intermediate result reads and writes that originally had to be performed through the global memory 101 (VRAM) are moved to the local memory 107 of the computational core 106. Compared to the global memory 101, the local memory 107 is on-chip storage with extremely low access latency, much faster than accessing the off-chip global memory 101. Secondly, access to the local memory 107 is a local behavior of the computational core 106 and does not occupy the bus bandwidth of the system-level global memory 101. In addition, the capacity of the local memory 107 is typically small (e.g., tens of KB to several MB), much smaller than that of the global memory 101. Therefore, the method in this disclosure reduces the pressure on video memory bandwidth, eliminates unnecessary synchronization and read / write latency, thereby improving the overall computational efficiency of large matrix multiplication and addition operations.
[0051] Figure 4 This is a flowchart illustrating the breakdown of large matrix multiplication and addition operations provided in this disclosure. (For example...) Figure 4 As shown, large matrix multiplication and addition operations are operations that multiply source matrix A and source matrix B; Break down large matrix multiplication and addition operations into one or more subsets of tasks, including: Step S401: Decompose the large matrix multiplication and addition operation into an operation that multiplies at least one row of source matrix A with at least one column of source matrix B, to obtain one or more medium-granularity computing tasks. Step S402: Divide each medium-granularity computing task into one or more computing subtasks according to the maximum processing capacity of the computing core; Step S403: Divide all the computational subtasks obtained from the same medium-granularity computation task into corresponding task subsets.
[0052] Taking a large 128×128 matrix multiplication and addition operation received by command processor 102, requiring the computation of sub-block D[0-15:0-15] in the result matrix D as an example, the decomposition process for the large matrix multiplication and addition operation is as follows: Based on the size of sub-block D[0-15:0-15] in the result matrix D, command processor 102 determines that the 128×128 large matrix multiplication and addition operation needs to be decomposed into K dimensions, dividing the K dimensions into 0-63 and 64-127. Then, based on rows 0-15 of source matrix A and columns 0-15 of source matrix B, all relevant computational subtasks are determined: Calculate subtask 1: A[0-15:0-63]×B[0-63:0-15]; Calculate subtask 2: A[0-15:64-127]×B[64-127:0-15].
[0053] Command processor 102 combines computation subtask 1, computation subtask 2, and possible addition tasks of C[0-15:0-15] into a task subset.
[0054] Figure 5 This is a flowchart of the two-level scheduling method provided in this disclosure. Figure 5 As shown, all computational subtasks in the task subset are scheduled to the same computational core 106, including: Step S501: The command processor schedules a subset of tasks to the computing cluster containing the computing core. In step S502, the scheduling unit within the computing cluster schedules all computing subtasks in the task subset to the computing core.
[0055] This disclosure provides a two-level scheduling hierarchy, the specific process of which is as follows: Execute Level 1 scheduling: Command processor 102 selects a computing cluster from N computing clusters 103, and packages the entire task subset (including computing subtask 1 and computing subtask 2) into a medium-granularity task, which is then sent to the computing cluster via control bus 104.
[0056] Next, secondary scheduling is performed: the scheduling unit 105 inside the computing cluster 103 receives the task subset, selects a computing core 106 from its M computing cores 106, and schedules all computing subtasks in the task subset (first computing subtask 1, then computing subtask 2) sequentially and completely to the computing core 106 for execution by polling the idle status bit of the computing core 106.
[0057] If computational subtask 2 arrives while computational core 106 is executing computational subtask 1, scheduling unit 105 places it into the waiting queue of computational core 106.
[0058] After the designated computing core 106 completes the computing subtask 1, it writes the first computing result Partial_D1 to its local memory 107. In this disclosure, the local memory 107 is the shared memory of the computing core 106.
[0059] After completing computational subtask 1, the designated computational core 106 becomes idle, and then the scheduling unit 105 schedules computational subtask 2 to the computational core 106. When executing computational subtask 2, the computational core 106 reads Partial_D1 from Shared Memory and adds it to the result Partial_D2 of computational subtask 2.
[0060] Therefore, the computational core 106 specified in this disclosure completes the calculation of Partial_D2+Partial_D1 within itself to obtain the final D[0-15:0-15].
[0061] Based on the above two-level scheduling architecture and the use of Shared Memory, the command processor always ensures that all matrix operation tasks of a certain row of source matrix A and a certain column of source matrix B are scheduled to the same computing cluster. The computing core 106 uses its private Shared Memory to store the intermediate results of the previous matrix task. The computing core 106 directly reads the intermediate results stored in Shared Memory for accumulation calculation, which can reduce the access of video memory and improve computing efficiency.
[0062] In this disclosure, the computational subtasks in the task subset include: A first computational subtask with a first computational granularity and a second computational subtask with a second computational granularity; The first and second computational subtasks are scheduled together to the corresponding computational core 106.
[0063] The first computational granularity is a regular matrix size that matches the maximum processing power of the 106-core computing system; the second computational granularity is an irregular matrix size that does not match the maximum processing power of the 106-core computing system.
[0064] In practical AI and HPC (High Performance Computing) applications, the dimensions of matrices are not always the well-defined size (e.g., multiples of 16, 64, or 128) expected by hardware designers, such as in a 130×130 matrix multiplication task. A common approach to handling such irregular dimensions is to use padding techniques, filling 130 to 192 (the next multiple of 64) or 144 (the next multiple of 16), but this introduces a large amount of unnecessary computation. Another common approach is to process the well-defined portion (0-127) and the irregular portion (128-129) separately, but assigning them to different computation cores (106), and accumulating their respective results via VRAM.
[0065] In this disclosure, the normalized and irregular parts are assigned to the same computational core 106 for computation, and the specific process is as follows: Command processor 102 receives a large matrix multiplication and addition operation of 130×130 dimensions: 16×16×130.
[0066] Assume that the maximum processing capacity of computation cluster 103 is 16×16×64. To process 16×16×130, it is split into two computational subtasks with a granularity of 16×16×64 (normalized matrix size) and one computational subtask with a granularity of 16×16×2 (disnormalized matrix size), specifically: The calculation subtask X: A[0-15:0-63]×B[0-63:0-15] belongs to the first calculation subtask.
[0067] The calculation subtask Y: A[0-15:64-127]×B[64-127:0-15] belongs to the first calculation subtask.
[0068] The computational subtask Z: A[0-15:128-129]×B[128-129:0-15] belongs to the second computational subtask.
[0069] The aforementioned computational subtasks X, Y, and Z all belong to the same subset of tasks. In this disclosure, computational subtasks X, Y, and Z are scheduled together to the same computational core 106. The computational core 106 executes these three tasks sequentially: Execute the computation subtask X, obtain the computation result R_X, and store R_X into Shared Memory; Execute the computation subtask Y, obtain the computation result R_Y, read R_X stored in Share Memory, calculate R_Y+R_X, and store the computation result R_XY into Share Memory.
[0070] Execute the computation subtask Z to obtain the computation result R_Z, read R_XY stored in Share Memory, calculate R_Z+R_XY, and store the final result in Share Memory.
[0071] Figure 6 This is a schematic diagram of the computing task scheduling device provided in this disclosure. Figure 6 As shown, the device is applied to a parallel processing system comprising multiple computing cores configured with local memory. The device includes: The splitting module 601 is used to split a large matrix multiplication and addition operation into one or more task subsets, each task subset including at least one computational subtask. The scheduling module 602 is used to schedule each subset of tasks to the corresponding computing core 106; The computing module 603 is used to execute each computing subtask in the corresponding task subset using the computing core 106.
[0072] The computational task scheduling device provided in this disclosure can achieve... Figures 1 to 5 The various processes implemented in the method embodiments shown will not be described again here to avoid repetition.
[0073] Figure 7 This is a structural block diagram of the computing device provided in this disclosure. In some examples, the computing device 70 can be at least one of devices such as a smartphone, smartwatch, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and laptop computer. The computing device 70 has communication functions and can access wired or wireless networks. The computing device 70 can refer to one of multiple terminals, and those skilled in the art will understand that the number of such terminals can be more or less. It is understood that the computing device 70 undertakes the calculation and processing work of the technical solution of this disclosure, and this disclosure does not limit it in this respect.
[0074] like Figure 7 As shown, the computing device in this disclosure may include one or more of the following components: processor 710 and memory 720.
[0075] Optionally, the processor 710 connects various parts within the computing device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 720, and by calling data stored in the memory 720. Optionally, the processor 710 can be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 710 can integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), and baseband chip. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the touch screen; the NPU is used to implement Artificial Intelligence (AI) functions; and the baseband chip is used to handle wireless communication. It is understandable that the aforementioned baseband chip may not be integrated into the processor 710, but may be implemented using a separate chip.
[0076] The memory 720 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 720 may include a non-transitory computer-readable storage medium. The memory 720 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 720 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created according to the use of the computing device, etc.
[0077] In addition, those skilled in the art will understand that the structure of the computing device shown in the above figures does not constitute a limitation on the computing device. The computing device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the computing device may also include a display screen, camera assembly, microphone, speaker, radio frequency circuit, input unit, sensors (such as accelerometer, angular velocity sensor, light sensor, etc.), audio circuit, WiFi module, power supply, Bluetooth module, etc., which will not be described in detail here.
[0078] This disclosure also provides a computer-readable storage medium storing at least one instruction that is executed by a processor to implement the computing task scheduling method described in the above embodiments.
[0079] Those skilled in the art will recognize that the functions described in this disclosure in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium accessible to a general-purpose or special-purpose computer.
[0080] It should be noted that the technical solutions described in this disclosure can be combined arbitrarily as long as they do not conflict.
[0081] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A computational task scheduling method, applied to a parallel processing system comprising multiple computational cores configured with local memory, characterized in that, The method includes: Large matrix multiplication and addition operations are broken down into one or more task subsets, each of which includes at least one computational subtask. Schedule each of the aforementioned task subsets to the corresponding computing core; Each computational subtask in the corresponding task subset is executed using the computational core; Specifically, during the execution of the current computational subtask by the computational core, the first computational result of the previous computational subtask is read from the local memory configured in the computational core, and the first computational result is added to the second computational result of the current computational subtask to obtain a third computational result, which is then stored in the local memory.
2. The computational task scheduling method according to claim 1, characterized in that, The large matrix multiplication and addition operation is the operation of multiplying source matrix A and source matrix B; The step of splitting large matrix multiplication and addition operations into one or more task subsets includes: The large matrix multiplication and addition operation is broken down into an operation that multiplies at least one row of the source matrix A with at least one column of the source matrix B, resulting in one or more medium-granularity computing tasks. Each of the medium-granularity computing tasks is broken down into one or more sub-tasks based on the maximum processing capacity of the computing core. All the computational subtasks obtained from the same medium-granularity computation task are divided into corresponding task subsets.
3. The computational task scheduling method according to claim 1, characterized in that, The local memory is the shared memory of the computing core.
4. The computational task scheduling method according to claim 1, characterized in that, The step of scheduling all computational subtasks in the task subset to the same computational core includes: The command processor schedules the subset of tasks to a computing cluster containing the computing cores; The scheduling unit within the computing cluster schedules all computing subtasks in the task subset to the computing core.
5. The computational task scheduling method according to claim 1, characterized in that, The computational subtasks in the task subset include: A first computational subtask with a first computational granularity and a second computational subtask with a second computational granularity; The first computing subtask and the second computing subtask are scheduled together to the corresponding computing core.
6. The computational task scheduling method according to claim 5, characterized in that, The first computational granularity is a regularized matrix size that matches the maximum processing capability of the computational core; The second computational granularity is an irregular matrix size that does not match the maximum processing capacity of the computational core.
7. A computational task scheduling device, applied to a parallel processing system comprising multiple computational cores configured with local memory, characterized in that, The device includes: A splitting module is used to split large matrix multiplication and addition operations into one or more task subsets, wherein the task subsets include at least one computational subtask; The scheduling module is used to schedule each of the task subsets to the corresponding computing cores; A computing module is used to execute each computing subtask in the corresponding task subset using the computing core; Specifically, during the execution of the current computational subtask by the computational core, the first computational result of the previous computational subtask is read from the local memory configured in the computational core, and the first computational result is added to the second computational result of the current computational subtask to obtain a third computational result, which is then stored in the local memory.
8. A parallel processing system, characterized in that, include: Global memory; Multiple computing cores, each of which is equipped with local memory; as well as The computing task scheduling device as described in claim 7.
9. A computing device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the computing task scheduling method according to any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, which is executed by a processor to implement the computational task scheduling method as described in any one of claims 1-6.