Hybrid computing task scheduling method and device on TCU and CUDA core
By employing single-stream serial or multi-stream parallel scheduling in the hybrid computing task scheduling method on TCU and CUDA cores, and distributing the load according to the sparse matrix multiplication type and non-zero element threshold, the problems of decreased computing frequency and unbalanced load in hybrid computing on TCU and CUDA cores are solved, thereby improving the computational efficiency of sparse matrix multiplication and the utilization of GPU resources.
Patent Information
- Application Number
- CN202511246782.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing technologies employ hybrid computing strategies combining TCU and CUDA cores, which suffer from reduced computing frequency, resource contention, and unbalanced load, resulting in low overall computing efficiency.
By determining the theoretical execution time ratio and TCU utilization, a hybrid computing task scheduling method of single-stream serial scheduling or multi-stream parallel scheduling is adopted. Load distribution and balancing are performed based on sparse matrix multiplication type and non-zero element thresholds to optimize the task scheduling of TCU and CUDA core.
It significantly improves the overall computational throughput and GPU hardware resource utilization for sparse matrix multiplication tasks, and enhances the reliability and effectiveness of GPU operation.
Smart Images

Figure CN121070567A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a TCU and CUDA core hybrid computing task scheduling method and device. BACKGROUND
[0002] A graphics processing unit (GPU) is a kind of coprocessor for processing image and graphics operation, which is widely used in personal computers, workstations and some mobile devices. With the rapid development of artificial intelligence and scientific computing, sparse matrix multiplication (SpMM) and sample dense matrix multiplication (SDDMM) in sparse matrix multiplication gradually become the key computing core in large-scale data processing and analysis tasks, and are widely used in the fields of machine learning, scientific computing, recommendation system, etc. Especially in the graphics processing unit (GPU), these calculations usually involve two kinds of heterogeneous computing resources, namely CUDA core and tensor core (TCU).
[0003] At present, the existing GPU implementation scheme adopts a TCU and CUDA core hybrid computing strategy, and can schedule the computing tasks of the TCU and the CUDA core simultaneously by using a single kernel function according to the divided hybrid computing load, so as to improve the overall parallelism and resource utilization. However, this strategy faces two main problems. First, the actual measurement shows that the parallel execution of the TCU and the CUDA core may lead to a decrease in computing frequency (such as thermal frequency reduction or resource contention), and in some scenarios, the execution efficiency is actually lower than that of using only the TCU. Second, in the non-structured sparse scenario, the load between the TCU and the CUDA core is often difficult to be evenly distributed within the thread block, and due to the difference in computing capacity between the two types of computing units, the execution time in the thread bundle (warp) is difficult to align, which easily leads to the idle of part of the thread resources, thereby reducing the overall computing efficiency and still affecting the resource utilization of the GPU. SUMMARY
[0004] In view of this, the embodiments of the present application provide a TCU and CUDA core hybrid computing task scheduling method and device to eliminate or improve one or more defects in the prior art.
[0005] An aspect of the present application provides a TCU and CUDA core hybrid computing task scheduling method, comprising: determining a value of a preset theoretical execution time ratio according to a ratio between a CUDA core estimation execution time corresponding to a CUDA core in a graphics processing unit (GPU) and a tensor core TCU estimation execution time corresponding to a TCU in the graphics processing unit (GPU); wherein the CUDA core is used to perform matrix multiplication calculation for a CUDA core computing task in each computing task; the TCU is used to perform matrix multiplication calculation for a TCU computing task in each computing task; inputting the value of the theoretical execution time ratio into a preset TCU utilization rate prediction model to obtain a corresponding TCU utilization rate; if the TCU utilization rate is less than a preset TCU utilization rate threshold, sequentially and serially scheduling each CUDA core computing task and each TCU computing task in a single computing stream.
[0006] In some embodiments of the present application, further comprising: If the TCU utilization rate is equal to or greater than the TCU utilization rate threshold, determining a value of a preset atomic operation ratio according to a ratio between a total number of atomic operation instructions in an execution process of the CUDA core computing task and the TCU computing task and a total number of all computing operation instructions; If the value of the atomic operation ratio is greater than a preset atomic index threshold, sequentially and serially scheduling each CUDA core computing task and each TCU computing task in a single computing stream.
[0007] In some embodiments of the present application, further comprising: If the value of the atomic operation ratio is less than or equal to the atomic index threshold, decoupling each TCU computing task and each CUDA core computing task to different computing streams for parallel scheduling.
[0008] In some embodiments of the present application, the TCU utilization rate prediction model is shown in formula (1): U tcu (r)=38.9exp(-3.66r)+70 (1) Wherein, r represents the theoretical execution time ratio; U tcu (r) represents the TCU utilization rate.
[0009] In some embodiments of the present application, before the value of the theoretical execution time ratio is input into the preset TCU utilization rate prediction model, further comprising: Constructing an exponential decay model between the theoretical execution time ratio and the TCU utilization rate; wherein the exponential decay model is shown in formula (2): U tcu(r) = Aexp(-kr) + C (2) wherein A, k and C represent undetermined coefficients; The TCU actual utilization rates respectively obtained by testing at different values of the theoretical execution time ratio are taken as observation data, and each of the undetermined coefficients in the exponential decay model is determined to correspond to a value by function fitting, to obtain a corresponding TCU utilization rate prediction model.
[0010] In some embodiments of the present application, before determining the value of the preset theoretical execution time ratio between the execution time estimated according to the CUDA core corresponding to the CUDA core in the graphics processing unit GPU and the execution time estimated according to the TCU corresponding to the TCU in the graphics processing unit GPU, the method further comprises: determining the partition granularity and the non-zero element threshold corresponding to the sparse matrix multiplication to be calculated based on the type of the sparse matrix multiplication; wherein the type of the sparse matrix multiplication includes: sparse matrix and dense matrix multiplication SpMM and sample dense matrix multiplication SDDMM; performing load distribution on the sub-matrix corresponding to each window in the sparse matrix multiplication respectively according to the partition granularity and the non-zero element threshold, to obtain a calculation group corresponding to each window; wherein each calculation group includes: a CUDA core calculation block corresponding to the CUDA core, and / or a TCU block corresponding to the TCU; performing load balancing processing on each calculation group to obtain a plurality of calculation tasks corresponding to each calculation group; wherein each calculation task includes: a CUDA core calculation task corresponding to the CUDA core calculation block, and / or a TCU calculation task corresponding to the TCU block; so that the CUDA core performs matrix multiplication calculation for the CUDA core calculation task and / or the TCU performs matrix multiplication calculation for the TCU calculation task.
[0011] In some embodiments of the present application, the determination of the partition granularity and the non-zero element threshold corresponding to the sparse matrix multiplication to be calculated based on the type of the sparse matrix multiplication to be calculated comprises: if the type of the sparse matrix multiplication to be calculated is the SpMM, determining that the partition granularity corresponding to the SpMM is the non-zero element granularity, and obtaining a first non-zero element threshold corresponding to the SpMM; if the type of the sparse matrix multiplication to be calculated is the SDDMM, determining that the partition granularity corresponding to the SDDMM is the TCU block granularity, and obtaining a second non-zero element threshold corresponding to the SDDMM.
[0012] In some embodiments of the present application, the sparse matrix multiplication is divided into a plurality of windows, and each window corresponds to a sub-matrix; and the sub-matrix is divided into a plurality of TCU blocks according to the TCU block granularity, and each TCU block is divided into a plurality of CUDA core calculation blocks according to the CUDA core calculation block granularity, so as to obtain a plurality of calculation groups corresponding to each window. If the type of the sparse matrix multiplication to be calculated currently is the SpMM, whether the number of non-zero elements in each column of the sub-matrix corresponding to each window is equal to or greater than the first non-zero element threshold is judged based on the non-zero element granularity; for each sub-matrix, if there is a column in the sub-matrix in which the number of non-zero elements is equal to or greater than the first non-zero element threshold, the column in which the number of non-zero elements is equal to or greater than the first non-zero element threshold in the sub-matrix is compressed into a TCU block; if there is a column in the sub-matrix in which the number of non-zero elements is less than the first non-zero element threshold, each column in which the number of non-zero elements is less than the first non-zero element threshold in each sub-matrix is divided into a CUDA core calculation block, so as to obtain a plurality of calculation groups corresponding to each window. If the type of the sparse matrix multiplication to be calculated currently is the SDDMM, for each window, each column is sorted in descending order of the number of non-zero elements contained in each column in the sub-matrix corresponding to the window; based on the TCU block granularity, each sorted column in each window is divided into a TCU judgment unit, and whether the number of non-zero elements in each TCU judgment unit is equal to or greater than the second non-zero element threshold is judged; for each sub-matrix, if there is a TCU judgment unit in the sub-matrix in which the number of non-zero elements is equal to or greater than the second non-zero element threshold, each TCU judgment unit in which the number of non-zero elements is equal to or greater than the second non-zero element threshold in each sub-matrix is allocated as a TCU block; if there is a TCU judgment unit in the sub-matrix in which the number of non-zero elements is less than the second non-zero element threshold, each element in the TCU judgment unit in which the number of non-zero elements is less than the second non-zero element threshold in each sub-matrix is divided into a CUDA core calculation block, so as to obtain a plurality of calculation groups corresponding to each window.
[0013] In some embodiments of the present application, the load balancing processing is performed on each calculation group to obtain a plurality of calculation tasks corresponding to each calculation group. The preset splitting step is performed on each calculation group corresponding to each window to obtain a plurality of calculation tasks corresponding to each calculation group. The splitting step includes: if the calculation group contains the TCU block, whether the number of columns in the TCU block exceeds a first column number threshold is judged, and if yes, each TCU block in the calculation group is divided into a plurality of TCU block calculation tasks based on the first column number threshold. and if the CUDA core computing block is included in the computing group, determining whether the number of the CUDA core computing blocks in the computing group exceeds a second number threshold, if yes, dividing each CUDA core computing block in the computing group into a plurality of CUDA core computing tasks based on the second number threshold, and determining each CUDA core computing block in the CUDA core computing task whose number of CUDA core computing blocks exceeds a third number threshold as a long CUDA core computing block, and determining each CUDA core computing block in the CUDA core computing task whose number of CUDA core computing blocks does not exceed the third number threshold as a short CUDA core computing block; In some embodiments of the present application, correspondingly, the TCU and CUDA core hybrid computing task scheduling method further comprises: Based on a plurality of auxiliary arrays, record the split information corresponding to each computing task in each computing group of each window; The auxiliary array includes: A first array for recording the number of TCU blocks corresponding to each computing task in each window; A second array for recording the number of non-zero elements corresponding to each computing task in each window; A third array for recording the index of each computing task in the corresponding original window; A fourth array for recording the index of each computing task in the corresponding original row; A fifth array for recording whether each computing task is to be executed by an atomic operation.
[0014] In some embodiments of the present application, before determining the partition granularity and the non-zero element threshold corresponding to the sparse matrix multiplication based on the type of the current sparse matrix multiplication to be calculated, further comprising: Constructing a data access cost model corresponding to each of the sparse matrix and dense matrix multiplication SpMM and the sampled dense matrix multiplication SDDMM; wherein the access cost model is used to represent the data access cost ratio of CUDA core and TCU; determining the partition granularity of the SpMM based on the data access cost model corresponding to the SpMM, and determining the partition granularity of the SDDMM based on the data access cost model corresponding to the SDDMM.
[0015] Another aspect of the present application provides a TCU and CUDA core hybrid computing task scheduling device, comprising: a theoretical execution time ratio determination module configured to determine a preset theoretical execution time ratio value according to a ratio between an estimated execution time of a CUDA core in a graphics processing unit (GPU) and an estimated execution time of a tensor core (TCU) in a graphics processing unit (GPU); wherein the CUDA core is used to perform matrix multiplication calculation for each CUDA core computing task in each computing task; and the TCU is used to perform matrix multiplication calculation for each TCU computing task in each computing task; a TCU utilization rate prediction module configured to input the theoretical execution time ratio value into a preset TCU utilization rate prediction model to obtain a corresponding TCU utilization rate; a single-flow serial scheduling module configured to sequentially and serially schedule each CUDA core computing task and each TCU computing task in a single computing flow if the TCU utilization rate is less than a preset TCU utilization rate threshold.
[0016] A third aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the TCU and CUDA core hybrid computing task scheduling method.
[0017] A fourth aspect of the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the TCU and CUDA core hybrid computing task scheduling method.
[0018] A fifth aspect of the present application provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the TCU and CUDA core hybrid computing task scheduling method.
[0019] The TCU and CUDA core hybrid computing task scheduling method provided in the application determines the value of a preset theoretical execution time ratio according to the ratio between the estimated execution time of a CUDA core corresponding to the CUDA core in a graphics processing unit (GPU) and the estimated execution time of a tensor core (TCU) corresponding to the TCU in the graphics processing unit (GPU); wherein the CUDA core is used to perform matrix multiplication calculation on a CUDA core computing task in each computing task; the TCU is used to perform matrix multiplication calculation on a TCU computing task in each computing task; the value of the theoretical execution time ratio is input into a preset TCU utilization rate prediction model to obtain the corresponding TCU utilization rate; if the TCU utilization rate is less than a preset TCU utilization rate threshold, each CUDA core computing task and each TCU computing task is sequentially and serially scheduled in a single computing stream, the optimal execution mode of each computing task kernel function can be accurately and without additional profiling overhead, thereby significantly improving the computing throughput of the overall sparse matrix multiplication computing task and the utilization rate of the GPU hardware computing resource, and improving the operation reliability and effectiveness of the GPU.
[0020] Additional advantages, objects, and features of the application will be set forth in part in the description which follows, and in part will become apparent to those having ordinary skill in the art upon examination of the following or can be learned from practice of the application. The objects and other advantages of the application can be realized and attained by the structure particularly pointed out in the specification and claims hereof as well as the appended drawings.
[0021] It will be understood by those skilled in the art that the objects and advantages of the present application can be realized and attained by the structure particularly pointed out in the appended claims and carriable out in the following detailed description and claims. BRIEF DESCRIPTION OF DRAWINGS
[0022] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the principles of the application. The components in the drawings are not drawn to scale, but are merely intended to illustrate the principles of the application. To facilitate an understanding of some portions of the application, corresponding portions of the drawings can be exaggerated relative to other portions, namely, parts shown in the drawings can be shown disproportionately large to illustrate details thereof. In the drawings: Figure 1 The first flowchart of the TCU and CUDA core hybrid computing task scheduling method in an embodiment of the application.
[0023] Figure 2 The second flowchart of the TCU and CUDA core hybrid computing task scheduling method in an embodiment of the application.
[0024] Figure 3A schematic diagram of a TCU and CUDA core hybrid computing task scheduling process is provided for the present application.
[0025] Figure 4 A third schematic diagram of a TCU and CUDA core hybrid computing task scheduling method is provided for an embodiment of the present application.
[0026] Figure 5 An example schematic diagram of a corresponding original sparse matrix for sparse matrix multiplication in an embodiment of the present application is provided.
[0027] Figure 6 A schematic diagram of a TCU and CUDA core hybrid computing task scheduling method is provided for an application example of the present application.
[0028] Figure 7 A schematic diagram of load distribution for SpMM and SDDMM is provided for an application example of the present application.
[0029] Figure 8 A schematic diagram of load balancing for SpMM and SDDMM is provided for an application example of the present application.
[0030] Figure 9 A schematic diagram of the relationship between the theoretical execution time ratio r and TCU utilization rate in an application example of the present application. DETAILED DESCRIPTION
[0031] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be given to the present application in combination with embodiments and drawings. Herein, the exemplary embodiments of the present application and their descriptions are used to explain the present application, but not as a limitation to the present application.
[0032] It should be noted that, in order to avoid the present application being obscured by unnecessary details, only structures and / or processing steps closely related to the scheme according to the present application are shown in the drawings, and other details not closely related to the present application are omitted.
[0033] It should be emphasized that the term “comprises / comprising” is used herein to indicate the presence of a feature, element, step or component, but not to exclude the presence or addition of one or more other features, elements, steps or components.
[0034] It should be noted that, if not specifically stated, the term “connected” herein can not only mean direct connection, but also indirect connection with an intermediate.
[0035] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0036] CUDA core is the basic computing unit in GPU, which is used to perform parallel computing tasks, including floating point operations, integer operations and memory access, etc. As the basic unit of GPU for large-scale parallel computing, it is suitable for scientific computing, deep learning, graphics rendering and other fields. TCU is a new type of processing core that can perform mixed precision computing and dynamically adjust the computing power according to the reduction of precision, while maintaining accuracy and improving throughput.
[0037] The current method generally lacks a mixed computing task scheduling method for TCU and CUDA core, and lacks a clear execution mode, so even if the load is divided, it is difficult to fully release the computing potential of mixed resources.
[0038] Therefore, in order to solve the problems that the existing scheduling method of simultaneously scheduling the computing tasks of TCU and CUDA core with a single kernel function may lead to a decrease in computing frequency (such as thermal frequency reduction or resource contention), and in the non-structured sparse scene, the load between TCU and CUDA core is often difficult to be balanced within the thread block, and due to the difference in computing power between the two types of computing units, the execution time in the thread bundle (warp) is difficult to align, which is easy to cause part of the thread resources to be idle, thereby reducing the overall computing efficiency and affecting the resource utilization of GPU, the embodiments of the present application respectively provide a mixed computing task scheduling method on TCU and CUDA core, a mixed computing task scheduling device on TCU and CUDA core for executing the mixed computing task scheduling method on TCU and CUDA core, an entity device, a computer readable storage medium and a computer program product, which can predict the appropriate execution mode to realize efficient collaborative computing on TCU and CUDA core, and meet the actual needs of high-performance deep learning and scientific computing for sparse operator acceleration.
[0039] To alleviate the above single kernel function scheduling problem, the present application introduces a multi-stream parallel execution mode, which decouples the tasks of TCU and CUDA core into different computing streams and schedules them in parallel to improve resource utilization efficiency. This mode shows good scalability in tasks with high decoupling of computing paths. However, the present designers find from actual test results that the multi-stream scheduling mode also introduces new performance bottlenecks in some cases: for example, for operators such as SpMM that need to perform load partitioning by row or window, the number of atomic operation executions will increase significantly when the tasks are executed across multiple computing streams, resulting in additional synchronization overhead and reducing overall computing performance. In addition, the multi-stream parallel execution mode also has competition between computing resources, causing a decrease in computing frequency and limiting performance improvement. In contrast, the single-stream serial execution mode (scheduling TCU tasks and CUDA Core tasks in sequence in 1 computing stream) has lower parallelism in nature, but can effectively avoid the frequent atomic operation synchronization conflicts that occur in the multi-stream parallel execution mode, and significantly reduce the problem of computing frequency reduction caused by resource competition. Therefore, under certain load structure and computing conditions, the serial execution mode can achieve higher hardware resource utilization efficiency and better overall computing performance.
[0040] Therefore, based on solving the above problems of simultaneously scheduling TCU and CUDA Core computing tasks for single kernel functions, the present designers can further overcome the new technical problems caused by introducing the multi-stream parallel execution mode, and design a new hybrid computing task scheduling method that can accurately evaluate and determine which execution mode is most suitable for a certain specific hybrid computing task to achieve the best performance. The specific implementation is described in detail in the following embodiments.
[0041] Based on this, the present embodiment provides a TCU and CUDA core hybrid computing task scheduling method that can be implemented by a TCU and CUDA core hybrid computing task scheduling device, as shown in Figure 1 , which specifically includes the following content: Step 100: determining the value of a preset theoretical execution time ratio according to the ratio between the estimated execution time of the CUDA core corresponding to the CUDA core in the graphics processing unit GPU and the estimated execution time of the tensor core TCU corresponding to the TCU in the graphics processing unit GPU; wherein the CUDA core is used to perform matrix multiplication calculation for the CUDA core computing task in each computing task; and the TCU is used to perform matrix multiplication calculation for the TCU computing task in each computing task.
[0042] In the embodiments of the present application, each of the computing tasks comprises: a CUDA core computing task corresponding to the CUDA core computing block, and / or a TCU computing task corresponding to the TCU block, so that the CUDA core performs matrix multiplication calculation for the CUDA core computing task and / or the TCU performs matrix multiplication calculation for the TCU computing task. Before the mixed computing task scheduling is performed, the TCU and CUDA core mixed computing task scheduling device can first determine a partition granularity and a non-zero element threshold corresponding to the sparse matrix multiplication to be currently calculated based on a type of the sparse matrix multiplication; wherein the type of the sparse matrix multiplication comprises: a sparse matrix and dense matrix multiplication SpMM and a sampled dense matrix multiplication SDDMM; each sub-matrix corresponding to each window in the sparse matrix multiplication is respectively allocated a load according to the partition granularity and the non-zero element threshold, so as to obtain a computing group corresponding to each window; each of the computing groups is respectively subjected to load balancing processing, so as to obtain a plurality of computing tasks corresponding to each of the computing groups.
[0043] It should be noted that the sparse matrix refers to a matrix in which most elements are 0 (or a default value).
[0044] The sparse matrix and dense matrix multiplication SpMM (Sparse Matrix-Matrix Multiplication) can also be referred to as sparse matrix-dense matrix multiplication or sparse matrix and dense matrix multiplication operation, etc., wherein the mathematical definition of SpMM is C=AxB; A is a sparse matrix (most elements are 0); C and B both refer to dense matrices (all elements are non-zero), and the calculation core is only to calculate the product of the non-zero elements (Number of Non-Zero elements) in the sparse matrix A and the corresponding rows and columns of B, avoiding zero element operation.
[0045] The mathematical definition of the sampled dense matrix multiplication SDDMM (Dense-Dense Matrix Multiplication) is C=S (AxB), A and B are dense matrices; S is a sparse sampling matrix (mask matrix, which specifies the element position to be calculated); represents element-wise multiplication (Hadamard product), and the calculation core is only to calculate the non-zero position value of S in (AxB), and the other positions are 0.
[0046] In step 100, the present application designs an index, i.e. a theoretical execution time ratio, which can also be referred to as a theoretical execution time ratio index, which can be denoted as time_ratio, or simply r.
[0047] Among them, the total number of floating point operations allocated to the CUDA core L cudaWith the theoretical peak throughput F of CUDA cores cuda The ratio between (in TFLOPS) is used to calculate the estimated execution time T of the CUDA core. CUDA And based on the total number L of floating-point operations allocated to the TCU tcu Compared to the theoretical peak throughput F of TCU tcu The estimated execution time T of the TCU is calculated by the ratio between (in TFLOPS). TCU Then T CUDA With T TCU The ratio between them is used as the theoretical execution time ratio r.
[0048] Specifically, the formula for calculating the theoretical execution time ratio r (time_ratio) is as follows: time_ratio = T CUDA / T TCU =(L cuda / F cuda ) / (L tcu / F tcu ) Step 200: Input the value of the theoretical execution time ratio into the preset TCU utilization prediction model to obtain the corresponding TCU utilization.
[0049] It is understood that the TCU utilization prediction model refers to a model used for predicting TCU utilization.
[0050] Step 300: If the TCU utilization rate is less than the preset TCU utilization rate threshold, then each CUDA core computing task and each TCU computing task are sequentially scheduled in a single computing stream.
[0051] In step 300, it is determined whether the TCU utilization rate is less than a preset TCU utilization threshold. If so, it indicates that the TCU utilization rate is significantly low, and a serial execution mode should be directly adopted. That is, each CUDA core computing task and each TCU computing task are sequentially scheduled in a single computing stream. In other words, when the TCU utilization rate is significantly low, the optimal execution mode for each CUDA core computing task and each TCU computing task is single-stream serial. By accurately predicting and selecting this optimal execution mode without additional analysis overhead, the overall task's computing throughput and GPU hardware resource utilization efficiency can be significantly improved.
[0052] In a preferred embodiment, to achieve effective task scheduling, this application establishes a TCU utilization threshold U based on extensive experiments (including 500 test matrices). thr= 78%, to further improve the accuracy and reliability of determining the optimal execution mode when the TCU utilization is obviously low.
[0053] From the above description, the TCU and CUDA core hybrid computing task scheduling method provided by the embodiments of the application can accurately and without additional profiling overhead predict and select the optimal execution mode of each computing task kernel function, thereby significantly improving the overall sparse matrix multiplication computing task computing throughput and GPU hardware computing resource utilization, and improving the operation reliability and effectiveness of the GPU.
[0054] To further improve the effectiveness and accuracy of predicting and selecting the optimal execution mode of each computing task kernel function, in a TCU and CUDA core hybrid computing task scheduling method provided by an embodiment of the application, referring to Figure 2 and Figure 3 , the step 200 in the TCU and CUDA core hybrid computing task scheduling method specifically contains the following content: step 400: if the TCU utilization is equal to or greater than the TCU utilization threshold, determining the value of a preset atomic operation ratio according to the ratio between the total number of atomic operation instructions and the total number of all computing operation instructions in the execution process of the CUDA core computing task and the TCU computing task.
[0055] Specifically, the atomic operation ratio atomic_ratio reflects the degree of synchronization overhead caused by atomic operations in the task, and its calculation formula is: atomic_ratio = N_atomic / N_total Where N_atomic represents the total number of atomic operation instructions in the task execution process, and N_total represents the total number of all computing operation instructions. The lower the atomic_ratio index, the less synchronization overhead caused by atomic operations in the task, and at this time, the multi-stream parallel execution mode is more suitable for improving the overall computing performance.
[0056] Step 500: if the value of the atomic operation ratio is greater than a preset atomic index threshold, sequentially serially scheduling each CUDA core computing task and each TCU computing task in a single computing stream.
[0057] In step 500, if the value of the atomic operation ratio atomic_ratio is greater than a preset atomic index threshold A thr, it indicates that the computing task has a high synchronization overhead, and a serial execution mode should be adopted, i.e., each of the CUDA core computing tasks and each of the TCU computing tasks are sequentially and serially scheduled in a single computing stream. That is, when the computing task has a high synchronization overhead, the optimal execution mode of each of the CUDA core computing tasks and each of the TCU computing tasks is single-stream serial, and then the optimal execution mode is predicted and selected accurately without additional profiling overhead, so as to further significantly improve the computing throughput of the overall task and the utilization efficiency of GPU hardware resources.
[0058] In a preferred manner, in order to realize effective task scheduling, the present application establishes an atomic index threshold A thr = 0.6 according to a large number of experiments (including 500 test matrices), so as to further improve the accuracy and reliability of determining the optimal execution mode of the computing task with a high synchronization overhead.
[0059] On this basis, in order to further improve the effectiveness and accuracy of predicting and selecting the optimal execution mode of each computing task kernel function, in a TCU and CUDA core hybrid computing task scheduling method provided in an embodiment of the present application, referring to Figure 2 and Figure 3 , the step 300 in the TCU and CUDA core hybrid computing task scheduling method specifically comprises the following content: Step 600: If the value of the atomic operation ratio is less than or equal to the atomic index threshold, each of the TCU computing tasks and each of the CUDA core computing tasks are decoupled and scheduled in parallel in different computing streams.
[0060] In step 600, if the value of the atomic operation ratio atomic_ratio is less than or equal to the preset atomic index threshold A thr , it indicates that the computing task has a low synchronization overhead, and a multi-stream parallel execution mode should be adopted, i.e., each of the TCU computing tasks and each of the CUDA core computing tasks are decoupled and scheduled in parallel in different computing streams. That is, when the computing task has a low synchronization overhead, the optimal execution mode of each of the CUDA core computing tasks and each of the TCU computing tasks is multi-stream parallel, and then the optimal execution mode is predicted and selected accurately without additional profiling overhead, so as to further significantly improve the computing throughput of the overall task and the utilization efficiency of GPU hardware resources.
[0061] In order to further improve the accuracy and application reliability of TCU utilization rate prediction, in a TCU and CUDA core hybrid computing task scheduling method provided in an embodiment of the present application, the TCU utilization rate prediction model is shown in formula (1): Utcu (r) = 38.9exp(-3.66r) + 70 (1) wherein, r represents the theoretical execution time ratio; tcu (r) represents the TCU utilization rate.
[0062] On this basis, in order to further improve the application effectiveness and reliability of the above TCU utilization rate prediction model, in a TCU and CUDA core hybrid computing task scheduling method provided by an embodiment of the present application, referring to Figure 2 , the TCU and CUDA core hybrid computing task scheduling method can further include the following content before step 200 or step 100: Step 001: Construct an exponential decay model between the theoretical execution time ratio and the TCU utilization rate; wherein the exponential decay model is as shown in formula (2): tcu (r) = A exp(-kr) + C (2) wherein, A, k and C all represent undetermined coefficients; the undetermined coefficient A represents the additional increase amplitude of the TCU utilization rate compared with the saturation value in the initial state (very small CUDA core load), the undetermined coefficient k describes the speed of the TCU utilization rate decreasing with the increase of the CUDA core load ratio, and the undetermined coefficient C represents the saturation value (i.e. the stable lower limit utilization rate) of the TCU utilization rate when the CUDA core load ratio is too high.
[0063] Step 002: Take the TCU actual utilization rate respectively tested under different values of the theoretical execution time ratio as observation data, determine the respective values of each of the undetermined coefficients in the exponential decay model through function fitting, to obtain the corresponding TCU utilization rate prediction model.
[0064] Specifically, taking the NVIDIA RTX4090 GPU as an example, the TCU actual utilization rate under different theoretical execution time ratios r is tested by performing experiments on the RTX4090 GPU. The experimental data shows that when r gradually increases from 0.1 to 1, the utilization rate of the TCU decreases rapidly; and when r further increases to 6, the utilization rate almost tends to be stable. The present application takes the TCU actual utilization rate tested under different theoretical execution time ratios r as observation data, determines the undetermined coefficients (including A, k and C) of the model through function fitting, and thus obtains an accurate TCU utilization rate prediction model.
[0065] That is to say, the embodiments of the present application can accurately and without additional profiling overhead predict and select the optimal execution mode of each computing task by comprehensively considering the above two heuristic indicators and exponential decay functions, thereby significantly improving the overall task computing throughput and hardware resource utilization efficiency.
[0066] On this basis, the existing sparse matrix multiplication calculation methods mainly include the following implementation schemes: (1) Based on CUDA core computing sparse matrix multiplication, while CUDA core has strong programming flexibility and can effectively handle irregular access of sparse data, but the calculation peak performance is significantly lower than TCU, making it difficult to efficiently handle high-density computing tasks.
[0067] (2) Using TCU to accelerate sparse matrix calculation, but the processing effect on highly sparse data is limited, and there is a large amount of calculation redundancy and data management overhead, that is, TCU is suitable for regular and high-density matrix multiplication operations and can provide extremely high calculation peak performance, but when processing highly sparse and irregular data, a large amount of redundant calculation is easily generated, the resource utilization rate is reduced, and the ideal performance improvement cannot be achieved.
[0068] (3) Although some schemes propose a coarse-grained TCU and CUDA core hybrid computing strategy, there is a lack of explicit sparse load partitioning mechanism for TCU and CUDA core, and there are problems such as coarse partitioning granularity, load imbalance, and low data processing efficiency, which cannot effectively solve the problems of data redundancy and calculation efficiency, limiting the overall computing performance.
[0069] That is to say, the existing sparse matrix multiplication calculation method lacks an efficient cooperative mechanism for TCU and CUDA core on the GPU platform, especially how to dynamically and finely partition computing tasks and optimize resource allocation in SpMM and SDDMM tasks. Specifically, the technical problems existing in the existing sparse matrix multiplication calculation method are as follows: (1) The current sparse matrix multiplication (SpMM and SDDMM) acceleration method on GPU mainly uses a single type of computing resource, i.e., using CUDA core or TCU for calculation. But these two methods each have obvious defects, for example, when TCU processes unstructured sparse data, since the operands must be forced to align and fill a large number of zero values, significant calculation redundancy is generated, greatly reducing the computing performance; while CUDA core can flexibly adapt to data of different sparsity, but its computing power is limited, unable to fully tap the computing potential of GPU hardware. Thus, the application efficiency of sparse operators in high-performance computing and graph neural networks and other practical scenarios is limited.
[0070] (2) The prior art lacks a load partitioning strategy that can effectively guide fine-grained mixed scheduling of tasks between CUDA cores and TCUs. Existing sparse load partitioning methods have the problem of coarse granularity. They usually only partition tasks according to the average sparsity of the entire sparse matrix or the entire TCU block, ignoring the local non-uniformity of the sparse pattern within the matrix or block. This coarse-grained partitioning cannot accurately match the performance advantages of CUDA cores and TCUs according to task characteristics, limiting the full release of GPU hardware computing resources.
[0071] (3) Manual parameter tuning is costly and lacks universality: A few studies attempt to mix CUDA cores and TCUs, such as PCGCN and SparseTIR, but they require complex graph partitioning tools (such as METIS) or manual parameter tuning, making it difficult to automatically adapt to input data with different sparsity and sparsity distribution. The optimization process is complex and the effect is unstable.
[0072] Therefore, on the basis of being able to accurately and without additional profiling overheads to predict and select the optimal execution mode for each computing task, it is also urgent to propose an efficient TCU and CUDA core mixed load partitioning strategy to fully exploit the advantages of both resources and improve the overall efficiency and performance of sparse matrix related computations on GPU platforms, to further improve the utilization of GPU hardware computing resources.
[0073] Based on this, in the TCU and CUDA core hybrid computing task scheduling method provided by the embodiments of the present application, referring to Figure 4 , the TCU and CUDA core hybrid computing task scheduling method further specifically comprises the following content before step 100: Step 010: determining the partition granularity and non-zero element threshold corresponding to the sparse matrix multiplication based on the type of the current sparse matrix multiplication to be calculated; wherein the type of the sparse matrix multiplication includes sparse matrix and dense matrix multiplication SpMM and sample dense matrix multiplication SDDMM.
[0074] It can be understood that the TCU and CUDA core hybrid computing task scheduling device can be set in the GPU as a sparse matrix multiplication accelerator for receiving the current sparse matrix multiplication calculation instruction or request, and obtaining the original sparse matrix corresponding to the sparse matrix multiplication from the instruction or request, and lacking the type to which it belongs SpMM or SDDMM, and then extracting the partition granularity and non-zero element threshold corresponding to the sparse matrix multiplication from the corresponding relationship data between the pre-stored types, partition granularities and non-zero element thresholds of the sparse matrix multiplication in the local.
[0075] The division granularity is used to represent a unit division granularity of load distribution on the original sparse matrix corresponding to the sparse matrix multiplication. The non-zero element threshold is used to represent a threshold of comparison between the number of non-zero elements in the unit division granularity.
[0076] The division granularity and the non-zero element threshold can be transmitted to the TCU and CUDA core hybrid computing task scheduling device for storage after being artificially specified. The non-zero element threshold can be set according to different GPU models, performance, and response speed, and can be determined after a limited number of experiments according to actual application scenarios.
[0077] In order to further improve the effectiveness and reliability of the load distribution of the subsequent step 200 for the sparse matrix multiplication, the division granularity can also be generated in advance according to the data access cost ratio of the CUDA core and the TCU, which will be described in detail in subsequent embodiments.
[0078] Step 020: According to the division granularity and the non-zero element threshold, each sub-matrix corresponding to each window in the sparse matrix multiplication is respectively load distributed to obtain a calculation group corresponding to each window.
[0079] In one or more embodiments of the present application, the size of the window can be obtained by artificially specifying according to actual application requirements. The size of the window is used to represent the number of rows of the original sparse matrix corresponding to the sparse matrix multiplication. For example, if the original sparse matrix corresponding to the sparse matrix multiplication has 9 rows of elements, and the size of the window is 3, the matrix can be divided into three sub-matrices corresponding to each window. The first window sub-matrix is the first three rows of the original sparse matrix, the second window sub-matrix is the middle three rows of the original sparse matrix, and the third window sub-matrix is the last three rows of the original sparse matrix. If the number of rows of the original sparse matrix corresponding to the sparse matrix multiplication cannot be divided by the size of the window, the number of rows of the sub-matrix of the last window is less than the size of the window.
[0080] In step 020, according to the division granularity and the non-zero element threshold, the load distribution of each sub-matrix corresponding to each window in the sparse matrix multiplication refers to: judging the sub-matrix corresponding to each window in the sparse matrix multiplication according to the division granularity in turn, and then comparing whether the non-zero elements in each judging unit exceed the non-zero element threshold. If yes, the judging unit is distributed to the TCU block corresponding to the TCU in the GPU; if no, the judging unit is distributed to the CUDA core computing block corresponding to the CUDA core in the GPU.
[0081] Therefore, in one window, the load distribution result (TCU block or CUDA core computing block) corresponding to each judging unit in the unique corresponding sub-matrix of the window constitutes the calculation group corresponding to the window. Since the load distribution results of each judging unit in the same calculation group may be consistent or inconsistent, the calculation group may only contain TCU blocks, may only contain CUDA core computing blocks, or may contain both TCU blocks and CUDA core computing blocks.
[0082] After completing the work load distribution, the next problem to be considered is how to uniformly map the distributed tasks to the thread blocks of the parallel system. For the distributed work load, some windows may contain too many TCU blocks or too long CUDA core computing blocks, so the windows need to be split by the following step 300 to ensure the balance of the load.
[0083] In an example, referring to Figure 5 , for the original coefficient matrix containing 8 rows and 12 columns of elements, if the window size is 4, the original coefficient matrix is divided into 2 windows, i.e. the first window and the second window each corresponding to a sub-matrix. If the original coefficient matrix is the original coefficient matrix corresponding to the SpMM, each of the sub-matrices of each window contains one judging unit of non-zero original columns, and the columns processed by the TCU can be compressed into a matrix composed of multiple groups of non-zero original columns, i.e. TCU blocks. For example, in the sub-matrix corresponding to the first window W0, the 3rd, 4th, 6th and 7th columns can form a 4x4 matrix, i.e. a TCU block; the 0th, 1st, 2nd, 5th, 8th, 9th and 10th columns are divided into CUDA core computing blocks; and the 11th column does not enter the judging unit to be distributed because it does not contain non-zero elements.
[0084] Step 030: load balancing is performed on each of the calculation groups respectively, to obtain a plurality of calculation tasks corresponding to each of the calculation groups respectively; wherein each of the calculation tasks comprises: a CUDA core calculation task corresponding to the CUDA core calculation block, and / or a TCU calculation task corresponding to the TCU block; so that the CUDA core performs matrix multiplication calculation on the CUDA core calculation task and / or the TCU performs matrix multiplication calculation on the TCU calculation task.
[0085] In step 030, if each CUDA core calculation block in the calculation group obtained in step 020 needs to be data split, each CUDA core calculation block in the calculation group is divided into a plurality of CUDA core calculation tasks respectively, and each CUDA core calculation task comprises at least one CUDA core calculation block; if each CUDA core calculation block in the calculation group obtained in step 020 does not need to be split, each CUDA core calculation block in the calculation group belongs to the same CUDA core calculation task.
[0086] In step 030, if each TCU block in the calculation group obtained in step 020 needs to be data split, each TCU block in the calculation group is divided into a plurality of TCU calculation tasks respectively, and each TCU calculation task comprises at least one TCU block; if each TCU block in the calculation group obtained in step 020 does not need to be split, each TCU block in the calculation group belongs to the same TCU calculation task.
[0087] It should be noted that the CUDA core performs matrix multiplication calculation on the CUDA core calculation task and / or the TCU performs matrix multiplication calculation on the TCU calculation task specifically refers to: (1) If the plurality of calculation tasks corresponding to the sparse matrix multiplication only comprise CUDA core calculation tasks, each of the CUDA core calculation tasks is mapped to a corresponding thread block, so that the CUDA core performs matrix multiplication calculation on the CUDA core calculation task.
[0088] (2) If the plurality of calculation tasks corresponding to the sparse matrix multiplication only comprise TCU calculation tasks, each of the TCU calculation tasks is mapped to a corresponding thread block, so that the TCU performs matrix multiplication calculation on the TCU calculation task.
[0089] (3) If the sparse matrix multiplication corresponds to a plurality of calculation tasks including CUDA core calculation tasks and TCU calculation tasks, each of the CUDA core calculation tasks and each of the TCU calculation tasks are mapped to a corresponding thread block, so that the CUDA core performs matrix multiplication calculation for the CUDA core calculation task and the TCU performs matrix multiplication calculation for the TCU calculation task.
[0090] As can be seen from the above description, the TCU and CUDA core hybrid computing task scheduling method provided by the embodiments of the application can also achieve efficient collaborative allocation of sparse computing load between the TCU and the CUDA core, realize fine division of sparse load between the CUDA core and the TCU under the GPU architecture, give full play to the advantages of heterogeneous computing resources, thereby significantly improving the overall performance of sparse operators such as SpMM and SDDMM, making full use of the performance of GPU hardware, reducing calculation redundancy and manual parameter adjustment cost, meeting the actual needs of high-performance deep learning and scientific computing for sparse operator acceleration, and thus effectively improving the efficiency of sparse matrix multiplication calculation and the utilization rate of GPU hardware computing resources, improving the operation reliability and effectiveness of the GPU.
[0091] In order to further improve the effectiveness and calculation efficiency of SpMM hybrid heterogeneous computing, in the TCU and CUDA core hybrid computing task scheduling method provided by the embodiments of the application, step 010 of the TCU and CUDA core hybrid computing task scheduling method specifically includes the following content: Step 011: If the type of the sparse matrix multiplication to be calculated is the SpMM, it is determined that the division granularity corresponding to the SpMM is the non-zero element granularity, and a first non-zero element threshold corresponding to the SpMM is obtained.
[0092] In order to further improve the effectiveness and reliability of load allocation for SpMM, in the TCU and CUDA core hybrid computing task scheduling method provided by the embodiments of the application, step 020 of the TCU and CUDA core hybrid computing task scheduling method specifically includes the following content: Step 021: Based on the non-zero element granularity, it is respectively judged whether the number of non-zero elements in each column in the sub-matrix corresponding to each window is equal to or greater than the first non-zero element threshold.
[0093] Step 022: for each of the sub-matrices, if there is a column in the sub-matrix in which the number of non-zero elements is equal to or greater than the first non-zero element threshold value, the column in the sub-matrix equal to or greater than the first non-zero element threshold value is compressed into a TCU block; if there is a column in the sub-matrix in which the number of non-zero elements is less than the first non-zero element threshold value, each of the columns in the sub-matrix less than the first non-zero element threshold value is divided into a CUDA core calculation block, to obtain a calculation group corresponding to each window respectively.
[0094] In step 022, the column in which there is a non-zero element is judged by the judging unit.
[0095] In order to further improve the effectiveness and calculation efficiency of the SDDMM hybrid heterogeneous computing, in the TCU and CUDA core hybrid computing task scheduling method provided in the embodiment of the application, step 010 in the TCU and CUDA core hybrid computing task scheduling method specifically contains the following contents: Step 012: if the type of the sparse matrix multiplication to be calculated at present is the SDDMM, it is determined that the division granularity corresponding to the SDDMM is the TCU block granularity, and a second non-zero element threshold value corresponding to the SDDMM is obtained.
[0096] In order to further improve the effectiveness and reliability of the load allocation to the SDDMM, in the TCU and CUDA core hybrid computing task scheduling method provided in the embodiment of the application, step 020 in the TCU and CUDA core hybrid computing task scheduling method specifically contains the following contents: Step 023: for each of the windows, each column in the sub-matrix corresponding to the window is sorted in descending order of the number of non-zero elements contained in each column.
[0097] Step 024: based on the TCU block granularity, each of the sorted columns in each of the windows is divided into a TCU judging unit, and it is judged whether the number of non-zero elements in each of the TCU judging units is equal to or greater than the second non-zero element threshold value.
[0098] For example, the sorted columns in each of the windows are combined according to the size of the preset TCU block to form a TCU block as a judging unit, i.e., a TCU judging unit.
[0099] Step 025: for each of the sub-matrices, if the number of the non-zero elements in the sub-matrix is equal to or greater than the second non-zero element threshold, the TCU judging unit is allocated as a TCU block; if the number of the non-zero elements in the sub-matrix is less than the second non-zero element threshold, each element in the TCU judging unit less than the second non-zero element threshold in each of the sub-matrices is divided into a CUDA core computing block, to obtain a respective computing group corresponding to each window.
[0100] To further improve the effectiveness and reliability of load balancing for sparse matrix multiplication, in a TCU and CUDA core hybrid computing task scheduling method provided in an embodiment of the present application, step 030 in the TCU and CUDA core hybrid computing task scheduling method specifically contains the following content: Step 031: a preset splitting step is performed on each of the computing groups corresponding to each of the windows, to obtain a plurality of computing tasks corresponding to each of the computing groups; wherein the splitting step includes: if the computing group contains the TCU block, it is determined whether the number of columns in the TCU block exceeds a first column number threshold, if yes, each TCU block in the computing group is divided into a plurality of TCU block computing tasks based on the first column number threshold; and if the computing group contains the CUDA core computing block, it is determined whether the number of the CUDA core computing blocks in the computing group exceeds a second number threshold, if yes, each CUDA core computing block in the computing group is divided into a plurality of CUDA core computing tasks based on the second number threshold, and each CUDA core computing block in the CUDA core computing task whose number of CUDA core computing blocks exceeds a third number threshold is determined as a long CUDA core computing block, and each CUDA core computing block in the CUDA core computing task whose number of CUDA core computing blocks does not exceed the third number threshold is determined as a short CUDA core computing block.
[0101] It can be understood that in step 031, if the number of columns in the TCU block does not exceed the first column number threshold, the TCU block is not split and is directly determined as a TCU computing task; if the number of columns in the CUDA core computing block does not exceed the second column number threshold, the CUDA core computing block is not split and is directly determined as a CUDA core computing task.
[0102] In an example, the first column number threshold T s = 4; and the second column number threshold C sShort_len = 5; the third column number threshold Short len = 2.
[0103] In order to further improve the application effectiveness and reliability of long CUDA core computing blocks, short CUDA core computing blocks and TCU computing tasks, and improve the efficiency and convenience of data query in the load distribution and balancing process, in the TCU and CUDA core hybrid computing task scheduling method provided in the embodiment of the application, step 030 in the TCU and CUDA core hybrid computing task scheduling method further specifically contains the following content: Step 032: based on a plurality of auxiliary arrays, record the split information corresponding to each of the computing tasks in each of the computing groups split by each window; wherein the auxiliary array includes: a first array for recording the number of TCU blocks corresponding to each of the computing tasks in each of the windows; a second array for recording the number of non-zero elements corresponding to each of the computing tasks in each of the windows; a third array for recording the index of each of the computing tasks in the corresponding original window; a fourth array for recording the index of each of the computing tasks in the corresponding original row; and a fifth array for recording whether each of the computing tasks is to be executed by an atomic operation.
[0104] Step 033: map the TCU computing tasks, long CUDA core computing blocks and short CUDA core computing blocks to different computing resources through three CUDA streams.
[0105] In an example, the first array can be denoted as WindowOffset; the second array can be denoted as RowOffset; the third array can be denoted as CurWindow; the fourth array can be denoted as CurRow; and the fifth array can be denoted as Atomic.
[0106] In order to further improve the application effectiveness and reliability of the division granularity corresponding to sparse matrix multiplication, in the TCU and CUDA core hybrid computing task scheduling method provided in the embodiment of the application, step 010 in the TCU and CUDA core hybrid computing task scheduling method further specifically contains the following content before step 010: Step 002: construct a data access cost model corresponding to each of sparse matrix and dense matrix multiplication SpMM and sample dense matrix multiplication SDDMM; wherein the access cost model is used to represent the data access cost ratio of CUDA core and TCU.
[0107] Step 003: determine the division granularity of the SpMM based on the data access cost model corresponding to the SpMM, and determine the division granularity of the SDDMM based on the data access cost model corresponding to the SDDMM.
[0108] To further illustrate the above embodiments, the application also provides a specific application example of a TCU and CUDA core hybrid computing task scheduling method, and proposes an efficient hybrid load partitioning strategy. The strategy considers the non-zero element density of sub-blocks in the sparse matrix and the data reusability characteristics on different computing resources when partitioning tasks, accurately allocates sparse computing tasks to the most suitable computing unit (CUDA core or TCU), and realizes the optimal balance of performance and hardware adaptability. The application proposes a hybrid computing task scheduling method based on double indicators (time_ratio and atomic_ratio) and an exponential decay model, which predicts the suitable execution mode (multi-stream parallel or serial execution) to realize efficient collaborative computing on TCU and CUDA core. The core goal of the application is to realize fine partitioning of sparse load between CUDA core and TCU under GPU architecture, fully utilize the advantages of heterogeneous computing resources, thereby significantly improving the overall performance of sparse operators such as SpMM and SDDMM, reducing computing redundancy and manual parameter tuning cost, and meeting the actual needs of high-performance deep learning and scientific computing for sparse operator acceleration.
[0109] The application example proposes a novel hybrid load partitioning strategy for efficient guidance of collaborative allocation of sparse computing load between TCU and CUDA core. In sparse operations such as SpMM and SDDMM, the performance bottleneck is usually caused by data access to dense matrices, and there are significant differences between TCU and CUDA core in terms of data reuse mode and theoretical computing performance. Therefore, in order to fully utilize the hardware potential and efficiently process sparse computing tasks, the allocation strategy of the application example considers two key dimensions: data reusability and actual performance of specific sparse tasks on different hardware resources.
[0110] Specifically, referring to Figure 6The application example of the application firstly determines the load division granularity according to the sparse task type (SpMM or SDDMM), and then judges whether the proportion of non-zero elements in the task or TCU block exceeds a preset threshold to determine the specific sparse load division. The selection of the threshold is determined according to the peak performance difference of the TCU and the CUDA core hardware resources. When the proportion of non-zero elements is higher than the threshold, it indicates that the sparse task is relatively dense, and is more suitable for processing by the TCU with higher computing performance; on the contrary, if the proportion of non-zero elements is lower than the threshold, it indicates that the task is relatively sparse, and is more suitable for processing by the CUDA core with better sparse flexibility. After the preliminary task division is completed, the application example of the application further adjusts through a hybrid computing task scheduling method based on double indicators and an exponential decay model. This method can accurately predict the performance of different computing paths, dynamically balance the atomic operation overhead and the computing resource utilization, and realize the fine matching of the task granularity and the hardware architecture. In addition, by reasonably setting the resource decomposition threshold, resource competition is avoided to the maximum extent, and the overall computing efficiency is improved. Overall, this fine load distribution strategy effectively reduces the computing redundancy caused by forced data alignment and zero value padding in unstructured sparse computing, and the efficient hybrid computing task scheduling method fully utilizes the advantages of heterogeneous computing resources, and significantly improves the overall computing performance and efficiency.
[0111] The hybrid computing task scheduling method on the TCU and the CUDA core provided by the application example of the application specifically includes the following contents: (I) Determining the division granularity In terms of data multiplicity, compared with the CUDA core, the TCU has unique architectural characteristics, that is, after being loaded into the register, the operands A1 (i.e., the sparse TCU block) and B1 (i.e., the dense TCU block) can be multiplexed multiple times in a single matrix multiplication and accumulation (MMA) instruction. To simplify the analysis, the application example of the application defines the data access cost as the overhead of loading data from the storage hierarchy, without distinguishing the data sources (such as global memory or cache).
[0112] 1) For SpMM, the main data access cost comes from loading the dense TCU block. The data access cost ratio R of the corresponding CUDA core and the TCU for SpMM Spmm can be expressed as: In formula (3), NNZ1 represents the number of non-zero elements in the sparse TCU block A1; m, n, and k are the dimensions of the MMA operand on the TCU.
[0113] For the CUDA core, each non-zero element is processed individually, so the data access cost is NNZ x n.
[0114] For TCU, each row of dense TCU block B1 is loaded into the register only once and then reused by multiple non-zeros in the same non-zero element.
[0115] Therefore, when NNZ1>k, TCU can reduce the data access cost by times.
[0116] Furthermore, the application examples use density p to represent the density of sparse TCU block A1, and substituting NNZ1=mkp into the above formula (1) further simplifies to m p, i.e., the average number of non-zeros of all non-zeros in the sparse TCU block. Intuitively, a higher density vector can obtain more benefits from data reuse when processing on TCU.
[0117] Therefore, for SpMM, the application examples allocate the sparse workload to TCU and CUDA core in non-zero column vector granularity (m x 1).
[0118] 2) For SDDMM, both input TCU blocks A2 and B2 are dense matrices. The data access cost ratio R sddmm of the corresponding CUDA core to TCU of SDDMM can be represented as: NNZ2 in formula (4) represents the number of non-zeros in the sparse TCU block C.
[0119] When using CUDA core, each non-zero element in the sparse TCU block C needs to access a row of TCU block A and a column of TCU block B respectively, so the data access cost is 2 x NNZ2 x k.
[0120] On TCU, TCU blocks A and B are loaded only once and can be reused by the MMA instruction.
[0121] Therefore, when , TCU can reduce the data access cost to times. Unlike SpMM, the formula of SDDMM cannot be further simplified, and intuitively, the more non-zeros in the sparse TCU block C, the more data reuse advantages TCU can obtain.
[0122] Therefore, for SDDMM, the application examples allocate the sparse workload to TCU and CUDA core in TCU block granularity (m x n).
[0123] (II) Load allocation The above analysis determines the load allocation granularity of different operations. In practice, it is easy to meet R Spmm >1 and R sddmmThe condition of >1 means that the vector or block with low data access cost should be processed on TCU first. However, considering only the data access cost is not enough. Although the theoretical peak performance of TCU is much higher than that of CUDA core, the lower R Spmm or R sddmm will cause a large amount of computational redundancy on TCU, because TCU can process a large number of unnecessary zero elements, resulting in a decrease in actual performance. Therefore, to ensure the actual performance of TCU on the basis of lower data access overhead, a higher number of non-zero elements is necessary. However, the actual performance cannot be determined in advance, so the present application uses a threshold adjuster to guide load distribution. Based on the previously determined distribution granularity, when the column vector NNZ of SpMM or the TCU block NNZ of SDDMM exceeds the threshold, the vector or block is allocated to TCU, otherwise it is processed by the CUDA core. Since the actual performance of TCU can be estimated by the theoretical peak performance and p, the optimal threshold is highly related to the hardware architecture, and less related to the specific matrix.
[0124] Referring to Figure 7 , in the sparse load distribution examples for SpMM and SDDMM, for SpMM, the present example sets the column vector threshold to 2, and counts the NNZ of non-zero column vectors in each window. If NNZ≥2, the vector is allocated to TCU, otherwise it is allocated to the CUDA core. The vectors allocated to TCU are usually compressed into TCU blocks (4x4), and the remaining vectors processed by the CUDA core are used to fill zero vectors to complete the TCU block. For SDDMM, the present example sets the TCU block threshold to 4, and the non-zero elements in each window are arranged in descending order of NNZ. The densest vector is compressed into the TCU block first. If the TCU block NNZ≥4, the block is allocated to TCU, otherwise it is allocated to the CUDA core. In fact, the present example uses mma.m16n8k4(TF32), mma.m16n8k8(FP16) instructions to process SpMM, and mma.m16n8k8(TF32), mma.m16n8k16(FP16) instructions to process SDDMM. Combined with the exchange transpose strategy, the present example uses 8x1 vector granularity for SpMM and 8x16 TCU block granularity for SDDMM. Overall, the present example theoretically analyzes the core problem of load distribution on heterogeneous computing resources, and provides clear guidance for accurate hybrid computing load distribution.
[0125] (III) Load balancing After the workload allocation is done, the next question to consider is how to evenly map these allocated tasks to the thread blocks of the parallel system. For the allocated workload, some windows can contain too many TCU blocks or long CUDA core computation blocks, so the windows need to be split to ensure the balance of the load. However, in the hybrid computation of SpMM, if the workload in a window is split, then each split part needs an atomic operation to accumulate. Therefore, the primary goal of the application instance design is to reduce the atomic operation overhead as much as possible due to window splitting. This design is guided by two main principles: the application instance observes that if the number of TCU blocks or CUDA core computation blocks in a window is small, splitting it can not bring enough performance gain, so as to make up for the additional overhead of atomic operations. Therefore, the application instance sets a clear standard to decide whether the window needs to be split.
[0126] The example strictly limits the splitting within a single window to avoid cross-window splitting to further reduce the overhead of atomic operations. Figure 8 Examples of window splitting are shown, including TCU blocks (such as 2x2) and CUDA core computation blocks, and specified splitting conditions (such as T s = 4 and C s = 5). For the CUDA core computation blocks, the example adopts the long-short computation block division method, and sets Short len = 2 as the standard.
[0127] The example summarizes three window splitting cases: (1) For window W0, since both TCU blocks and long CUDA core computation blocks need to be split, atomic operations are needed for all segments in the window; (2) For window W1, since the CUDA core computation blocks need to be split, the TCU blocks in the same window also need to be atomic operated. Similarly, the split CUDA core computation blocks also need to perform atomic operations; (3) For window W2 and window W3, these windows only contain a single type of workload and do not meet the splitting condition, so no atomic operation is needed.
[0128] In addition, the application example also introduces multiple auxiliary arrays to record the split information: WindowOffset and RowOffset record the number of TCU blocks and non-zero elements in each split part within each window; CurWindow and CurRow track the index of each split part in the original window and row; Atomic indicates whether each split part needs to perform atomic operations. Overall, the strategy of the application example effectively balances the load while significantly reducing the atomic operation overhead in hybrid computing.
[0129] (iv) Hybrid computing task scheduling method on TCU and CUDA Core In order to dynamically select between multi-stream parallel execution mode and serial execution mode to obtain optimal performance, the application example proposes a hybrid computing task scheduling method on TCU and CUDA Core (by designing two heuristic indicators to accurately evaluate and determine which execution mode is most suitable for a specific hybrid computing task to achieve optimal performance, which includes: 1) theoretical execution time ratio indicator (denoted as time_ratio, abbreviated as r); 2) atomic operation ratio indicator (denoted as atomic_ratio).
[0130] The application example proposes a hybrid computing task scheduling method on TCU and CUDA Core, which includes the following 5 steps: 1) Define the theoretical execution time ratio time_ratio: This indicator is used to measure the execution time proportion of CUDA Core and TCU under a specific task.
[0131] 2) Define the atomic operation ratio atomic_ratio: This indicator reflects the degree of synchronization overhead caused by atomic operations in the task.
[0132] 3) TCU utilization rate prediction model selection TCU and CUDA Core parallel execution can cause the computing frequency to drop (such as thermal throttling or resource contention), resulting in reduced TCU actual utilization. In order to accurately depict the impact of different computing load proportions on TCU utilization rate during parallel execution, the application example selects the exponential decay model as shown in the aforementioned formula (2) based on the experimental observed data characteristics (such as Figure 9 as shown) and the theoretical execution time ratio r defined by the application example.
[0133] 4) Determine the TCU utilization rate prediction model to be determined sparse Taking the NVIDIA RTX4090 GPU as an example, the actual utilization of the TCU under different theoretical execution time ratios r is tested by performing experiments on the RTX4090 GPU. The experimental data show that when r gradually increases from 0.1 to 1, the utilization of the TCU rapidly decreases; and when r further increases to 6, the utilization almost tends to be stable. In the application example, the actual utilization of the TCU under different theoretical execution time ratios r obtained by testing is taken as observation data, the undetermined coefficients (including A, k and C) of the model are determined by function fitting, so as to obtain an accurate TCU utilization prediction model as shown in the above formula (1). The fitting model is highly consistent with the experimental data, and effectively describes the change rule of the TCU utilization with the load ratio. This model not only facilitates the formulation and optimization of subsequent hardware scheduling strategies, but also provides a theoretical basis for in-depth analysis of the resource utilization of mixed computing load on GPU heterogeneous hardware.
[0134] (5) Task scheduling mode selection strategy In order to realize effective task scheduling, the application example establishes a TCU utilization threshold U thr= 78% and an atomic index threshold A thr = 0.6 according to a large number of experiments (including 500 test matrices). First, the index r is calculated according to the defined theoretical execution time ratio formula; The obtained execution time ratio r is substituted into the fitted exponential decay function to calculate the predicted TCU utilization U TCU (r). If U TCU (r) < U thr , it indicates that the TCU utilization is obviously low, and the serial execution mode should be directly used; If U TCU (r) >= U thr , the atomic operation ratio atomic_ratio of the task is further calculated and compared with the threshold A thr . If atomic_ratio > A thr , it indicates that the task has high synchronization overhead, and the serial execution mode should be used; otherwise, it indicates that the task has low synchronization overhead, and the multi-stream parallel execution mode should be used.
[0135] In summary, the application example can accurately and without additional profiling overhead predict and select the optimal execution mode of each computing task kernel function by comprehensively considering the above two heuristic indexes and the exponential decay function, thereby significantly improving the computing throughput and hardware resource utilization efficiency of the overall task.
[0136] (Five) Application scenario The fine mixed load division and execution mode prediction strategy proposed in the application example is particularly suitable for sparse computing scenarios in graph neural networks (GNN). In a typical GNN inference or training process, the node feature aggregation stage involves large-scale sparse matrix multiplication (SpMM) operations and has highly heterogeneous load characteristics, including strong unstructured sparsity and uneven local density distribution. Traditional scheduling strategies based on a single execution mode cannot fully adapt to this heterogeneity, which can easily lead to low resource utilization and computing bottlenecks. By introducing the feature extraction mechanism and execution mode prediction model described in the application example, the sparse computing tasks in each layer of the GNN can select a more suitable execution mode (multi-stream parallel execution or serial execution) according to the actual graph structure, achieving a high degree of matching between load and hardware resources. Experiments show that compared with DGL, the method of the application example achieves a geometric mean speedup of 1.57 times (up to 1.89 times) on GCN and a geometric mean speedup of 2.9 times (up to 3.9 times) on AGNN.
[0137] In addition, in addition to the unstructured sparse computing scenario in graph neural networks, the application example is also widely applicable in large language model (LLM) inference. In recent years, large language models have widely adopted weight pruning techniques for model compression. The pruned large models also involve large-scale sparse matrix multiplication (SpMM) operations, and high-accuracy large model pruning techniques often introduce unstructured sparsity, making it difficult to balance the computing load between CUDA cores and TCUs. The mixed load division and mixed computing method proposed in the application example can adaptively select the optimal execution path according to the distribution characteristics of the pruned sparse weights, efficiently schedule the unstructured sparse operators, and significantly improve the inference performance of the pruned sparse large model.
[0138] Therefore, whether in the sparse adjacency multiplication in GNN or in the weight pruning scenario in large language models, this method can effectively improve the throughput and resource utilization, showing good generality and scalability.
[0139] (Six) Experimental test Referring to Table 1, the experimental results on the NVIDIA H100 GPU show that, in the test of 500 sparse matrices, the TCU and CUDA core hybrid computing method proposed by the application instance can cover more optimal matrices than using only CUDA core or only TCU, and can significantly improve the performance. Specifically, the hybrid computing in SpMM is 1.59 times and 1.22 times higher than using only CUDA core and only TCU, respectively, and in SDDMM, it is 1.97 times and 1.21 times higher, respectively, with a maximum speedup of 10.38 times. These results verify that the hybrid computing method proposed by the application instance can greatly improve the computing performance of sparse matrix operators on TCU and CUDA core heterogeneous computing resources.
[0140] Table 1 The hybrid load partitioning strategy provided by the application instance, i.e., a new task partitioning method, accurately and dynamically allocates sparse matrix operators to CUDA cores or TCUs for execution based on the non-zero element density of subblocks in the sparse matrix and data reuse characteristics, to improve resource matching, reduce computational redundancy, and significantly improve computing performance. A high-efficiency hybrid computing task scheduling method is also provided for a TCU and CUDA Core hybrid architecture. The core idea of the application instance is to accurately analyze and predict the load ratio and computing efficiency of different computing resources based on the explicit theoretical execution time ratio index (denoted as time_ratio, abbreviated as r) and atomic operation ratio index (atomic_ratio), and an exponential decay model fitted, to determine whether to use single-stream serial execution or multi-stream parallel execution mode.
[0141] The hybrid load partitioning strategy proposed by the application instance can accurately allocate sparse operator tasks to suitable CUDA cores or TCUs for execution based on the density distribution of sparse matrix subblocks and the characteristics of computing resources, effectively solving the problem of overly rough task allocation in the prior art, which leads to underutilization of computing resources and reduced performance. The application instance achieves better matching of sparse tasks and computing unit characteristics through fine-grained task partitioning, reducing computational redundancy caused by task-resource mismatch. A hybrid computing task scheduling method for TCU and CUDA Core hybrid architecture is proposed, which can automatically determine and select the optimal scheduling strategy (multi-stream parallel execution or serial execution) before running, significantly improving the overall performance and stability of hybrid computing. Compared with the traditional hybrid computing scheme that relies only on single-core functions, the application instance achieves dynamic adaptation of computing paths and hardware characteristics through execution mode matching, improving parallel efficiency, and is particularly suitable for heterogeneous computing tasks in non-structured sparse scenarios with large structural differences.
[0142] In addition, one possible alternative of the application example is to use a coarse-grained random task partitioning method. Instead of performing a detailed analysis of the sparse matrix, a simpler partitioning strategy is used, such as directly cutting the sparse matrix into multiple roughly uniform sub-matrices or sub-tasks of fixed size or at random, and then mapping these sub-tasks to a single kernel function or multiple computing streams for execution. Although this method is simpler to implement, it often cannot accurately partition mixed sparse loads due to the lack of consideration of the local density characteristics of the matrix and the adaptability of hardware resources, resulting in a large amount of computational redundancy and memory access conflicts. In addition, unreasonable task scheduling can also cause waste of thread resources and decrease in utilization of execution units, thereby limiting further improvement of overall performance. Compared with the above simple and rough partitioning scheme, the application example proposes a more detailed mixed load partitioning and scheduling strategy. Based on reasonable partitioning of sparse computing loads, the strategy further selects the optimal execution mode through a hybrid computing task scheduling method of double indicators (time_ratio and atomic_ratio) and an exponential decay model, maps the sub-tasks to appropriate computing paths to achieve efficient parallel execution. This method not only significantly reduces computational redundancy and memory access conflicts, but also fully utilizes the computing advantages of CUDA Core and TCU, thereby more effectively releasing the collaborative computing potential of GPU heterogeneous resources. Therefore, although using coarse-grained random task partitioning and simple execution strategy is a feasible alternative, it is far inferior to the technical solution proposed by the application example in terms of performance, computational efficiency and resource utilization.
[0143] From the software level, the application also provides a TCU and CUDA core hybrid computing task scheduling device for executing all or part of the TCU and CUDA core hybrid computing task scheduling method. The TCU and CUDA core hybrid computing task scheduling device specifically includes the following contents: A theoretical execution time ratio determination module is configured to determine a value of a preset theoretical execution time ratio according to a ratio between an estimated execution time of a CUDA core corresponding to a CUDA core in a graphics processing unit (GPU) and an estimated execution time of a tensor core (TCU) corresponding to a TCU in the graphics processing unit (GPU). The CUDA core is used to perform matrix multiplication calculation for CUDA core computing tasks in each computing task. The TCU is used to perform matrix multiplication calculation for TCU computing tasks in each computing task. A TCU utilization rate prediction module is configured to input the value of the theoretical execution time ratio into a preset TCU utilization rate prediction model to obtain a corresponding TCU utilization rate. The single-flow serial scheduling module is configured to serially schedule each of the CUDA core computing tasks and each of the TCU computing tasks in a single computing flow if the TCU utilization is less than the preset TCU utilization threshold.
[0144] The TCU and CUDA core hybrid computing task scheduling device provided in the embodiments of the present application can be specifically used to execute the processing procedure of the embodiments of the TCU and CUDA core hybrid computing task scheduling method described above, and the functions thereof will not be repeated here. Please refer to the detailed description of the embodiments of the TCU and CUDA core hybrid computing task scheduling method described above.
[0145] The part of the TCU and CUDA core hybrid computing task scheduling device performing TCU and CUDA core hybrid computing task scheduling can be completed in the GPU.
[0146] As can be seen from the above description, the TCU and CUDA core hybrid computing task scheduling device provided in the embodiments of the present application can accurately and without additional profiling overhead predict and select the optimal execution mode of each computing task kernel function, thereby significantly improving the computing throughput of the overall sparse matrix multiplication computing task and the utilization rate of the GPU hardware computing resources, and improving the operation reliability and effectiveness of the GPU.
[0147] The embodiments of the present application also provide an electronic device, which can include a processor, a memory, a receiver and a transmitter. The processor is configured to execute the TCU and CUDA core hybrid computing task scheduling method mentioned in the embodiments described above. The processor and the memory can be connected through a bus or other means, and the connection through the bus is taken as an example. The receiver can be connected with the processor and the memory through wired or wireless means.
[0148] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, or combinations thereof.
[0149] The memory, as a non-transitory computer readable storage medium, can be configured to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the TCU and CUDA core hybrid computing task scheduling method in the embodiments of the present application. The processor can execute various functions and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, implement the TCU and CUDA core hybrid computing task scheduling method in the above method embodiments.
[0150] The memory can include a program storage area and a data storage area. The program storage area can store an operating system and application programs required by at least one function. The data storage area can store data created by the processor and the like. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0151] The one or more modules are stored in the memory and, when executed by the processor, perform the TCU and CUDA core hybrid computing task scheduling method in the embodiments.
[0152] In some embodiments of the present application, a user equipment can include a processor, a memory and a transceiver unit, which can include a receiver and a transmitter. The processor, the memory, the receiver and the transmitter can be connected through a bus system. The memory is configured to store computer instructions, and the processor is configured to execute the computer instructions stored in the memory to control the transceiver unit to transceive signals.
[0153] As an implementation manner, the functions of the receiver and the transmitter in the present application can be implemented by a transceiver circuit or a transceiver dedicated chip. The processor can be implemented by a dedicated processing chip, a processing circuit or a general-purpose chip.
[0154] As another implementation manner, the server provided by the embodiments of the present application can be implemented by using a general-purpose computer. That is, program codes for implementing the functions of the processor, the receiver and the transmitter are stored in the memory, and the general-purpose processor implements the functions of the processor, the receiver and the transmitter by executing the codes in the memory.
[0155] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps of the foregoing hybrid computing task scheduling method of TCU and CUDA core.
[0156] The embodiment of the present application further provides a computer program product, which comprises a computer program. The computer program is executed by a processor to implement the steps of the foregoing hybrid computing task scheduling method of TCU and CUDA core.
[0157] Those skilled in the art should understand that the exemplary components, systems and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software or a combination thereof. The exact implementation depends on the specific application and design constraints imposed on the overall system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are program or code segments used to perform the required tasks. The program or code segments can be stored in a machine readable medium or transmitted through a data signal carried in a carrier wave in a transmission medium or communication link.
[0158] It should be noted that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of well-known methods are omitted. In the above embodiments, several specific steps are described and shown as examples. However, the method processes of the present application are not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications and additions, or change the order of the steps, after understanding the spirit of the present application.
[0159] In the present application, the features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or in combination with or instead of features of other embodiments.
[0160] The above descriptions are only the preferred embodiments of the present application, and are not intended to limit the present application. The embodiments of the present application can be variously changed and modified by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for scheduling a hybrid computing task on a TCU and a CUDA core, characterized in that, The method comprises the following steps: determining a preset theoretical execution time ratio value according to a ratio between a CUDA core estimated execution time corresponding to a CUDA core in a graphics processing unit (GPU) and a tensor core (TCU) estimated execution time corresponding to a TCU in the graphics processing unit (GPU); wherein the CUDA core is used to perform matrix multiplication calculation for a CUDA core calculation task in each calculation task; the TCU is used to perform matrix multiplication calculation for a TCU calculation task in each calculation task; inputting the theoretical execution time ratio value into a preset TCU utilization rate prediction model to obtain a corresponding TCU utilization rate; if the TCU utilization rate is less than a preset TCU utilization rate threshold, sequentially and serially scheduling each CUDA core calculation task and each TCU calculation task in a single calculation stream. 2.The method of claim 1, wherein, Further comprising: if the TCU utilization rate is equal to or greater than the TCU utilization rate threshold, determining a preset atomic operation ratio value according to a ratio between a total number of atomic operation instructions in an execution process of the CUDA core calculation task and the TCU calculation task and a total number of all calculation operation instructions; if the atomic operation ratio value is greater than a preset atomic index threshold, sequentially and serially scheduling each CUDA core calculation task and each TCU calculation task in a single calculation stream. 3.The method of claim 2, wherein, Further comprising: if the atomic operation ratio value is less than or equal to the atomic index threshold, decoupling each TCU calculation task and each CUDA core calculation task into different calculation streams for parallel scheduling.
4. The method of claim 1, wherein the TCU and CUDA core hybrid computing task scheduling method is characterized in that, The TCU utilization rate prediction model is shown in formula (1): U tcu (r) = 38.9 exp(-3.66r) + 70 (1) wherein r represents the theoretical execution time ratio; U tcu (r) represents the TCU utilization.
5. The method of claim 4, wherein, Before the step of inputting the theoretical execution time ratio value into the preset TCU utilization rate prediction model, the method further comprises the following steps: constructing an exponential decay model between the theoretical execution time ratio and the TCU utilization rate; wherein the exponential decay model is shown in formula (2): U tcu (r) = A exp(-kr) + C (2) wherein A, k and C represent to-be-determined coefficients; using TCU actual utilization rates respectively tested under different theoretical execution time ratio values as observation data, determining respective values of each to-be-determined coefficient in the exponential decay model through function fitting to obtain a corresponding TCU utilization rate prediction model.
6. The method of scheduling a computing task on a hybrid of TCU and CUDA core according to any one of claims 1 to 5, characterized in that, Before the step of determining the preset theoretical execution time ratio value according to a ratio between a CUDA core estimated execution time corresponding to a CUDA core in a graphics processing unit (GPU) and a tensor core (TCU) estimated execution time corresponding to a TCU in the graphics processing unit (GPU), the method further comprises the following steps: determining a corresponding division granularity and a non-zero element threshold of a sparse matrix multiplication according to a type of the sparse matrix multiplication to be calculated; wherein the type of the sparse matrix multiplication includes a sparse matrix and dense matrix multiplication (SpMM) and a sampled dense matrix multiplication (SDDMM); According to the division granularity and the non-zero element threshold, each sub-matrix corresponding to each window in the sparse matrix multiplication is respectively allocated load to obtain a calculation group corresponding to each window; wherein each calculation group comprises a CUDA core calculation block corresponding to the CUDA core and / or a TCU block corresponding to the TCU; each calculation group is respectively subjected to load balancing processing to obtain a plurality of calculation tasks corresponding to each calculation group; wherein each calculation task comprises a CUDA core calculation task corresponding to the CUDA core calculation block and / or a TCU calculation task corresponding to the TCU block; so that the CUDA core performs matrix multiplication calculation for the CUDA core calculation task and / or the TCU performs matrix multiplication calculation for the TCU calculation task.
7. The method of claim 6, wherein the TCU and CUDA core hybrid computing task scheduling method is characterized in that, The type of the current sparse matrix multiplication to be calculated is determined to determine the division granularity and the non-zero element threshold corresponding to the sparse matrix multiplication, comprising: If the type of the current sparse matrix multiplication to be calculated is the SpMM, the division granularity corresponding to the SpMM is determined to be the non-zero element granularity, and the first non-zero element threshold corresponding to the SpMM is obtained; If the type of the current sparse matrix multiplication to be calculated is the SDDMM, the division granularity corresponding to the SDDMM is determined to be the TCU block granularity, and the second non-zero element threshold corresponding to the SDDMM is obtained.
8. The method of claim 7, wherein the TCU and CUDA core hybrid computing task scheduling method is characterized in that, According to the division granularity and the non-zero element threshold, each sub-matrix corresponding to each window in the sparse matrix multiplication is respectively allocated load to obtain a calculation group corresponding to each window, comprising: If the type of the current sparse matrix multiplication to be calculated is the SpMM, whether the number of non-zero elements in each column in each sub-matrix corresponding to each window is equal to or greater than the first non-zero element threshold is respectively judged based on the non-zero element granularity; for each sub-matrix, if there is a column in the sub-matrix whose number of non-zero elements is equal to or greater than the first non-zero element threshold, the column in the sub-matrix whose number of non-zero elements is equal to or greater than the first non-zero element threshold is compressed into a TCU block; if there is a column in the sub-matrix whose number of non-zero elements is less than the first non-zero element threshold, each column in the sub-matrix whose number of non-zero elements is less than the first non-zero element threshold is divided into a CUDA core calculation block to obtain a calculation group corresponding to each window; If the type of the sparse matrix multiplication to be calculated currently is the SDDMM, for each window, each column in the window is sorted in descending order of the number of non-zero elements contained in each column in the sub-matrix corresponding to the window; based on the TCU block granularity, each sorted column in each window is divided into a TCU judgment unit, and it is judged whether the number of non-zero elements in each TCU judgment unit is equal to or greater than the second non-zero element threshold; for each sub-matrix, if there is a TCU judgment unit in which the number of non-zero elements is equal to or greater than the second non-zero element threshold, the TCU judgment unit equal to or greater than the second non-zero element threshold in each sub-matrix is allocated as a TCU block; if there is a TCU judgment unit in which the number of non-zero elements is less than the second non-zero element threshold, each element in the TCU judgment unit less than the second non-zero element threshold in each sub-matrix is divided into a CUDA core calculation block, to obtain a calculation group corresponding to each window respectively.
9. The method of claim 6, wherein the TCU and CUDA core hybrid computing task scheduling method is characterized in that, The load balancing processing is performed on each calculation group respectively to obtain a plurality of calculation tasks corresponding to each calculation group respectively, including: a preset splitting step is performed on the calculation group corresponding to each window respectively to obtain a plurality of calculation tasks corresponding to each calculation group respectively; wherein, the splitting step includes: if the calculation group contains the TCU block, it is judged whether the number of columns in the TCU block exceeds a first column threshold, if yes, each TCU block in the calculation group is divided into a plurality of TCU block calculation tasks based on the first column threshold; and if the calculation group contains the CUDA core calculation block, it is judged whether the number of CUDA core calculation blocks in the calculation group exceeds a second number threshold, if yes, each CUDA core calculation block in the calculation group is divided into a plurality of CUDA core calculation tasks based on the second number threshold, and each CUDA core calculation block in the CUDA core calculation task whose number of CUDA core calculation blocks exceeds a third number threshold is determined as a long CUDA core calculation block, and each CUDA core calculation block in the CUDA core calculation task whose number of CUDA core calculation blocks does not exceed the third number threshold is determined as a short CUDA core calculation block; Correspondingly, the TCU and CUDA core hybrid computing task scheduling method further includes: based on a plurality of auxiliary arrays, the splitting information corresponding to each calculation task obtained by splitting in each calculation group of each window is recorded; wherein, the auxiliary array includes: a first array for recording the number of TCU blocks corresponding to each calculation task in each window; a second array for recording the number of non-zero elements corresponding to each calculation task in each window; a third array for recording the index of each calculation task in the original window corresponding to each calculation task respectively. a fourth array for recording indexes of each of the computing tasks in the respective original row; a fifth array for recording whether each of the computing tasks is to perform an atomic operation.
10. The method of claim 6, wherein the TCU and CUDA core hybrid computing task scheduling method is characterized by, Before determining the partition granularity and the non-zero element threshold corresponding to the sparse matrix multiplication based on the type of the current sparse matrix multiplication to be calculated, the method further comprises: constructing a data access cost model corresponding to each of a sparse matrix and dense matrix multiplication (SpMM) and a sampled dense matrix multiplication (SDDMM); wherein the access cost model is used to represent a data access cost ratio of a CUDA core and a TCU; determining the partition granularity of the SpMM based on the data access cost model corresponding to the SpMM, and determining the partition granularity of the SDDMM based on the data access cost model corresponding to the SDDMM.
11. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method for scheduling a computing task on a TCU and a CUDA core according to any one of claims 1 to 10 when executing the computer program.
Citation Information
Patent Citations
Method and device for processing vector operator based on Tensor Core and computer equipment
CN119578573A
Task scheduling method and device, equipment and medium
CN119759543A
Hash-based sparse matrix vector multiplication optimization method and device
CN119884572A
High-performance SpMM kernel implementation method based on dense tensor core
CN120045827A
Method, device, and computer program product for assigning tasks to dedicated processing resources
US20200133735A1