Acceleration method and device based on irregular sparse model, electronic equipment and medium
By splitting the sparse model and generating attribute information, dynamically scheduling to appropriate computing units, the problems of waste of resources and low performance of irregular sparse models are solved, and efficient sparse acceleration is achieved, suitable for edge computing and large language models.
Patent Information
- Application Number
- CN202510332831.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-20
AI Technical Summary
When processing irregular sparse models, the prior art cannot accurately identify sub-blocks suitable for sparse calculations or intensive calculations, resulting in waste of resources and low performance, and failure to effectively utilize hardware resources, especially in large-scale data processing to create computing bottlenecks.
By inputting the sparse model into the target compiler, splitting it into sparse sub-blocks, generating attribute information of each sparse sub-block, and fusing and mapping it to appropriate computing units based on the attribute information, dynamic resource scheduling and end-to-end optimization are achieved.
It significantly improves the computing efficiency of sparse models, reduces resource waste, improves memory efficiency and end-to-end inference throughput, and is suitable for edge computing and efficient sparse acceleration of large language models.
Smart Images

Figure CN120336834A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to computer technology, an acceleration method, device, electronic device, and medium based on an irregular sparse model Background Art
[0002] With the wide application of deep neural networks (DNNs) in edge computing and large language models (LLMs), accelerating the processing efficiency of irregular sparse models has become a key challenge. Currently, sparse model acceleration is mainly achieved through the following two types of methods: static optimization of sparse formats and compilation-directed sparse-aware optimization. The above-mentioned static optimization of sparse formats usually refers to using fixed sparse coding formats (such as CSR, COO) and a single computing paradigm (such as the pure sparse computing kernel of cuSPARSE), and improving regularity through data rearrangement or block strategies. The above-mentioned compilation-directed sparse-aware optimization usually refers to using traditional frameworks (such as TVM-Sparse, SparseTIR) to achieve multi-format hybrid coding and operator fusion, but still relies on static sparse block partitioning strategies.
[0003] However, the above methods often have the following technical problems when dealing with irregular sparse models (such as the sparse weight matrix of Llama-2): Sparse acceleration schemes (such as Sputnik, ASpT) do not model the performance of sparse sub-blocks in different sparse rate regions, resulting in the inability to accurately identify sub-blocks suitable for sparse or dense computing in irregular distribution scenarios, causing waste of computing resources; traditional sparse coding (such as CSR, COO) uses a unified format for irregularly distributed matrices, resulting in low memory efficiency of sparse sub-blocks due to redundant storage; existing frameworks (such as cuSPARSE) map the entire sparse matrix to a single type of computing unit, resulting in the overall performance being slowed down by invalid zero-value calculations when dense sub-blocks occupy sparse computing units, or waste of resources due to empty calculations when sparse sub-blocks use dense units; traditional batch processing schemes (such as SparseTIR) do not consider cross-batch load correlations. When dealing with large-scale data, there are often computational bottlenecks where computing units are continuously allocated high-density sub-blocks, resulting in a decrease in inference throughput.
[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the concept of the present disclosure. Therefore, it may include information that does not form the prior art known to ordinary skilled in the art in this country. Summary of the Invention
[0005] The content part of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure propose an acceleration method, apparatus, electronic device, and medium based on an irregular sparse model to explain one or more of the technical problems mentioned in the above background art section.
[0007] In a first aspect, some embodiments of the present disclosure propose an acceleration method based on an irregular sparse model. The method includes: inputting a sparse model into a target compiler, where the target compiler includes a layer intermediate representation, a block layer intermediate representation, and an operator layer intermediate representation; in the layer intermediate representation, splitting the sparse weight matrix of the sparse model into sparse sub-blocks to obtain a split sparse weight matrix, where the split sparse weight matrix includes each sparse sub-block; based on a sparse operator performance model, generating attribute information of each sparse sub-block in the split sparse weight matrix; in the block layer intermediate representation, performing a fusion process on each sparse sub-block according to the attribute information of each sparse sub-block to obtain a fused sparse weight matrix, where each sparse block in the fused sparse weight matrix includes each target fusion sparse block and each unfused sparse sub-block, and each target fusion sparse block corresponds to attribute information; in the operator layer intermediate representation, mapping each sparse block to a computing unit according to the attribute information corresponding to each sparse block to obtain a mapping result; generating a processing task corresponding to the sparse model according to the mapping result, and executing the processing task.
[0008] In a second aspect, some embodiments of the present disclosure propose an acceleration apparatus based on an irregular sparse model. The apparatus includes: an input unit configured to input a sparse model into a target compiler, where the target compiler includes a layer intermediate representation, a block layer intermediate representation, and an operator layer intermediate representation; a splitting unit configured to split the sparse weight matrix of the sparse model into sparse sub-blocks in the layer intermediate representation to obtain a split sparse weight matrix, where the split sparse weight matrix includes each sparse sub-block; a first generating unit configured to generate attribute information of each sparse sub-block in the split sparse weight matrix based on a sparse operator performance model; a fusion unit configured to perform a fusion process on each sparse sub-block according to the attribute information of each sparse sub-block in the block layer intermediate representation to obtain a fused sparse weight matrix, where each sparse block in the fused sparse weight matrix includes each target fusion sparse block and each unfused sparse sub-block, and each target fusion sparse block corresponds to attribute information; a mapping unit configured to map each sparse block to a computing unit according to the attribute information corresponding to each sparse block in the operator layer intermediate representation to obtain a mapping result; a second generating unit configured to generate a processing task corresponding to the sparse model according to the mapping result, and execute the processing task.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect above.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect above.
[0011] In a fifth aspect, some embodiments of the present disclosure provide a computer program product including a computer program, which, when executed by a processor, implements the method described in any implementation manner of the first aspect above.
[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the acceleration method based on the irregular sparse model of some embodiments of the present disclosure, dynamic resource scheduling and end-to-end collaborative optimization are realized, and the computing efficiency of the sparse model is significantly improved. Specifically, the following problems often exist in traditional sparse acceleration methods: The performance of sparse sub-blocks in different sparse rate regions is not modeled, resulting in the inability to accurately identify sub-blocks suitable for sparse computing or dense computing in the irregular distribution scenario, causing waste of computing resources; A unified format is adopted for the irregular distribution matrix, resulting in low memory efficiency due to redundant storage of sparse sub-blocks; The entire sparse matrix is mapped to a single type of computing unit, resulting in different types of sub-blocks in the sparse matrix not being mapped to the corresponding computing units, causing low performance improvement and waste of resources; The cross-batch load correlation is not considered, resulting in a decrease in end-to-end inference throughput. Based on this, in the acceleration method based on the irregular sparse model of some embodiments of the present disclosure, first, the sparse model is input into the target compiler, where the above-mentioned target compiler includes a layer intermediate representation, a block layer intermediate representation, and an operator layer intermediate representation. Thus, the sparse model can be used as a data source. Next, in the above-mentioned layer intermediate representation, the sparse weight matrix of the above-mentioned sparse model is split into sparse sub-blocks to obtain the split sparse weight matrix, where the above-mentioned split sparse weight matrix includes each sparse sub-block. Thus, the sparse weight matrix in the above-mentioned sparse model can be divided for subsequent processing. Then, based on the sparse operator performance model, the attribute information of each sparse sub-block in the above-mentioned split sparse weight matrix is generated. Thus, by modeling the performance of sparse sub-blocks in different sparse rate regions, a sparse operator performance model can be obtained to facilitate predicting the attribute information of each sparse sub-block and reducing waste of computing resources. Then, in the above-mentioned block layer intermediate representation, according to the attribute information of each sparse sub-block, the above-mentioned sparse sub-blocks are fused to obtain the fused sparse weight matrix, where each sparse block in the above-mentioned fused sparse weight matrix includes each target fused sparse block and each unfused sparse sub-block, and each of the above-mentioned target fused sparse blocks corresponds to attribute information. Thus, the above-mentioned sparse sub-blocks are fused and the format encoding is converted to facilitate reducing the redundant storage of the above-mentioned fused sparse weight matrix and thus improving the memory efficiency. Then, in the above-mentioned operator layer intermediate representation, according to the attribute information corresponding to each sparse block, each sparse block is mapped to a computing unit to obtain a mapping result. Thus, the above-mentioned sparse blocks can be mapped to different computing units according to their types to obtain a mapping result, improving the overall performance improvement and reasonably allocating resources. Finally, according to the above-mentioned mapping result, a processing task corresponding to the above-mentioned sparse model is generated, and the above-mentioned processing task is executed. Thus, the above-mentioned mapping result and the above-mentioned sparse operator performance model can be used to optimize the model and process the task, thereby reflecting the acceleration effect.In addition, when processing computing tasks across batches, high-load sparse sub-blocks and low-load sparse sub-blocks in adjacent sparse weight matrices can be cross-mapped to different computing units. Therefore, the load correlation between different batches can be improved and the throughput of end-to-end inference can be increased. As a result, it can adapt to different sparse patterns and hardware environments, showing significant advantages in large model inference scenarios, providing an efficient sparse acceleration solution for edge computing and large language model deployment, significantly enhancing the applicability of the model in mobile and cloud inference scenarios, and breaking through the benefit boundary of traditional unstructured sparse acceleration designs. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that elements and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of an acceleration method based on an unstructured sparse model according to the present disclosure;
[0015] Figure 2 is a schematic structural diagram of some embodiments of an acceleration method device based on an unstructured sparse model according to the present disclosure;
[0016] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure.
[0017] Figure 4 is a bar chart of multi-dimensional performance comparison of some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should also be noted that, for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0020] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0021] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The following will detail this disclosure with reference to the accompanying drawings and in conjunction with embodiments.
[0024] Figure 1 Flow 100 of some embodiments of an acceleration method based on an unstructured sparse model according to this disclosure is shown. The acceleration method based on the unstructured sparse model includes the following steps:
[0025] Step 101, input the sparse model into the target compiler.
[0026] In some embodiments, an execution entity (such as a computing device) of an acceleration method based on an irregular sparse model may input the sparse model into a target compiler. Among them, the above-mentioned target compiler includes a layer intermediate representation, a block layer intermediate representation, and an operator layer intermediate representation. The above-mentioned target compiler can be a compilation tool that supports multi-level intermediate representations and is specifically used to process acceleration tasks based on irregular sparse models. Through the above-mentioned layer intermediate representation, the above-mentioned block layer intermediate representation, and the above-mentioned operator layer intermediate representation, the weight matrix of the sparse model can be split, fused, and mapped, and finally an efficient processing task is generated and executed. The above-mentioned layer intermediate representation is used to describe the overall structure of the model and is used for high-level optimization. The above-mentioned block layer intermediate representation is used to describe the local structure of the model and is used for middle-level optimization. The above-mentioned operator layer intermediate representation is used to describe the specific calculations of the model and is used for low-level optimization. The above-mentioned sparse model can be an irregular sparse model (such as a DNN). The above-mentioned irregular sparse model is usually a mathematical model through unstructured sparsity. The above-mentioned unstructured sparsity usually refers to pruning any element in the matrix of the model to obtain an irregular matrix. The above-mentioned pruning operation usually sets any element in the matrix to 0, thereby reducing the complexity and computational amount of the model while trying to maintain the performance of the model. In practice, the above-mentioned execution entity can input the irregular sparse DNN model into the target compiler according to the file format of the model (such as ONNX, TensorFlow SavedModel, or PyTorch) for subsequent processing.
[0027] Step 102, in the layer intermediate representation, split the sparse weight matrix of the sparse model into sparse sub-blocks to obtain the split sparse weight matrix.
[0028] In some embodiments, the above-mentioned execution entity may split the sparse weight matrix of the above-mentioned sparse model into sparse sub-blocks in the above-mentioned layer intermediate representation to obtain the split sparse weight matrix. Among them, the above-mentioned split sparse weight matrix includes each sparse sub-block. The weight matrix of the above-mentioned sparse model is usually a matrix form used to describe model parameters. Each of the above-mentioned sparse sub-blocks is usually each local matrix divided from the above-mentioned sparse weight matrix, and each of the above-mentioned local matrices can be used as a sparse sub-block of the above-mentioned sparse weight matrix. In practice, the above-mentioned execution entity can split the above-mentioned sparse weight matrix into each sparse sub-block according to the dimension of the preset sub-block (such as 2×2 or 4×4). And record the positions of each of the above-mentioned sparse sub-blocks in the above-mentioned sparse weight matrix. Since the positions of each of the above-mentioned sparse sub-blocks in the above-mentioned sparse weight matrix have not changed, the sparse weight matrix split into each sparse sub-block can be determined as the split sparse weight matrix.
[0029] Step 103: Based on the sparse operator performance model, generate the attribute information of each sparse sub-block in the split sparse weight matrix.
[0030] In some embodiments, the above-mentioned execution entity may generate the attribute information of each sparse sub-block in the split sparse weight matrix based on the sparse operator performance model. Among them, the sparse operator performance model may be a structured table including various candidate matrix feature information. The various candidate matrix feature information usually includes the shape, threshold, and fitting execution time of each candidate matrix. The various candidate matrices may be matrices with various pre-set shapes (such as 256×256, 128×128, 64×64, etc.). The thresholds of the various candidate matrices usually include the dense threshold and the sparse threshold of each candidate matrix. The thresholds of the various candidate matrices can be used to determine the type of each matrix. When the sparsity rate of the matrix is greater than the sparse threshold, it can be obtained that the type of the matrix is the sparse type; when the sparsity rate of the matrix is between the sparse threshold and the dense threshold, it can be obtained that the type of the matrix is the swing type; when the sparsity rate of the matrix is less than the dense threshold, it can be obtained that the type of the matrix is the dense type. The above types include the dense type, the swing type, and the sparse type. The fitting execution time may be the time predicted for each candidate matrix to perform matrix multiplication at each sparsity rate. The above matrix multiplication includes dense calculation and sparse calculation. The dense calculation refers to operating on all elements one by one in matrix multiplication. The sparse calculation refers to operating only on non-zero elements in matrix multiplication. The fitting execution time of the various candidate matrices may include the fitting dense calculation time and the fitting sparse calculation time of each candidate matrix. The fitting dense calculation time of the various candidate matrices may be the actual execution time for each candidate matrix to perform dense calculation. The fitting sparse calculation time of the various candidate matrices may be the time predicted for each candidate matrix to perform sparse calculation at different sparsity rates. The fitting sparse calculation time of the various candidate matrices is usually represented by the fitting parameters corresponding to each candidate matrix. The fitting parameters corresponding to the various candidate matrices represent the relationship between the sparsity rate of each candidate matrix and the time for sparse calculation. The above attribute information may include the sparsity rate, shape, type, coordinates, and predicted execution time of each sparse sub-block. The sparsity rate usually refers to the proportion of zero elements in the matrix. The predicted execution time is usually the time predicted for each sparse block to perform matrix multiplication. The predicted execution time includes the dense calculation predicted execution time and the sparse calculation predicted execution time. The dense calculation predicted execution time may be the actual execution time for each sparse block to perform dense calculation on the candidate matrix corresponding to the sparse operator performance model. The sparse calculation predicted execution time may be the fitting sparse calculation time of each sparse block on the candidate matrix corresponding to the sparse operator performance model.
[0031] In some alternative implementation manners of some embodiments, the above-mentioned execution entity may generate attribute information of each sparse sub-block based on a sparse operator performance model through the following steps:
[0032] Step 1: Obtain the actual execution time data of matrix multiplication of each candidate matrix at each sparsity rate. Among them, the above-mentioned actual execution time data may be the actual time taken for matrix multiplication of each candidate matrix obtained at each sparsity rate. The above-mentioned actual execution time data includes the actual execution time of sparse calculation and the actual execution time of dense calculation. In practice, the above-mentioned execution entity may implement sparse calculation through Sputnik. Dense calculation may be implemented through cuBlas. Then, the actual execution time data of each of the above-mentioned candidate matrices at each sparsity rate may be obtained respectively. For example, the above-mentioned execution entity may generate 6 groups of candidate matrices with different shapes, and take 23 sparsity rates on average in the sparsity rate range [0.55, 0.99]. Then, the above-mentioned 6 groups of candidate matrices with different shapes may be respectively subjected to sparse calculation and dense calculation with a preset feature matrix (generally a dense matrix) at 23 sparsity rates. Finally, the execution time of the above-mentioned 6 groups of candidate matrices with different shapes for the two matrix operations at the above-mentioned 23 sparsity rates is obtained as the actual execution time data of each candidate.
[0033] Step 2: Generate a sparse operator performance model based on the above-mentioned actual execution time data.
[0034] Step 3: Generate the attribute information of each sparse sub-block according to the sparse operator performance model. In practice, the above-mentioned execution entity can identify the dimensions of each sparse sub-block (such as 256×256 or 1024×1024), and use the dimensions of each sparse sub-block as the shape of each sparse sub-block. Then, it can identify the proportion of zero elements in each sparse sub-block, and use the proportion of zero elements in each sparse sub-block as the sparsity rate of each sparse sub-block. In the above-mentioned sparse operator performance model, retrieve the corresponding candidate matrices of each sparse sub-block according to the shape of each sparse sub-block. Next, obtain the attribute information of each sparse sub-block according to the characteristic information of the corresponding candidate matrices. For example, for a sparse sub-block with a shape of (256×256), the above-mentioned execution entity can retrieve a candidate matrix with a shape of 256×256 in the above-mentioned sparse operator performance model. Then, according to the sparsity rate of this sparse sub-block (such as 0.9) and the thresholds of the candidate matrix with a shape of 256×256 (the dense threshold is 0.675 and the sparse threshold is 0.835), obtain the type of this sparse sub-block as the sparse type. Next, use the position of this sparse sub-block in the above-mentioned split sparse weight matrix as the coordinates of this sparse sub-block. Then, use the fitting dense calculation time corresponding to the candidate matrix with a shape of 256×256 as the dense calculation predicted execution time of this sparse sub-block. Next, perform a quadratic polynomial solution on the sparsity rate of this sparse sub-block and the fitting parameters corresponding to the candidate matrix with a shape of 256×256, and obtain the sparse calculation predicted execution time of this sparse sub-block. Integrate the dense calculation predicted execution time of this sparse sub-block and the sparse calculation predicted execution time of this sparse sub-block into the predicted execution time of this sparse sub-block. Finally, integrate the sparsity rate, shape, type, coordinates, and predicted execution time of this sparse sub-block into the attribute information of this sparse sub-block. Similarly, the attribute information of each sparse sub-block in the above-mentioned split sparse weight matrix can be generated.
[0035] In some optional implementation manners of some embodiments, the above-mentioned execution entity can generate a sparse operator performance model based on the above-mentioned actual execution time data through the following steps:
[0036] First step, based on the above actual execution time data, generate the fitting execution time of each of the above candidate matrices. In practice, the above execution entity can use the polynomial curve fitting algorithm to perform polynomial solution on the actual execution time of sparse calculation of each of the above candidate matrices and the sparse rate of each corresponding candidate matrix, so as to obtain the fitting sparse calculation time of each candidate matrix. The actual execution time of dense calculation of each candidate matrix can be used as the fitting dense calculation time of each candidate matrix. Integrate the fitting sparse calculation time of each of the above candidate matrices with the fitting dense calculation time of each candidate matrix to obtain the fitting execution time of each candidate matrix. For example, for any candidate matrix of a certain shape, the above execution entity takes the above 23 sparse rates as independent variables, and takes the actual execution time of sparse calculation corresponding to the above 23 sparse rates of the candidate matrix as the dependent variable. Construct a polynomial from the above independent variable and the above dependent variable. Then use the least squares method to solve the fitting parameters of the polynomial, that is, the fitting sparse calculation time of the candidate matrix. Then, take the actual execution time of dense calculation of the candidate matrix as the predicted execution time of sparse calculation of the candidate matrix. Similarly, the fitting execution time of each of the above candidate matrices can be obtained.
[0037] Second step, based on the fitting execution time of each of the above candidate matrices, generate the threshold of each candidate matrix. In practice, the above execution entity can perform quadratic polynomial solution on the actual execution time of dense calculation of each of the above candidate matrices and the above fitting parameters respectively to obtain the dense threshold of each candidate matrix. Then, according to the performance advantage coefficient (such as 0.8), perform quadratic polynomial solution on the actual execution time of dense calculation of each candidate matrix combined with the performance advantage coefficient and the fitting parameters of each candidate matrix respectively to obtain the sparse threshold of each candidate matrix. Integrate the dense threshold of each candidate matrix and the sparse threshold of each candidate matrix into the threshold of each candidate matrix.
[0038] Third step, according to the threshold of each of the above candidate matrices and the fitting execution time of each of the above candidate matrices, generate a sparse operator performance model. In practice, the above execution entity can integrate the shape of each of the above candidate matrices, the threshold of each of the above candidate matrices, and the fitting execution time of each of the above candidate matrices into a Table model as the sparse operator performance model.
[0039] Step 104, in the block-level intermediate representation, according to the attribute information of each sparse sub-block, perform fusion processing on each sparse sub-block to obtain a fused sparse weight matrix.
[0040] In some embodiments, the above-mentioned execution entity may perform a fusion process on each of the sparse sub-blocks according to the attribute information of each of the sparse sub-blocks in the above-mentioned block-level intermediate representation to obtain a fused sparse weight matrix. Among them, each sparse block in the above-mentioned fused sparse weight matrix includes each target fused sparse block and each unfused sparse sub-block, and each of the above-mentioned target fused sparse blocks corresponds to attribute information. The above-mentioned fusion process generally refers to fusing each of the sparse sub-blocks into each sparse block according to a preset fusion condition, and each fused sparse block will be obtained. The above-mentioned fusion process may include homogeneous block fusion and non-homogeneous block fusion. The above-mentioned preset fusion condition may include a homogeneous block fusion condition and a non-homogeneous block fusion condition. The above-mentioned homogeneous block fusion condition generally means that the type of the sparse block after homogeneous block fusion should be the same as the types of its corresponding two sparse sub-blocks. The above-mentioned non-homogeneous block fusion condition generally means that the type of the sparse block after non-homogeneous block fusion should be the same as the type of the sparse sub-block with a higher sparsity rate among its corresponding two sparse sub-blocks. Each of the above-mentioned fused sparse blocks includes each sparse block after homogeneous block fusion and each sparse block after non-homogeneous block fusion. Each of the above-mentioned sparse blocks after homogeneous block fusion may include each target sparse block after homogeneous block fusion and each sparse block after homogeneous block fusion that does not meet the homogeneous block fusion condition. Each of the above-mentioned sparse blocks after non-homogeneous block fusion may include each target sparse block after non-homogeneous block fusion and each sparse block after non-homogeneous block fusion that does not meet the non-homogeneous block fusion condition. Each of the above-mentioned target fused sparse blocks may include each target sparse block after homogeneous block fusion and each target sparse block after non-homogeneous block fusion. Each of the above-mentioned target sparse blocks after homogeneous block fusion refers to each sparse block after homogeneous block fusion that meets the above-mentioned homogeneous block fusion condition. Each of the above-mentioned target sparse blocks after non-homogeneous block fusion refers to each sparse block after non-homogeneous block fusion that meets the above-mentioned non-homogeneous block fusion condition. Each of the above-mentioned unfused sparse sub-blocks generally refers to each sparse sub-block corresponding to each fused sparse block that does not meet the above-mentioned preset fusion condition. In practice, the above-mentioned execution entity may fuse each of the sparse sub-blocks according to the attribute information of each of the sparse sub-blocks to obtain each fused sparse block. Determine each fused sparse block that meets the above-mentioned preset fusion condition as each target fused sparse block. Integrate each of the above-mentioned target fused sparse blocks and each sparse sub-block corresponding to each fused sparse block that does not meet the above-mentioned preset fusion condition into a fused sparse weight matrix.
[0041] In some optional implementation manners of some embodiments, the above-mentioned execution entity may perform a fusion process on each of the sparse sub-blocks according to the attribute information of each of the sparse sub-blocks in the above-mentioned block-level intermediate representation through the following steps to obtain a fused sparse weight matrix:
[0042] Step 1: According to the attribute information of each of the above sparse sub - blocks, perform homogeneous block fusion on adjacent sparse sub - blocks of the same type among the above sparse sub - blocks to obtain sparse blocks after homogeneous block fusion. Among them, the above homogeneous block fusion generally refers to integrating two adjacent sparse sub - blocks of the same type in the split sparse weight matrix into one matrix, that is, the sparse block after homogeneous block fusion. The attribute information of each of the above sparse blocks after homogeneous block fusion is also determined by the attribute information of the two adjacent sparse sub - blocks corresponding to each of the above sparse blocks after homogeneous block fusion. In practice, the above execution entity can merge two adjacent sparse sub - blocks of the same type in the split sparse weight matrix to obtain sparse blocks after homogeneous block fusion.
[0043] Step 2: For each sparse block after homogeneous block fusion obtained, perform the following steps:
[0044] Sub-step 1: Generate the attribute information of the sparse block after homogeneous block fusion based on the attribute information of the adjacent sparse sub-blocks corresponding to the sparse block after homogeneous block fusion. Among them, the attribute information of the sparse block after homogeneous block fusion includes the sparsity rate, shape, type, coordinates, and predicted execution time of the sparse block after homogeneous block fusion. In practice, the above-mentioned execution entity can obtain the sparsity rate of the sparse block after homogeneous block fusion through the proportion of zero elements in the sparse block after homogeneous block fusion. And, based on the shapes and coordinates of the two sparse sub-blocks corresponding to the sparse block after homogeneous block fusion, obtain the shape and coordinates of the sparse block after homogeneous block fusion. When the sparsity rate and shape of the sparse block after homogeneous block fusion are known, the type and predicted execution time of the sparse block after homogeneous block fusion can be obtained through the above-mentioned sparse operator performance model. Integrate the sparsity rate, shape, type, coordinates, and predicted execution time of the above-mentioned sparse block after homogeneous block fusion into the attribute information of the sparse block after homogeneous block fusion. For example, the above-mentioned execution entity can perform homogeneous block fusion on any two adjacent sparse sub-blocks of the same type. The attribute information of any two adjacent sparse sub-blocks of the same type is respectively [sparsity rate: 0.84, block type: sparse type, coordinates (256, 0), block shape (256×256)]; [sparsity rate: 0.84, block type: sparse type, coordinates (256, 256), block shape (256×256)]. It can be obtained that some attribute information of the sparse block after homogeneous block fusion obtained by performing homogeneous block fusion on any two adjacent sparse sub-blocks of the same type includes [sparsity rate: 0.84, coordinates (256, 0), block shape (256×512)]. According to the above-mentioned partial attribute information of the sparse block after homogeneous block fusion, combined with the above-mentioned sparse operator performance model, the type and predicted execution time of the obtained sparse block after homogeneous block fusion can be further obtained. The specific implementation method can refer to the steps of generating the attribute information of each sparse sub-block based on the sparse operator performance model in step 103, which will not be elaborated here. Integrate the above-mentioned partial attribute information with the type and predicted execution time of the obtained sparse block after homogeneous block fusion into the attribute information of the sparse block after homogeneous block fusion.
[0045] Sub-step 2: In response to the attribute information of the sparse block after homogeneous block fusion satisfying the homogeneous block fusion condition, determine the sparse block after homogeneous block fusion as the target sparse block after homogeneous block fusion. In practice, when the type of the sparse block after homogeneous block fusion is the same as the types of the two adjacent sparse sub-blocks of the same type, the sparse block after homogeneous block fusion can be determined as the target sparse block after homogeneous block fusion.
[0046] Step 3: Perform non - homogeneous block fusion on adjacent sparse sub - blocks of different types in each of the above sparse sub - blocks to obtain sparse blocks after non - homogeneous block fusion. Among them, the above non - homogeneous block fusion generally refers to integrating two adjacent sparse sub - blocks of different types in the split sparse weight matrix into one matrix, that is, the sparse block after non - homogeneous block fusion. The attribute information of each sparse block after non - homogeneous block fusion is also determined by the attribute information of the two adjacent sparse sub - blocks corresponding to each sparse block after non - homogeneous block fusion. In practice, the above - mentioned execution entity can merge any two adjacent sparse sub - blocks of different types in the split sparse weight matrix to obtain sparse blocks after non - homogeneous block fusion.
[0047] Step 4: For each sparse block after non - homogeneous block fusion obtained, perform the following steps:
[0048] Sub - step 1: Generate the attribute information of the sparse block after non - homogeneous block fusion according to the attribute information of the adjacent sparse sub - blocks corresponding to the sparse block after non - homogeneous block fusion. Among them, the attribute information of the sparse block after non - homogeneous block fusion includes the sparsity rate, shape, type, coordinates, and predicted execution time of the sparse block after non - homogeneous block fusion. In practice, the above - mentioned execution entity can obtain the sparsity rate of the sparse block after non - homogeneous block fusion, and obtain the shape and coordinates of the sparse block after non - homogeneous block fusion according to the shapes and coordinates of the two sparse sub - blocks corresponding to the sparse block after non - homogeneous block fusion. The specific implementation method can refer to the steps of generating partial attribute information of the sparse block after homogeneous block fusion above, and will not be elaborated here. When the sparsity rate of the sparse block after non - homogeneous block fusion and the shape of the sparse block after non - homogeneous block fusion are known, the type and predicted execution time of the sparse block after non - homogeneous block fusion can be obtained through the above - mentioned sparse operator performance model. The specific implementation method can refer to the steps of generating the attribute information of each sparse sub - block based on the sparse operator performance model in step 103 above, and will not be elaborated here. Finally, the sparsity rate, shape, type, coordinates, and predicted execution time of the sparse block after non - homogeneous block fusion can be integrated into the attribute information of the sparse block after non - homogeneous block fusion.
[0049] Sub - step 2: In response to the attribute information of the sparse block after non - homogeneous block fusion satisfying the non - homogeneous block fusion condition, determine the sparse block after non - homogeneous block fusion as the target sparse block after non - homogeneous block fusion. In practice, when the type of the sparse block after non - homogeneous block fusion is the same as the type of the sparse sub - block with a higher sparsity rate among the two sparse sub - blocks of different types corresponding to the sparse block after non - homogeneous block fusion, the above - mentioned execution entity can determine the sparse block after non - homogeneous block fusion as the target sparse block after non - homogeneous block fusion. For example, when a sparse sub - block of the swing type is fused with a sparse sub - block of the sparse type and the type of the sparse block after non - homogeneous block fusion is the sparse type, the sparse block after non - homogeneous block fusion is determined as the target sparse block after non - homogeneous block fusion.
[0050] Step 5: Generate a fused sparse weight matrix based on the sparse blocks after fusing the above homogeneous blocks and the sparse blocks after fusing the above heterogeneous blocks. In practice, the above execution entity may integrate the respective sparse sub-blocks corresponding to the above target fused sparse blocks and the fused sparse blocks that do not meet the above preset fusion conditions into a fused sparse weight matrix according to the corresponding attribute information.
[0051] Optionally, the above execution entity may also perform the following steps:
[0052] First step: Obtain the attribute information of each sparse block in the above fused sparse weight matrix. Among them, each sparse block in the above fused sparse weight matrix includes each target fused sparse block and the above un-fused sparse sub-blocks. The attribute information of each sparse block includes the sparsity rate, shape, type, coordinates, and predicted execution time of each sparse block in the above fused sparse weight matrix. In practice, the above execution entity may use the above sparse operator performance model to extract the attribute information of each sparse block in the above fused sparse weight matrix according to the shape and sparsity rate of each sparse block. The specific implementation method may refer to the steps of generating the attribute information of each sparse sub-block based on the sparse operator performance model in the above step 103, which will not be elaborated here.
[0053] Second step: According to the attribute information of each sparse block in the above fused sparse weight matrix, perform format encoding on each matrix in the above fused sparse weight matrix. Among them, the above format encoding generally refers to switching each matrix in the above fused sparse weight matrix to a coding format that occupies less storage. In practice, the above execution entity may encode each sparse block with a sparse type in the attribute information using a row or column compression method (such as CSR or CSC) or a block compression method (such as Block-CSR or ELLPACK). Each sparse block in the above fused sparse weight matrix with a dense type in the attribute information may be encoded using a native array method or quantization compression (such as INT8 quantization). Each sparse block in the above fused sparse weight matrix with a swaying type in the attribute information may be encoded using a hybrid encoding, that is, generating two copies of sparse and dense formats at the same time to facilitate subsequent adaptation to multiple computing units.
[0054] Step 105: In the operator layer intermediate representation, map each sparse block to a computing unit according to the corresponding attribute information of each sparse block to obtain a mapping result.
[0055] In some embodiments, the above-mentioned execution entity may map each of the above-mentioned sparse blocks to a computing unit in the above-mentioned operator layer intermediate representation according to the attribute information corresponding to each of the above-mentioned sparse blocks, and obtain a mapping result. Among them, the above-mentioned computing unit is usually a pre-set functional module for performing computing tasks. The above-mentioned computing unit includes a dense computing unit and a sparse computing unit. The above-mentioned mapping result is the mapping relationship information between each sparse block and the computing unit. Among them, the above-mentioned dense computing unit is usually a computing module for processing dense matrix calculations. The above-mentioned sparse computing unit is usually a computing module for processing sparse matrix calculations. The above-mentioned mapping relationship information includes mapping sparse blocks of the sparse type to the sparse computing unit, mapping sparse blocks of the dense type to the dense computing unit, and mapping sparse blocks of the sway type to the computing unit with a smaller load among the two computing units. The load of the above-mentioned computing unit can be represented by the predicted execution time of each sparse block on the computing unit. In practice, the above-mentioned execution entity may adjust the mapping relationship between each sparse block and the computing unit, optimizing the total load of the computing unit.
[0056] In some optional implementation manners of some embodiments, the above-mentioned execution entity may map each of the above-mentioned sparse blocks to a computing unit in the above-mentioned operator layer intermediate representation according to the attribute information corresponding to each sparse block through the following steps, and obtain a mapping result:
[0057] Step 1, for each of the above-mentioned sparse blocks, perform the following steps:
[0058] Sub-step 1, in response to the type characterized by the attribute information of the above-mentioned sparse block being the sparse type, map the above-mentioned sparse block to the sparse computing unit. Among them, the above-mentioned sparse computing unit is usually a computing module that supports direct calculation in a compressed format (such as CSR or Block-CSR) or zero-value skipping. In practice, the above-mentioned execution entity may map the above-mentioned sparse block with the sparse type in the attribute information to the above-mentioned sparse computing unit (such as the Sparse Tensor Core of NVIDIA A100, the Intel DLBoost sparse instruction set, or Google Sparsity Core).
[0059] Sub-step 2, in response to the type characterized by the attribute information of the above-mentioned sparse block being the dense type, map the above-mentioned sparse block to the dense computing unit. Among them, the above-mentioned dense computing unit is usually a computing module that maximizes the computing throughput through continuous memory access and vectorized computing (such as SIMD). In practice, the above-mentioned execution entity may map the above-mentioned sparse block with the dense type in the attribute information to the above-mentioned dense computing unit (such as the CUDA Core of the GPU, the AVX-512 unit of the CPU, or the MXU of the TPU).
[0060] Sub-step 3: In response to the fact that the attribute information representation type of the above sparse block is the swing type, map the above sparse block to a computing unit that meets the preset load condition. Among them, the above preset load condition is usually the computing unit with the lower load among the sparse computing unit load and the dense computing unit load. The computing unit load can be the sum of the predicted execution times of the sparse blocks mapped to the computing unit as the computing unit load. In practice, when mapping the above sparse block with the swing type in the mapping attribute information, the above execution entity dynamically identifies the current sparse computing unit load and the dense computing unit load. Then, the above sparse block with the swing type in the attribute information can be mapped to one of the computing units with the lower load among the above sparse computing unit and the above dense computing unit.
[0061] Step 2: In response to mapping the computing units simultaneously for the sparse weight matrices after fusion between adjacent batches on the computing unit, rearrange the mapping order of each sparse block in the sparse weight matrices after fusion between adjacent batches with respect to the computing unit to obtain inter-matrix adjustment information. Among them, the above-mentioned inter-matrix adjustment information can be the mapping order of each sparse block and the computing unit in the merged matrix. In practice, since the sparse weight matrices after fusion are used for adjacent batches during task generation, and in each batch, each sparse block in each sparse weight matrix after fusion is sequentially mapped to the computing unit. The above-mentioned execution entity can merge the sparse weight matrices after fusion in two batches. And adjust the mapping order of each sparse block in each sparse weight matrix after fusion with respect to the computing unit. So that when each sparse sub-block in the two batches is sequentially mapped, the load difference between each mapping to the computing unit is kept small, and the inter-matrix adjustment information is obtained. For example, name two identical sparse weight matrices after fusion as Matrix A and Matrix B. Among them, the above-mentioned two identical sparse weight matrices after fusion both include 4 sparse blocks. It is known that the first two sparse blocks (Sparse Block 1, Sparse Block 2) in Matrix A are mapped to the dense computing unit, the computing load of Sparse Block 1 is high, and the computing load of Sparse Block 2 is low. The last two sparse blocks (Sparse Block 3, Sparse Block 4) are mapped to the sparse computing unit, the computing load of Sparse Block 3 is high, and the computing load of Sparse Block 4 is low. Correspondingly, the first two sparse blocks (Sparse Block 5, Sparse Block 6) in Matrix B are mapped to the dense computing unit, the computing load of Sparse Block 5 is high, and the computing load of Sparse Block 6 is low. The last two sparse blocks (Sparse Block 7, Sparse Block 8) are mapped to the sparse computing unit, the computing load of Sparse Block 7 is high, and the computing load of Sparse Block 8 is low. After fusion, since Sparse Block 1 and Sparse Block 5 will be processed simultaneously, it will cause the computing unit load to be too high. The order of the sparse blocks in Matrix B can be adjusted so that Sparse Block 1 and Sparse Block 6 are processed together on the dense computing unit, Sparse Block 2 and Sparse Block 5 are processed on the dense computing unit, Sparse Block 3 and Sparse Block 8 are processed on the sparse computing unit, and Sparse Block 4 and Sparse Block 7 are processed on the sparse computing unit, ensuring that the load difference between each mapping to the computing unit is small.
[0062] Step 3: Determine the mapping relationship information, the above-mentioned inter-matrix adjustment information, and the attribute information of each sparse block as the mapping result. Among them, the above-mentioned mapping result is a set of sparse computing resource scheduling strategies and execution process optimization schemes. The above-mentioned sparse computing resource scheduling strategy can be to reduce the utilization rate difference by dynamically mapping the computing unit. The above-mentioned execution process optimization scheme can be to rearrange the sparse blocks across matrices to reduce the over-high load of the technical unit. In practice, the above-mentioned execution entity can use the sparse computing resource scheduling strategy and the execution process optimization scheme as the mapping result.
[0063] Step 106: Generate a processing task corresponding to the sparse model according to the mapping result, and execute the processing task.
[0064] In some embodiments, the above-mentioned execution entity may generate a processing task corresponding to the above-mentioned sparse model according to the above-mentioned mapping result, and execute the above-mentioned processing task.
[0065] In some optional implementation manners of some embodiments, the above-mentioned execution entity may generate a processing task corresponding to the above-mentioned sparse model according to the above-mentioned mapping result, and execute the above-mentioned processing task through the following steps:
[0066] First step: Insert a block-level intermediate representation of the compiler stack between the layer and operator layers of the compiler stack. Among them, the block-level intermediate representation of the compiler stack is used to carry information and format conversion. The information carried by the block-level intermediate representation of the compiler stack includes the above-mentioned sparse block execution time information, the above-mentioned attribute information corresponding to each sparse block, and the above-mentioned mapping result. In practice, first, the above-mentioned execution entity may divide the input sparse matrix into blocks of 128×128. And record the sparsity rate of each sparse block (such as 92%, indicating that 92% of the elements in this sparse block are zero), execution time prediction (for example, it takes 1.2 milliseconds to process using a sparse computing core and 2.1 milliseconds for a dense core), and hardware mapping strategy (such as allocating this sparse block to a dedicated sparse computing unit of the GPU). Through the newly added BlockLevelPass module, extract the attributes of these blocks (such as shape, non-zero element distribution) from the layer, and use the above-mentioned sparse operator performance model to predict the processing time of different computing modes. Then, make a format decision according to the sparsity rate. For example, convert a sparse block with a high sparsity rate (such as >90%) from the COO format (storing non-zero elements' row, column, and value through triples) to the CSR format (optimizing storage through compressed row pointers and column indices).
[0067] Second, in the block-level intermediate representation in the above compiler stack, adjust the shapes of the above-mentioned fused sparse blocks according to the hardware characteristic information to obtain the fused sparse blocks with adjusted shapes. Among them, the above-mentioned hardware characteristic information is usually the physical constraints and performance characteristics of the target hardware. The above-mentioned hardware characteristic information may include warp, memory alignment, and shared memory capacity. The above-mentioned fused sparse block usually combines multiple adjacent small sparse blocks into a larger block. The above-mentioned shape adjustment usually refers to modifying the dimensions of the block or padding zero elements according to hardware constraints. The above-mentioned shape adjustment may include column alignment and row alignment. In practice, in the block-level intermediate representation of the TVM compiler, according to the hardware characteristic information (such as the warp size of 32 and the memory alignment requirement of 128 bytes for NVIDIA A100 GPU), perform the above-mentioned shape adjustment on the above-mentioned fused sparse blocks (such as combining 4 adjacent 64×64 blocks into a 128×128 block). If the actual non-zero data area of any fused block is 128 rows × 93 columns, pad zero elements to 128×128 to align the block size with the hardware memory access granularity. At the same time, update the index array in CSR format (compressed sparse row format). The virtual index of the padded column can be included by expanding the col_idx array (such as marking the padded column as -1), and the offset of the row_ptr array can be adjusted to reflect the row boundary after padding. Finally, output the fused sparse blocks with adjusted shapes. The number of block columns can also be aligned with the hardware access granularity (such as 128 columns) through shape adjustment to ensure that each thread accesses adjacent elements to ensure that all threads within the same warp access consecutive and aligned memory addresses.
[0068] Step 3: In the operator layer of the above compiler stack, combine the above sparse block execution time information and the attribute information corresponding to each of the above sparse blocks to generate dynamic scheduling information for each fused sparse block with adjusted shapes. Among them, the above dynamic scheduling information includes a sparse block priority queue, a computing unit status monitoring task, and a real-time load feedback task. The above sparse block priority queue is usually a task queue sorted according to optimization objectives, which determines the calculation order of blocks and the hardware allocation priority. The sorting rule of the above sparse block priority queue can be to preferentially process sparse blocks with high sparsity rates and low execution times. The above computing unit status monitoring task is usually a background task that periodically collects hardware resource usage data and can provide real-time input for dynamic scheduling. The monitoring metrics of the above computing unit status monitoring task can include streaming multiprocessor utilization, video memory bandwidth (memory access throughput), and GPU temperature. The above streaming multiprocessor is usually the core computing unit of the GPU, which includes multiple CUDA Cores, shared memory, and a scheduler. The above real-time load feedback task is usually an algorithm module that dynamically adjusts task allocation according to monitoring data. When the above real-time load feedback task detects that the utilization rate of a certain streaming multiprocessor is > 90% for 3 consecutive cycles, it migrates 2 sparse blocks (such as Block_3 and Block_5) in its queue to an idle streaming multiprocessor (such as an idle streaming multiprocessor with a utilization rate of 45%). When the actual execution time of the above sparse block exceeds the predicted value by 20% (such as a predicted value of 1.2 ms and an actual value of 1.5 ms), the above real-time load feedback task reduces the priority of its subsequent similar blocks by one level. In the operator layer of the TVM compiler, a sparse block priority queue is constructed, a computing unit status monitoring task is started, and a real-time load feedback task is executed for the above block-level intermediate representation (including a 128×128 block shape, CSR format, sparsity rate of 92%, and predicted execution time of 1.2 ms / dense core of 2.1 ms) to obtain dynamic scheduling information.
[0069] In the fourth step, sparse computing cores and dense computing cores are generated based on the above dynamic scheduling information. Among them, the above sparse computing cores can be kernel functions optimized for sparse formats (such as CSR). The above dense computing cores can be kernel functions optimized for dense data, and use vectorized instructions (such as SIMD) to batch process continuous memory data. In practice, the above execution entity performs index parsing, multiplication and addition calculations, and hardware optimization on the sparse blocks (row_ptr, col_idx, data) in CSR format and the dense input matrix to obtain sparse computing cores. Among them, the above hardware optimization usually refers to using the Tensor Core of the GPU (supporting structured sparse computing) and shared memory to cache frequently accessed data. The above shared memory usually refers to the on-chip high-speed cache shared by GPU threads within the sparse block, which is used to reduce the global memory access latency (such as caching the frequently accessed part of the input matrix). Then, vectorized calculations and memory optimization are performed on the dense format blocks (continuous memory arrays) and the input matrix to obtain dense computing cores. Among them, the above vectorized calculation usually refers to using the WMMA (Warp Matrix MultiplyAccumulate) instruction of the CUDA Core to batch process 128×128 matrix multiplications. The above memory optimization usually refers to reducing the global memory access by caching the input matrix blocks through shared memory.
[0070] Step 5: Through the inter-core communication interface, perform instruction-level interleaving on the index processing unit of the above sparse computing core and the vectorized computing unit of the above dense computing core to form a collaborative computing core. Among them, the above inter-core communication interface is usually a mechanism for exchanging data between different computing cores in a GPU, which includes shared memory, global memory atomic operations, semaphores, etc. The above index processing unit usually refers to the module in the sparse core that parses the sparse format (such as row_ptr and col_idx of CSR), and is used to locate the positions of non-zero elements. The above vectorized computing unit is usually a module in the dense core that batch processes data using SIMD instructions (such as WMMA of Tensor Core). The above instruction-level interleaving usually refers to coordinating the timing of different computing units through synchronization instructions (such as __syncthreads()) so that the two execute alternately to form a pipeline. The above collaborative computing core is usually a hybrid kernel function that integrates sparse index parsing and dense computing. The above atomic operation usually refers to an uninterruptible read and write operation (such as atomicAdd for accumulating semaphores). The above semaphore usually refers to a synchronization flag variable (such as updating the semaphore after the sparse core finishes writing data). In practice, the above execution entity can, on an NVIDIA A100 GPU, through the inter-core communication interface of CUDA (such as shared memory and atomic operations), perform instruction-level interleaving on the index processing unit of the sparse computing core (parsing row_ptr and col_idx of the CSR format) and the vectorized computing unit of the dense computing core (using WMMA instructions of Tensor Core) to form a unified collaborative computing core.
[0071] Step 6: Generate the execution time of the above collaborative computing core corresponding to the above processing task according to the above collaborative computing core. In practice, the above execution entity can, through the above sparse operator performance model and hardware actual measurement calibration (using nvprof to collect the number of GPU instruction cycles and memory transactions), combined with task parameters (sparse rate 95%, block size 128×128) and hardware configuration (such as 108 SMs of A100, 1555GB / s memory bandwidth), obtain the execution time of the collaborative computing core. For example, first superimpose the theoretical time consumption of each stage (such as 0.3ms parsing + 0.15ms loading + 0.2ms computing = 0.65ms), and then correct according to the parallelism (such as the GPU processes 4 blocks simultaneously, and the equivalent single-block time consumption is 0.65ms / 4≈0.16ms), and finally output the calibrated execution time (such as the actual measurement is 0.17ms±5% error).
[0072] Step 7: Generate visualization information according to the above execution time. In practice, the above execution entity can sort out the execution time of the above collaborative computing core corresponding to the above processing task, and collect the speedup ratios of each model and this method at different sparse rates, and use visualization tools (such as Matplotlib or Seaborn) to generate visualization information (such as icons or curves).
[0073] In the eighth step, the above visualization information is displayed. In practice, the above execution entity can use Matplotlib or Seaborn through a Python library to generate a static chart and perform color coding on the visualization information. The static chart is then displayed.
[0074] The above first step to the eighth step are an inventive point of the embodiments of the present disclosure, which solves the technical problem of "existing solutions (such as TVM-Sparse) only support static sparse format compilation, lack the ability of runtime sparsity rate awareness and computational paradigm switching, resulting in the inability to achieve end-to-end optimization from high-level model sparsification to low-level hardware instructions". The difficulties in meeting the requirements of efficient sparse matrix calculation in the prior art are as follows: Existing sparse matrix calculation solutions have deficiencies in dynamic sparsity rate awareness, flexible switching of computational paradigms, and implementation of end-to-end optimization, making it difficult to effectively adapt to data distributions with different sparsity rates. At the same time, there are defects in hardware resource utilization and load balancing, resulting in low computational efficiency and the inability to fully meet the requirements of efficient sparse matrix calculation. If the above factors are solved, the effects of improving the computational efficiency of sparse matrices, optimizing hardware resource utilization, and achieving end-to-end performance optimization can be achieved. To achieve this effect, the present disclosure adopts a sparse matrix calculation method based on runtime sparsity rate awareness and computational paradigm switching, inserts block-level intermediate representations to carry the execution time information and attribute information of sparse blocks; adjusts the shape of fused sparse blocks according to hardware characteristics; generates dynamic scheduling information by combining the execution time information and attribute information of sparse blocks; generates sparse computing cores and dense computing cores based on the dynamic scheduling information; forms a cooperative computing core through the instruction-level interleaving of sparse computing cores and dense computing cores via an inter-core communication interface; generates the execution time of the processing tasks corresponding to the cooperative computing core; generates visualization information and displays it. Thus, end-to-end optimization from high-level model sparsification to low-level hardware instructions can be achieved. Thereby, it can effectively sense the runtime sparsity rate, flexibly switch computational paradigms, optimize all aspects of sparse matrix calculation, improve computational efficiency and hardware resource utilization rate, and thus meet the requirements of efficient sparse matrix calculation.
[0075] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the acceleration method based on the irregular sparse model of some embodiments of the present disclosure, dynamic resource scheduling and end-to-end collaborative optimization are achieved, significantly improving the computing efficiency of the sparse model. Specifically, the traditional sparse acceleration methods often have the following problems: They do not model the performance of sparse sub-blocks in different sparse rate regions, resulting in the inability to accurately identify sub-blocks suitable for sparse or dense computing in the irregular distribution scenario, causing waste of computing resources; They use a unified format for the irregular distribution matrix, resulting in low memory efficiency due to redundant storage of sparse sub-blocks; They map the entire sparse matrix to a single type of computing unit, resulting in different types of sub-blocks in the sparse matrix not being mapped to the corresponding computing units, causing low performance improvement and waste of resources; They do not consider the cross-batch load correlation, resulting in a decrease in the end-to-end inference throughput. Based on this, in the acceleration method based on the irregular sparse model of some embodiments of the present disclosure, first, the sparse model is input into the target compiler, where the above-mentioned target compiler includes a layer intermediate representation, a block layer intermediate representation, and an operator layer intermediate representation. Thus, the sparse model can be used as a data source. Next, in the above-mentioned layer intermediate representation, the sparse weight matrix of the above-mentioned sparse model is split into sparse sub-blocks to obtain the split sparse weight matrix, where the above-mentioned split sparse weight matrix includes each sparse sub-block. Thus, the sparse weight matrix in the above-mentioned sparse model can be divided for subsequent processing. Then, based on the sparse operator performance model, the attribute information of each sparse sub-block in the above-mentioned split sparse weight matrix is generated. Thus, by modeling the performance of sparse sub-blocks in different sparse rate regions, a sparse operator performance model can be obtained to facilitate predicting the attribute information of each sparse sub-block and reducing waste of computing resources. Then, in the above-mentioned block layer intermediate representation, according to the attribute information of each sparse sub-block, the above-mentioned sparse sub-blocks are fused to obtain the fused sparse weight matrix, where each sparse block in the above-mentioned fused sparse weight matrix includes each target fused sparse block and each unfused sparse sub-block, and each of the above-mentioned target fused sparse blocks corresponds to attribute information. Thus, the above-mentioned sparse sub-blocks are fused and the format encoding is converted to facilitate reducing the redundant storage of the above-mentioned fused sparse weight matrix and thus improving the memory efficiency. Then, in the above-mentioned operator layer intermediate representation, according to the attribute information corresponding to each sparse block, the above-mentioned sparse blocks are mapped to the computing units to obtain the mapping result. Thus, the above-mentioned sparse blocks can be mapped to different computing units according to their types to obtain the mapping result, improving the overall performance improvement and reasonably allocating resources. Finally, according to the above-mentioned mapping result, a processing task corresponding to the above-mentioned sparse model is generated, and the above-mentioned processing task is executed. Thus, the above-mentioned mapping result and the above-mentioned sparse operator performance model can be used to optimize the model and process the task, thereby reflecting the acceleration effect.In addition, when processing computing tasks across batches, middle-high load sparse sub-blocks and low load sparse sub-blocks in adjacent sparse weight matrices can be cross-mapped to different computing units. Therefore, the load correlation between different batches can be improved, and the throughput of end-to-end inference can be increased. As a result, it can adapt to different sparse patterns and hardware environments, showing significant advantages in large model inference scenarios, providing an efficient sparse acceleration solution for edge computing and large language model deployment, significantly enhancing the applicability of the model in mobile and cloud inference scenarios, and breaking through the benefit boundary of traditional unstructured sparse acceleration design.
[0076] Further refer to Figure 4 , which is a multi-dimensional performance comparison bar chart according to an embodiment of the present disclosure.
[0077] As Figure 4As shown, it presents the performance comparison and analysis results of different sparse acceleration algorithms in large language model inference tasks. From the figure, it can be seen that this group of bar charts compares the acceleration effects of 7 acceleration algorithms (including cuBLAS, SparseTIR, and the method "Ours" proposed by the applicant) for 5 mainstream models (Bert-base, Bert-large, GPT-2, Llama-2, Bart-large) at sparsity rates of 50% - 90% through a multi-dimensional arrangement. At the top of the chart, the corresponding sparsity rate is clearly marked in the title of each sub-chart (such as "90% Sparsity"). The five horizontally arranged sub-charts intuitively show the differences in the speedup ratios of different models at the same sparsity rate. The X-axis of each sub-chart lists the specific model names such as BERT, GPT-2, etc., and the Y-axis uniformly uses "Speedup Ratio" to quantify the acceleration effect, with the baseline being the original un-accelerated performance (speedup ratio of 1.0 times). Different algorithm types are distinguished by bar bodies of different colors, and the "Ours" method proposed by the researcher is specifically marked with a red bar body, forming a visual contrast focus. From the specific data distribution, in the scenario of 90% high sparsity rate (the leftmost sub-chart), the "Ours" method reaches a peak speedup of 3.75 times on the Llama-2 model, and its red bar body is significantly higher than the adjacent SparseTIR (blue, speedup ratio of about 2.8 times) and cuBLAS (gray, speedup ratio of about 1.5 times). As the sparsity rate decreases to 50% (the rightmost sub-chart), the speedup ratios of all algorithms show a downward trend. For example, the speedup ratio of "Ours" on Llama-2 drops to 2.3 times, but it still maintains a performance advantage over other algorithms. It is particularly worth noting that the traditional algorithm cuBLAS (gray bar body) has a speedup ratio always lower than 2.0 times at each sparsity rate, and there is a negative optimization phenomenon where the speedup ratio is less than 1.0 on large models such as Bert-large. The influence of the model architecture on the acceleration effect is revealed through vertical comparison in the chart: Llama-2 based on the Transformer-XL architecture (the rightmost model in each sub-chart) shows the highest speedup ratio at all sparsity rates, while the acceleration gain of the basic BERT model (Bert-base) is relatively limited. This difference may stem from the synergistic effect between the large model parameter quantity and sparse computing optimization. In addition, the legend box in the upper right corner of the task bar clearly marks the color coding of the 7 algorithms, which helps to quickly identify the performance distribution characteristics of different algorithm series.
[0078] Generally speaking, this group of comparison figures systematically verifies the effectiveness of the sparse acceleration algorithm in LLM inference optimization. By means of hierarchical visualization, it reveals the laws of algorithm performance varying with model scale and sparsity rate, providing data support for the joint optimization of algorithm selection and sparse strategy in actual deployment. For example, in the 50% sparse scenario with high-precision requirements, the "SparseTIR + Ours" combination scheme can be preferentially selected, while in the 90% sparse scenario with high compression rate, the "Ours" method alone can achieve optimal acceleration.
[0079] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an acceleration device based on an irregular sparse model. These device embodiments correspond to Figure 1 the method embodiments shown, and the device can be specifically applied to various electronic devices.
[0080] As Figure 2 shown, some embodiments of the acceleration device 200 based on an irregular sparse model include: an input unit 201, a splitting unit 202, a first generating unit 203, a fusion unit 204, a mapping unit 205, and a second generating unit 206. Among them, the input unit 201 is configured to input a sparse model into a target compiler, where the target compiler includes a layer intermediate representation, a block layer intermediate representation, and an operator layer intermediate representation; the splitting unit 202 is configured to split the sparse weight matrix of the sparse model into sparse sub-blocks in the layer intermediate representation to obtain a split sparse weight matrix, where the split sparse weight matrix includes each sparse sub-block; the first generating unit 203 is configured to generate attribute information of each sparse sub-block in the split sparse weight matrix based on a sparse operator performance model; the fusion unit 204 is configured to perform a fusion process on each sparse sub-block according to the attribute information of each sparse sub-block in the block layer intermediate representation to obtain a fused sparse weight matrix, where each sparse block in the fused sparse weight matrix includes each target fused sparse block and each unfused sparse sub-block, and each target fused sparse block corresponds to attribute information; the mapping unit 205 is configured to map each sparse block to a computing unit according to the attribute information corresponding to each sparse block in the operator layer intermediate representation to obtain a mapping result; the second generating unit 206 is configured to generate a processing task corresponding to the sparse model according to the mapping result and execute the processing task.
[0081] It can be understood that the units described in the device 200 and referring to Figure 1Correspond to each step in the described method. Thus, the operations, features, and beneficial effects described above for the method also apply to the apparatus 200 and the units included therein, and will not be elaborated herein.
[0082] Reference is made below to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0083] As Figure 3 shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0084] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 shows the electronic device 300 having various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had. Figure 3 Each block shown in
[0085] Specifically, according to some embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0086] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0087] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0088] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: input a sparse model into a target compiler, where the above target compiler includes a layer intermediate representation, a block layer intermediate representation, and an operator layer intermediate representation; in the above layer intermediate representation, split the sparse weight matrix of the above sparse model into sparse sub-blocks to obtain a split sparse weight matrix, where the above split sparse weight matrix includes each sparse sub-block; generate attribute information of each sparse sub-block in the above split sparse weight matrix based on a sparse operator performance model; in the above block layer intermediate representation, perform a fusion process on each of the above sparse sub-blocks according to the attribute information of each of the above sparse sub-blocks to obtain a fused sparse weight matrix, where each sparse block in the above fused sparse weight matrix includes each target fused sparse block and each unfused sparse sub-block, and each of the above target fused sparse blocks corresponds to attribute information; in the above operator layer intermediate representation, map each of the above sparse blocks to a computing unit according to the attribute information corresponding to each of the above sparse blocks to obtain a mapping result; generate a processing task corresponding to the above sparse model according to the above mapping result, and execute the above processing task.
[0089] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0091] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an input unit, a splitting unit, a first generating unit, a fusing unit, a mapping unit, and a second generating unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the input unit can also be described as "the unit that inputs a sparse model into a target compiler".
[0092] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0093] Some embodiments of the present disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the above acceleration methods based on an irregular sparse model.
[0094] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An acceleration method based on an irregular sparse model, comprising: Inputting the sparse model into a target compiler, where the target compiler includes a layer intermediate representation, a block layer intermediate representation, and an operator layer intermediate representation; In the layer intermediate representation, splitting the sparse weight matrix of the sparse model into sparse sub-blocks to obtain a split sparse weight matrix, where the split sparse weight matrix includes each sparse sub-block; Generating attribute information of each sparse sub-block in the split sparse weight matrix based on a sparse operator performance model; In the block layer intermediate representation, performing a fusion process on each sparse sub-block according to the attribute information of each sparse sub-block to obtain a fused sparse weight matrix, where each sparse block in the fused sparse weight matrix includes each target fused sparse block and each unfused sparse sub-block, and each target fused sparse block corresponds to attribute information; In the operator layer intermediate representation, mapping each sparse block to a computing unit according to the attribute information corresponding to each sparse block to obtain a mapping result; Generating a processing task corresponding to the sparse model according to the mapping result, and executing the processing task.
2. The method according to claim 1, wherein The generating attribute information of each sparse sub-block in the split sparse weight matrix based on the sparse operator performance model includes: Obtaining actual execution time data of matrix multiplication of each candidate matrix at each sparsity rate; Generating a sparse operator performance model based on the actual execution time data; Generating attribute information of each sparse sub-block according to the sparse operator performance model.
3. The method according to claim 2, wherein, The generating a sparse operator performance model based on the actual execution time data includes: Generating a fitted execution time of each candidate matrix based on the actual execution time data; Generating a threshold of each candidate matrix based on the fitted execution time of each candidate matrix; Generating a sparse operator performance model according to the threshold of each candidate matrix and the fitted execution time of each candidate matrix.
4. The method according to claim 1, wherein, The performing a fusion process on each sparse sub-block according to the attribute information of each sparse sub-block in the block layer intermediate representation to obtain a fused sparse weight matrix includes: Performing homogeneous block fusion on adjacent sparse sub-blocks of the same type in each sparse sub-block according to the attribute information of each sparse sub-block to obtain each homogeneous block fused sparse block; For each obtained homogeneous block fused sparse block, perform the following steps: Generating attribute information of the homogeneous block fused sparse block according to the attribute information of the adjacent sparse sub-blocks corresponding to the homogeneous block fused sparse block; In response to the attribute information of the homogeneous block fused sparse block satisfying the homogeneous block fusion condition, determining the homogeneous block fused sparse block as a target homogeneous block fused sparse block; performing non-homogeneous block fusion on adjacent sparse sub-blocks of different types in each sparse sub-block to obtain each non-homogeneous block fused sparse block; For each obtained non-homogeneous block fused sparse block, perform the following steps: Generate the attribute information of the sparse block after heterogeneous block fusion according to the attribute information of the adjacent sparse sub-blocks corresponding to the sparse block after heterogeneous block fusion; In response to the attribute information of the sparse block after heterogeneous block fusion satisfying the heterogeneous block fusion condition, determine the sparse block after heterogeneous block fusion as the target sparse block after heterogeneous block fusion; Generate a fused sparse weight matrix according to the sparse blocks after each homogeneous block fusion and the sparse blocks after each heterogeneous block fusion.
5. The method according to claim 4, wherein The method further includes: Obtain the attribute information of each sparse block in the fused sparse weight matrix, where each sparse block in the fused sparse weight matrix includes each target fused sparse block and each unfused sparse sub-block; Perform format encoding on each sparse block in the fused sparse weight matrix according to the attribute information of each sparse block in the fused sparse weight matrix.
6. The method according to claim 1, wherein, In the intermediate representation of the operator layer, mapping each sparse block to a computing unit according to the attribute information corresponding to each sparse block, and obtaining a mapping result, including: For each sparse block in each of the sparse blocks, perform the following steps: In response to the attribute information representation type of the sparse block being the sparse type, map the sparse block to a sparse computing unit; In response to the attribute information representation type of the sparse block being the dense type, map the sparse block to a dense computing unit; In response to the attribute information representation type of the sparse block being the swing type, map the sparse block to a computing unit that satisfies a preset load condition; In response to mapping the fused sparse weight matrix between adjacent batches to the computing unit simultaneously on the computing unit, rearrange the mapping order of each sparse block in the fused sparse weight matrix between adjacent batches to the computing unit to obtain matrix adjustment information; Determine the mapping relationship information, the matrix adjustment information, and the attribute information of each sparse block as the mapping result.
7. An acceleration device based on an irregular sparse model, including: An input unit configured to input a sparse model into a target compiler, where the target compiler includes an intermediate representation of a layer, an intermediate representation of a block layer, and an intermediate representation of an operator layer; A splitting unit configured to split the sparse weight matrix of the sparse model into sparse sub-blocks in the intermediate representation of the layer to obtain a split sparse weight matrix, where the split sparse weight matrix includes each sparse sub-block; A first generating unit configured to generate the attribute information of each sparse sub-block in the split sparse weight matrix based on a sparse operator performance model; A fusion unit configured to perform fusion processing on each sparse sub-block according to the attribute information of each sparse sub-block in the intermediate representation of the block layer to obtain a fused sparse weight matrix, where each sparse block in the fused sparse weight matrix includes each target fused sparse block and each unfused sparse sub-block, and each target fused sparse block corresponds to attribute information; A mapping unit, configured to map each of the sparse blocks to a computing unit according to the attribute information corresponding to each sparse block in the intermediate representation of the operator layer, so as to obtain a mapping result; A second generating unit, configured to generate a processing task corresponding to the sparse model according to the mapping result, and execute the processing task.
8. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Sparse matrix vector multiplication parallel optimization method and system based on CSR storage format
CN119598081A
Sparse Matrix-Vector Multiplication on Graphics Processor Units
US20110078226A1
Accelerator for sparse-dense matrix multiplication
US20190042542A1
Apparatus for acquiring depth image, method for fusing depth images, and terminal device
US20230042846A1
Generating sparse neural networks
US20240152407A1
Cited By
Irregular matrix multiplication-oriented dual-core collaborative parallel computing method
CN122086547A