Method and device for generating calculation model for sparse matrix multiplication
By obtaining the location of all zero data blocks in sparse matrix multiplication and training a deep neural network to generate a computing model, the problem of high computational complexity of sparse matrix in the prior art is solved, and the calculation complexity is reduced.
Patent Information
- Application Number
- CN202510282859.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, when processing sparse matrices, the AlphaTensor algorithm fails to fully utilize the sparseness, resulting in high computational complexity.
By obtaining the location of all zero data blocks in the matrix dataset and combining deep neural network training, a computational model for sparse matrix multiplication is generated, taking into account the computational savings brought by sparseness.
Reduces the number of multiplications required for matrix multiplication and reduces the computational complexity.
Smart Images

Figure CN120256796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method and apparatus for generating a computational model for sparse matrix multiplication. Background Art
[0002] A matrix is an important basic concept in mathematics. An M×N matrix is a rectangular array formed by arranging M rows and N columns of elements. Currently, matrix multiplication is one of the very important operations in data calculation when a graphics processing unit (GPU) performs image processing, when input trajectory analysis is performed in handwriting recognition, and / or when input audio analysis is performed in speech recognition.
[0003] In the prior art, matrix multiplication includes the Strassen algorithm and the AlphaTensor algorithm. Among them, the Strassen algorithm, as a classic matrix multiplication optimization method, can reduce the computational complexity from O(n 3 ) to O(n 2.81 ) by reducing the number of multiplication operations, but its optimization effect still has limitations; the AlphaTensor algorithm adopts deep reinforcement learning technology, enhances the flexibility of algorithm discovery, and reduces the computational complexity to O(n 2.778 ). However, the AlphaTensor algorithm does not consider the characteristics of sparse matrices, so it may not be able to fully exert its advantages when dealing with sparse matrices. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0005] To this end, an object of the present invention is to propose a method for generating a computational model for sparse matrix multiplication. This method considers the computational savings brought by sparsity during the training process of the computational model through the sparse information of the first positions of all-zero data blocks in the first matrix and the second positions of all-zero data blocks in the second matrix in the matrix dataset, thereby reducing the number of multiplications required for matrix multiplication and reducing the computational complexity.
[0006] Another object of the present invention is to propose an apparatus for generating a computational model for sparse matrix multiplication.
[0007] To achieve the above object, an embodiment of one aspect of the present invention proposes a method for generating a computational model for sparse matrix multiplication, including:
[0008] Obtain a matrix dataset for training;
[0009] Determine the first positions of all-zero data blocks in the first matrix and the second positions of all-zero data blocks in the second matrix in the matrix dataset;
[0010] Update the network parameters of the initial deep neural network based on the first matrix, the second matrix, the first position, and the second position;
[0011] Repeat the above steps until the training stop condition is reached, obtain the target deep neural network, and determine the target deep neural network as the calculation model for sparse matrix multiplication.
[0012] The method for generating a calculation model for sparse matrix multiplication according to the embodiments of the present invention may further have the following additional technical features:
[0013] Further, the determining the first position of the all-zero data block in the first matrix and the second position of the all-zero data block in the second matrix in the matrix dataset includes:
[0014] Determine the target dimension for dividing the matrix into data blocks;
[0015] Based on the target dimension, divide the first matrix and the second matrix respectively to obtain corresponding first target data blocks and second target data blocks;
[0016] Determine the first position of the all-zero data block in the first target data block;
[0017] Determine the second position of the all-zero data block in the second target data block.
[0018] Further, the updating the network parameters of the initial deep neural network based on the first matrix, the second matrix, the first position, and the second position includes: updating the network parameters of the initial deep neural network based on the first target data block, the second target data block, the first position, and the second position.
[0019] Further, the initial deep neural network includes a policy network and a value network; the updating the network parameters of the initial deep neural network based on the first target data block, the second target data block, the first position, and the second position includes:
[0020] Determine a state set and an action set, and initialize the state score, the policy network, and the value network;
[0021] Randomly select an initial state from the state set;
[0022] Select a policy action through the policy network, and update the initial state to the current state based on the policy action;
[0023] Determine whether there is a new coefficient vector pair based on Monte Carlo tree search;
[0024] If there is a new pair of coefficient vectors, update the state score and update the value network;
[0025] Repeat the above steps until the multiplication operation of the first target data block and the second target data block is completed, and obtain the target state score;
[0026] Based on the target state score, update the network parameters of the policy network and the value network.
[0027] Further, the step of if there is a new pair of coefficient vectors, update the state score and update the value network includes:
[0028] Increment the state score by 1;
[0029] Determine whether the new pair of coefficient vectors corresponds to an all-zero data block;
[0030] If the new coefficient vector corresponds to an all-zero data block, decrement the state score by 1;
[0031] Update the value network.
[0032] Further, the stop training condition includes:
[0033] The initial deep neural network converges, and / or the number of iterations of the network parameters reaches a preset number.
[0034] To achieve the above object, another embodiment of the present invention proposes a generation device for a calculation model for sparse matrix multiplication, and the device includes:
[0035] An acquisition module, configured to acquire a matrix data set for training;
[0036] A first determination module, configured to determine a first position of an all-zero data block in a first matrix and a second position of an all-zero data block in a second matrix in the matrix data set;
[0037] An update module, configured to update the network parameters of the initial deep neural network based on the first matrix, the second matrix, the first position, and the second position;
[0038] A second determination module, repeats the above steps until the stop training condition is reached, obtains a target deep neural network, and determines the target deep neural network as a calculation model for sparse matrix multiplication.
[0039] The method and device for generating a calculation model for sparse matrix multiplication proposed by the present invention consider the sparse information of the first positions of all-zero data blocks in the first matrix and the second positions of all-zero data blocks in the second matrix in the matrix dataset, taking into account the computational savings brought by sparsity during the training process of the calculation model, thereby reducing the number of multiplications required for matrix multiplication and lowering the computational complexity.
[0040] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0041] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, in which:
[0042] Figure 1 is a flowchart of a method for generating a calculation model for sparse matrix multiplication according to an embodiment of the present invention;
[0043] Figure 2 is a schematic structural diagram of a device for generating a calculation model for sparse matrix multiplication according to an embodiment of the present invention. Detailed Embodiments
[0044] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0045] Among them, the Strassen algorithm reduces the number of operations required for 2×2×2 matrix multiplication to seven, thereby reducing the computational complexity from O(n 3 ) to O(n 2.81 ). On this basis, the AlphaTensor algorithm enhances the flexibility of algorithm discovery using deep reinforcement learning. Specifically, in the multiply-accumulate calculation process of the AlphaTensor algorithm, a 3D tensor Tn is defined to represent an n×n×n matrix multiplication, and Tn is decomposed into R first-order terms to obtain an algorithm that requires R scalar multiplications, and its form is as follows:
[0046]
[0047] where represents the outer product, and u(r), v(r), and w(r) are coefficient vector pairs representing the multiply-accumulate process, with a dimension of n×n×R.
[0048] Exemplarily, assume matrix multiplication C = A × B, where the dimensions of A, B, and C are all n × n. First, intermediate results m1 to mR are calculated through the elements of matrices A and B and vectors u and v, that is:
[0049]
[0050] where r = 1, …, R, a1 is an element of matrix A, and b1 is an element of matrix B.
[0051] Then, matrix C is obtained based on vectors m and u, that is:
[0052]
[0053] where i = 1, …, n 2 .
[0054] Based on the above description, the present invention retrains the AlphaTensor algorithm to support sparse matrix multiplication, proposes a method for generating a computational model for sparse matrix multiplication, integrates sparse matrix constraints into the AlphaTensor algorithm for retraining, thereby reducing the number of multiplications in the generated computational model.
[0055] First, a method for generating a computational model for sparse matrix multiplication according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0056] Figure 1 FIG. is a flowchart of a method for generating a computational model for sparse matrix multiplication according to an embodiment of the present invention.
[0057] As Figure 1 shown, the method for generating a computational model for sparse matrix multiplication includes the following steps:
[0058] Step S1, obtain a matrix data set for training;
[0059] In an embodiment of the present invention, the above matrix data set for training may be a large sparse matrix. And, in an embodiment of the present invention, the above matrix data set for training may also be the matrix dimensions corresponding to matrix multiplication with a known number of matrix multiplications, so as to be used for adjusting the network parameters of the subsequent model. Exemplarily, the matrix size corresponding to matrix multiplication is (3, 3, 3), and the known number of matrix multiplications for the corresponding optimal method is 23.
[0060] Step S2, determine the first position of the all-zero data block in the first matrix and the second position of the all-zero data block in the second matrix in the matrix data set;
[0061] Among them, in one embodiment of the present invention, when performing multiplication operations on a large-scale matrix, the large-scale matrix needs to be divided into data blocks for operations. And, in one embodiment of the present invention, when the large-scale matrix is a sparse matrix, it is necessary to determine the first position of the all-zero data blocks in the first matrix and the second position of the all-zero data blocks in the second matrix in the matrix dataset to reduce the number of required multiplications.
[0062] Specifically, in one embodiment of the present invention, the method for determining the first position of the all-zero data blocks in the first matrix and the second position of the all-zero data blocks in the second matrix in the matrix dataset may include the following steps:
[0063] Step S21, determine the target dimension for dividing the matrix into data blocks;
[0064] Step S22, based on the target dimension, divide the first matrix and the second matrix respectively to obtain the corresponding first target data blocks and second target data blocks;
[0065] Step S23, determine the first position of the all-zero data blocks in the first target data blocks;
[0066] Step S24, determine the second position of the all-zero data blocks in the second target data blocks.
[0067] Among them, in one embodiment of the present invention, the above-mentioned target dimension can be set as needed. For example, when the matrix dimension is 6×6, the corresponding target dimension is 3×3.
[0068] Step S3, update the network parameters of the initial deep neural network based on the first matrix, the second matrix, the first position, and the second position;
[0069] In one embodiment of the present invention, after determining the first position of the all-zero data blocks in the first matrix and the second position of the all-zero data blocks in the second matrix in the matrix dataset through the above steps, the network parameters of the initial deep neural network can be updated based on the first matrix, the second matrix, the first position, and the second position.
[0070] Among them, in one embodiment of the present invention, the method for updating the network parameters of the initial deep neural network based on the first matrix, the second matrix, the first position, and the second position may include: updating the network parameters of the initial deep neural network based on the first target data blocks, the second target data blocks, the first position, and the second position.
[0071] In one embodiment of the present invention, the above-mentioned initial deep neural network includes a policy network and a value network.
[0072] Specifically, in an embodiment of the present invention, the method for updating the network parameters of the initial deep neural network based on the first matrix, the second matrix, the first position, and the second position may include the following steps:
[0073] Step S31, determine the state set and the operation set, and initialize the state score, the policy network, and the value network;
[0074] Step S32, randomly select an initial state from the state set;
[0075] Step S33, select a policy action through the policy network, and update the initial state to the current state based on the policy action;
[0076] Step S34, determine whether there is a new coefficient vector pair based on Monte Carlo tree search;
[0077] Step S35, if there is a new coefficient vector pair, update the state score and update the value network;
[0078] Step S36, repeat the above steps S33 to S35 until the multiplication operation of the first target data block and the second target data block is completed to obtain the target state score;
[0079] Step S37, update the network parameters of the policy network and the value network based on the target state score.
[0080] Wherein, in an embodiment of the present invention, the above initial state S represents the difference between the current calculation algorithm and the correct algorithm; the operation set M defines a set of allowed operations, and each operation corresponds to an operation in matrix multiplication. And, in an embodiment of the present invention, the above policy network selects the next best operation, and the value network evaluates the current state score. When it is determined that there is a new coefficient vector pair u and v, it indicates that a matrix multiplication operation has been successfully executed.
[0081] And, in an embodiment of the present invention, the method for updating the state score and updating the value network if there is a new coefficient vector pair may include the following steps:
[0082] Step 1, increment the state score by 1;
[0083] Step 2, determine whether the new coefficient vector pair corresponds to an all-zero data block;
[0084] Step 3, if the new coefficient vector corresponds to an all-zero data block, decrement the state score by 1;
[0085] Step 4, update the value network.
[0086] Further, in an embodiment of the present invention, the method for determining whether the new coefficient vector pair corresponds to an all-zero data block may include: determining whether the new coefficient vector pair u and v satisfy or If it satisfies or then the new coefficient vector corresponds to an all-zero data block, that is, the selected multiplication strategy does not increase the multiplication calculation load, and the state score is decremented by 1; otherwise, no other operation is performed.
[0087] Further, in an embodiment of the present invention, assuming that the matrix obtained by multiplying the above first matrix and the second matrix is of dimension m×n×p, the above method, while minimizing the state score, can correspond to minimizing the number of multiplications required for the m×n×p matrix multiplication operation.
[0088] It should be noted that, in an embodiment of the present invention, the above method, through the sparse information of the first position of the all-zero data block in the first matrix and the second position of the all-zero data block in the second matrix in the matrix dataset, considers the computational savings brought by sparsity on the basis of the existing AlphaTensor algorithm, thereby reducing the number of multiplications required for matrix multiplication and reducing the computational complexity.
[0089] Step S4, repeat the above steps until the training stop condition is reached, obtain the target deep neural network, and determine the target deep neural network as the calculation model for sparse matrix multiplication.
[0090] Wherein, in an embodiment of the present invention, the above training stop condition may include: the initial deep neural network converges, and / or the number of iterations of the network parameters reaches a preset number.
[0091] And, in an embodiment of the present invention, the above preset number can be set as needed. By way of example, assume the preset number is 50.
[0092] Further, in an embodiment of the present invention, after determining the target deep neural network through the above steps, sparse matrix multiplication operations can be performed through the target neural network.
[0093] Table 2 is a comparative analysis table of the computational complexity of the present invention proposed in the embodiments of the present invention and the currently established optimal matrix multiplication algorithms in different dimensions. Among them, the above check is based on an assumption that there is only one zero element in the entire multiplication task, and it is speculated that when applied to a sparser matrix, the present invention may further reduce the computational complexity. And, previous research efforts have effectively reduced the required number of multiplications to a very low threshold, but the integration of sparse information in the present invention further reduces the computational complexity, and this phenomenon is particularly obvious in specific dimension configurations, such as
[0094] In (4, 4, 5) and (5, 5, 5) shown in Table 2, the number of matrix multiplications required by the present invention is lower than that of the existing known optimal method.
[0095]
[0096] According to the method for generating a calculation model for sparse matrix multiplication proposed by an embodiment of the present invention, a matrix data set for training is obtained; the first position of the all-zero data block in the first matrix and the second position of the all-zero data block in the second matrix in the matrix data set are determined; the network parameters of the initial deep neural network are updated based on the first matrix, the second matrix, the first position, and the second position; the above steps are repeated until the training stop condition is reached to obtain a target deep neural network, and the target deep neural network is determined as the calculation model for sparse matrix multiplication. Thus, the present invention considers the computational savings brought by sparsity during the training process of the calculation model through the sparse information of the first position of the all-zero data block in the first matrix and the second position of the all-zero data block in the second matrix in the matrix data set, thereby reducing the number of multiplications required for matrix multiplication and reducing the computational complexity.
[0097] Next, a device for generating a calculation model for sparse matrix multiplication proposed by an embodiment of the present invention will be described with reference to the accompanying drawings.
[0098] Figure 2 It is a schematic structural diagram of a device for generating a calculation model for sparse matrix multiplication according to an embodiment of the present invention.
[0099] As Figure 2 shown, the device 10 for generating a calculation model for sparse matrix multiplication includes: an acquisition module 201, a first determination module 202, an update module 203, and a second determination module 204, wherein,
[0100] The acquisition module 201 is used to obtain a matrix data set for training;
[0101] The first determination module 202 is used to determine the first position of the all-zero data block in the first matrix and the second position of the all-zero data block in the second matrix in the matrix data set;
[0102] The update module 203 is used to update the network parameters of the initial deep neural network based on the first matrix, the second matrix, the first position, and the second position;
[0103] The second determination module 204 repeats the above steps until the training stop condition is reached to obtain a target deep neural network, and determines the target deep neural network as the calculation model for sparse matrix multiplication.
[0104] Further, the first determination module 202 is specifically configured to:
[0105] Determine the target dimension for dividing the matrix into data blocks;
[0106] Based on the target dimension, divide the first matrix and the second matrix respectively to obtain corresponding first target data blocks and second target data blocks;
[0107] Determine the first positions of the all-zero data blocks in the first target data blocks;
[0108] Determine the second positions of the all-zero data blocks in the second target data blocks.
[0109] Further, the update module 203 is specifically configured to:
[0110] Update the network parameters of the initial deep neural network based on the first target data blocks, the second target data blocks, the first positions, and the second positions.
[0111] Further, the initial deep neural network includes a policy network and a value network; the update module 203 is further configured to:
[0112] Determine a state set and an action set, and initialize a state score, the policy network, and the value network;
[0113] Randomly select an initial state from the state set;
[0114] Select a policy action through the policy network, and update the initial state to the current state based on the policy action;
[0115] Determine whether there is a new coefficient vector pair based on Monte Carlo tree search;
[0116] If there is a new coefficient vector pair, update the state score and update the value network;
[0117] Repeat the above steps until the multiplication operation of the first target data blocks and the second target data blocks is completed to obtain a target state score;
[0118] Based on the target state score, update the network parameters of the policy network and the value network.
[0119] Further, the update module 203 is further configured to:
[0120] Increment the state score by 1;
[0121] Determine whether the new coefficient vector pair corresponds to an all-zero data block;
[0122] If the new coefficient vector corresponds to an all-zero data block, decrement the state score by 1;
[0123] Update the value network.
[0124] Furthermore, the above-mentioned stopping training conditions include:
[0125] The initial deep neural network converges, and / or the number of iterations of network parameters reaches a preset number.
[0126] According to the generating device of the calculation model for sparse matrix multiplication proposed by the embodiment of the present invention, a matrix data set for training is obtained; the first position of the all-zero data block in the first matrix in the matrix data set and the second position of the all-zero data block in the second matrix are determined; the network parameters of the initial deep neural network are updated based on the first matrix, the second matrix, the first position, and the second position; the above steps are repeated until the stopping training condition is reached to obtain a target deep neural network, and the target deep neural network is determined as the calculation model for sparse matrix multiplication. Thus, the present invention considers the computational savings brought by sparsity in the training process of the calculation model through the sparse information of the first position of the all-zero data block in the first matrix and the second position of the all-zero data block in the second matrix in the matrix data set, thereby reducing the number of multiplications required for matrix multiplication and reducing the computational complexity.
[0127] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0128] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0129] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for generating a computational model for sparse matrix multiplication, characterized in that, The method includes: Obtaining a matrix dataset for training; Determining a first position of an all-zero data block in a first matrix and a second position of an all-zero data block in a second matrix in the matrix dataset; Updating network parameters of an initial deep neural network based on the first matrix, the second matrix, the first position, and the second position; Repeating the above steps until a training stop condition is reached, obtaining a target deep neural network, and determining the target deep neural network as a calculation model for sparse matrix multiplication.
2. The method according to claim 1, wherein The determining a first position of an all-zero data block in a first matrix and a second position of an all-zero data block in a second matrix in the matrix dataset includes: Determining a target dimension for dividing the matrix into data blocks; Dividing the first matrix and the second matrix respectively based on the target dimension to obtain corresponding first target data blocks and second target data blocks; Determining a first position of an all-zero data block in the first target data block; Determining a second position of an all-zero data block in the second target data block.
3. The method according to claim 2, characterized in that, The updating network parameters of an initial deep neural network based on the first matrix, the second matrix, the first position, and the second position includes: Updating network parameters of an initial deep neural network based on the first target data block, the second target data block, the first position, and the second position.
4. The method according to claim 3, wherein The initial deep neural network includes a policy network and a value network; the updating network parameters of an initial deep neural network based on the first target data block, the second target data block, the first position, and the second position includes: Determining a state set and an action set, and initializing a state score, a policy network, and a value network; Randomly selecting an initial state from the state set; Selecting a policy action through the policy network and updating the initial state to the current state based on the policy action; Determining whether there is a new coefficient vector pair based on Monte Carlo tree search; If there is a new coefficient vector pair, updating the state score and updating the value network; Repeating the above steps until the multiplication operation of the first target data block and the second target data block is completed, obtaining a target state score; Updating network parameters of the policy network and the value network based on the target state score.
5. The method according to claim 4, wherein The if there is a new coefficient vector pair, updating the state score and updating the value network includes: Adding 1 to the state score; Determining whether the new coefficient vector pair corresponds to an all-zero data block; If the new coefficient vector corresponds to an all-zero data block, subtracting 1 from the state score; Updating the value network.
6. The method according to claim 1, wherein The training stop condition includes: The initial deep neural network converges, and / or, the number of iterations of the network parameters reaches a preset number of times.
7. An apparatus for generating a computational model for sparse matrix multiplication, characterized in that, The device includes: An obtaining module, configured to obtain a matrix dataset for training; A first determining module, configured to determine a first position of an all-zero data block in a first matrix and a second position of an all-zero data block in a second matrix in the matrix dataset; An update module, configured to update network parameters of an initial deep neural network based on the first matrix, the second matrix, the first position, and the second position; A second determination module, configured to repeat the above steps until a stop training condition is met, obtain a target deep neural network, and determine the target deep neural network as a calculation model for sparse matrix multiplication.
8. The device according to claim 7, characterized in that, The first determination module is specifically configured to: Determine a target dimension for dividing a matrix into data blocks; Based on the target dimension, divide the first matrix and the second matrix respectively to obtain corresponding first target data blocks and second target data blocks; Determine a first position of all-zero data blocks in the first target data blocks; Determine a second position of all-zero data blocks in the second target data blocks.
9. An electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-6.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1-6 is implemented.