Method and apparatus for processing matrix multiplication data, electronic device and storage medium
By determining optimal matrix combinations and partitioning based on storage space and bandwidth information, the method addresses the memory access challenges in matrix multiplication, improving efficiency and AI chip performance.
Patent Information
- Application Number
- JP2025038294
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-11
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-11
AI Technical Summary
Existing technologies face challenges in efficiently processing matrix multiplication data due to high memory access overhead, particularly in deep learning operations, where the general matrix multiplication operator is memory access-intensive and requires optimization to improve execution efficiency.
The method involves determining I matrix combinations from a plurality of matrices, calculating target sub-scale information, and optimizing matrix partitioning based on storage space capacities and bandwidth information to minimize memory access overhead, thereby selecting a target matrix combination that meets predetermined access memory resource overhead conditions.
This approach reduces memory access overhead, improves the execution efficiency of matrix multiplication operations, and enhances the performance of artificial intelligence chips by optimizing memory resource utilization.
Smart Images

Figure 2025087884000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to the field of chip technology. More specifically, the present disclosure provides a method, apparatus, electronic device, and storage medium for processing matrix multiplication data.
Background Art
[0002] With the development of artificial intelligence technology, the operators of deep learning models can be adjusted based on the hardware resources of artificial intelligence chips.
Summary of the Invention
[0003] The present disclosure provides a data processing apparatus, method, device, and storage medium.
[0004] According to one aspect of the present disclosure, a method for processing matrix multiplication data is provided. The method includes determining I matrix combinations from a plurality of matrices corresponding to a matrix multiplication operation, where each of the I matrix combinations includes a first matrix and a second matrix, the first matrix corresponds to a first storage space, the second matrix corresponds to a second storage space, and I is an integer greater than or equal to 1; determining first target sub-scale information for each of the I matrix combinations based on the scale information of the second matrix in each of the I matrix combinations and the capacity of the second storage space in each of the I matrix combinations, where the first target sub-scale information is related to the first matrix and the second matrix; determining second target sub-scale information for each of the I matrix combinations based on the capacity of the first storage space in each of the I matrix combinations and the first target sub-scale information for each of the I matrix combinations, where the second target sub-scale information is related to the first matrix and a third matrix, and the third matrix corresponds to the matrix multiplication operation; determining target partitioning information for each of the I matrix combinations based on the second target sub-scale information for each of the I matrix combinations and the scale information of the first matrix in each of the I matrix combinations; determining a target matrix combination from the I matrix combinations based on the bandwidth information of a computing unit that executes the matrix multiplication operation and the target partitioning information for each of the I matrix combinations, where the access memory resource overhead corresponding to the target matrix combination satisfies a predetermined access memory resource overhead condition.
[0005] According to another aspect of the present disclosure, a processing device for matrix multiplication data is provided. The device determines I matrix combinations from a plurality of matrices corresponding to matrix multiplication operations, where each of the I matrix combinations includes a first matrix and a second matrix. The first matrix corresponds to a first storage space, the second matrix corresponds to a second storage space, and I is an integer greater than or equal to 1. A first determination module, based on the scale information of the second matrix in each of the I matrix combinations and the capacity of the second storage space in each of the I matrix combinations, determines the first target sub-scale information for each of the I matrix combinations, where the first target sub-scale information is related to the first matrix and the second matrix. A second determination module, based on the capacity of the first storage space in each of the I matrix combinations and the first target sub-scale information in each of the I matrix combinations, determines the second target sub-scale information for each of the I matrix combinations, where the second target sub-scale information is related to the first matrix and the third matrix, and the third matrix is a third determination module corresponding to the matrix multiplication operation. A fourth determination module determines the target division information for each of the I matrix combinations based on the second target sub-scale information in each of the I matrix combinations and the scale information of the first matrix in each of the I matrix combinations. Based on the bandwidth information of the computing unit that executes the matrix multiplication operation and the target division information for each of the I matrix combinations, a target matrix combination is determined from the I matrix combinations, where the access memory resource overhead corresponding to the target matrix combination satisfies a predetermined access memory resource overhead condition.
[0006] According to another aspect of the present disclosure, an electronic device including the data processing device according to the present disclosure is provided.
[0007] According to another aspect of the present disclosure, an electronic device including at least one processor and a memory communicatively connected to the at least one processor is provided. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, which cause a computer to execute the method according to the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program that, when executed by a processor, implements the method according to the present disclosure.
[0010] It should be understood that the content described in this part is not intended to indicate the key points or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description.
Brief Description of the Drawings
[0011] The drawings are for better understanding of the present invention and do not limit the present disclosure.
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0012] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included, which are merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of known functions and structures are omitted in the following description.
[0013] In fields such as scientific computing and deep learning, a large number of linear filtering operations can be used, and the corresponding operator is generally a General Matrix Multiplication (GEMM) operator. In the operations related to deep learning models, the time of the general matrix multiplication operator is large. The general matrix multiplication operator is a memory access-intensive operator. Improving the utilization rate of the access memory can improve the execution efficiency of the general matrix multiplication operator and also help to provide the performance of related artificial intelligence chips.
[0014] The general matrix multiplication operator can be related to the multiplicand matrix A, the multiplier matrix B, and the result matrix C. The calculation process of the general matrix multiplication operator can be realized as follows.
[0015]
Equation
[0016] In some embodiments, based on the capacity of the Level 1 Cache (L1 Cache), one or both of the two input matrices and the result matrix can be decomposed to obtain a plurality of sub-matrices.
[0017] For example, based on the capacity of the level-1 cache unit for the multiplicand matrix A, the multiplicand matrix A can be decomposed into a plurality of sub-matrices. The data volume of each sub-matrix may match the capacity of the level-1 cache unit for the multiplicand matrix A. The scale of the sub-matrix may be l_m×l_k. l_m may be 1 or more and m or less, and l_k may be 1 or more and k or less. Therefore, in the row dimension, the number of times the multiplicand matrix A is decomposed may be t m as well.
[0018]
Number
[0019]
Number
[0020] FIG. 1 is a schematic diagram of the access memory of a plurality of matrices related to a matrix multiplication operation according to an embodiment of the present disclosure.
[0021] As shown in FIG. 1, the multiplicand matrix 110 may be decomposed into four sub-matrices. The scale of the sub-matrix 111 of the multiplicand matrix 110 may be, for example, l_m×l_k. The multiplicand matrix 110 may be the above-mentioned multiplicand matrix A. The multiplier matrix 120 may be decomposed into two sub-matrices. The scale of the sub-matrix 121 of the multiplier matrix 120 may be l_m×n. When performing the matrix multiplication operation, the scale of the matrix multiplication operation executed each time may be [l_m, n, l_k]. t m *t k After performing the matrix multiplication operation t
[0022] In the embodiment shown in FIG. 1, when determining the scale of the sub - matrix 121 of the multiplier matrix 120, the sub - matrix 111 was considered, but the capacity of the level - 1 cache unit for the multiplier matrix 120 was not considered. That is, the data volume of the sub - matrix 121 may be larger than the capacity of the level - 1 cache unit of the multiplier matrix 120. This causes redundant access memory for the sub - matrix 121 when performing the matrix multiplication operation.
[0023] When decomposing the multiplicand matrix 110 based on the capacity of the level - 1 cache unit for the multiplicand matrix 110, as shown in FIG. 1, in the processes of the first access memory LS111, the second access memory LS112, the third access memory LS113, and the fourth access memory LS114, the sub - matrices of the multiplicand matrix 110 may not be repeatedly accessed. Therefore, the access memory amount indication value LS_A of the multiplicand matrix 110
Number
[0024] The multiplier matrix 120 may be accessed t m - 1 times repeatedly. When the rows of the sub - matrices of the multiplicand matrix 110 change, the data in the level - 1 cache unit for the multiplier matrix 120 can be multiplexed to reduce the access memory amount. As shown in FIG. 1, in the processes of the first access memory LS111 and the second access memory LS112, some data of the sub - matrix 121 may be multiplexed. In the processes of the third access memory LS113 and the fourth access memory L114, some data of the sub - matrix 122 may be multiplexed. Therefore, the access memory amount indication value LS_B of the multiplier matrix 120
Number
[0025] The l1b_size may be the capacity of the level-1 cache unit for the multiplier matrix 120. The max() may be a function that takes the maximum value.
[0026] The result matrix 130 can be accessed by repeating t k -1 times. When the columns of the sub-matrix of the multiplicand matrix 110 change, the data in the level-1 cache unit for the result matrix 130 can be multiplexed to reduce the access memory amount. As shown in FIG. 1, in the process of the second access memory LS132 and the third access memory L133, some data related to the sub-matrix 132 may be multiplexed. Therefore, the access memory amount indication value LS_C of the result matrix 130 is [Number] It may be.
[0027] The l1c_size may be the capacity of the level-1 cache unit for the result matrix 130.
[0028] As described above, based on the capacity of the level-1 cache unit of the multiplicand matrix A, the multiplicand matrix A is decomposed, and based on the sub-matrix of the multiplicand matrix A, the multiplier matrix is decomposed. However, in some embodiments, the first matrix may be decomposed using the capacity of the level-1 cache unit of the first matrix, and the second matrix may be decomposed using the sub-matrix of the first matrix. The first matrix may be any of the multiplicand matrix A, the multiplier matrix B, and the result matrix C. The second matrix may be any matrix other than the first matrix.
[0029] By using different matrices as the first matrix and the second matrix, the sum of a plurality of access memory amount indication values can be determined. Based on the matrix decomposition method corresponding to the sum of the minimum access memory amount indication values, access memory resources can be saved. However, if different matrices are used as the first matrix and matrix decomposition is performed based on the capacity of the first-level cache unit of the first matrix, as the matrix scale increases, the number of decomposition methods increases, and it is necessary to calculate a large number of access memory amount indication values, and it is necessary to calculate the sum of a large number of access memory amount indication values. The overhead of computing resources is large and the time cost is high.
[0030] Based on this, in order to further improve the efficiency of matrix multiplication operations, the present disclosure provides a method for processing matrix multiplication data, which will be described below.
[0031] FIG. 2 is a schematic flowchart of a method for processing matrix multiplication data according to an embodiment of the present disclosure.
[0032] As shown in FIG. 2, method 200 may include operation S210 to operation S250.
[0033] In operation S210, I matrix combinations are determined from a plurality of matrices corresponding to the matrix multiplication operation.
[0034] In an embodiment of the present disclosure, each matrix combination in the I matrix combinations may include a first matrix and a second matrix. I may be an integer greater than or equal to 1. For example, when the plurality of matrices are three matrices, I may be 6. The fact that the number of matrices is 3 is only an example. The plurality of matrices may be two, four or more matrices. Also, for example, the three matrices may be the multiplicand matrix A, the multiplier matrix B, and the result matrix C respectively. Also, for example, for the first matrix combination in the I matrix combinations, the first matrix may be the multiplicand matrix A, and the second matrix may be the multiplier matrix B. The second matrix combination may be the multiplicand matrix A, and the second matrix may be the result matrix C. Hereinafter, the first matrix combination will be referred to for description.
[0035] In an embodiment of the present disclosure, the first matrix may correspond to a first storage space, and the second matrix may correspond to a second storage space. For example, when the first matrix is the multiplicand matrix A, the first storage space may be a first-level cache unit for the multiplicand matrix A. When the second matrix is the multiplier matrix B, the second storage space may be a first-level cache unit for the multiplier matrix B described above.
[0036] In operation S220, based on the scale information of the second matrix of each of the I matrix combinations and the capacity of the second storage space of each of the I matrix combinations, the first target sub-scale information of each of the I matrix combinations is determined.
[0037] In an embodiment of the present disclosure, the first target sub-scale information is related to the first matrix and the second matrix. For example, the scale information of the second matrix may include the scale of the second matrix. Taking the first matrix combination as an example, the scale of the first matrix may be m×k, and the scale of the second matrix may be k×n. The first target sub-scale information may be related to k. Taking the second matrix combination as an example, the scale of the first matrix may be m×k, and the scale of the second matrix (the result matrix C above) may be m×n. The first target sub-scale information may be related to m.
[0038] In an embodiment of the present disclosure, the first target sub-scale information can be determined based on the sub-scale information of the second matrix that has no relation to the first matrix. For example, taking the first matrix combination as an example, the sub-scale information of the second matrix includes an initial sub-scale value n that has no relation to the first matrix. Based on the initial sub-scale value n and the capacity of the second storage space, the first target sub-scale information can be determined.
[0039] In an embodiment of the present disclosure, based on the first target sub-scale information, the data amount of the sub-matrix of the second matrix may be less than or equal to the capacity of the second storage space. For example, for each matrix combination, the data amount of the sub-matrix of the second matrix may be equal to the capacity of the second storage space.
[0040] In operation S230, based on the capacity of the first storage space of each of the I matrix combinations and the first target sub-scale information of each of the I matrix combinations, the second target sub-scale information of each of the I matrix combinations is determined.
[0041] In an embodiment of the present disclosure, the second target sub-scale information is related to the first matrix and the third matrix, and the third matrix corresponds to a matrix multiplication operation. For example, taking the first matrix combination as an example, the third matrix may be the result matrix C. The second target sub-scale information may be related to m. Based on the first target scale sub-information and the capacity of the first storage space of the first matrix combination, the second target sub-scale information of the first matrix combination can be determined.
[0042] In operation S240, based on the second target sub-scale information of each of the I matrix combinations and the scale information of the first matrix of each of the I matrix combinations, the target segmentation information of each of the I matrix combinations is determined.
[0043] In an embodiment of the present disclosure, the target segmentation information can indicate the number of times the first matrix is divided in a dimension not related to the second matrix. For example, taking the first matrix combination as an example, the row dimension of the first matrix is not related to the second matrix, and the target segmentation information of the first matrix combination may indicate the number of times the first matrix is divided in the row dimension.
[0044] In operation S250, based on the bandwidth information of the computing unit that executes the matrix multiplication operation and the target segmentation information of each of the I matrix combinations, a target matrix combination is determined from the I matrix combinations.
[0045] In an embodiment of the present disclosure, the access memory resource overhead corresponding to the target matrix combination satisfies a preset access memory resource overhead condition. For example, based on the target division information of each of the I matrix combinations, the data volume of the input sub-matrices of the two input matrices in each matrix combination can be determined. Based on the bandwidth information of the computing unit and the data volume of the input sub-matrices of each matrix combination, the access memory resource overhead of each matrix combination can be determined. The access memory resource overhead condition may be that the access memory resource overhead is less than or equal to a preset access memory resource overhead threshold value.
[0046] According to an embodiment of the present disclosure, by determining the division method of the first matrix based on the capacity of the second storage space, the computing resources required to determine the target matrix combination can be reduced, the access memory amount can be effectively reduced, and in particular, it is possible to avoid the repeated access of the sub-matrices of the first matrix and the sub-matrices of the second matrix to the access memory, improve the execution efficiency of the matrix multiplication operator, and thus help improve the efficiency of the data processing device.
[0047] The method of the present disclosure has been described above. Hereinafter, the method of the present disclosure will be further described with reference to FIGS. 3A-3B.
[0048] FIG. 3A is a schematic diagram of matrix multiplication data according to an embodiment of the present disclosure.
[0049] As shown in FIG. 3A, the matrix multiplication data may include a multiplicand matrix 310, a multiplier matrix 320, and a result matrix 330. The scale of the multiplicand matrix 310 may be m×k, the scale of the multiplier matrix 320 may be k×n, and the scale of the result matrix 330 may be m×n.
[0050] In some embodiments, in some forms of the above operation S210, the I matrix combinations may further include a third matrix combination, a fourth matrix combination, a fifth matrix combination, and a sixth matrix combination. For the third matrix combination, the first matrix may be the multiplier matrix 320, and the second matrix may be the multiplicand matrix 310. For the fourth matrix combination, the first matrix may be the multiplier matrix 320, and the second matrix may be the result matrix 330. For the fifth matrix combination, the first matrix may be the result matrix 330, and the second matrix may be the multiplicand matrix 310. For the sixth matrix combination, the first matrix may be the result matrix 330, and the second matrix may be the multiplier matrix 320.
[0051] In an embodiment of the present disclosure, the scale information of the first matrix includes a first initial sub-scale value that has no relation to the second matrix, and the scale information of the second matrix includes a second initial sub-scale value that has no relation to the first matrix. For example, taking the first matrix as the multiplicand matrix 310 and the second matrix as the multiplier matrix 320 as an example, the first initial sub-scale value may be m, and the second initial sub-scale value may be n. To reduce the duplicate access memory to the multiplier matrix 320, referring to the above formula 4, t m = 1 or n * l_k can be set to be less than or equal to l1b_size. To reduce the duplicate access memory to the result matrix 330, referring to the above formula 5, t k = 1 can be set. However, setting t m = 1 or t k = 1 cannot efficiently divide the multiplicand matrix 310. Thus, to reduce the duplicate access memory, n * l_k can be set to be less than or equal to l1b_size. Also, to minimize the duplicate access memory to the result matrix 330, the value of t k should be as small as possible. That is (referring to the above formula 2), l_k can be as large as possible. Therefore, n * l_k = l1b_size can be set.
[0052] Based on this, in some embodiments, the first target sub-scale information includes a first target sub-scale value. In some embodiments of the above operation S220, determining the first target sub-scale information for each of the I matrix combinations based on the scale information of the second matrix and the capacity of the second storage space for each of the I matrix combinations includes determining the first target sub-scale value based on the ratio of the capacity of the second storage space to the second initial sub-scale value. Taking the first matrix as the multiplicand matrix 310 and the second matrix as the multiplier matrix 320 as an example, the first target sub-scale value large_k may be l_k when n*l_k = l1b_size. As another example, the first target sub-scale value large_k of the first matrix combination may be determined by the following formula.
[0053] [Number] n may be the second initial sub-scale value.
[0054] Also, in some embodiments, the second target sub-scale information may include a second target sub-scale value. In some embodiments of the above operation S230, determining the second target sub-scale information for each of the I matrix combinations based on the capacity of the first storage space and the first target sub-scale information for each of the I matrix combinations includes determining the second target sub-scale value based on the ratio of the capacity of the first storage space to the first target sub-scale value. For example, the second target sub-scale value large_m can be determined by the following formula.
[0055] [Number] l1a_size may be the capacity of the first-level cache unit for the multiplicand matrix 310.
[0056] Next, in some embodiments of the above operation S240, determining the target division information for each of the I matrix combinations based on the second target sub-scale information for each of the I matrix combinations and the scale information of the first matrix for each of the I matrix combinations includes determining the target division information based on the ratio of the first initial sub-scale value to the second target sub-scale value. The target division information may include a target division parameter value. The target division parameter value T m can be determined by the following formula.
[0057]
Equation
[0058] FIG. 3B is a schematic diagram of a plurality of sub-matrices according to an embodiment of the present disclosure.
[0059] As shown in FIG. 3B, taking the above first matrix combination as an example, based on the above first target sub-scale value large_k and the second target sub-scale value large_m, the sub-matrix 311 of the multiplicand matrix 310 can be determined. The multiplicand matrix 310 may be divided into, for example, six sub-matrices. Based on the second initial sub-scale value n and the first target sub-scale value large_k, the sub-matrix 321 of the multiplier matrix 320 can be determined. The multiplier matrix 320 may include, for example, three sub-matrices. Therefore, the result matrix 330 may include, for example, two sub-matrices. The scale of the sub-matrix 331 of the result matrix 330 may be large_m×n.
[0060] As shown in FIG. 3B, in the processes of the first access memory LS311, the second access memory LS312, the third access memory LS313, the fourth access memory LS314, the fifth access memory LS315, and the sixth access memory LS316 to the multiplicand matrix 310, the multiplicand matrix 310 is not repeatedly accessed. In the processes of the first access memory LS321, the second access memory LS323, and the third access memory LS325 to the multiplier matrix 320, the multiplier matrix 320 is not repeatedly accessed. Also, in the process of performing the matrix multiplication operation, the data in the one-level cache unit for the multiplier matrix 320 is multiplexed. For example, the sub-matrix 321 of the multiplier matrix 320 is multiplied by the sub-matrices 311 and 312 of the multiplicand matrix 310 respectively, that is, the sub-matrix 321 is multiplexed once.
[0061] As shown in FIG. 3B, the result matrix 330 needs to be repeatedly accessed. To determine the sub-matrix 331 of the result matrix 330, based on the global memory unit, the first access memory LS331, the fourth access memory LS334, and the fifth access memory LS335 can be executed. To determine the sub-matrix 332 of the result matrix 330, based on the global memory unit, the second access memory LS332, the third access memory LS333, and the sixth access memory LS336 can be executed. After multiplying the above-mentioned sub-matrix 311 and sub-matrix 321, a result sub-matrix can be obtained. Next, the first access memory LS331 is executed, and the result sub-matrix is stored from the one-level cache unit for the result matrix 330 to the global memory unit. Next, the resource overhead of the access memory can be determined.
[0062] In some embodiments, in some embodiments of the above operation S250, determining a target matrix combination from I matrix combinations based on the bandwidth information of the computing unit that performs the matrix multiplication operation and the target partitioning information of each of the I matrix combinations includes determining the access memory resource overhead of each of the I matrix combinations based on the bandwidth information of the computing unit that performs the matrix multiplication operation and the target partitioning information of each of the I matrix combinations.
[0063] In an embodiment of the present disclosure, the bandwidth information includes the memory bandwidth between the computing unit and the memory unit. For example, the memory unit may be a global memory unit.
[0064] In an embodiment of the present disclosure, determining the access memory resource overhead of each of the I matrix combinations based on the bandwidth information of the computing unit that performs the matrix multiplication operation and the target partitioning information of each of the I matrix combinations includes determining a first access memory time resource overhead based on the data volume of the first matrix, the data volume of the second matrix, the data volume of the third matrix, and the memory bandwidth. An initial access memory time resource overhead is determined based on the target partitioning information, the initial sub-data volume of the third matrix, and the memory bandwidth. The initial sub-data volume may be the data volume for storing the third matrix in the global memory unit.
[0065] For example, the total data volume can be determined based on the data volume of the first matrix, the data volume of the second matrix, and the data volume of the third matrix. The first access memory time resource overhead is determined based on the ratio of the total data volume to the memory bandwidth. The access memory resource overhead of the matrix combination can be determined based on the first access memory time resource overhead and the initial access memory time resource overhead. As another example, the access memory resource overhead total_time(a_m) of the first matrix combination can be determined by the following formula.
[0066]
Number
Number
[0067] With reference to the first matrix combination, some methods for determining the access memory resource overhead in the present disclosure have been described. Since the method for determining the access memory resource overhead based on the second to sixth matrix combinations is the same as or similar to the method for determining the access memory resource overhead based on the first matrix combination, the description is omitted here.
[0068] According to the embodiments of the present disclosure, the memory bandwidth between the computing unit and the memory unit can be fully utilized, and repeated access memory can be performed only on the third matrix, reducing the number of access memory times and contributing to the improvement of access memory efficiency.
[0069] As described above, it will be understood that some data of the third matrix that requires repeated access memory is stored in the global memory unit. However, the present disclosure is not limited thereto, and some data of the third matrix that requires repeated access memory may be stored in multiple levels of cache units, which will be described below.
[0070] In some embodiments, the third matrix is stored in at least two levels of cache units among a plurality of levels of cache units. The at least two levels of cache units include a first cache unit and a second cache unit. For example, taking the first matrix combination as an example, the first cache unit may be a first-level cache unit for the result matrix C. The second cache unit may be a second-level cache unit (Level2Cache, L2Cache) related to the first-level cache unit. It should be understood that the first cache unit may be a zero-level cache (level0cache, L0cache). When the first cache unit is a zero-level cache, the second cache unit may be a first-level cache unit. In another example, the at least two levels of cache units may further include a third cache unit. When the first cache unit is a first-level cache unit, the third cache unit may be a third-level cache unit (Level3Cache, L3Cache).
[0071] In an embodiment of the present disclosure, the bandwidth information between the computing unit and the memory unit may further include at least one cache bandwidth between a plurality of levels of cache units. For example, taking the first cache unit being a first-level cache unit as an example, the at least one cache bandwidth may include a cache bandwidth B 2 between the second-level cache unit and the first-level cache unit, and a cache bandwidth B 3 between the third-level cache unit and the second-level cache unit.
[0072] Next, based on the target segmentation information, the target sub-data volume of the third matrix, and the cache bandwidth, the second access memory time resource overhead can be determined. The target sub-data volume may be the amount of data for storing the third matrix in the target cache unit. For example, the target sub-data volume may be one or more. When the data volume of the third matrix is large, a part of the data of the third matrix can be stored in the three-level cache unit.
[0073] In the embodiments of the present disclosure, determining the second access memory time resource overhead based on the target segmentation information, the target sub-data volume of the third matrix, and the cache bandwidth may include determining the sub-access memory time resource overhead based on the ratio of the target sub-data volume of the third matrix to the cache bandwidth. The second access memory time resource overhead is determined based on the product of the sub-access memory time resource overhead and the target segmentation parameter value. Next, different from determining the access memory resource overhead of the matrix combination based on the first access memory time resource overhead and the initial access memory time resource overhead, the access memory resource overhead of the matrix combination can be determined based on the first access memory time resource overhead and the second access memory time resource overhead. For example, the access memory resource overhead total_time(A_m) can be determined by the following formula.
[0074] [Number] m l2 is the target sub-data volume in which the third matrix is stored in the second cache unit, and m l3 is the target sub-data volume in which the third matrix is stored in the third cache unit, and m lj is the target sub-data volume in which the third matrix is stored in the memory unit at level j. j may be an integer greater than 2. As can be understood, m l3may be 0, m lj may be 0. It is understood that the memory unit at the j level may be a global memory unit. B j may be the bandwidth between the memory unit at the j level and the memory unit at the 3 level.
[0075] According to the embodiments of the present disclosure, a part of the third data that is repeatedly accessed and memorized can be stored in memory units of multiple levels, the bandwidth of each level of memory unit in the memory units of multiple levels can be fully utilized, and the efficiency of the computing unit can be further improved.
[0076] In the above, several methods for determining the access memory resource overhead have been described. Hereinafter, several methods for determining the target matrix combination will be described.
[0077] In some embodiments, in some embodiments of the above operation S250, determining the target matrix combination from I matrix combinations includes determining the target matrix combination from the I matrix combinations based on the access memory resource overhead of each of the I matrix combinations.
[0078] In the embodiments of the present disclosure, it is possible to determine whether the access memory resource overhead of the I matrix combination meets a predetermined access memory resource overhead condition. The predetermined access memory resource overhead condition includes at least one of that the access memory resource overhead corresponding to the target matrix combination is below a predetermined access memory resource overhead threshold, and that among the I matrix combinations, the access memory resource overhead corresponding to the target matrix combination is the smallest. For example, the matrix combination with the smallest access memory resource overhead may be used as the target matrix combination. As another example, as the target matrix combination, a matrix combination in which any access memory resource overhead is smaller than a predetermined access memory resource overhead threshold may be used.
[0079] Having described several forms of determining the target matrix combination above, the method of the present disclosure will be further described below with reference to the target matrix combination.
[0080] In some embodiments, the method 200 may further include performing a matrix multiplication operation based on the first target sub-scale information and the second target sub-scale information of the target matrix combination.
[0081] In an embodiment of the present disclosure, the third matrix may be the result matrix of the matrix multiplication operation. The method may further include determining a plurality of first sub-matrices of the first matrix of the target matrix combination based on the first target sub-scale information and the second target sub-scale information of the target matrix combination. Based on the first target sub-scale information of the target matrix combination, a plurality of second sub-matrices of the second matrix of the target matrix combination are determined. A matrix multiplication operation is performed based on the plurality of first sub-matrices and the plurality of second sub-matrices to obtain a plurality of result sub-matrices. Based on the plurality of result sub-matrices, a third matrix is obtained. For example, taking the first matrix combination above as an example, based on the first target sub-scale value large_k and the second target sub-scale value large_m of the first matrix combination, a plurality of first sub-matrices may be determined from the multiplicand matrix. The scale of the first sub-matrix may be large_m×large_k. Based on the first target sub-scale value large_k and the second initial sub-scale value n, a plurality of second sub-matrices may be determined from the multiplier matrix. The scale of the second sub-matrix may be large_k×n. Next, the first sub-matrix and the second sub-matrix are multiplied to obtain a result sub-matrix. For the plurality of result sub-matrices, addition and stitching operations are performed to obtain a result matrix. It should be understood that the manner of performing the matrix multiplication operation based on the second matrix combination is the same as or similar to the manner of performing the matrix multiplication operation based on the first matrix combination. Details are not elaborated here.
[0082] In an embodiment of the present disclosure, the second matrix may be the result matrix of a matrix multiplication operation. The method may further include determining a plurality of first sub-matrices of a first matrix of a target matrix combination based on first target sub-scale information and second target sub-scale information of the target matrix combination. Based on the second target sub-scale information of the target matrix combination, determine a plurality of third sub-matrices of a third matrix of the target matrix combination. Execute a matrix multiplication operation based on the plurality of first sub-matrices and the plurality of third sub-matrices to obtain a plurality of result sub-matrices. Obtain a second matrix based on the plurality of result sub-matrices.
[0083] In an embodiment of the present disclosure, the first matrix may be the result matrix of a matrix operation. The method may further include determining a plurality of second sub-matrices of a second matrix of a target matrix combination based on first target sub-scale information of the target matrix combination. Based on the second target sub-scale information of the target matrix combination, determine a plurality of third sub-matrices of a third matrix of the target matrix combination. Execute a matrix multiplication operation based on the plurality of second sub-matrices and the plurality of third sub-matrices to obtain a plurality of result sub-matrices. Obtain a first matrix based on the plurality of result sub-matrices.
[0084] It can be understood that the manner of executing the matrix multiplication operation based on the third matrix combination to the sixth matrix combination is the same as or similar to the manner of executing the matrix operation based on the first matrix combination. Details are not described here to avoid redundancy.
[0085] Above, the method for processing matrix multiplication data of the present disclosure has been described. Next, the apparatus of the present disclosure will be described.
[0086] FIG. 4 is a schematic diagram of an apparatus for processing matrix multiplication data according to an embodiment of the present disclosure.
[0087] As shown in FIG. 4, the apparatus 400 may include a first determination module 410, a second determination module 420, a third determination module 430, a fourth determination module 440, and a fifth determination module 450.
[0088] The first determination module 410 determines I matrix combinations from a plurality of matrices corresponding to matrix multiplication operations. Each matrix combination in the I matrix combinations includes a first matrix and a second matrix. The first matrix corresponds to a first storage space, the second matrix corresponds to a second storage space, and I is an integer greater than or equal to 1.
[0089] The second determination module 420 determines the first target sub-scale information for each of the I matrix combinations based on the scale information of the second matrix in each of the I matrix combinations and the capacity of the second storage space in each of the I matrix combinations. The first target sub-scale information is related to the first matrix and the second matrix.
[0090] The third determination module 430 determines the second target sub-scale information for each of the I matrix combinations based on the capacity of the first storage space in each of the I matrix combinations and the first target sub-scale information for each of the I matrix combinations. The second target sub-scale information is related to the first matrix and the third matrix, and the third matrix corresponds to the matrix multiplication operation.
[0091] The fourth determination module 440 determines the target division information for each of the I matrix combinations based on the second target sub-scale information for each of the I matrix combinations and the scale information of the first matrix in each of the I matrix combinations.
[0092] The fifth determination module 450 determines a target matrix combination from the I matrix combinations based on the bandwidth information of the calculation unit that executes the matrix multiplication operation and the target division information for each of the I matrix combinations. The access memory resource overhead corresponding to the target matrix combination satisfies a predetermined access memory resource overhead condition.
[0093] In some embodiments, the scale information of the first matrix includes a first initial sub-scale value that is not related to the second matrix, the scale information of the second matrix includes a second initial sub-scale value that is not related to the first matrix, the first target sub-scale information includes a first target sub-scale value, and the second target sub-scale information includes a second target sub-scale value.
[0094] In some embodiments, the second determination module further determines a first target sub-scale value based on a ratio of the capacity of the second memory space to the second initial sub-scale value.
[0095] In some embodiments, the third determination module further determines a second target sub-scale value based on a ratio of the capacity of the first memory space to the first target sub-scale value.
[0096] In some embodiments, the fourth determination module further determines target segmentation information based on a ratio of the first initial sub-scale value to the second target sub-scale value.
[0097] In some embodiments, the third matrix is stored in at least two levels of cache units among a plurality of levels of cache units, and the at least two levels of cache units include a first cache unit and a second cache unit.
[0098] In some embodiments, the fifth determination module includes: a first determination sub-module that determines the access memory resource overhead of each of the I matrix combinations based on the bandwidth information of the calculation unit that executes the matrix multiplication operation and the target segmentation information of each of the I matrix combinations; and a second determination sub-module that determines a target matrix combination from the I matrix combinations according to the access memory resource overhead of each of the I matrix combinations.
[0099] In some embodiments, the bandwidth information includes the memory bandwidth between the computing unit and the memory unit and at least one cache bandwidth between multiple levels of cache units. The first determination sub-module includes a first determination unit, a second determination unit, and a third determination unit. The first determination unit determines a first access memory time resource overhead based on the data volume of the first matrix, the data volume of the second matrix, the data volume of the third matrix, and the memory bandwidth. The second determination unit determines a second access memory time resource overhead based on the target segmentation information, the target sub-data volume of the third matrix, and the cache bandwidth. The target sub-data volume is the data volume stored in the target cache unit by the third matrix, and the target cache unit is a cache unit other than the first cache unit in the multiple levels of cache units. The third determination unit determines the access memory resource overhead of the matrix combination based on the first access memory time resource overhead and the second access memory time resource overhead.
[0100] In some embodiments, the first determination unit includes a first determination subunit for determining the total data volume based on the data volume of the first matrix, the data volume of the second matrix, and the data volume of the third matrix, and a second determination subunit for determining the first access memory time resource overhead based on the ratio of the total data volume to the memory bandwidth.
[0101] In some embodiments, the target segmentation information includes a target segmentation parameter value. The second determination unit includes a third determination subunit for determining a sub-access memory time resource overhead based on the ratio of the target sub-data volume of the third matrix to the cache bandwidth, and a fourth determination subunit for determining the second access memory time resource overhead based on the product of the sub-access memory time resource overhead and the target segmentation parameter value.
[0102] In some embodiments, it further includes an execution module that executes a matrix multiplication operation based on the first target sub-scale information and the second target sub-scale information of the target matrix combination.
[0103] In some embodiments, the predetermined access memory resource overhead condition includes at least one of the following: the access memory resource overhead corresponding to the target matrix combination is less than or equal to a predetermined access memory resource overhead threshold; and among the I matrix combinations, the access memory resource overhead corresponding to the target matrix combination is minimized.
[0104] Above, the device of the present disclosure has been described. Next, an electronic device including the device will be described.
[0105] FIG. 5 is a schematic diagram of an electronic device according to an embodiment of the present disclosure.
[0106] As shown in FIG. 5, the electronic device 50 may include a matrix multiplication data processing device 500 according to the present disclosure. The matrix multiplication data processing device 500 may be, for example, the device 400 described above.
[0107] In the technical solution of the present disclosure, any processing such as the collection, storage, use, processing, transmission, provision, and disclosure of such user personal information complies with the provisions of relevant laws and does not violate public order and good customs.
[0108] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.
[0109] FIG. 6 shows an exemplary block diagram for implementing an exemplary electronic device 600 of an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may further represent various forms of mobile devices, such as, for example, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The members, their connections and relationships, and their functions shown herein are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0110] As shown in FIG. 6, the device 600 includes a computing unit 601, which can execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 can further store various programs and data necessary for the operation of the device 600. The computing unit 601, the ROM 602, and the RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0111] A plurality of components in the device 600 are connected to the I / O interface 605 and include an input unit 606, such as a keyboard, a mouse, etc., an output unit 607, such as various types of displays, speakers, etc., a storage unit 608, such as a magnetic disk, an optical disk, etc., and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 enables the device 600 to exchange information / data with other devices via a computer network, such as the Internet, and / or various telecommunication networks.
[0112] The computing unit 601 may be various general-purpose and / or dedicated processing modules having processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a GPU (Graphics Processing Unit), various dedicated artificial intelligence (AI) computing chips, computing units for various machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes processing with each of the methods described above, such as a data processing method. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly included in a machine-readable medium such as the storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed into the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the data processing method described above may be executed. Alternatively, in another embodiment, the computing unit 601 may be configured to execute the data processing method in any other suitable form (e.g., via firmware).
[0113] The various embodiments of the systems and techniques described in this specification may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be implemented in one or more computer programs, which may be executed and / or interpreted in a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, and which receives data and instructions from, and transmits data and instructions to, a memory system, at least one input device, and at least one output device.
[0114] The program code for implementing the methods of the present disclosure may be created in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a dedicated computer, or other programmable data processing apparatus, such that, when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the device, partially on the device, partially on the device as an independent software package, and partially on a remote device or entirely on a remote device or server.
[0115] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use in or in combination with an instruction execution system, apparatus, or electronic device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or electronic devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections by one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] To provide for interaction with a user, the computer may implement the systems and techniques described herein, the computer having a display device (e.g., a CRT (cathode ray tube) display or an LCD (liquid crystal display)) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball), whereby the user can provide input to the computer through the keyboard and the pointing device. Other kinds of devices may further provide for interaction with the user; for example, feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form (including voice input, speech input, or tactile input).
[0117] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser, where a user can interact with embodiments of the systems and techniques described herein via the graphical user interface or the network browser), or a computing system that includes any combination of such background components, middleware components, or front-end components. The components of the system can be connected to each other by digital data communication in any form or medium (e.g., a communication network). Exemplary communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0118] The computer system may include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by a computer program running on the corresponding computer and having a client-server relationship.
[0119] It should be understood that the various forms of flow shown above may be used, and the steps may be sorted, added, or deleted again. For example, each step described in the present invention may be executed in parallel, sequentially, or in a different order, and the present specification is not limited herein as long as the desired results of the disclosed technical solutions can be achieved.
[0120] The foregoing specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and alternatives can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.
Claims
1. A method for processing matrix multiplication data, comprising the steps of: determining I matrix combinations from a plurality of matrices corresponding to a matrix multiplication operation, where each of the I matrix combinations includes a first matrix and a second matrix, the first matrix corresponds to a first storage space, the second matrix corresponds to a second storage space, and I is an integer equal to or greater than 1; determining first target sub-scale information for each of the I matrix combinations based on the size information of the second matrix for each of the I matrix combinations and a capacity of the second storage space for each of the I matrix combinations, where the first target sub-scale information is associated with the first matrix and the second matrix; determining second target sub-scale information of each of the I matrix combinations based on a capacity of the first storage space of each of the I matrix combinations and a first target sub-scale information of each of the I matrix combinations, where the second target sub-scale information is related to the first matrix and a third matrix, and the third matrix corresponds to the matrix multiplication operation; determining target division information for each of the I matrix combinations based on second target sub-scale information for each of the I matrix combinations and scale information for a first matrix for each of the I matrix combinations; determining a target matrix combination from the I matrix combinations according to bandwidth information of a computing unit that performs the matrix multiplication operation and target partition information of each of the I matrix combinations, where an access memory resource overhead corresponding to the target matrix combination satisfies a predetermined access memory resource overhead condition. How to process matrix multiplication data.
2. The first matrix scale information includes a first initial sub-scale value unrelated to the second matrix, the second matrix scale information includes a second initial sub-scale value unrelated to the first matrix, the first target sub-scale information includes a first target sub-scale value, and the second target sub-scale information includes a second target sub-scale value. The method of claim 1.
3. Determining first target sub-scale information of each of the I matrix combinations based on the second matrix scale information of each of the I matrix combinations and the second storage space capacity of each of the I matrix combinations includes: determining the first target sub-size value based on a ratio of a capacity of the second storage space to the second initial sub-size value. The method of claim 2.
4. Determining second target sub-scale information for each of the I matrix combinations based on the capacity of the first storage space for each of the I matrix combinations and first target sub-scale information for each of the I matrix combinations includes: determining the second target sub-size value based on a ratio of a capacity of the first storage space to the first target sub-size value. The method of claim 2.
5. Determining target division information of each of the I matrix combinations based on second target sub-scale information of each of the I matrix combinations and scale information of a first matrix of each of the I matrix combinations, determining the target division information based on a ratio of the first initial sub-scale value and the second target sub-scale value. The method of claim 2.
6. The third matrix is stored in at least two levels of cache units of a multi-level cache unit, the at least two levels of cache units including a first cache unit and a second cache unit. The method of claim 1.
7. determining a target matrix combination from the I matrix combinations based on bandwidth information of a computing unit that performs the matrix multiplication operation and target partition information for each of the I matrix combinations, determining an access memory resource overhead for each of the I matrix combinations based on bandwidth information of a computing unit that performs the matrix multiplication operation and target partition information for each of the I matrix combinations; determining a target matrix combination from the I matrix combinations based on an access memory resource overhead for each of the I matrix combinations. The method of claim 1.
8. the bandwidth information includes a storage bandwidth between the computing unit and a memory unit and at least one cache bandwidth between multiple levels of cache units; determining an access memory resource overhead for each of the I matrix combinations based on bandwidth information of a computing unit that performs the matrix multiplication operation and target partition information for each of the I matrix combinations, determining a first access memory time resource overhead based on the first matrix data amount, the second matrix data amount, the third matrix data amount, and the storage bandwidth; determining a second access memory time resource overhead based on the target partition information, a target sub-data amount of the third matrix, and the cache bandwidth, where the target sub-data amount is a data amount of the third matrix stored in a target cache unit, and the target cache unit is a cache unit other than the first cache unit among the multiple levels of cache units; determining an access memory resource overhead for the matrix combination based on the first access memory time resource overhead and the second access memory time resource overhead. The method of claim 7.
9. Determining a first access memory time resource overhead based on the first matrix data amount, the second matrix data amount, the third matrix data amount, and the storage bandwidth includes: determining a total amount of data based on the amount of data of the first matrix, the amount of data of the second matrix, and the amount of data of the third matrix; determining the first access memory time resource overhead based on a ratio of the total data amount to the storage bandwidth. The method according to claim 8.
10. The target division information includes a target division parameter value; determining a second access memory time resource overhead based on the target partition information, the target sub-data amount of the third matrix, and the cache bandwidth; determining a sub-access memory time resource overhead based on a ratio of a target sub-data amount of the third matrix and the cache bandwidth; determining the second access memory time resource overhead based on a product of the sub-access memory time resource overhead and the target partitioning parameter value. The method according to claim 8.
11. performing the matrix multiplication operation based on first target sub-scale information and second target sub-scale information of the target matrix combination. The method of claim 1.
12. The predetermined access memory resource overhead condition is: an access memory resource overhead corresponding to the target matrix combination is less than or equal to a predetermined access memory resource overhead threshold; In the I matrix combinations, an access memory resource overhead corresponding to the target matrix combination is minimal. The method of claim 1.
13. A processing device for matrix multiplication data, comprising: a first determination module for determining I matrix combinations from a plurality of matrices corresponding to a matrix multiplication operation, where each of the I matrix combinations includes a first matrix and a second matrix, the first matrix corresponding to a first storage space, the second matrix corresponding to a second storage space, and I is an integer equal to or greater than 1; A second determination module determines first target sub-scale information of each of the I matrix combinations based on the size information of the second matrix of each of the I matrix combinations and the capacity of the second storage space of each of the I matrix combinations, where the first target sub-scale information is associated with the first matrix and the second matrix; a third determination module for determining second target sub-scale information of each of the I matrix combinations based on the capacity of the first storage space of each of the I matrix combinations and the first target sub-scale information of each of the I matrix combinations, where the second target sub-scale information is related to the first matrix and a third matrix, and the third matrix corresponds to the matrix multiplication operation; a fourth determination module for determining target division information of each of the I matrix combinations based on second target sub-scale information of each of the I matrix combinations and scale information of a first matrix of each of the I matrix combinations; and a fifth determining module for determining a target matrix combination from the I matrix combinations based on bandwidth information of a computing unit that performs the matrix multiplication operation and target partition information of each of the I matrix combinations, where an access memory resource overhead corresponding to the target matrix combination satisfies a predetermined access memory resource overhead condition. A processing device for matrix multiplication data.
14. The first matrix scale information includes a first initial sub-scale value unrelated to the second matrix, the second matrix scale information includes a second initial sub-scale value unrelated to the first matrix, the first target sub-scale information includes a first target sub-scale value, and the second target sub-scale information includes a second target sub-scale value.
14. The apparatus of claim 13.
15. The second determination module further comprises: used to determine the first target sub-scale value based on a ratio of the capacity of the second storage space to the second initial sub-scale value.
15. The apparatus of claim 14.
16. The third decision module further comprises: and determining the second target sub-size value based on a ratio of the capacity of the first storage space to the first target sub-size value.
15. The apparatus of claim 14.
17. The fourth determination module further comprises: and determining the target division information based on a ratio between the first initial sub-scale value and the second target sub-scale value.
15. The apparatus of claim 14.
18. The third matrix is stored in at least two levels of cache units of a multi-level cache unit, the at least two levels of cache units including a first cache unit and a second cache unit.
14. The apparatus of claim 13.
19. The fifth determination module: a first determining submodule for determining an access memory resource overhead for each of the I matrix combinations based on bandwidth information of a computing unit that performs the matrix multiplication operation and target partition information for each of the I matrix combinations; and a second determining submodule for determining a target matrix combination from the I matrix combinations based on an access memory resource overhead of each of the I matrix combinations.
14. The apparatus of claim 13.
20. the bandwidth information includes a storage bandwidth between the computing unit and a memory unit and at least one cache bandwidth between multiple levels of cache units; The first determination submodule: a first determining unit for determining a first access memory time resource overhead based on the data amount of the first matrix, the data amount of the second matrix, the data amount of the third matrix, and the storage bandwidth; a second determination unit for determining a second access memory time resource overhead based on the target partition information, a target sub-data amount of the third matrix, and the cache bandwidth, where the target sub-data amount is a data amount of the third matrix stored in a target cache unit, and the target cache unit is a cache unit other than the first cache unit among the multiple levels of cache units; and a third determining unit for determining an access memory resource overhead of the matrix combination based on the first access memory time resource overhead and the second access memory time resource overhead.
20. The apparatus of claim 19.
21. The first determination unit comprises: a first determining subunit for determining a total data amount based on the data amount of the first matrix, the data amount of the second matrix, and the data amount of the third matrix; and a second determining subunit for determining the first access memory time resource overhead based on a ratio of the total data amount and the storage bandwidth.
21. The apparatus of claim 20.
22. The target division information includes a target division parameter value; The second determination unit, a third determining sub-unit for determining a sub-access memory time resource overhead based on a ratio between a target sub-data amount of the third matrix and the cache bandwidth; and a fourth determining subunit for determining the second access memory time resource overhead based on a product of the sub access memory time resource overhead and the target division parameter value.
21. The apparatus of claim 20.
23. and an execution module for performing the matrix multiplication operation based on first target sub-scale information and second target sub-scale information of the target matrix combination.
14. The apparatus of claim 13.
24. The predetermined access memory resource overhead condition is: an access memory resource overhead corresponding to the target matrix combination is less than or equal to a predetermined access memory resource overhead threshold; In the I matrix combinations, an access memory resource overhead corresponding to the target matrix combination is minimal.
14. The apparatus of claim 13.
25. Including a device according to any one of claims 13 to 24 electronic equipment.
26. At least one processor; a memory in communication with the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the method of any one of claims 1 to 12. electronic equipment.
27. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to carry out a method according to any one of claims 1 to 12. A non-transitory computer-readable storage medium.
28. A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Data processing device and method, electronic equipment and storage medium
CN116382593A
Information processor, matrix operation method, and matrix operation program
JP2020013412A
Method for optimizing matrix multiplication operation on system on chip, and related product
US20240028666A1