Irregular sparse matrix multiplication method and hardware architecture for Transformer model

By load-balanced RCBA arrangement of irregular sparse matrices of Transformer model and designing efficient operation arrays and data stream processing methods, the problems of irregular sparse matrix multiplication calculation in the prior art are solved, and efficient matrix multiplication operations and low-power edge deployment are realized.

CN115357850BActive Publication Date: 2025-05-09SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210818335.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-05-09
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

When handling irregular sparse matrix multiplication, existing Transformer accelerators have unbalanced calculation load, random storage memory access, and high index matching overhead, resulting in high power consumption and frequent data movement, making it difficult to deploy on the edge end with limited hardware resources.

Method used

A Transformer model multiplication operation method is proposed. By rearranging the weight matrix into a load-balanced RCBA arrangement format, and designing the operation array, including multiplication units and merging units, adopting a simple index allocation mechanism and high multiplexing data flow design, reducing intermediate result storage and movement.

Benefits of technology

It realizes efficient calculation of sparse and dense matrix multiplication, reduces index matching overhead and intermediate result cache overhead, improves data multiplexing and overall energy efficiency, and is suitable for deploying Transformer models at the edge end.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115357850B_ABST
    Figure CN115357850B_ABST
Patent Text Reader

Abstract

The present invention provides a Transformer model irregular sparse matrix multiplication operation method and hardware architecture; wherein the method is: setting an operation array: the operation array includes N operation groups; each operation group includes a multiplication unit, a distributor and a merging unit; rearranging the row and column order of the sparse weight matrix to generate a load-balanced weight matrix RCBA arrangement format; inputting into the operation array: in each operation group, the multiplication unit multiplies the non-zero value of the corresponding row module with the input matrix element corresponding to the row index of the row module one by one to obtain the multiplication result; the distributor distributes the multiplication result to each merging unit for merging according to the column index of each column module. This method can meet the multiplication acceleration of irregular sparse weight matrix and dense input matrix, and has the characteristics of low intermediate result mobile cache overhead, simple index mechanism and high reuse rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of integrated circuit design, and more specifically, to a Transformer model irregular sparse matrix multiplication operation method and hardware architecture. Background Art

[0002] Natural language processing technology is a bridge for humans to interact with machines using language. With the rapid development of deep learning, various models for natural language processing have emerged one after another, and their performance has been continuously improved. Especially after the Transformer model based on the self-attention mechanism was proposed, various Transformer-based models have achieved leading performance in most natural language processing tasks.

[0003] However, the high performance of the Transformer model comes at the cost of a huge model and massive parameters. When the Transformer model is deployed on the edge, its massive parameters bring about huge storage and computing power consumption, making it difficult to apply it on the edge with limited hardware resources. In order to accelerate the deployment of the Transformer model on the edge, it is necessary to design a dedicated accelerator for the computing and storage aspects of the Transformer model.

[0004] The design of the Transformer accelerator compresses the model algorithmically and performs a proprietary custom architecture design on the hardware. Pruning is one of the most effective methods for compressing neural network models. After pruning, the weight matrix of the model changes from a dense matrix to a sparse matrix, which brings new challenges to hardware calculations. The pruned sparse matrix is ​​divided into regular sparse and irregular sparse weight matrices according to its regularity. When the irregular sparse matrix weights with the largest application range are calculated on the hardware, due to the completely irregular distribution of its non-zero weight values, the calculation load of the operation array is unbalanced, the random access of the storage unit, and the high index matching overhead are caused. The existing Transformer accelerators are divided into the operation and storage architecture design for accelerating the operation of specific regular sparse weight matrices with a low application range, the operation and storage architecture design for accelerating the operation of completely dense weight matrices, and the operation and storage architecture design for accelerating the operation of irregular sparse weight matrices. The Transformer accelerator that accelerates irregular sparse weights considers data flows where both the input and weights are completely sparse or dense. Ignoring that most of the Transformer model's calculations are multiplications of dense input matrices and sparse weight matrices, and ignoring the knowability of the computational load of this sparse dense matrix multiplication, it results in unnecessary data movement and index matching overhead in hardware design. At the same time, its data reusability is low, and the intermediate result mobile cache overhead is high, which increases the overall data movement power consumption. For the deployment of the Transformer model at the edge, power consumption is one of the most important considerations. Therefore, there is an urgent need to design an irregular sparse matrix multiplication operation method tailored to the characteristics of the Transformer computational load to improve the overall energy efficiency of the accelerator. Summary of the invention

[0005] In order to overcome the shortcomings and deficiencies in the prior art, the purpose of the present invention is to provide a Transformer model irregular sparse matrix multiplication operation method and hardware architecture; this method makes full use of the known nature of the computational load of the multiplication of the sparse weight matrix and the dense input matrix, which accounts for the vast majority of the calculations in the Transformer model, and can eliminate the complex index matching mechanism and overhead, reduce the storage and movement of intermediate calculation results, and has a high data reuse rate; it can meet the acceleration of the multiplication of the irregular sparse weight matrix and the dense input matrix of the Transformer model, and has the characteristics of low intermediate result mobile cache overhead, simple index mechanism and high reuse rate.

[0006] In order to achieve the above object, the present invention is implemented by the following technical scheme: a Transformer model irregular sparse matrix multiplication method, characterized in that it includes the following steps:

[0007] S1. Set the operation array: the operation array includes N operation groups; N is the vector-matrix multiplication parallelism; each operation group includes G1 multiplication units, a distributor and G2 merging units;

[0008] S2. Rearrange the order of rows and columns of the sparse weight matrix of the Transformer model to generate a load-balanced weight matrix RCBA (Row-column Balance Arrange) arrangement format; wherein the method for generating the weight matrix RCBA arrangement format is: count the number of non-zero elements in each row and column of the weight matrix respectively; divide the rows of the weight matrix into G1 row modules, so that the number of non-zero elements in each row module is roughly equal, and record the row index of each row module; divide the columns of the weight matrix into G2 column modules, so that the number of non-zero elements in each column module is roughly equal, and record the column index of each column module;

[0009] S3. Input the input matrix and weight matrix in RCBA arrangement format into the operation array: in each operation group, G1 row modules correspond to G1 multiplication units, and the G1 multiplication units respectively multiply the non-zero value of the corresponding row module with the input matrix elements corresponding to the row index of the row module one by one to obtain the multiplication result; the distributor distributes the multiplication results output by the G1 multiplication units to the G2 merging units according to the column index of each column module, and the G2 merging units merge the input multiplication results respectively; the results obtained by the G2 merging units are respectively output to the storage units.

[0010] Preferably, in said S2, the method for generating the weight matrix RCBA arrangement format further includes: rearranging the order of the allocated row modules and the input order of the non-zero values ​​in each row module to avoid input conflicts of the merging unit.

[0011] Preferably, in S2, the row module generation method is: rearrange according to the number of non-zero elements in each row, and cyclically distribute the non-zero elements in each row to each row module in turn, so that the number of non-zero elements in each row module is approximately equal;

[0012] The column module generation method is: rearrange according to the number of non-zero elements in each column, and cyclically distribute the non-zero elements in each column to each column module in turn, so that the number of non-zero elements in each column module is roughly equal.

[0013] Preferably, each multiplication unit comprises a multiplier, a multiplier cache, a multiplication result cache and a control module;

[0014] In S3, the operation method of the multiplication unit is: according to the corresponding relationship between G1 row modules and G1 multiplication units, the input allocation signal and the transfer signal are set; the control module allocates the non-zero value of the row module and the input matrix element input to the corresponding multiplier cache according to the input allocation signal, and selects whether to change the input of the multiplier according to the transfer signal, performs a multiplication once per cycle, and the multiplication result is cached in the corresponding multiplication result cache, and outputs a multiplication result per cycle.

[0015] Preferably, each merging unit includes an adder, an addition signal buffer, an addition intermediate result buffer, a merging result buffer, a quantization module and an input and output control unit;

[0016] In S3, the operation method of the merging unit is: the distributor sets the switching signal according to the column index of each column module; the input and output control unit distributes the input multiplication result to the addition intermediate result cache according to the switching signal, and determines whether the adder performs partial sum accumulation of the current column or partial sum accumulation of the new column. When the partial sum accumulation is completed, the result is stored in the corresponding merged result cache through the input and output control unit, and finally the result in the merged result cache is quantized by the quantization module to obtain the final output.

[0017] A hardware architecture for irregular sparse matrix multiplication operation of a Transformer model, characterized by comprising:

[0018] The weight matrix RCBA arrangement format generation module is used to rearrange the order of rows and columns of the sparse weight matrix of the Transformer model to generate a load-balanced weight matrix RCBA arrangement format; wherein, the weight matrix RCBA arrangement format generation method is: count the number of non-zero elements in each row and column of the weight matrix respectively; divide the rows of the weight matrix into G1 row modules, so that the number of non-zero elements in each row module is roughly equal, and record the row index of each row module; divide the columns of the weight matrix into G2 column modules, so that the number of non-zero elements in each column module is roughly equal, and record the column index of each column module;

[0019] A multiplication operation module is used to perform multiplication operations on an input matrix and a weight matrix RCBA arrangement format; the multiplication operation module includes N operation groups; N is the vector matrix multiplication parallelism; each operation group includes G1 multiplication units, a distributor and G2 merging units; in each operation group, G1 row modules correspond to G1 multiplication units, and the G1 multiplication units respectively multiply the non-zero values ​​of the corresponding row modules with the input matrix elements corresponding to the row index of the row module one by one to obtain the multiplication results; the distributor distributes the multiplication results output by the G1 multiplication units to the G2 merging units according to the column index of each column module, and the G2 merging units respectively merge the input multiplication results; the results obtained by the G2 merging units are respectively output to the storage units.

[0020] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0021] (1) The present invention focuses on matrix multiplication, which accounts for the vast majority of calculations in the Transformer model. In view of the multiplication characteristics of sparse matrices and dense matrices that are different from ordinary matrix multiplication, the proposed weight matrix RCBA arrangement format balances the load of each multiplier and adder, and has a high hardware computing load balancing performance.

[0022] (2) The index allocation mechanism proposed in the present invention directly allocates the multiplication output results through the input column index, avoiding the high-overhead index matching mechanism of irregular sparse matrices in the existing Transformer accelerator, avoiding the frequent movement and caching of index data, and reducing the power consumption of data movement.

[0023] (3) The present invention distributes the output results of the multiplication unit array to the merging unit array for accumulation in real time, avoiding the caching and matching of intermediate results, reducing the size of the on-chip cache, and reducing the area overhead. At the same time, since the intermediate results exist for a short time, a smaller cache overhead can be used to exchange for a larger reuse rate, reducing the power consumption of data movement.

[0024] (4) The intermediate results of the present invention are used immediately, and data is reused from the input matrix, weight matrix, and partial sum three dimensions. Adjusting the size of parameters A and N can make the operation array calculate the entire output matrix in parallel at the same time. The input and weight data only need to be written once from outside the operation array, and the output matrix only needs to be written out once. Data movement is minimized and the total power consumption of the operation is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is the overall architecture of the operation array of the irregular sparse matrix multiplication operation method of the Transformer model of the present invention;

[0026] Figure 2It is a multiplication unit architecture of a computing array of the irregular sparse matrix multiplication computing method of the Transformer model of the present invention;

[0027] Figure 3 It is a merging unit architecture of the operation array of the irregular sparse matrix multiplication operation method of the Transformer model of the present invention;

[0028] Figure 4 are the input matrix and weight matrix of the embodiment;

[0029] Figure 5 is a computing array architecture of an embodiment;

[0030] Figure 6(a) to Figure 6(c) Schematic diagram of the weight matrix RCBA arrangement format generation process of the embodiment;

[0031] Figure 7 Schematic diagram of parallel data stream allocation in an embodiment. DETAILED DESCRIPTION

[0032] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0033] Example

[0034] This embodiment provides a Transformer model irregular sparse matrix multiplication operation method, and its operation array architecture is as follows: Figure 1 As shown, the architecture of the multiplication unit and the merging unit in the operation array architecture is as follows Figure 2 and Figure 3 shown.

[0035] Assume that the input matrix is ​​I 2x8 , the weight matrix is ​​W 8x16 , where the subscripts are the dimensions of the matrix rows and columns, such as Figure 4 As shown, the number of operation groups N is set to 2; the number of multiplication units G1 in each operation group is 2, and the number of merging units G2 is 4.

[0036] The method of this embodiment includes the following steps:

[0037] S1. According to the RCBA format of the input matrix and the weight matrix, the operation array is set: the operation array includes N operation groups; N is the vector matrix multiplication parallelism; each operation group includes G1 multiplication units, a distributor and G2 merging units, and adopts a multiplier and adder separation architecture. In this embodiment, N = 2, G1 = 2, G2 = 4, so the operation array architecture is as follows Figure 5 shown.

[0038] S2. Rearrange the order of rows and columns of the sparse weight matrix of the Transformer model to generate a load-balanced weight matrix RCBA (Row-column Balance Arrange) arrangement format;

[0039] Among them, the method for generating the weight matrix RCBA arrangement format is:

[0040] First, count the number of non-zero elements in each row and column of the weight matrix;

[0041] Then, the rows of the weight matrix are divided into G1 row modules: the non-zero elements of each row are rearranged according to the number of non-zero elements in each row, and the non-zero elements of each row are cyclically distributed to each row module in turn, so that the number of non-zero elements in each row module is roughly equal, as shown in Figure 6(a), and the row index of each row module is recorded; the calculation load balance of the multiplication unit is achieved;

[0042] Afterwards, the columns of the weight matrix are divided into G2 column modules: they are rearranged according to the number of non-zero elements in each column, and the non-zero elements in each column are cyclically distributed to each column module in turn, so that the number of non-zero elements in each column module is roughly equal, as shown in Figure 6(b), and the column index of each column module is recorded; the computational load balance of the merging unit is achieved.

[0043] As shown in Figure 6(c), according to the data distribution method of the multiplication unit output and the merging unit input, in order to make the result of each distribution array a valid input for each merging unit, that is, the output result of each multiplication unit is distributed to different merging units according to the column index for merging, and to avoid input conflicts of the merging units, the order of each row module and the input order of non-zero values ​​in each row module need to be rearranged to obtain the final weight matrix RCBA arrangement format. Since the distributor only distributes one group of multiplication results to the merging unit in each cycle, it is necessary to ensure that the group of multiplication results can be distributed to each merging unit at the same time according to the distribution rules, that is, the column index of the multiplication result belongs to each merging group. As shown in Figure 6(b), merging units 1 and 2 are used as merging group 0, and merging units 3 and 4 are used as merging group 1. The first group of inputs in Figure 6(c) is the multiplication result obtained by non-zero weights 6 and 1, and the corresponding column index values ​​are 1 and 2 respectively. According to the allocation method of Figure 6(b), column index values ​​1 and 2 both belong to the computational load of merge group 1, resulting in that merge group 0 has no corresponding input in this cycle while merge group 1 has two inputs, resulting in unbalanced inputs. Therefore, the order of multipliers input to the multiplication unit is adjusted to adjust the input of the merge unit. As shown in Figure 6(c), non-zero values ​​6 and 1 are adjusted to 6 and 2, and non-zero values ​​6 and 2 belong to merge groups 0 and 1 respectively, thereby obtaining balanced merge unit inputs. Similarly, non-zero values ​​7 and 2 are adjusted to 7 and 3, 8 and 3 are adjusted to 8 and 4, and so on. The order of multiplier inputs is rearranged to obtain the multiplier input order of Figure 6(c).

[0044] S3. Input the input matrix and weight matrix in RCBA arrangement format into the operation array: in each operation group, G1 row modules correspond to G1 multiplication units, and the G1 multiplication units respectively multiply the non-zero value of the corresponding row module with the input matrix elements corresponding to the row index of the row module one by one to obtain the multiplication result; the distributor distributes the multiplication results output by the G1 multiplication units to the G2 merging units according to the column index of each column module, and the G2 merging units merge the input multiplication results respectively; the results obtained by the G2 merging units are respectively output to the storage units.

[0045] Specifically, each multiplication unit includes a multiplier, a multiplier buffer, a multiplication result buffer and a control module. The control module is responsible for the input multiplier allocation and the control of the output valid signal. The multiplication unit generates a partial sum result obtained by multiplying a single input non-zero value with the weight matrix row one by one. In S3, the operation method of the multiplication unit is: according to the corresponding relationship between G1 row modules and G1 multiplication units, the input allocation signal and the transfer signal are set; the control module allocates the non-zero value of the row module and the input matrix element input to the corresponding multiplier buffer according to the input allocation signal, and selects whether to change the input of the multiplier according to the transfer signal. Multiplication is performed once per cycle, and the multiplication result is cached in the corresponding multiplication result buffer, and a multiplication result is output per cycle.

[0046] Each merging unit includes an adder, a switching signal buffer, an addition intermediate result buffer, a merging result buffer, a quantization module and an input-output control unit. The input-output control unit controls input allocation, output validity, interaction with the intermediate result buffer, and valid output after quantization by the quantization module. The merging unit is responsible for merging the multiplication results output by the multiplication unit obtained by the distributor; the final result obtained by the merging unit is output to the storage unit. In S3, the operation method of the merging unit is: the distributor sets the switching signal according to the column index of each column module; the input-output control unit distributes the input multiplication result to the addition intermediate result buffer according to the switching signal, determines whether the adder performs partial sum accumulation of the current column or partial sum accumulation of the new column, and when the partial sum accumulation is completed, the result is stored in the corresponding merging result buffer through the input-output control unit, and finally the result in the merging result buffer is quantized by the quantization module to obtain the final output.

[0047] The distributor is located between the multiplication unit and the merging unit, and is responsible for distributing the multiplication results output by the multiplication unit to the corresponding merging unit. The multiplication part and are distributed by the column index corresponding to the input weight matrix. The multiplication results output by the multiplication unit are arranged according to the distribution index values ​​in the weight matrix RCBA arrangement format described in the invention. Each index value corresponds to a merging unit. The distributor distributes the multiplication results calculated in each cycle to each merging unit in turn according to the index, so that the merging unit load is balanced and has high utilization.

[0048] The present invention adopts parameterized parallel data flow, merges parts of vectors one by one and reduces the intermediate result cache. The parallelism of vector matrix multiplication is 2, that is, the operation array performs two vector matrix multiplications at the same time; the number of rows calculated by each multiplication unit and the number of columns calculated by each merging unit are 4; the total number of multiplication units and merging units are 4 and 8 respectively, all multiplication units and merging units are divided into 2 groups, and the number of multiplication units and merging units in each group is 2 and 4; the input matrix dimension is 2x8, and the weight matrix dimension is 8x16. By adjusting the size of the data flow allocation parameter to adjust the allocation balance rate and the number of parallel calculations, data movement is minimized.

[0049] According to the area resource constraints and throughput requirements of the design, determine the appropriate operation array multiplication unit and merging unit. Group the multiplication units and merging units in the designed operation array. Each multiplication unit and merging group calculates a vector-matrix multiplication in parallel. According to the number of rows of the input matrix and the total number of multiplication units and merging units, select the appropriate number of groups N so that the operation array can simultaneously calculate the entire matrix multiplication in parallel. The number of groups N is the parallelism of the vector-matrix multiplication. Merging the parts of the vectors one by one reduces the intermediate result cache.

[0050] According to the total number of multiplication units and merging units and the size of the number of groups, the number of multiplication units and merging units in each group is 2 and 4 respectively. According to the calculated matrix dimension and the number of multiplication units and merging units in each group, the merged column load size of each merging unit is calculated to be 4.

[0051] Each group of multiplication units and merging units calculates a vector-matrix multiplication with a parallelism of 2, that is, the operation array simultaneously calculates and executes 2 vector-matrix multiplications. By increasing the parallelism N, that is, the number of vectors calculated simultaneously by the operation array each time, the weight matrix reuse can be improved and the computing power consumption can be reduced.

[0052] After the input elements are assigned to the corresponding multiplication units, the two multiplication units calculate the partial sum of the vector-matrix multiplication results, which are then distributed to the corresponding four merging units by column index through the distributor for accumulation to obtain the final calculation result.

[0053] The present invention adopts a multi-dimensional data reuse strategy to reuse data from three perspectives: input matrix, weight matrix and output part, striving to minimize data movement and reduce the total power consumption of matrix multiplication operations.

[0054] First, the input data is reused inside the multiplication unit. For each element in the column of the input matrix, it is fixed as an input inside the multiplier of the multiplication unit, and the corresponding weight matrix rows are input in turn and multiplied with the element to achieve input reuse.

[0055] The output partial sum of the multiplier is accumulated inside the merging unit until the partial sum of all dimensions in the column is accumulated to obtain the final output and then output to the storage unit. The output is multiplexed in the merging unit.

[0056] Each operation group in the operation array calculates a vector-matrix multiplication in parallel, where the weight matrix is ​​shared among the various operation groups. The weight matrix is ​​written from outside the operation array and multicast to each group, realizing weight reuse, reducing the frequent data movement of the weight matrix between the operation array and the storage unit, and reducing operation power consumption.

[0057] In order to implement the above-mentioned Transformer model irregular sparse matrix multiplication operation method, this embodiment also provides a hardware architecture for Transformer model irregular sparse matrix multiplication operation, including:

[0058] The weight matrix RCBA arrangement format generation module is used to rearrange the order of rows and columns of the sparse weight matrix of the Transformer model to generate a load-balanced weight matrix RCBA arrangement format; wherein, the weight matrix RCBA arrangement format generation method is: count the number of non-zero elements in each row and column of the weight matrix respectively; divide the rows of the weight matrix into G1 row modules, so that the number of non-zero elements in each row module is roughly equal, and record the row index of each row module; divide the columns of the weight matrix into G2 column modules, so that the number of non-zero elements in each column module is roughly equal, and record the column index of each column module;

[0059] A multiplication operation module is used to perform multiplication operations on an input matrix and a weight matrix RCBA arrangement format; the multiplication operation module includes N operation groups; N is the vector matrix multiplication parallelism; each operation group includes G1 multiplication units, a distributor and G2 merging units; in each operation group, G1 row modules correspond to G1 multiplication units, and the G1 multiplication units respectively multiply the non-zero values ​​of the corresponding row modules with the input matrix elements corresponding to the row index of the row module one by one to obtain the multiplication results; the distributor distributes the multiplication results output by the G1 multiplication units to the G2 merging units according to the column index of each column module, and the G2 merging units respectively merge the input multiplication results; the results obtained by the G2 merging units are respectively output to the storage units.

[0060] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A Transformer model irregular sparse matrix multiplication method, characterized by: The steps include: S1. Set the operation array: the operation array includes N operation groups; N is the vector-matrix multiplication parallelism; each operation group includes G1 multiplication units, a distributor and G2 merging units; S2. Rearrange the order of rows and columns of the sparse weight matrix of the Transformer model to generate a load-balanced weight matrix RCBA arrangement format; wherein the method for generating the weight matrix RCBA arrangement format is: count the number of non-zero elements in each row and column of the weight matrix respectively; divide the rows of the weight matrix into G1 row modules, so that the number of non-zero elements in each row module is roughly equal, and record the row index of each row module; divide the columns of the weight matrix into G2 column modules, so that the number of non-zero elements in each column module is roughly equal, and record the column index of each column module; S3, input the input matrix and the weight matrix RCBA arrangement format into the operation array: in each operation group, G1 row modules correspond to G1 multiplication units, and the G1 multiplication units respectively multiply the non-zero value of the corresponding row module with the input matrix element corresponding to the row index of the row module one by one to obtain the multiplication result; the distributor distributes the multiplication results output by the G1 multiplication units to the G2 merging units according to the column index of each column module, and the G2 merging units merge the input multiplication results respectively; the results obtained by the G2 merging units are respectively output to the storage unit; In S2, the row module generation method is: rearrange according to the number of non-zero elements in each row, and cyclically distribute the non-zero elements in each row to each row module in turn, so that the number of non-zero elements in each row module is roughly equal; The column module generation method is: rearrange according to the number of non-zero elements in each column, and cyclically distribute the non-zero elements in each column to each column module in turn, so that the number of non-zero elements in each column module is roughly equal.

2. The irregular sparse matrix multiplication method of the Transformer model according to claim 1, characterized in that: In S2, the method for generating the weight matrix RCBA arrangement format further includes: rearranging the order of the allocated row modules and the input order of the non-zero values ​​in each row module to avoid input conflicts in the merging unit.

3. The irregular sparse matrix multiplication method of the Transformer model according to claim 1, characterized in that: Each multiplication unit includes a multiplier, a multiplier cache, a multiplication result cache and a control module; In S3, the operation method of the multiplication unit is: according to the corresponding relationship between G1 row modules and G1 multiplication units, the input allocation signal and the transfer signal are set; the control module allocates the non-zero value of the row module and the input matrix element input to the corresponding multiplier cache according to the input allocation signal, and selects whether to change the input of the multiplier according to the transfer signal, performs a multiplication once per cycle, and the multiplication result is cached in the corresponding multiplication result cache, and outputs a multiplication result per cycle.

4. The irregular sparse matrix multiplication method of the Transformer model according to claim 1, characterized in that: Each merging unit includes an adder, an addition signal buffer, an addition intermediate result buffer, a merging result buffer, a quantization module and an input and output control unit; In S3, the operation method of the merging unit is: the distributor sets the switching signal according to the column index of each column module; the input and output control unit distributes the input multiplication result to the addition intermediate result cache according to the switching signal, and determines whether the adder performs partial sum accumulation of the current column or partial sum accumulation of the new column. When the partial sum accumulation is completed, the result is stored in the corresponding merged result cache through the input and output control unit, and finally the result in the merged result cache is quantized by the quantization module to obtain the final output.

5. A hardware architecture for irregular sparse matrix multiplication of a Transformer model, characterized by: include: The weight matrix RCBA arrangement format generation module is used to rearrange the order of rows and columns of the sparse weight matrix of the Transformer model to generate a load-balanced weight matrix RCBA arrangement format; wherein, the weight matrix RCBA arrangement format generation method is: count the number of non-zero elements in each row and column of the weight matrix respectively; divide the rows of the weight matrix into G1 row modules, so that the number of non-zero elements in each row module is roughly equal, and record the row index of each row module; divide the columns of the weight matrix into G2 column modules, so that the number of non-zero elements in each column module is roughly equal, and record the column index of each column module; A multiplication operation module is used to perform multiplication operation on the input matrix and the weight matrix RCBA arrangement format; the multiplication operation module includes N operation groups; N is the vector matrix multiplication parallelism; each operation group includes G1 multiplication units, a distributor and G2 merging units; in each operation group, G1 row modules correspond to G1 multiplication units, and the G1 multiplication units respectively multiply the non-zero values ​​of the corresponding row modules with the input matrix elements corresponding to the row index of the row module one by one to obtain the multiplication results; the distributor distributes the multiplication results output by the G1 multiplication units to the G2 merging units according to the column index of each column module, and the G2 merging units respectively merge the input multiplication results; the results obtained by the G2 merging units are respectively output to the storage units; In the weight matrix RCBA arrangement format generation module, the row module generation method is: rearrange according to the number of non-zero elements in each row, and cyclically distribute the non-zero elements in each row to each row module in turn, so that the number of non-zero elements in each row module is roughly equal; The column module generation method is: rearrange according to the number of non-zero elements in each column, and cyclically distribute the non-zero elements in each column to each column module in turn, so that the number of non-zero elements in each column module is roughly equal.

Citation Information

Patent Citations

  • Transformer neural network-based model compression method and matrix multiplication module

    CN113486298A

  • Low-voltage distribution network phase-household relation identification method based on matrix completion

    CN113839384A