Transform network attention matrix sparse processing system based on non-zero highest order detection

By designing the Transformer network attention matrix sparse processing system based on non-zero highest bit detection, the problem of high cost and high latency of attention matrix sparse hardware in the Transformer network is solved, and the low-cost and low-latency matrix sparse is achieved, which improves computing efficiency and storage capabilities.

CN120068959APending Publication Date: 2025-05-30SOUTHEAST UNIV +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510084150.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The hardware implementation of attention matrix sparse hardware in Transformer network is costly and has a large processing delay, making it difficult to meet the challenges of computing resources and storage requirements.

Method used

A Transformer network attention matrix sparse processing system based on non-zero highest bit detection is designed, including a non-zero highest bit detection module, a maximum value search module and a mask generation module. Through these modules, the sparse mask is gradually generated and matrix sparse is carried out.

Benefits of technology

Low-cost and low-latency matrix sparseness is achieved, significantly reducing time complexity, simplifying circuit design, and improving computing efficiency and storage capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068959A_ABST
    Figure CN120068959A_ABST
Patent Text Reader

Abstract

The invention discloses a Transform network attention matrix sparse processing system based on non-zero highest bit detection. The Transform network attention matrix sparse processing system comprises a non-zero highest bit detection module, a maximum value search module and a mask generation module. The non-zero highest-bit detection module is used for carrying out non-zero highest-bit detection on all data in the quantized attention matrix column by column, and enabling all the data to be downward approximate to an index of 2; the maximum value search module is used for carrying out maximum value search on numerical values of each row in the matrix column by column; and the mask generation module is used for generating sparse masks column by column according to the maximum value information represented by the non-zero highest bit and a preset threshold value, and multiplying the mask matrix by the original attention matrix element by element to complete matrix sparseness. According to the invention, a low-cost and low-delay hardware implementation method can be provided for matrix sparseness after the Softmax function in the Transform network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of neural network accelerator design, and particularly relates to a sparse processing system for the attention matrix of a Transformer network based on non-zero most significant bit detection. Background Art

[0002] Due to its unique self-attention mechanism, the Transformer network has superior performance in dealing with the context understanding problem of long-distance dependencies. Dense matrix calculations are the core operations of the Transformer network, and performing these calculations requires a large amount of computing resources, which may conflict with hardware resource limitations. In addition, the Transformer network has a large number of model parameters and a large number of redundant parameters, which makes the storage and loading of the model on devices a challenge.

[0003] To address these problems, model compression techniques such as sparsity and quantization are mainly adopted at present to reduce the volume of the model, remove redundant parameters that contribute less to the model performance, and design hardware accelerators based on this. Sparsity refers to the process of removing or reducing weights and connections in a neural network that contribute less to the model performance by utilizing zero or near-zero values in the matrix, thereby reducing the model volume and improving the computing efficiency. The sparsity in the Transformer network mainly occurs in the process of applying the Softmax function to the attention matrix. Therefore, it is necessary to sparsify the attention matrix after being processed by the Softmax function, and this usually requires a relatively complex circuit design on hardware. Summary of the Invention

[0004] In order to reduce the hardware implementation cost of sparsifying the matrix in the Transformer network and reduce the processing delay, the present invention proposes a sparse processing system for the attention matrix of a Transformer network based on non-zero most significant bit detection, achieving low-cost and low-latency matrix sparsity.

[0005] The present invention discloses a sparse processing system for the attention matrix of a Transformer network based on non-zero most significant bit detection, which includes a non-zero most significant bit detection module, a maximum value search module, and a mask generation module connected in sequence;

[0006] The non-zero most significant bit detection module is used to perform non-zero most significant bit detection on all data in the quantized attention matrix column by column, approximate all data downward to an exponent of 2, and obtain a matrix with the same size as the input matrix, where each element can reflect the numerical size of the original matrix;

[0007] The maximum value search module is used to search for the maximum value of each row in the matrix column by column to obtain the maximum value of each row in the entire matrix;

[0008] The mask generation module is used to generate a sparse mask column by column according to the maximum value information represented by the non-zero most significant bit and a preset threshold, obtaining a mask matrix with the same size as the original matrix, and multiplying the mask matrix and the original attention matrix element by element to complete matrix sparsification.

[0009] Further, the non-zero most significant bit detection module includes m LOD units, each LOD unit is composed of combinational logic, and from left to right, each time it converts the data of m rows and 1 column in the matrix block; after conversion, the value only contains a single-bit 1 in binary, and the rest of the bits are all 0, which is expressed by the formula:

[0010]

[0011] In the formula, X in represents the data of m rows and 1 column in the attention matrix; X out represents the output data processed by the LOD unit.

[0012] Further, the maximum value search module includes m search units, each search unit is composed of a comparator, a multiplexer and a cache unit; the maximum value search module receives the input matrix data of m rows and 1 column column by column, compares it with the current maximum value stored in the cache unit, if the input value is greater than the current maximum value, updates the maximum value in the cache unit by controlling the multiplexer, until all column data are input, and outputs the maximum value of each row; the formula of the maximum value search module is expressed as:

[0013]

[0014] In the formula, X ij represents the matrix data of the i-th row and j-th column, and Max i represents the maximum value of the i-th row.

[0015] Further, the mask generation module includes m mask generation units and a sparsification unit; each mask generation unit shifts the obtained maximum value to the right by T bits according to a preset threshold T and sends it to a subtractor, subtracts the shifted value with a fixed value 9^′b100000000, then performs a bitwise AND with the current input value and sends it to the first combinational logic unit, and finally sends it to the second combinational logic unit for bitwise OR. If the result is 1, the mask is set to 1, otherwise it is set to 0, generating the mask for the corresponding row in the matrix block; the formula of the mask generation unit is expressed as:

[0016] M i = Max i >> (T)

[0017] N i= (9'b1 0000 0000) - M i

[0018] O ij = N i & X ij

[0019]

[0020] Wherein, M i represents the data shifted right by T bits, N i represents the output data of the subtractor, O ij represents the output data of the first combinational logic unit, X ij represents the matrix data of the i-th row and j-th column, Mask ij represents the mask corresponding to the matrix data of the i-th row and j-th column;

[0021] The sparsification unit performs an element-wise multiplication operation on the mask generated by the mask generation unit, which has the same size as the original matrix block, and the original matrix block to obtain a sparsified matrix block.

[0022] The beneficial effects of the present invention are as follows:

[0023] First, in the Transformer network attention matrix sparse processing system based on non-zero most significant bit detection of the present invention, the threshold for mask generation is defined as the ratio of each input value to the maximum value in its row. Compared with the traditional Top-K method, using this threshold significantly reduces the time complexity in the process of generating the mask.

[0024] Second, in the Transformer network attention matrix sparse processing system based on non-zero most significant bit detection of the present invention, the LOD module converts the matrix values into the form of powers of 2, enabling full utilization of the numerical characteristics of binary in the process of mask generation, with simple circuit implementation and advantages in both power consumption and area. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic diagram of the sparsified attention matrix in the calculation process of the Transformer network;

[0026] Figure 2 In (a) and (b) are schematic diagrams of the embodiment architecture and the corresponding calculation process including the solution of the present invention;

[0027] Figure 3 is a schematic diagram of the process of the solution of the present invention;

[0028] Figure 4 In (a) and (b) are schematic diagrams of the function and circuit structure of the LOD module in the present invention;

[0029] Figure 5 It is a schematic diagram of the maximum value search module in the present invention;

[0030] Figure 6 It is a schematic diagram of the mask generation module in the present invention. Specific embodiments

[0031] The following embodiments can enable those skilled in the art to understand the present invention more comprehensively, but do not limit the present invention in any way.

[0032] The present invention discloses a sparse processing system for the attention matrix of a Transformer network based on non-zero most significant bit detection. The system includes a non-zero most significant bit detection (LOD) module, a maximum value search module, and a mask generation module connected in sequence.

[0033] The LOD module is used to perform non-zero most significant bit detection on all data in the matrix; specifically, the LOD module includes several LOD units, each LOD unit is implemented by combinational logic, the LOD module processes m rows and 1 column of data in parallel each time, after processing is completed, it processes the data of the next column, after the m rows of data are processed, it switches to new m rows of data and continues to process until the entire matrix is processed. The specific processing process is to perform non-zero most significant bit detection on all data in the matrix, approximate all data downward to the exponent of 2, and obtain a matrix with the same size as the input matrix, where each element can reflect the numerical size of the original matrix.

[0034] The maximum value search module is used to search for the maximum value in each row of the matrix; specifically, the maximum value search module is composed of several maximum value search units, each maximum value search unit includes a comparator, a multiplexer, and a maximum value cache module, and each unit processes one row of data in the matrix; the maximum value search module processes m rows and 1 column of data in parallel each time, compares the input data with the maximum value in the cache unit, if the input value is larger, updates the maximum value in the cache unit through the multiplexer, and then processes the data of the next column. After the m rows of data are processed, the maximum values of the m rows are obtained, and then it switches to new m rows and continues to search until the maximum value of each row in the entire matrix is obtained.

[0035] The mask generation module is used to generate a mask according to the maximum value information represented by the non-zero highest bit and a preset threshold. Specifically, the mask generation module includes several mask generation units. Each mask generation unit includes a shift unit, a subtractor, a combinational logic unit, and a cache unit, and can process one row of data in the matrix. The mask generation module processes m rows and 1 column of data in parallel each time. After being processed by the LOD module, the relative magnitude relationship between the values is also a multiple of 2 at this time. Therefore, according to the preset multiple threshold, it can be determined whether the input needs to be sparse through simple logical operations, and is correspondingly marked as 0 or 1 as the generated sparse mask. After the processing is completed, the data of the next column is processed. After the m rows of data are processed, it switches to new m rows of data for continuous processing until the entire matrix is processed. At this time, a mask matrix with the same size as the original matrix can be obtained, and then it is multiplied element by element with the original matrix to complete matrix sparsification.

[0036] Figure 1 It is a schematic diagram of the calculation process of one layer of the Transformer network involved in the present invention. The input and output sizes and calculation processes of different layers in the network are exactly the same. Among them, the attention matrix exists in the calculation of the attention sublayer. Based on this calculation process, Figure 2 It shows the hardware structure of a Transformer accelerator adopting the matrix sparsification hardware implementation method disclosed in the present invention, which mainly includes a matrix calculation module, a non-linear calculation module, an approximate calculation module, a cache module, and the matrix sparsification module described in the present invention. In this embodiment, according to the different number of calculation layers, the input matrix can be either the input of the entire network or the output result from the previous layer of the network, and the size is 64*768. After being processed by the approximate calculation module, the corresponding attention matrix size is 64*64*12, and sparsity is generated row by row. Therefore, the matrix is divided into 16*64*4*12, where 16 corresponds to the parallel degree of the subsequent sparsification module, and a matrix block with a size of 16*64 is sparsified each time.

[0037] Figure 3 It is a schematic diagram of the calculation process of the matrix sparsification hardware implementation method described in the present invention. The input each time is a quantized matrix block with a size of 16*64. First, the non-zero highest bit detection is performed on it, and all the values in the matrix block are converted into exponents of 2. Then, based on this, the maximum value search is performed row by row to obtain the maximum value in each row and send it to the mask generation module to generate a mask according to the preset threshold. The overall hardware implementation scheme includes three modules: the non-zero highest bit detection (LOD) module, the maximum value search module, and the mask generation module;

[0038] Figure 4This is a schematic diagram of the LOD module function in the present invention. Through this module, all the values in the matrix block are converted into exponents of 2. The LOD module is composed of 16 LOD units in total. Each LOD unit is composed of combinational logic. From left to right, each time the data of 16 rows and 1 column in the matrix block is converted. After conversion, the value in binary only contains a single-bit 1 and the remaining bits are all 0. The formula is expressed as:

[0039]

[0040] In this embodiment, the data bit width is 8 bits. Therefore, the converted data contains 1, 2, 4, 8, 16, 32, 64, 128 in decimal.

[0041] Figure 5 This is a schematic diagram of the maximum value search module in the present invention, which is composed of 16 search units in total. Each search unit is composed of a comparator, a multiplexer and a cache unit. From left to right, each time the matrix data of 16 rows and 1 column is input into the module. The current maximum value and the input value are sent to the comparator for comparison. If the input value is greater than the current maximum value, the maximum value is updated by controlling the multiplexer until all 64 columns of data are input, and the maximum value of each row is obtained. A total of 16 search units obtain 16 * 1 maximum values. The formula is expressed as:

[0042]

[0043] (i = 1, 2, 3,..., 16; j = 1, 2, 3,..., 64).

[0044] Figure 6 This is a schematic diagram of the mask generation module in the present invention, which is composed of 16 mask generation units in total. From left to right, each time the matrix block data of 16 rows and 1 column is input into the module. Each mask generation unit generates a mask for the corresponding row in the matrix block. First, according to the pre-set threshold T, the maximum value obtained by the LOD module is shifted to the right by T bits, and then sent to the subtractor. The fixed value (9'b1 0000 0000) is used to subtract the shifted value, and then the shifted value and the input value at this time are sent to the combinational logic unit 1 for bitwise AND, and finally sent to the combinational logic unit 2 for bitwise OR. If the result is 1, the mask is set to 1, otherwise it is set to 0. The formula is expressed as:

[0045] M i =Max i >>(T)

[0046] N i =(9'b1 0000 0000)-M i

[0047] O ij =N i &Xij

[0048]

[0049] The matrix blocks are input into the module column by column to generate a mask with the same size as the original matrix blocks. Subsequently, by performing an element-wise multiplication operation on the generated mask and the original matrix blocks, the sparsified matrix blocks are obtained. After performing the above operations on all matrix blocks, a sparse matrix and a mask matrix with a size of 64×64×12 are finally obtained, both of which are consistent with the size of the original input matrix. Next, the generated mask matrix and sparse matrix are respectively transmitted to the read-write control module and the matrix calculation module to complete subsequent calculations.

[0050] In summary, the present invention provides a hardware implementation method for sparse attention matrices of a Transformer network based on non-zero most significant bit detection. By defining the mask generation threshold as the ratio of each value to the maximum value in its row, the time complexity is significantly reduced. In addition, the LOD module is adopted to convert the values in the matrix into the form of powers of 2, so as to make full use of binary numerical features during the mask generation process, simplifying the circuit design. Low-cost and low-latency matrix sparsification is achieved.

[0051] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A Transformer network attention matrix sparse processing system based on non-zero highest bit detection, characterized in that: The system comprises a non-zero highest bit detection module, a maximum value search module and a mask generation module connected in sequence; The non-zero highest bit detection module is used to perform non-zero highest bit detection on all data in the quantized attention matrix column by column, approximate all data downward to the exponent of 2, and obtain a matrix with the same size as the input matrix, in which each element can reflect the numerical size of the original matrix; The maximum value search module is used to search the maximum value of each row of values ​​in the matrix column by column to obtain the maximum value of each row in the entire matrix; The mask generation module is used to generate a sparse mask column by column based on the maximum value information represented by the non-zero highest bit and a preset threshold, to obtain a mask matrix with the same size as the original matrix, and to multiply the mask matrix by the original attention matrix element by element to complete matrix sparseness.

2. The Transformer network attention matrix sparse processing system based on non-zero highest bit detection according to claim 1 is characterized in that: The non-zero highest bit detection module includes m LOD units, each of which is composed of combinational logic. From left to right, the data of m rows and 1 column in the matrix block are converted each time; the converted value contains only a single bit of 1 in binary, and the rest of the bits are 0, and the formula is expressed as: Where, X in Represents the data of m rows and 1 column in the attention matrix; X out Represents the output data after being processed by the LOD unit.

3. The Transformer network attention matrix sparse processing system based on non-zero highest bit detection according to claim 1 is characterized in that: The maximum value search module includes m search units, each of which is composed of a comparator, a multiplexer and a cache unit; the maximum value search module receives the input matrix data of m rows and 1 column column by column, and compares it with the current maximum value stored in the cache unit. If the input value is greater than the current maximum value, the maximum value in the cache unit is updated by controlling the multiplexer until all column data are input and the maximum value of each row is output; the formula of the maximum value search module is expressed as: Where, X ij Represents the matrix data of the i-th row and j-th column, Max i Represents the maximum value of the i-th row.

4. The Transformer network attention matrix sparse processing system based on non-zero highest bit detection according to claim 1 is characterized in that: The mask generation module includes m mask generation units and a sparse unit; each mask generation unit shifts the obtained maximum value rightward by T bits according to a preset threshold value T and sends it to a subtractor, subtracts the shifted value from a fixed value 9^'b10000 0000, and then sends it to the first combinational logic unit for bitwise AND with the current input value, and finally sends it to the second combinational logic unit for bitwise OR. If the result is 1, the mask is set to 1, otherwise it is set to 0, and the mask of the corresponding row in the matrix block is generated; the formula of the mask generation unit is expressed as: M i =Max i >>(T) N i =(9′b1 0000 0000)-M i O ij =N i &X ij Where M i Indicates data shifted right by T bits, N i Represents the subtractor output data, O ij represents the output data of the first combinational logic unit, X ij Represents the matrix data of row i and column j, Mask ij Represents the mask corresponding to the matrix data of the i-th row and j-th column; The sparse unit obtains a sparse matrix block by performing element multiplication operation on the mask generated by the mask generating unit with the same size as the original matrix block and the original matrix block.

Citation Information

Cited By

  • Large language model softmax function hardware acceleration circuit and method

    CN120952078A

  • Tensor core component of processor and processor

    CN121029685A