Matrix multiplication hardware architecture
The proposed matrix multiplication hardware architecture addresses the inefficiencies of existing FSB architectures by employing a multi-level tree topology and a cascaded DSP48 chain, resulting in reduced resource consumption and optimized timing.
Patent Information
- Application Number
- JP2024163930
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-09-20
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2044-09-20
AI Technical Summary
Existing Flexible Sparse Block (FSB) hardware computing architectures for matrix multiplication suffer from low computational efficiency and high resource consumption on Field-Programmable Gate Arrays (FPGAs), particularly due to inefficient addition tree structures and excessive resource utilization.
A matrix multiplication hardware architecture featuring a reduction network with a multi-level tree topology and a cascaded chain of digital signal processing units (DSP48), where each DSP48 has four input ports for sparse and dense matrix data and an output port connected to the next DSP48, optimizing resource usage and timing.
This architecture significantly reduces resource consumption and optimizes timing by changing the addition tree to an addition chain adapted to the DSP48 structure, multiplexing the post-adder, and converting sign-bit extension to zero filling.
Smart Images

Figure 2025097277000001_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital signal processing, and specifically relates to a matrix multiplication hardware architecture.
Background Art
[0002] In large language models based on the Transformer algorithm, matrix-matrix multiplication (General Matrix Multiplication, GEMM) is widely applied in technical fields such as the solution of complex physical systems, the calculation of current distribution in circuits and the analysis of engineering problems, the processing of multi-dimensional data, the analysis of social networks, movie recommendation systems, and traffic planning and management. This is the most important and time-consuming arithmetic operation. In order to reduce the computational load and improve the computational efficiency, it is necessary to adopt optimization calculation methods such as sparsification and design an efficient hardware architecture to speed up the matrix multiplication operation. Sparsification has become a widely used GEMM acceleration method, and related dedicated hardware architectures have been realized.
[0003] For a matrix with flexible sparsity at the block level, the existing Flexible Sparse Block (FSB) hardware computing architecture consists of multiple multipliers and a dynamically expandable reduction network. Here, the dynamically expandable reduction network is composed of an adder tree and configuration logic. As shown in Figure 1, Figure 1(a) shows the process of reading an array from the storage of the flexible sparse block hardware computing architecture, which includes steps such as defining a memory address, loading array data, reading array data, and processing array data. In the step of defining the memory address, a register for storing binary code is adopted. Figure 1(b) shows the decode input process of the flexible sparse block hardware computing architecture. The decode input means decoding the input signal or data so as to extract useful information or data therefrom, which includes receiving input data, analyzing the data format, decoding the data, and filtering or converting the decoded data by a selector. Figure 1(c) shows the reduction network configuration process of the flexible sparse block hardware computing architecture, which realizes data transmission and optimization based on a preset network topology. Figure 1(d) shows the data transfer process of the flexible sparse block hardware computing architecture, which also realizes hierarchical multiplexed time-sharing data transfer based on a preset network topology, and it is necessary to consider factors such as the use of hardware resources and energy consumption. Figure 1(e) shows the calculation and reduction process of the flexible sparse block hardware computing architecture, which uses a vector-matrix multiplication and a reduction method based on a tree structure. Figure 1(f) shows the cumulative addition process of the flexible sparse block hardware computing architecture.
[0004] According to the current sparsity of the blocks, each layer's reduction node is configured to perform an addition or forward transmission function, that is, the selector selects and outputs either the adder result or the bit-splicing result.
[0005] However, the hardware computing architectures in related technologies such as FSB have the following drawbacks.
[0006] The hardware optimization for Field-Programmable Gate Array (FPGA) is insufficient, which is manifested in the low computational efficiency of the addition tree structure and the large resource consumption of the FPGA. When the hardware parallelism is p, the computing unit requires p multipliers and (p / 2^1 + p / 2^2 + p / 2^3 +... + 1) adders. To speed up the inference process of large language models, it is necessary to repeatedly configure multiple computing units to achieve high computational efficiency, which consumes a large amount of the limited Look Up Table (LUT) resources of the FPGA.
Summary of the Invention
[0007] In response to the problems in the related art, the present invention provides a matrix multiplication hardware architecture, which can significantly save resources and optimize timing.
[0008] Embodiments of the present invention provide a matrix multiplication hardware architecture, including a reduction network including a multi-level tree topology formed by a plurality of reduction network nodes each including a data selector and two computing paths, A digital signal processing unit DSP48 chain cascaded by a plurality of digital signal processing units DSP48, wherein the output ends of adjacent digital signal processing units DSP48 are respectively connected to two calculation paths of the same reduction network node in the first-level tree topology, and the outputs of the two calculation paths are connected to the reduction network node in the upper-level tree topology through a data selector.
[0009] Furthermore, the digital signal processing unit DSP48 includes four input ports for receiving sparse matrix data and dense matrix data, and an output port connected to an adjacent cascaded digital signal processing unit DSP48.
[0010] Furthermore, each of the four input ports is input port B, input port A, and input port D for receiving sparse matrix data and dense matrix data, input port C for being connected to the output end of the upper-level digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain, and output port P for being connected to the input end of the next-level digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain.
[0011] Furthermore, inside the digital signal processing unit DSP48, a pre-adder, a post-adder, a plurality of sets of logic circuits provided on the input side and the output side of the pre-adder, and a plurality of sets of logic circuits provided on the input side and the output side of the post-adder are provided.
[0012] Furthermore, the plurality of sets of logic circuits are connected to two output ports of the digital signal processing unit DSP48 and are logic circuits for connecting to the pre-adder, and It is connected to other output ports of the digital signal processing unit DSP48, used to connect to a post-adder, and is a logic circuit for simultaneously connecting the post-adder to the output end of the pre-adder, and includes a logic circuit used at the output end of the digital signal processing unit DSP48.
[0013] Furthermore, the digital signal processing unit DSP48 is configured to be able to simultaneously calculate a plurality of 8-bit multiplications.
[0014] Furthermore, the parallelism of the digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain is an integer multiple of 4, and each parallelism corresponds to matrix multiplication calculations in various sparse format matrices.
[0015] Furthermore, the two calculation paths are respectively used for addition operations and splicing functions on the input data of the reduction network node.
[0016] Furthermore, the data selector is used to select and output one of the two calculation paths based on a preset selection signal.
[0017] Furthermore, the preset selection signal of the data selector is determined based on the sparsity of the input matrix, the parallelism of the digital signal processing unit DSP48, and the depth in the multi-level tree topology where the reduction network node is located.
[0018] Some of the other optional features and technical effects of the embodiments of the present invention are described below, and some will become apparent by reading this specification.
[0019] Compared with the prior art, the present invention has the following beneficial technical effects.
[0020] The present invention provides a matrix multiplication hardware architecture, including a reduction network including a multi-level tree topology formed by a plurality of reduction network nodes each including a data selector and two calculation paths, and a digital signal processing unit DSP48 chain cascaded by a plurality of digital signal processing units DSP48. The output ends of adjacent digital signal processing units DSP48 are respectively connected to two calculation paths of the same reduction network node in the first-level tree topology, and the outputs of the two calculation paths are connected to the reduction network node in the upper-level tree topology through a data selector. In the present application, the addition tree of the FSB is changed to an addition chain adapted to the DSP48 structure, thereby multiplexing the post adder of the DSP48 and improving the hardware utilization rate. At the same time, the hardware architecture of the present application can change the sign extension of the upper bits to zero filling, thereby greatly saving resources and optimizing timing.
Brief Description of the Drawings
[0021] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the elements shown in the drawings are not limited to the scale shown in the drawings, and the same or similar reference numerals in the drawings indicate the same or similar elements. Here,
Figure 1
Figure 2
Modes for Carrying Out the Invention
[0022] In order to make the object, technical solution, and advantages of the present invention clearer, hereinafter, the present invention will be described in more detail in combination with the embodiments and the drawings. Here, the exemplary embodiments of the present invention and their descriptions are for explaining the present invention, but not for limiting the present invention.
[0023] As used herein, the term "comprising" and variations thereof mean non-limiting inclusion, i.e., it means "including but not limited to". The term "or" means "and / or" unless otherwise specified. The term "based on" means "at least partially based on". The terms "one exemplary embodiment" and "one example" mean "at least one exemplary embodiment". The term "another example" means "at least one another example". The terms such as "first", "second", etc. may refer to different or the same things. Hereinafter, there may also be other explicit definitions and implicit definitions.
[0024] FIG. 2 shows a matrix multiplication hardware architecture in an embodiment of the present invention. As shown in FIG. 2, the matrix multiplication hardware architecture according to the present invention includes a reduction network including a multi-level tree topology formed by a plurality of reduction network nodes each including a data selector and two calculation paths, a digital signal processing unit DSP48 chain cascaded by a plurality of digital signal processing units DSP48, wherein the output ends of the adjacent digital signal processing units DSP48 are respectively connected to two calculation paths of the same reduction network node in the first-level tree topology, and the outputs of the two calculation paths are connected to a reduction network node in a higher-level tree topology through a data selector.
[0025] Note that the multi-level tree topology described in this embodiment is a local area network topology similar to the bus topology, which is composed of a tree structure and has the characteristics of a tree structure. In the tree topology, the tree network may include branches, and each branch may include a plurality of reduction network nodes. The tree topology is an extended form of the bus topology, and the transmission medium is a branched circuit that is not closed.
[0026] The tree topology has a root node and each branch node, is suitable for a hierarchical structure, and is very suitable for a hierarchical management system with priorities and grades. The characteristics of the tree topology are the same as those of the bus topology. One site can send data and other sites can receive it. Furthermore, the tree topology has strong foldability and can effectively protect wiring investment.
[0027] In the multi-level tree topology structure and the digital signal processing unit DSP48 cascade in this embodiment, as shown in FIG. 2, there are a plurality of multiplexers (MUXs) used for data signal processing and transmission between the upper and lower level digital signal processing units DSP48 and the reduction network nodes. Specifically, the multiplexer (MUX) is a multiplexer for integrating the input multiplex lines of multiple signals into one output line. The multiplexer (MUX) has a specific set of input terminals, and each input terminal may have one or more input signals or no signal. It includes a specific set of input terminals, a selection terminal for which it is necessary to pre-select the input signal in advance, and a set of output terminals. The multiplexer (MUX) has only one output port, and all signals need to be integrated into this one output port. The multiplexer (MUX) outputs only the selected input signal and ignores the rest. The main operating principle of the multiplexer (MUX) is that when there is a signal input at the input terminal, based on the input selection signal, the corresponding input signal is integrated into the output terminal, and the signals of other input terminals are all ignored. Therefore, the multiplexer (MUX) can effectively integrate multiple signals, thereby saving the resources of the output line and reducing the cost of the system.
[0028] In this embodiment, the digital signal processing unit DSP48 includes four input ports for receiving sparse matrix data and dense matrix data, and an output port connected to an adjacent cascaded digital signal processing unit DSP48. Specifically, the four input ports are respectively input port B, input port A, and input port D for receiving sparse matrix data and dense matrix data, An input port C for connecting to the output end of the upper-level digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain, and an output port P for connecting to the input end of the next-level digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain.
[0029] More specifically, in this embodiment, DSP48 is adopted, the width of its input port B is 18 bits, the width of input port A is 30 bits, and input port D is the data port of the pre-adder, with a width of 25 bits.
[0030] In this embodiment, inside the digital signal processing unit DSP48, there are a pre-adder, a post-adder, a plurality of sets of logic circuits provided on the input side and output side of the pre-adder, and a plurality of sets of logic circuits provided on the input side and output side of the post-adder. Specifically, the plurality of sets of logic circuits include logic circuits connected to two output ports of the digital signal processing unit DSP48 for connecting to the pre-adder, and logic circuits connected to another output port of the digital signal processing unit DSP48, used for connecting to the post-adder, and simultaneously connecting the post-adder to the output end of the pre-adder, and logic circuits used at the output end of the digital signal processing unit DSP48.
[0031] Specifically, a logic circuit generally includes components such as an input interface, an arithmetic unit, control logic, and an output interface. The input interface receives data signals input from the outside and converts them into a format suitable for internal operations. The arithmetic unit is the core part of the digital signal processing unit DSP48, including a multiplier, an adder, a shifter, etc., and is used to execute various digital signal processing algorithms. The control logic controls the workflow of the arithmetic unit to ensure the correct operation sequence and result output. The output interface outputs the processed data signals to an external device or memory. In addition, the logic circuit of the digital signal processing unit DSP48 needs to consider the interface and communication protocol with the external device to ensure correct communication and data transmission with the external device.
[0032] In this embodiment, the digital signal processing unit DSP48 is configured to be able to calculate multiple 8-bit multiplications simultaneously. Generally, multiple multipliers capable of performing multiple 8-bit multiplication operations simultaneously are integrated inside the digital signal processing unit DSP48. When performing a multiplication operation, the digital signal processing unit DSP48 multiplies the input data by the coefficients of the multipliers respectively and accumulates the results. Since the multipliers operate in parallel, multiple multiplication operations can be processed simultaneously, thereby realizing the simultaneous calculation of multiple 8-bit multiplications.
[0033] In this embodiment, the parallelism of the digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain is an integer multiple of 4, and each parallelism corresponds to matrix multiplication calculations in various sparse format matrices. Specifically, the digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain may be composed of 4, 8, or 16 parallel DSPs. That is, the parallelism of the digital signal processing unit DSP48 is 4, 8, or 16. The calculation units for each DSP parallelism may correspond to matrix multiplication operations in multiple sparse format matrices. For example, the calculation unit with a DSP parallelism of 16 can correspond to sparse matrix calculations of 1:16, 2:16, 4:16, and 8:16 in total, and at the same time, it can also complete the multiplication of dense matrices.
[0034] In this embodiment, the two calculation paths are respectively used for the addition operation and the joining function for the input data of the reduction network node. Specifically, the addition operation of the calculation path has the roles of an accumulation function and a filter function.
[0035] Regarding the accumulation function, in digital signal processing, accumulation is a common operation for calculating the sum or average value of signals. The digital signal processing unit DSP48 can perform a cumulative operation on the input data through an addition operation to obtain a desired result.
[0036] Regarding the filtering function, the addition operation plays an important role in digital filters. The digital signal processing unit DSP48 can realize signal filtering by adding the input data and the filter coefficients, and can remove noise or extract specific frequency components.
[0037] The splicing function of the calculation path has the roles of data merging and resolution improvement.
[0038] Regarding data merging, multiple input data can be combined into a single larger data block by the splicing function. This is very useful when processing segment signals or when it is necessary to combine multiple signals into one signal. Through the splicing operation, the digital signal processing unit DSP48 can process longer data sequences, thereby improving the processing efficiency.
[0039] Regarding the improvement of resolution, the resolution of the data can be increased by splicing multiple input data. This is particularly important in applications such as image processing and audio processing. By splicing multiple 8-bit data, data with a higher number of bits can be obtained, thereby improving the accuracy and quality of the processing.
[0040] Specifically, in this embodiment, the data width is expanded after passing through the reduction network node. Before the digital signal processing unit DSP48 performs matrix multiplication calculation, the configuration of the internal data stream of the digital signal processing unit DSP48 is performed according to the sparsity of the matrix by the selection signal. Therefore, in the present application, it is possible to flexibly respond to various sparsity calculations.
[0041] In this embodiment, the data selector is used to select and output one of two operation paths based on a preset selection signal. The data selector is used in the digital signal processing unit DSP48 to optionally receive and process multiple data as needed, realizing the multiplexed time-division transfer and logical control of the data, thereby improving the processing efficiency and functional flexibility of the digital signal processing unit DSP48. Specifically, it is used for data selection function, data time-division transmission, and logical control.
[0042] Regarding the data selection function, the data selector can select one specified from a set of input signals based on the given input address code and send it to the output terminal. As a result, the digital signal processing unit DSP48 can receive and process multiplexed data arbitrarily as needed.
[0043] Regarding data time-division transmission, during the process of multiplexed data transmission, the data selector can select any one of them as needed. As a result, the digital signal processing unit DSP48 can achieve multiplexed time-division transmission of data and improve data processing efficiency.
[0044] Regarding logic control, as part of the logic control, the data selector can realize specific logic functions by selecting different input signals, which is very useful in digital signal processing and can support the digital signal processing unit DSP48 to realize various complex logical operations and controls.
[0045] In this embodiment, the preset selection signal of the data selector is determined based on the sparsity of the input matrix, the parallelism of the digital signal processing unit DSP48, and the depth in the multi-level tree topology where the reduction network node exists. Note that the depth in the multi-level tree topology means the number of levels where the reduction network node is located in the multi-level tree topology.
[0046] As shown in Figure 2, taking the arithmetic unit composed of four digital signal processing units DSP48 as an example, the four digital signal processing units DSP48 include the digital signal processing unit DSP48 0, the digital signal processing unit DSP48 1, the digital signal processing unit DSP48 2, and the digital signal processing unit DSP48 3. The arithmetic operation for a set of 4*4 row-column input data with a sparsity of 2:4 is as follows.
[0047] The four 8-bit data of A, B, C, and D from the sparse matrix are sent to the input port B of the digital signal processing units DSP48 0, DSP48 1, DSP48 2, and DSP48 3 respectively. The four 8-bit data of a, b, c, and d from the dense matrix are sent to the input port A of the digital signal processing units DSP48 0, DSP48 1, DSP48 2, and DSP48 3 respectively. Similarly, the four 8-bit data of a’, b’, c’, and d’ from the dense matrix are sent to the input port D of the digital signal processing units DSP48 0, DSP48 1, DSP48 2, and DSP48 3 respectively.
[0048] One digital signal processing unit DSP48 (A + D)*B + C = (A*B + C1) + (D*B + C2), that is, it can calculate two multiplication operations, where the data of input port A and input port D are obtained from different big model input sequence length (seq_len) dimensions.
[0049] The output result from the output port P of the upper-level digital signal processing unit DSP48 is transmitted to the input port C of the lower-level digital signal processing unit DSP48. The output result of the output port P of the digital signal processing unit DSP48 is composed of two parts, including the output result and the 0 splice configuration, which are (A*B + C1) and (D*B + C2) respectively. For every two digital signal processing units DSP48, these two output signals are sent to the first-level splicer. The splicer zero-fills the input signal to 32 bits, and the output result of the first-level splicer is sent to the second-level splicer and spliced into a 64-bit output result.
[0050] The computational architecture of the digital signal processing unit DSP48 with configurable sparsity in this embodiment includes two parts: a DSP cascade chain that can be reconfigured during execution and a configurable reduction network. The adder tree of the FSB can be changed to an adder chain adapted to the structure of the DSP48, thereby multiplexing the post-adder of the DSP48, improving the hardware utilization rate. At the same time, the sign-bit extension of the upper bits can be changed to fill 0, thereby significantly saving resources and optimizing the timing.
[0051] In this specification, multiple embodiments of the present invention have been described. For the sake of simplicity, the description of each embodiment is not complete, and the same or similar features or parts between each embodiment may be omitted. In this specification, "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that they apply to at least one embodiment or example according to the present invention, but not all embodiments. The above terms do not necessarily mean the same embodiment or example. A person skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.
[0052] The exemplary systems and methods of the present invention have been specifically shown and described with reference to the above-described embodiments, but these are only examples of the best modes for implementing these systems and methods. A person skilled in the art can understand that various changes can be made to the embodiments of these systems and / or methods without departing from the spirit and scope of the present invention defined in the claims when implementing these systems and / or methods.
Claims
1. 1. A matrix multiplication hardware architecture, comprising: a reduction network including a multi-level tree topology formed by a plurality of reduction network nodes, each of which includes a data selector and two computation paths; a digital signal processing unit DSP48 chain cascaded by a plurality of digital signal processing units DSP48, in which the output ends of two adjacent digital signal processing units DSP48 are respectively connected to two calculation paths of the same reduction network node in the first level tree topology, and the outputs of the two calculation paths are connected to the reduction network node in the upper level tree topology via a data selector; 1. A matrix multiplication hardware architecture comprising:
2. The digital signal processing unit DSP 48 includes: four input ports for receiving sparse and dense matrix data; an output port connected to an adjacent cascaded digital signal processing unit DSP 48; 2. The matrix multiplication hardware architecture of claim 1, comprising:
3. Each of the four input ports is an input port B, an input port A and an input port D for receiving sparse matrix data and dense matrix data; an input port C for connection to the output of a higher level digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain; and is an output port P for connection to the input end of the next level digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain.
3. The matrix multiplication hardware architecture of claim 2.
4. The digital signal processing unit DSP48 includes a pre-adder, a post-adder, a plurality of sets of logic circuits provided on the input side and the output side of the pre-adder, and a plurality of sets of logic circuits provided on the input side and the output side of the post-adder. A matrix multiplication hardware architecture according to claims 1 to 3.
5. The plurality of sets of logic circuits include a logic circuit connected to two output ports of the digital signal processing unit DSP 48 for connecting to the pre-adder; a logic circuit connected to another output port of the digital signal processing unit DSP 48, used to connect the post-adder and simultaneously connect the post-adder to the output of the pre-adder; A logic circuit used at the output end of the digital signal processing unit DSP 48; 5. The matrix multiplication hardware architecture of claim 4, comprising:
6. The digital signal processing unit DSP 48 is configured to be able to calculate multiple 8-bit multiplications simultaneously. A matrix multiplication hardware architecture according to any one of claims 1 to 3 and 5.
7. The parallelism of the digital signal processing unit DSP48 in the digital signal processing unit DSP48 chain is an integer multiple of 4, and each parallelism corresponds to matrix multiplication calculations of various sparse formats. A matrix multiplication hardware architecture according to any one of claims 1 to 3 and 5.
8. The two computation paths are used for the addition operation and the splice function on the input data of the reduction network node, respectively. A matrix multiplication hardware architecture according to any one of claims 1 to 3 and 5.
9. The data selector is used to select and output one of two calculation paths based on a preset selection signal. A matrix multiplication hardware architecture according to any one of claims 1 to 3 and 5.
10. The preset selection signal of the data selector is determined based on the sparseness of the input matrix, the parallelism of the digital signal processing unit DSP 48, and the depth in the multi-level tree topology in which the reduction network node is located. A matrix multiplication hardware architecture according to any one of claims 1 to 3 and 5.
Citation Information
Patent Citations
Systems and methods for mapping executable models to programmable logic device resources
US10114917B1
An improved hardware primitive for implementations of deep neural networks
WO2020215124A1