Matrix multiplication acceleration circuit architecture and operation method thereof

By designing a multi-level chained multiply-accumulate array and a delay register, the problems of low data caching and bandwidth utilization in matrix multiplication operations in existing technologies are solved, achieving efficient matrix multiplication operations and flexible PE configuration, thereby improving the computing performance of AI processors.

CN120848838APending Publication Date: 2025-10-28AIVATECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510965884.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing AI processors suffer from problems such as large data cache space, low PE input bandwidth utilization, increased pipeline computation latency, and inflexible array reconfigurability in large-scale matrix multiplication operations.

Method used

A multi-level chained multiply-accumulate array is adopted, which combines feature FIFO, weight FIFO and delay register. The weight vector data is loaded through delay processing and multiply-accumulate operation is performed in the multiply-accumulate. The independent accumulator memory and transfer memory are used for ping-pong operation to realize flexible matrix multiplication operation.

Benefits of technology

It improves the throughput of matrix multiplication operations, reduces input/output load, and enhances the flexibility and configurability of the PE array to adapt to matrix operation needs of different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848838A_ABST
    Figure CN120848838A_ABST
Patent Text Reader

Abstract

The invention provides a matrix multiplication acceleration circuit architecture and an operation method thereof, and the circuit architecture comprises a multi-stage chain type multiply-accumulator array which comprises i multiply-accumulators which are connected stage by stage; the i feature FIFOs are in one-to-one correspondence connection with the multiply-accumulators through feature element registers, the kth feature FIFO is used for caching the kth row of feature elements of the feature matrix, k is equal to 1, 2,..., i, and the feature element registers can load the corresponding next feature element only after keeping caching the current feature element in the configured i continuous periods; the weight FIFO is used for caching weight vector data corresponding to the weight matrix in sequence and is directly connected with the first multiply-accumulator; and each time delay register is used for delaying the weight vector data output by the weight FIFO by one period, so that the weight vector data is loaded to the hth multiply accumulator after being delayed by h-1 periods through the previous (h-1) time delay registers, and h is equal to 1, 2,..., (i-1). The multiply accumulator is high in utilization rate, flexible in configurability and high in reusability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a matrix multiplication acceleration circuit architecture and its operation method. Background Technology

[0002] In the field of artificial intelligence computing, matrix operation efficiency is a major factor affecting the performance of AI accelerators. For example, the operational principle of convolutional neural networks, commonly used in AI, is as follows: Figure 1 As shown, the Input Tensor represents the input feature tensor (C is the number of channels in the input feature tensor, and H / W are the dimensions of the input feature tensor), Filters represent the convolution kernels (K is the number of convolution kernels, and R / S is the size of the convolution kernel), and the Output Tensor represents the output feature tensor (K is the number of channels in the output feature tensor, corresponding to the number of convolution kernels, and R / Q is the size of the output feature tensor). In the process of accelerating computation in convolutional neural networks, tensor convolution operations can be converted into matrix operations to improve the computational efficiency of the accelerator by performing matrix multiplication.

[0003] According to the rules of matrix multiplication, the multiplication and accumulation operation between the feature matrix X and the weight matrix W can be expressed as a general expression: Where i represents the number of rows in the feature matrix X and the number of columns in the weight matrix W, and j represents the number of columns in the feature matrix X and the number of rows in the weight matrix W. The feature matrix X is a third-order feature matrix. The weight matrix W is a third-order weight matrix. For example, in calculation The calculation process is shown in Table 1 below:

[0004] Table 1

[0005]

[0006] Existing AI processors generally employ a systolic array architecture to implement large-scale matrix or tensor multiplication operations. Specifically, as shown in Table 2, in the classic systolic array third-order matrix multiplication calculation based on data broadcasting, cl1 to cl9 represent the 1st to 9th cycles. Each feature element of the row vector in the feature matrix must be maintained for i (in the example, i=3) operation cycles before the next feature element can be input into the corresponding multiply-accumulate (PE). After the weight matrix W is vectorized, the same weight element is loaded into each PE in each cycle. This scheme requires a large amount of data cache space and PE input bandwidth, and the data reuse rate is low. Furthermore, the weight vector data is broadcast to each PE unit for multiply-accumulate operations, which limits the overall IO bandwidth utilization of the array as the PE array size increases.

[0007] Table 2

[0008]

[0009] To address these shortcomings, existing technologies have proposed solutions such as on-chip caching, data reuse architecture, data stream rectification, near-memory computing, and in-memory computing to improve array computing efficiency. However, these solutions still face challenges such as increased pipeline computing latency, reduced PE utilization, and inflexible array reconfigurability or configurability as the PE array scales up. Summary of the Invention

[0010] In view of the shortcomings of the prior art, the purpose of this invention is to provide an improved matrix multiplication acceleration circuit architecture and its operation method.

[0011] In order to achieve the above object, the present invention adopts the following technical solutions:

[0012] In a first aspect, the present invention provides a matrix multiplication acceleration circuit architecture for implementing multiplication operations between an i×j order feature matrix and a j×i order weight matrix, the circuit architecture comprising:

[0013] A multi-level chained multiply-accumulate array, the multi-level chained multiply-accumulate array comprising i multiply-accumulates connected in stages;

[0014] The i feature FIFOs are configured in a stepwise manner. Each feature FIFO is connected to the multiply-accumulator through a feature element register. The k-th feature FIFO is used to cache the feature element of the k-th row of the feature matrix, k = 1, 2, ..., i. The feature element register can only load the next feature element after it has kept the current feature element cached for i consecutive cycles.

[0015] A weighted FIFO is used to sequentially cache the weight vector data corresponding to the weight matrix and is directly connected to the first multiply-accumulate unit.

[0016] There are i-1 delay registers, each of which is used to delay the weight vector data output by the weight FIFO by one cycle, so that the weight vector data is loaded into the h-th multiply-accumulate after being delayed for h-1 cycles by the first h-1 delay registers, where h = 1, 2, ..., i-1.

[0017] Furthermore, the multiply-accumulate includes:

[0018] A multiplier, wherein the first input of the multiplier is connected to the feature element register to load the corresponding feature element, and the second input of the multiplier is connected to the weight FIFO or delay register to load the corresponding weight vector data;

[0019] An adder, wherein the first input terminal of the adder is connected to the output terminal of the multiplier;

[0020] An accumulator memory includes an i-th layer of storage units, with its top-level storage unit connected to the second input terminal of the adder and its bottom-level storage unit connected to the output terminal of the adder. This allows the adder to add the multiplication result output by the multiplier to the data in the top-level storage unit, store the addition result in the bottom-level storage unit, and move the data in the (g+1)-th layer of storage units of the accumulator memory to the g-th layer of storage units, where g = 1, 2, ..., i-1.

[0021] Furthermore, the multiply-accumulate also includes:

[0022] The transfer memory includes an i-level storage unit, and its top-level storage unit is connected to an external memory.

[0023] Specifically, after the adder performs i×j addition operations, the accumulator memory and the transfer memory are interchanged, so that the accumulator memory becomes the new transfer memory and the transfer memory becomes the new accumulator memory.

[0024] Furthermore, the multiply-accumulate also includes:

[0025] An input selection unit is used to switch the output of the adder to the underlying storage unit of the new accumulator when the accumulator memory and the transfer memory are interchanged.

[0026] The output selection unit is used to switch the second input terminal of the adder to the top-level storage unit of the new accumulator when the accumulator memory and the transfer memory are interchanged, and at the same time switch the top-level storage unit of the new transfer memory to the external memory.

[0027] Furthermore, each of the feature element registers is connected to a data selector via the corresponding multiply-accumulate unit. The data selector is used to control the feature element register to keep the current feature element cached for i consecutive cycles.

[0028] Furthermore, the data selector is also used to control the i feature FIFOs to delay sequentially by one cycle and begin loading the corresponding feature elements into the corresponding multiply-accumulate.

[0029] Secondly, the present invention provides a computational method for the matrix multiplication acceleration circuit architecture as described above, the method comprising:

[0030] The i feature FIFOs are controlled to be delayed by one cycle in sequence, and the corresponding feature elements are loaded into the corresponding multiply-accumulate in a pulsating manner in column order. The feature element register enables the same feature element to be loaded into the corresponding multiply-accumulate within i consecutive cycles.

[0031] The weight FIFO is controlled to output the weight vector data in a pulsed manner. The weight vector data is loaded into the h-th multiply-accumulate unit after being delayed for h-1 cycles by the first h-1 delay registers, where h = 1, 2, ..., i. This allows each multiply-accumulate unit to perform multiply-accumulate operations on the feature elements and weight vector data loaded in its respective cycle.

[0032] Furthermore, before controlling the weight FIFO to output the weight vector data in a pulsed manner, the method further includes: converting the weight matrix into the corresponding weight vector data, and caching the weight vector data sequentially into the weight FIFO.

[0033] Furthermore, the multiply-accumulate unit includes a multiplier, an adder, and an accumulation memory, and the accumulation memory includes i layers of storage units, with each layer of storage units having an initial value of empty;

[0034] The multiply-accumulate unit performs multiply-accumulate operations on the feature elements and weight vector data loaded in the corresponding period through the following steps:

[0035] The feature elements and weight vector data loaded in the corresponding period are multiplied by a multiplier to obtain the multiplication result.

[0036] The multiplication result obtained in the corresponding cycle is added to the current data in the top-level storage unit of the accumulator memory by the adder, and the addition result is stored in the bottom-level storage unit of the accumulator memory. At the same time, the data in the (g+1)th level storage unit of the accumulator memory is moved to the g-th level storage unit, where g = 1, 2, ..., i-1.

[0037] Furthermore, the multiply-accumulate unit also includes a transfer memory, and the transfer memory includes i layers of storage units, each of which has an initial value of empty;

[0038] After the adder performs i×j addition operations, the accumulator memory and the transfer memory are swapped so that the accumulator memory becomes the new transfer memory and the transfer memory becomes the new accumulator memory. At the same time, the data in the top-level storage unit of the new transfer memory is transferred to the external memory.

[0039] By adopting the above technical solution, the present invention has the following beneficial effects: The present invention adopts a multi-level chained multiply-accumulate array structure, and by using a delay register to delay the loaded weights, it realizes the loading of weight data into the multi-level chained multiply-accumulate array through a single weight vector loading method, thereby reducing the input and output load of the multi-level chained multiply-accumulate array; at the same time, it can greatly improve the overall throughput when performing block operations on large-scale AI computation matrices; and due to the independent design of the multiply-accumulate and based on time-sharing accumulation operation, multiply-accumulates of different sizes can be flexibly configured according to the matrix size to adapt to the matrix operation requirements of different sizes, giving full play to the reusable characteristics of data and operation units. Attached Figure Description

[0040] Figure 1 This is a schematic diagram illustrating the operational principle of existing convolutional neural networks.

[0041] Figure 2 This is a schematic diagram of the matrix multiplication acceleration circuit architecture of the present invention;

[0042] Figure 3 This is a schematic diagram of the multiply-accumulator used in this invention;

[0043] Figure 4 This is an example diagram illustrating the data organization and computation cycle of third-order matrix multiplication in this invention;

[0044] Figure 5 This is a schematic diagram of the workflow of the third-order matrix multiplication array multiplication and accumulation in this invention. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0046] The terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0047] Example 1

[0048] This embodiment provides a matrix multiplication acceleration circuit architecture for implementing multiplication operations between an i×j order feature matrix X and a j×i order weight matrix W. For example... Figure 2 As shown, the circuit architecture includes: a multi-level chained multiply-accumulate array (PE), i feature FIFOs configured in stages, a weight FIFO, i-1 delay registers (D_REG), i feature element registers (F_REG), and i data selectors (MUX).

[0049] The following example uses i as 3, combined with... Figure 2 Each of the above components will be described in detail:

[0050] exist Figure 2 In the example shown, the multi-level chained multiply-accumulate (PE) array includes i (i=3) multiply-accumulate units connected in stages, namely PE1, PE2, and PE3. Each multiply-accumulate unit is used to receive the corresponding feature element and weight vector data from the corresponding feature element register F_REG and weight FIFO, respectively, to perform multiply-accumulate operations, and outputs the multiply-accumulate result to the storage unit MEM within a preset calculation cycle. Multiple multi-level chained multiply-accumulate (PE) arrays can form a PE matrix to perform multiple block matrix multiplication operations.

[0051] exist Figure 2 In the example shown, three feature FIFOs are configured sequentially: feature FIFO_1, feature FIFO_2, and feature FIFO_3. Each feature FIFO is connected to a corresponding multiply-accumulate (PE) via a feature element register F_REG. The k-th feature FIFO caches the feature elements of the k-th row of the feature matrix X in column order, where k = 1, 2, ..., i. Multiple feature FIFOs are used to preload feature elements into their respective multiply-accumulate (PE) in parallel. The feature element register F_REG maintains the cache of the current feature element for i consecutive cycles before loading the next feature element (i.e., the next column of the same row). By configuring the cycle number i of different feature element registers F_REG, feature elements can be loaded into their corresponding multiply-accumulate (PE) according to matrix calculation rules.

[0052] exist Figure 2 In the example shown, a weighted FIFO is used to cache the weight vector data corresponding to the weight matrix W in order. Before caching, the weight matrix W needs to be converted into a vector structure row by row in column order. For example, when... When the weight vector is converted row by row into the following vector form: w11, w12, w13, w21, w22, w23, w31, w32, w33, the converted weight vector data is then input into the weight FIFO in sequence.

[0053] exist Figure 2In the example shown, the multiply-accumulate units (PEs) at each stage are linked together via delay registers (D_REG), and the first multiply-accumulate unit (PE) is directly connected to the weight FIFO. Each delay register (D_REG) delays the weight vector data output from the weight FIFO by one cycle, ensuring that the weight vector data is loaded into the h-th multiply-accumulate unit (PE) after being delayed by h-1 cycles through the first h-1 delay registers (D_REG), where h = 1, 2, ..., i-1. This aligns the weight vector data with the corresponding feature element multiplication and accumulation operations. Delay registers are used because, in the initial loading state, each row of feature elements in the first column vector of the feature matrix needs to be loaded into the corresponding multiply-accumulate unit sequentially with a one-cycle delay. According to the matrix multiplication data organization method of this invention, the input weight vector data also needs to be delayed by one cycle sequentially. Therefore, setting delay registers (D_REG) to control the delay of the weight vector data avoids complex software control. It should be noted that the cycle in this invention refers to an instruction cycle, clock cycle, or operation cycle, etc. After the initial loading state, the feature elements in the feature FIFO are loaded into the corresponding multiply-accumulate units (PEs) in parallel according to the column order of the operation instructions. At the same time, the weight vector data is also loaded into the corresponding multiply-accumulate units (PEs) from bottom to top according to the synchronous operation instructions using a sequential delay method, and the multiply-accumulate operation is performed with the feature elements loaded in the corresponding PE.

[0054] exist Figure 2 In the example shown, each feature element register F_REG is connected to a data selector MUX between itself and the corresponding multiply-accumulate unit PE. Each data selector MUX is used to control the corresponding feature element register F_REG to maintain the current feature element cache for i consecutive cycles according to the corresponding control signal SEL. In addition, before entering the initial loading state, each data selector MUX can also be used to control the i feature FIFOs to delay sequentially by one cycle before starting to load the corresponding feature element into the corresponding multiply-accumulate unit PE.

[0055] In this embodiment, the structure of the multiply-accumulate PE is as follows: Figure 3 As shown, it includes: a multiplier, an adder, an accumulator memory, and a transfer memory. The accumulator memory and the transfer memory are ping-pong memories, used to cache intermediate accumulation results and intermediate accumulation output results of the adder, respectively, and to switch functions based on arithmetic instructions.

[0056] Specifically, the first input of the multiplier is connected to the corresponding feature element register F_REG to load the corresponding feature element, and the second input of the multiplier is connected to the weight FIFO or the corresponding delay register D_REG to load the corresponding weight vector data. The first input of the adder is connected to the output of the multiplier. The accumulator memory P_SUMA includes i-level storage units P_SUMA0 to P_SUMAi-1, and its top-level storage unit P_SUMA0 is connected to the second input of the adder, and its bottom-level storage unit P_SUMAi-1 is connected to the output of the adder. This allows the adder to add the multiplication result output by the multiplier to the data in the top-level storage unit P_SUMA0, store the addition result in the bottom-level storage unit P_SUMAi-1, and update the data in each storage unit. That is, the data in the (g+1)th level storage unit of the accumulator memory P_SUMA is moved to the g-th level storage unit, where g = 1, 2, ..., i-1. The transfer memory P_SUMB also includes i-level storage units, and its top-level storage unit P_SUMB0 is connected to the external memory MEM. After the adder performs i×j addition operations, the accumulator P_SUMA and the transfer memory P_SUMB are swapped (i.e., a ping-pong operation is performed) so that the accumulator becomes the new transfer memory and the transfer memory becomes the new accumulator.

[0057] Preferably, in this embodiment, the accumulator P_SUMA and the transfer memory P_SUMB can be single-port or dual-port memories, and the data refresh and function switching of the accumulator P_SUMA and the transfer memory P_SUMB can be realized by the operation instructions based on the configuration signal.

[0058] When the corresponding feature elements and weight vector data enter PE, they are first sent to the multiplier for multiplication. The multiplication result is then sent to the adder and added to the data in the top-level storage unit P_SUMA0 of the accumulator P_SUMA. The result of the addition is sent to the bottom-level storage unit P_SUMAi-1 of the accumulator. Each time the adder performs an operation, all data is refreshed, i.e., moved forward by one address (the data in the (g+1)th level storage unit of the accumulator P_SUMA is moved to the gth level storage unit, where g = 1, 2, ..., i-1), until it is in the top-level storage unit P_SUMA0. The above operation is repeated until the operation is completed after i×j cycles.

[0059] After the current matrix operation is completed, the result needs to be output. At this time, the memory performs a ping-pong operation, that is, it switches between the accumulation memory P_SUMA and the transfer memory P_SUMB, which was originally used to output the matrix operation result. At this time, the transfer memory P_SUMB is converted into the accumulation memory P_SUMA. This structure decouples the storage of intermediate results of the accumulation operation from the storage of the final result output, so that the PE can perform multiplication and accumulation operations more independently and uninterruptedly, without waiting for the final accumulation result output delay. At the same time, the PE unit can be flexibly configured to perform complex matrix multiplication operations according to configuration instructions.

[0060] In addition, such as Figure 3 As shown, the multiply-accumulate (PE) in this embodiment may further include an input selection unit and an output selection unit. The input selection unit is used to switch the output of the adder to the lower-level storage unit of the new accumulator memory when the accumulator memory and the transfer memory are interchanged. The output selection unit is used to switch the second input of the adder to the top-level storage unit of the new accumulator memory, and simultaneously switch the top-level storage unit of the new transfer memory to external memory when the accumulator memory and the transfer memory are interchanged. This achieves the switching of input and output between the accumulator memory and the transfer memory.

[0061] Example 2

[0062] This embodiment provides a computation method for the aforementioned matrix multiplication acceleration circuit architecture. The method includes: controlling i feature FIFOs to be sequentially delayed by one cycle (sequentially delayed by one cycle means that the first feature FIFO is not delayed, the second feature FIFO is delayed by one cycle relative to the first feature FIFO, the third feature FIFO is delayed by one cycle relative to the second feature FIFO, and so on), and starting to load the corresponding feature elements into the corresponding multiply-accumulate unit PE in a pulsating manner in column order. The feature element register F_REG loads the same feature element into the corresponding multiply-accumulate unit PE within i consecutive cycles. The corresponding multiply-accumulate PE can only load the next feature element after the i-th cycle has ended; at the same time, the weight FIFO is controlled to output the weight vector data in a pulsating manner (the weight matrix W is converted into the corresponding weight vector data in advance, and the weight vector data is cached into the weight FIFO in order). The weight vector data is loaded into the h-th multiply-accumulate PE after being delayed for h-1 cycles by the first h-1 delay registers D_REG, where h = 1, 2, ..., i, so that each multiply-accumulate PE performs multiply-accumulate operations on the feature element and weight vector data loaded in the corresponding cycle.

[0063] Specifically, in the initial loading state, the feature matrix X inputs the first column of feature elements into the multi-level chained PE array row by row with a one-cycle delay, and maintains pulsating loading of feature elements after the initial loading state is completed. Pulsating loading here means that the corresponding feature FIFO loads one feature element into the corresponding multiply-accumulate in column order every i cycles.

[0064] like Figure 4 As shown, taking third-order matrix multiplication as an example, in the initial loading state, the feature elements x11, x21, and x31 of the first column of each row are sequentially loaded into the corresponding multiply-accumulate units PE1, PE2, and PE3 after a one-cycle delay. During cycle cl3, the multi-level chained PE array is in a fully loaded state, indicating that the initial loading state is complete. Afterward, the feature elements of each row continue to be loaded in a pulsating manner until the corresponding row data is fully loaded. Specifically, in the initial loading state, the feature column vector can be loaded into the multi-level PE array row by row after a one-cycle delay via software or delay operation commands, or the delay can be controlled by configuring the data selector MUX data path. After the initial feature column vector data loading is completed, the configuration parameters of the data selector MUX data path are restored, ensuring that the currently input feature element maintains a cycle number of i.

[0065] In this embodiment, after the weight matrix W is converted into weight vector data column by column, the weight vector data input to the multi-level chained PE array is pulsed through the corresponding delay register D_REG for h-1 cycles and then input to the h-th multiply-accumulate PE, where h = 1, 2, ..., i. That is, each weight vector data is directly input to the first multiply-accumulate PE1, then after a one-cycle delay through delay register D_REG1, it is input to the second multiply-accumulate PE2, then after a one-cycle delay through delay register D_REG2, it is input to the third multiply-accumulate PE3 (delayed by two cycles relative to PE1), and so on. Then, in the corresponding multiply-accumulate, the weight vector data and the corresponding feature elements are multiplied and accumulated.

[0066] The reason for the delay processing in this embodiment is that each calculation of the weight vector data corresponds one-to-one with the feature elements. For example, the feature elements x12, x22, and x23 in the second column of the feature matrix X need to be calculated in consecutive periods with the weight vector data W21 in the first column of the second row of the weight matrix. Therefore, after delaying the feature elements x12, x22, and x23 in each row of the second column of the feature matrix by one period, the weight vector data W21 needs to be delayed accordingly through the delay register D_REG. Figure 4In the cl4 period, the feature element x12 of the first row and second column vector of the feature matrix is ​​multiplied with the weight element w21 in PE1. Since the feature elements of each row of the second column vector of the feature matrix are delayed by one period, in the cl5 period, the weight element W21 is multiplied with the feature element x22 of the second row and second column vector of the feature matrix in PE2 after a delay of one period of D_REG1. Similarly, in the cl6 period, the weight element W21 is multiplied with the feature element x32 of the third row and second column vector of the feature matrix in PE2 after a delay of two periods of D_REG1 and D_REG2.

[0067] Based on the above, this invention adopts a multi-level chained PE array structure. By using a delay register D_REG to delay the loaded weights, weight data can be loaded into the multi-level chained PE array through a single weight vector loading method, thereby reducing the IO load of the multi-level chained PE array.

[0068] Preferably, in this embodiment, the data path of the data selector MUX can be controlled by a preset configuration signal, so that the feature element currently loaded in the feature element register is cyclically maintained within i cycles, thereby allowing the same feature element to continuously perform multiplication operations with the corresponding weight element within i cycles. For example... Figure 4 As shown, each element in the eigenvector of each column of the third-order feature matrix needs to be multiplied with the corresponding weight element for i cycles before the next element is loaded into the corresponding PE. For example, x11, x21, and x31 constitute the column vector of the feature matrix. x11 is loaded into PE1 within i = 3 cycles to perform multiplication with w11, w12, and w13 respectively; x21 is loaded into PE2 within i = 3 cycles to perform multiplication with w11, w12, and w13 respectively; and x31 is loaded into PE3 within i = 3 cycles to perform multiplication with w11, w12, and w13 respectively. It should be noted that in the initial calculation state of the multi-level chained PE array, the column vector of the feature matrix loaded for the first time needs to be delayed by one cycle before being loaded into the corresponding PE. The specific delay operation can be configured by software or by configuring the input selection path MUX state of the feature element register F_REG during the first loading to achieve the delay effect.

[0069] In summary, during the initial loading state, the column vectors of the feature matrix are loaded into the corresponding PEs by delaying one cycle for each row. The weight vector data is loaded into the multi-level chained PE array in a serial delayed loading manner and multiplied with the corresponding feature elements to obtain the final accumulated calculation result.

[0070] In this embodiment, the multiply-accumulate is as follows: Figure 3As shown, the system includes a multiplier, an adder, an accumulator memory P_SUMA, and a transfer memory P_SUMB. The accumulator memory P_SUMA and the transfer memory P_SUMB each contain i layers of storage units, with each layer initially empty. When the multiply-accumulate unit performs a multiplication-accumulation operation, it first multiplies the feature elements and weight vector data loaded in the corresponding cycle using the multiplier to obtain the multiplication result. Then, the adder adds the multiplication result obtained in the corresponding cycle to the current data in the top-level storage unit P_SUMA0 of the accumulator memory, and stores the addition result in the bottom-level storage unit P_SUMAi-1 of the accumulator memory. Simultaneously, the data in the (g+1)th layer of storage unit of the accumulator memory is moved to the g-th layer, where g = 1, 2, ..., i-1. Repeat the above operation until the adder performs i×j addition operations (i.e., the feature elements of column j in the same row are multiplied with the corresponding weight vector data in i consecutive periods, and the corresponding multiplication results are accumulated). After one matrix multiplication operation is completed, the accumulator memory P_SUMA and the transfer memory P_SUMB are swapped so that the accumulator memory becomes the new transfer memory and the transfer memory becomes the new accumulator memory. At the same time, the data in the top-level storage unit of the new transfer memory is transferred to the external memory MEM.

[0071] The specific multiply-accumulate operation performed by each multiply-accumulate unit is as follows:

[0072] First, each feature element of the first column vector of the feature matrix is ​​multiplied by the corresponding weight vector data in the corresponding multiply-accumulator to obtain the first multiplication result. The first multiplication result is added to the data in the top-level storage unit P_SUMA0 of the corresponding accumulator to obtain the first accumulation result. The first accumulation result is sent to the bottom-level storage unit P_SUMAi-1 of the accumulator. After i cycles, the first accumulation result will be transferred to the top-level storage unit P_SUMA0 of the accumulator.

[0073] In the (i+1)th cycle, each feature element of the second column vector of the feature matrix is ​​multiplied by the corresponding weight vector data in the corresponding multiply-accumulator to obtain the second multiplication result. The second multiplication result is then added to the first accumulation result stored in the top-level storage unit P_SUMA0 of the corresponding accumulation memory to obtain the second accumulation result. The second accumulation result is then sent to the bottom-level storage unit P_SUMAi-1 of the accumulation memory, and the storage area data is refreshed, i.e., moved forward by one address.

[0074] Similarly, after multiple feature matrix column cycles and feature element retention cycles (i.e., after i×j cycles), the final multiplication and accumulation result of the current row feature element and the corresponding weight vector data is output.

[0075] Then, the accumulator P_SUMA and the transfer memory P_SUMB are switched. The original accumulator P_SUMA performs the final accumulation result output operation to save the result to the MEM memory; the original transfer memory P_SUMB is converted into an accumulator memory as an intermediate accumulation result memory to save the intermediate multiplication and accumulation results in subsequent operation cycles.

[0076] The following example uses the aforementioned third-order feature matrix X and weight matrix W, combined with... Figure 5 Further details on the operation of each multiply-accumulate unit:

[0077] In cycle cl1, x11 and w11 are multiplied in PE1 to obtain the multiplication result y11_1, and then added to the data in the initial state accumulation memory P_SUMA0. The addition result is sent to P_SUMA2 (there is no data in the initial state accumulation memory P_SUMA0, so the first accumulation result y11 is still y11_1), and data update is performed.

[0078] After i = 3 cycles, the data x12 in the next column of the row where x11 is located in cycle cl4 is loaded into PE and multiplied with the corresponding weight w21 in PE1 to obtain the multiplication result y11_2. Then, it is accumulated with the previous intermediate accumulation result y11_1 (i.e., y11 = y11_1 + y11_2). The accumulation result is sent to the bottom layer of the accumulation memory (P_SUMA2) and the data is updated.

[0079] In cycle cl7, the data x13 in the next column of the row where x12 is located is loaded into PE and multiplied with the corresponding weight w31 in PE1 to obtain the multiplication result y11_3. At this time, the intermediate accumulation result data (y11_1+y11_2) in the bottom layer of the accumulation memory P_SUMA2 in cycle cl4 has been sent to the top layer of the accumulation memory P_SUMA0 after 3 cycles, and is accumulated with the current cycle multiplication result y11_3 (y11=y11_1+y11_2+y11_3). The result y11 is sent to the top layer of the accumulation memory (P_SUMA2) and data update is performed.

[0080] In cycle cl9, each element x11, x12, and x13 performs a multiplication-accumulation operation with its corresponding weight data within three cycle hold periods. In PE1, the accumulation memory P_SUMA is converted to the transfer memory P_SUMB in cycle cl10, and the corresponding final accumulation result is output to the external MEM memory. Similarly, in cycle cl10, the eigenvector and weight vector in the second row of the feature matrix perform a multiplication-accumulation operation in PE2, and the final result is output. In cycle cl11, the eigenvector and weight vector in the third row of the feature matrix perform a multiplication-accumulation operation in PE3, and the final result is output.

[0081] After the above steps, the multiplication and accumulation results y11, y12, y13, ..., y33 shown in Table 1 above can be obtained.

[0082] This invention uses a 3rd-order matrix example with i=j=3 to facilitate the description and explanation of the entire matrix multiplication array operation process and the multi-level chained PE array design scheme. In actual implementation, the number of multiply-accumulators in the multi-level chained PE array can be set to 8 / 16 / 32 / 64, etc., and multiple multi-level chained PE arrays can also be combined to achieve parallel matrix operation, so as to be suitable for block parallel multiplication operation of ultra-large matrix multiplication. Moreover, this kind of ultra-large matrix block operation can be automatically processed by the compilation software without introducing additional work. That is, the multi-level chained matrix array of this invention can be used for matrix operations of different i×j orders.

[0083] This invention employs a multi-level chained multiply-accumulate array structure. By using a delay register to delay the loaded weights, it achieves the loading of weight data into the multi-level chained multiply-accumulate array via a single weight vector loading method, thereby reducing the input and output load of the multi-level chained multiply-accumulate array. Simultaneously, it significantly improves the overall throughput when performing block operations on large-scale AI computation matrices. Furthermore, due to the independent design of the multiply-accumulate units and their time-sharing accumulation operation, multiply-accumulate units of different sizes can be flexibly configured according to the matrix size to adapt to the needs of matrix operations of different scales, fully leveraging the reusability of data and computation units.

[0084] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these are merely illustrative examples, and the scope of protection of the present invention is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but all such changes and modifications fall within the scope of protection of the present invention.

Claims

1. A matrix multiplication acceleration circuit architecture for implementing multiplication operations between an i×j order feature matrix and a j×i order weight matrix, characterized in that, The circuit architecture includes: A multi-level chained multiply-accumulate array, the multi-level chained multiply-accumulate array comprising i multiply-accumulates connected in stages; The i feature FIFOs are configured in a stepwise manner. Each feature FIFO is connected to the multiply-accumulator through a feature element register. The k-th feature FIFO is used to cache the feature element of the k-th row of the feature matrix, k = 1, 2, ..., i. The feature element register can only load the next feature element after it has kept the current feature element cached for i consecutive cycles. A weighted FIFO is used to sequentially cache the weight vector data corresponding to the weight matrix and is directly connected to the first multiply-accumulate unit. There are i-1 delay registers, each of which is used to delay the weight vector data output by the weight FIFO by one cycle, so that the weight vector data is loaded into the h-th multiply-accumulate after being delayed for h-1 cycles by the first h-1 delay registers, where h = 1, 2, ..., i-1.

2. The matrix multiplication acceleration circuit architecture as described in claim 1, characterized in that, The multiply-accumulator includes: A multiplier, wherein the first input of the multiplier is connected to the feature element register to load the corresponding feature element, and the second input of the multiplier is connected to the weight FIFO or delay register to load the corresponding weight vector data; An adder, wherein the first input terminal of the adder is connected to the output terminal of the multiplier; An accumulator memory includes an i-th layer of storage units, with its top-level storage unit connected to the second input terminal of the adder and its bottom-level storage unit connected to the output terminal of the adder. This allows the adder to add the multiplication result output by the multiplier to the data in the top-level storage unit, store the addition result in the bottom-level storage unit, and move the data in the (g+1)-th layer of storage units of the accumulator memory to the g-th layer of storage units, where g = 1, 2, ..., i-1.

3. The matrix multiplication acceleration circuit architecture as described in claim 2, characterized in that, The multiply-accumulator also includes: The transfer memory includes an i-level storage unit, and its top-level storage unit is connected to an external memory. Specifically, after the adder performs i×j addition operations, the accumulator memory and the transfer memory are interchanged, so that the accumulator memory becomes the new transfer memory and the transfer memory becomes the new accumulator memory.

4. The matrix multiplication acceleration circuit architecture as described in claim 3, characterized in that, The multiply-accumulator also includes: An input selection unit is used to switch the output of the adder to the underlying storage unit of the new accumulator when the accumulator memory and the transfer memory are interchanged. The output selection unit is used to switch the second input terminal of the adder to the top-level storage unit of the new accumulator when the accumulator memory and the transfer memory are interchanged, and at the same time switch the top-level storage unit of the new transfer memory to the external memory.

5. The matrix multiplication acceleration circuit architecture as described in claim 1, characterized in that, Each of the feature element registers is connected to a corresponding multiply-accumulate unit via a data selector, which controls the feature element register to maintain a cache of the current feature element for i consecutive cycles.

6. The matrix multiplication acceleration circuit architecture as described in claim 5, characterized in that, The data selector is also used to control the i feature FIFOs to delay sequentially by one period and start loading the corresponding feature elements into the corresponding multiply-accumulate.

7. A computational method for a matrix multiplication acceleration circuit architecture as described in any one of claims 1-6, characterized in that, The method includes: The i feature FIFOs are controlled to be delayed by one cycle in sequence, and the corresponding feature elements are loaded into the corresponding multiply-accumulate in a pulsating manner in column order. The feature element register enables the same feature element to be loaded into the corresponding multiply-accumulate within i consecutive cycles. The weight FIFO is controlled to output the weight vector data in a pulsed manner. The weight vector data is loaded into the h-th multiply-accumulate unit after being delayed for h-1 cycles by the first h-1 delay registers, where h = 1, 2, ..., i. This allows each multiply-accumulate unit to perform multiply-accumulate operations on the feature elements and weight vector data loaded in its respective cycle.

8. The calculation method as described in claim 7, characterized in that, Before controlling the weight FIFO to output the weight vector data in a pulsed manner, the method further includes: converting the weight matrix into the corresponding weight vector data, and caching the weight vector data sequentially into the weight FIFO.

9. The calculation method as described in claim 7, characterized in that, The multiply-accumulate unit includes a multiplier, an adder, and an accumulation memory, and the accumulation memory includes i layers of storage units, each of which has an initial value of empty; The multiply-accumulate unit performs multiply-accumulate operations on the feature elements and weight vector data loaded in the corresponding period through the following steps: The feature elements and weight vector data loaded in the corresponding period are multiplied by a multiplier to obtain the multiplication result. The multiplication result obtained in the corresponding cycle is added to the current data in the top-level storage unit of the accumulator memory by the adder, and the addition result is stored in the bottom-level storage unit of the accumulator memory. At the same time, the data in the (g+1)th level storage unit of the accumulator memory is moved to the g-th level storage unit, where g = 1, 2, ..., i-1.

10. The calculation method as described in claim 9, characterized in that, The multiply-accumulate unit also includes a transfer memory, which comprises i layers of storage units, each with an initial value of empty. After the adder performs i×j addition operations, the accumulator memory and the transfer memory are swapped so that the accumulator memory becomes the new transfer memory and the transfer memory becomes the new accumulator memory. At the same time, the data in the top-level storage unit of the new transfer memory is transferred to the external memory.