A method for implementing data local caching in systolic array computing

By implementing local data cache and pre-alignment technology in AI chips, the problem of bandwidth limitation is solved, data transmission speed and processing capabilities are improved, and the computing power and bandwidth utilization of the chip are improved.

CN118012340BActive Publication Date: 2025-06-17SHANGHAI QINGWEI INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410087715.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-22
Publication Date
2025-06-17
Estimated Expiration
2044-01-22

AI Technical Summary

Technical Problem

When existing AI chips process and transmit data, bandwidth limitations have become an important factor restricting the improvement of computing power, resulting in limited data transmission speed and processing capabilities.

Method used

A method is proposed to realize the local cache of data in pulsating array calculation. The local cache module selects data according to the generated specific address, and pre-aligns the data and caches it, ensuring that the data is written at most 16 points per clock cycle.

Benefits of technology

Through data multiplexing and bandwidth optimization, bandwidth utilization during chip computing is improved, and design complexity and routing difficulty are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118012340B_ABST
    Figure CN118012340B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of systolic arrays, and specifically discloses a method for realizing data local caching in systolic array computing. This method uses the local cache module as an intermediate module between the control and computing arrays, enabling the systolic array computing module to read and write data blocks cached in the local cache module according to the calculation results of the transformed matrix outer product and matrix inner product, realizing data reuse in the chip computing process and effectively improving the bandwidth utilization rate in the chip computing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of systolic arrays, and particularly relates to a method for realizing data local caching in systolic array computing. Background Art

[0002] With the continuous popularization of AI applications, more and more industries and fields have achieved the implementation of AI. Currently, the demand for AI computing power shows a trend of rapid growth. Whether it is in the fields of autonomous driving, intelligent manufacturing, financial services, scientific computing, etc., powerful AI computing power is required to support the analysis and processing of these data.

[0003] In order to meet the increasingly high requirements for AI computing power, AI computing power not only faces challenges in terms of computing speed, computing accuracy, etc. Since the bandwidth directly affects the data transmission speed when an AI chip processes and transmits data, and further affects the data processing ability of the AI chip, data bandwidth has also become an important factor restricting the further improvement of the computing power of AI chips. Summary of the Invention

[0004] In order to solve at least one of the problems mentioned in the above background art, the present invention proposes a method for realizing data local caching in systolic array computing.

[0005] A method for realizing data local caching in systolic array computing includes the steps of:

[0006] Step S1, storing data in a shared storage array module according to a predefined storage specification.

[0007] Step S2, reading the data in the shared storage array module by an address generation module and generating a specific address according to the specific type of the data.

[0008] Step S3, selecting data by a local cache module according to the generated specific address, and caching the data after pre-aligning it.

[0009] Wherein, the local cache module includes a feature register block and a weight register block. Each register block is divided into ping-pong sub-blocks for supporting simultaneous reading and writing.

[0010] The feature register block is used to store the feature data in the computing process and is composed of 2x16x64 19-bit registers. Among them, 64 19-bit registers form 1 register row, 16 register rows form 1 ping-pong sub-block, and 2 ping-pong sub-blocks are respectively used for simultaneously reading and writing feature data.

[0011] The weight register block is used to store the weight data in the calculation process and consists of 2x64x64 19-bit registers. Among them, 64 19-bit registers form 1 register row, 64 register rows form 1 ping-pong sub-block, and 2 ping-pong sub-blocks are each used to read and write weight data simultaneously;

[0012] Both the feature register block and the weight register block support data storage of fp16, bf16, tf32, and int8 types, and both support 6 compression storage methods including 8-bit integer 4-channel, 8-bit integer 8-channel, 8-bit integer 16-channel, 8-bit integer 32-channel, 8-bit integer 64-channel, and 8-bit integer 128-channel;

[0013] Pre-align the data, and the process is as follows:

[0014] First, determine the data that needs to be cached, and obtain the data type and the number of channels of this data;

[0015] Then, calculate the number of cache points of this data according to the obtained data type and the number of channels;

[0016] Next, the feature register and the weight register in the local cache module execute data writing according to the data writing rule. The specific writing rule is:

[0017] According to the calculated number of data cache points, select the corresponding data cache points of this data to execute data writing respectively. Among them, the calculation method of the number of data cache points is:

[0018]

[0019] p_num ∈ [2 n , n ∈ [1, 4]

[0020]

[0021] In the formula, p_num is the number of data cache points, Bit_size is the size of the data type, and C_num is the number of channels.

[0022] If the calculated number of cache points is greater than or equal to 16, then when the number of cache points of the data in the feature register and the weight register is equal to 16, execute the data writing instruction;

[0023] If the calculated number of cache points is equal to 16, then when the number of cache points of the data in the feature register and the weight register is equal to the calculated number of cache points, execute the data writing instruction.

[0024] Regardless of whether it is training or inference, the number of valid data points at an address in each clock cycle is different. Directly writing to the register block increases the complexity of the design and the difficulty of backend placement and routing. To solve this problem, the present invention proposes that the local cache module selects data according to the generated specific address, pre-aligns the data, and then caches it to ensure that at most 16 data points are written to the activation register block and the weight register block in each clock cycle.

[0025] Step S4, the cached data is pushed by the local cache module to the systolic array computing module, and the systolic array computing module performs systolic array computing on the pushed data. Specifically, the systolic array computing includes three types of outer product matrix computing and three types of inner product matrix computing.

[0026] Among them, the three types of outer product matrix computing are respectively transposed convolution computing, left and right matrix transposition computing, and left matrix transposition computing. Specifically,

[0027] The process of transposed convolution computing is as follows: within 64 clock cycles, the ping-pong sub-block of the feature register block is switched every 16 clock cycles, for a total of 4 ping-pong sub-block switches; when a ping-pong sub-block switch is performed, the matrix outer product computing of the feature map W dimension and the feature map H dimension is executed, and the result is accumulated with the result of the previous matrix outer product computing in the feature map W dimension and the feature map H dimension.

[0028] The process of left and right matrix transposition computing is as follows: within 64 clock cycles, the ping-pong sub-block of the feature register block is switched every 16 clock cycles, for a total of 4 ping-pong sub-block switches; when a ping-pong sub-block switch is performed, the matrix outer product computing on the column dimension of the right matrix is executed, and the result is accumulated with the result of the previous matrix outer product computing in the column dimension of the right matrix.

[0029] The process of left matrix transposition computing is as follows: within 64 clock cycles, the ping-pong sub-block of the feature register block is switched every 16 clock cycles, for a total of 4 ping-pong sub-block switches; when a ping-pong sub-block switch is performed, the matrix outer product computing on the column dimension of the right matrix is executed, and the result is accumulated with the result of the previous matrix outer product computing in the row dimension of the right matrix.

[0030] The three types of inner product matrix computing are respectively convolution computing, conventional matrix computing, and right matrix transposition computing. Specifically,

[0031] The process of convolution computing is as follows: within 64 clock cycles, the ping-pong sub-block of the feature register block is switched every 16 clock cycles, for a total of 4 ping-pong sub-block switches; when a ping-pong sub-block switch is performed, the dot product computing on the C dimension of the feature map is executed, and the result is accumulated with the result of the previous dot product computing in the C dimension of the feature map.

[0032] Conventional matrix calculation: within 64 clock cycles, the ping-pong sub-block of the feature register block is switched every 16 clock cycles, for a total of 4 ping-pong sub-block switches; when a ping-pong sub-block switch is performed, the matrix inner product calculation in the column dimension of the left matrix is executed, and the result is accumulated in the column dimension of the left matrix with the result of the previous matrix inner product calculation.

[0033] Right matrix transpose calculation: within 64 clock cycles, the ping-pong sub-block of the feature register block is switched every 16 clock cycles, for a total of 4 ping-pong sub-block switches; when a ping-pong sub-block switch is performed, the matrix inner product calculation in the row direction of the right matrix is executed, and the result is accumulated in the row dimension of the right matrix with the result of the previous matrix inner product calculation.

[0034] When traditional AI chips perform calculations during the training or inference process, such as NPU, they mostly use the method of performing dot product operations on eigenvalues and weights. During the calculation process, the data stream's access to the memory is unidirectional. The present invention proposes to use the local cache module as an intermediate module between the control and calculation arrays, enabling the systolic array calculation module to read and write data blocks cached in the local cache module according to the calculation results of the transformed matrix outer product and matrix inner product, realizing data reuse in the chip calculation process and effectively improving the bandwidth utilization rate during the chip calculation process.

[0035] Step S5: The calculation result processing module post-processes the calculation results of the systolic array calculation module and returns the post-processed calculation results to the shared memory array module.

[0036] The present invention proposes a method for realizing data caching in the local cache during systolic array calculation. Compared with the existing technologies, it has the following beneficial effects:

[0037] The present invention proposes to use the local cache module as an intermediate module between the control and calculation arrays, enabling the systolic array calculation module to read and write data blocks cached in the local cache module according to the calculation results of the transformed matrix outer product and matrix inner product, realizing data reuse in the chip calculation process and effectively improving the bandwidth utilization rate during the chip calculation process.

[0038] The present invention proposes that the local cache module selects data according to the generated specific address, pre-aligns the data and then caches it, ensuring that at most 16 points of data are written into the activation register block and the weight register block per clock cycle, reducing the complexity of chip design and the difficulty of backend wiring. Brief Description of the Drawings

[0039] Figure 1 is the flow schematic diagram of the present invention;

[0040] Figure 2It is a schematic diagram of the process of caching pulsating array calculation data locally in an embodiment of the present invention. Detailed implementation manners

[0041] In order to make the objectives and features of the present invention more obvious and understandable, the following describes the technical solution in detail through embodiments in combination with the accompanying drawings.

[0042] In this embodiment, Figure 2 the SPM in Figure 2 is a shared storage array module, the udma is an address generation module, the Localbuffer is a local cache module, the PEA is a pulsating array calculation module, and the POST is a calculation result processing module. According to Figure 1 the system architecture shown in

[0043] and

[0044] the process shown in

[0045] a method for implementing data local caching in pulsating array calculation is as follows in the implementation process:

[0046] Step S1: Store the data in the shared storage array module according to the agreed storage specifications.

[0047] Step S2: The address generation module reads the data in the shared storage array module and generates a specific address according to the specific type of the data.

[0048] Step S3: The local cache module selects the data according to the generated specific address, pre-aligns the data and then caches it.

[0049] Among them, the local cache module includes a feature register block and a weight register block. Each register block is divided into ping-pong sub-blocks to support simultaneous reading and writing.

[0050] The feature register block is used to store the feature data in the calculation process and is composed of 2x16x64 19-bit registers. Among them, 64 19-bit registers form 1 register row, 16 register rows form 1 ping-pong sub-block, and 2 ping-pong sub-blocks are respectively used to simultaneously read and write feature data.

[0048] The weight register block is used to store the weight data in the calculation process and is composed of 2x64x64 19-bit registers. Among them, 64 19-bit registers form 1 register row, 64 register rows form 1 ping-pong sub-block, and 2 ping-pong sub-blocks are respectively used to simultaneously read and write weight data.

[0049] Both the feature register block and the weight register block support the storage of data of types fp16, bf16, tf32, and int8, and both support 6 compression storage methods including 8-bit integer 4 channels, 8-bit integer 8 channels, 8-bit integer 16 channels, 8-bit integer 32 channels, 8-bit integer 64 channels, and 8-bit integer 128 channels.

[0050] Pre-align the data, and the process is as follows:

[0051] First, determine the data to be cached, and obtain the data type and the number of channels of the data.

[0052] Then, calculate the number of cache points of the data according to the obtained data type and the number of channels.

[0053] Next, the feature register and weight register in the local cache module execute data writing according to the data writing rule. The specific writing rule is:

[0054] According to the calculated number of data cache points, select the corresponding data cache points of the data to execute data writing respectively. Among them, the calculation method of the number of data cache points is:

[0055]

[0056] p_num ∈ [2 n , n ∈ [1, 4]

[0057]

[0058] In the formula, p_num is the number of data cache points, Bit_size is the size of the data type, and C_num is the number of channels.

[0059] If the calculated number of cache points is greater than or equal to 16, when the number of cache points of the data in the feature register and weight register is equal to 16, execute the data writing instruction;

[0060] If the calculated number of cache points is equal to 16, when the number of cache points of the data in the feature register and weight register is equal to the calculated number of cache points, execute the data writing instruction.

[0061] Step S4, the cached data is pushed by the local cache module to the systolic array calculation module, and the systolic array calculation module performs systolic array calculation on the pushed data. Specifically, the systolic array calculation includes 3 kinds of outer product matrix calculations and 3 kinds of inner product matrix calculations.

[0062] Among them, the 3 kinds of outer product matrix calculations are respectively transposed convolution calculation, left - right matrix transpose calculation, and left matrix transpose calculation. Specifically,

[0063] The process of transposed convolution calculation is: within 64 clock cycles, perform the ping - pong sub - block switching of the feature register block every 16 clock cycles, for a total of 4 times of ping - pong sub - block switching; when performing a ping - pong sub - block switching, perform the matrix outer product calculation of the feature map W dimension and the feature map H dimension, and accumulate the result with the result of the previous matrix outer product calculation in the feature map W dimension and the feature map H dimension;

[0064] The calculation process of left - right matrix transposition is as follows: within 64 clock cycles, the ping - pong sub - block of the feature register block is switched every 16 clock cycles, for a total of 4 times of ping - pong sub - block switching; when a ping - pong sub - block switching occurs, the matrix outer product calculation in the column dimension of the right matrix is performed, and the result is accumulated in the column dimension of the right matrix with the result of the previous matrix outer product calculation.

[0065] The calculation process of left matrix transposition is as follows: within 64 clock cycles, the ping - pong sub - block of the feature register block is switched every 16 clock cycles, for a total of 4 times of ping - pong sub - block switching; when a ping - pong sub - block switching occurs, the matrix outer product calculation in the column dimension of the right matrix is performed, and the result is accumulated in the row dimension of the left matrix with the result of the previous matrix outer product calculation.

[0066] The three kinds of inner - product matrix calculations are convolution calculation, conventional matrix calculation, and right - matrix transposition calculation. Specifically,

[0067] The convolution calculation process is as follows: within 64 clock cycles, the ping - pong sub - block of the feature register block is switched every 16 clock cycles, for a total of 4 times of ping - pong sub - block switching; when a ping - pong sub - block switching occurs, the dot - product calculation in the C dimension of the feature map is performed, and the result is accumulated in the C dimension of the feature map with the result of the previous dot - product calculation.

[0068] Conventional matrix calculation: within 64 clock cycles, the ping - pong sub - block of the feature register block is switched every 16 clock cycles, for a total of 4 times of ping - pong sub - block switching; when a ping - pong sub - block switching occurs, the matrix inner - product calculation in the column dimension of the left matrix is performed, and the result is accumulated in the column dimension of the left matrix with the result of the previous matrix inner - product calculation.

[0069] Right - matrix transposition calculation: within 64 clock cycles, the ping - pong sub - block of the feature register block is switched every 16 clock cycles, for a total of 4 times of ping - pong sub - block switching; when a ping - pong sub - block switching occurs, the matrix inner - product calculation in the row direction of the right matrix is performed, and the result is accumulated in the row dimension of the right matrix with the result of the previous matrix inner - product calculation.

[0070] Step S5: The calculation result processing module post - processes the calculation results of the systolic array calculation module and returns the post - processed calculation results to the shared storage array module.

[0071] In this embodiment, optionally, the post - processing includes quantization processing and data - type transformation of the calculation results.

[0072] So far, according to the method disclosed in the present invention, one working process of the present invention has been completed.

[0073] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. And the term "including", "comprising" or any other variant thereof in this article is intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, article or method. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article or method including that element.

[0074] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0075] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for implementing local caching of data in systolic array computing, characterized in that: Includes steps: Step S1, storing the data in a shared storage array module according to the agreed storage specifications; Step S2, the address generation module reads the data in the shared storage array module and generates a specific address according to the specific type of the data; Step S3, the local cache module selects data according to the generated specific address, and pre-aligns the data before caching; The pre-aligning of data in step S3 includes the following steps: Step A1, determine the data to be cached, and obtain the data type and channel number of the data; Step A2, calculating the number of cache points of the data according to the acquired data type and number of channels; Step A3, the feature register and the weight register in the local cache module perform data writing according to the data writing rule; Among them, according to the calculated number of cache points, the corresponding data is selected and the data writing instruction is executed. The specific writing rules are: According to the calculated number of data cache points, select the data cache points corresponding to the data to perform data writing, wherein the number of data cache points is calculated as follows: p_num∈[2 n ],n∈[1,4] In the formula, p_num is the number of data cache points, Bit_size is the size of the data type, and C_num is the number of channels; If the calculated number of cache points is greater than or equal to 16, then when the number of cache points of the data in the feature register and the weight register is equal to 16, the data write instruction is executed; If the calculated number of cache points is equal to 16, then when the number of cache points of the data in the feature register and the weight register is equal to the calculated number of cache points, the data write instruction is executed; Step S4, the cached data is pushed by the local cache module to the systolic array computing module, and the systolic array computing module performs systolic array computing on the pushed data; Step S5, the calculation result processing module performs post-processing on the calculation result of the systolic array calculation module, and returns the post-processed calculation result to the shared storage array module.

2. A method for implementing local caching of data in systolic array computing according to claim 1, characterized in that: The local cache module described in step S3 includes a feature register block and a weight register block, wherein each register block is divided into ping-pong sub-blocks for supporting simultaneous reading and writing; The feature register block is used to store the feature data of the calculation process. It consists of 2x16x64 19-bit registers, where 64 19-bit registers form a register row, 16 register rows form a ping-pong block, and two ping-pong blocks are used to read and write feature data simultaneously. The weight register block is used to store the weight data of the calculation process. It consists of 2x64x64 19-bit registers, where 64 19-bit registers form a register row, 64 register rows form a ping-pong block, and two ping-pong blocks are each used to read and write weight data simultaneously.

3. A method for implementing local caching of data in systolic array computing according to claim 2, characterized in that: The feature register block and weight register block both support fp16, bf16, tf32, and int8 types of data storage, and both support 6 types of compressed storage methods: 8-bit integer 4 channels, 8-bit integer 8 channels, 8-bit integer 16 channels, 8-bit integer 32 channels, 8-bit integer 64 channels, and 8-bit integer 128 channels.

4. A method for implementing local caching of data in systolic array computing according to claim 1, characterized in that: The systolic array calculations performed on the pushed data in step S4 specifically include three outer product matrix calculations and three inner product matrix calculations.

5. A method for implementing local caching of data in systolic array computing according to claim 4, characterized in that: The three outer product matrix calculations are reverse convolution calculation, left and right matrix transposition calculation, and left matrix transposition calculation. The reverse convolution calculation process is: In 64 clock cycles, the ping-pong sub-block switching of the feature register block is performed once every 16 clock cycles, for a total of 4 ping-pong sub-block switchings; When a ping-pong block switch is performed, the matrix outer product calculation of the feature map W dimension and the feature map H dimension is performed, and the result of the previous matrix outer product calculation is accumulated on the feature map W dimension and the feature map H dimension; The calculation process of left and right matrix transposition is: In 64 clock cycles, the ping-pong sub-block switching of the feature register block is performed once every 16 clock cycles, for a total of 4 ping-pong sub-block switchings; When a ping-pong block switch is performed, the matrix outer product calculation on the right matrix column dimension is performed, and the result of the last matrix outer product calculation is accumulated on the right matrix column dimension; The calculation process of left matrix transpose is: In 64 clock cycles, the ping-pong sub-block switching of the feature register block is performed once every 16 clock cycles, for a total of 4 ping-pong sub-block switchings; When a ping-pong sub-block switch is performed, a matrix outer product calculation is performed on the column dimension of the right matrix, and the result of the previous matrix outer product calculation is accumulated on the row dimension of the right matrix.

6. A method for implementing local caching of data in systolic array computing according to claim 4, characterized in that: The three inner product matrix calculations are convolution calculation, regular matrix calculation, and right matrix transposition calculation, where: The convolution calculation process is: In 64 clock cycles, the ping-pong sub-block switching of the feature register block is performed once every 16 clock cycles, for a total of 4 ping-pong sub-block switchings; When a ping-pong block is switched, the dot product calculation on the C dimension of the feature map is performed, and the result of the last dot product calculation is accumulated on the C dimension of the feature map; General matrix calculations: In 64 clock cycles, the ping-pong sub-block switching of the feature register block is performed once every 16 clock cycles, for a total of 4 ping-pong sub-block switchings; When a ping-pong block switch is performed, the matrix inner product calculation on the left matrix column dimension is performed, and the result of the previous matrix inner product calculation is accumulated on the left matrix column dimension; Right matrix transpose calculation: In 64 clock cycles, the ping-pong sub-block switching of the feature register block is performed once every 16 clock cycles, for a total of 4 ping-pong sub-block switchings; When a ping-pong sub-block switch is performed, a matrix inner product calculation is performed in the row direction of the right matrix, and the result of the previous matrix inner product calculation is accumulated in the row dimension of the right matrix.

Citation Information

Patent Citations

  • A universal convolutional neural network accelerator based on a one-dimensional pulsation array

    CN109934339A