Flexible convolution operation accelerator based on systolic array

By introducing flexible cache modules and pulsating array structures into the convolution operation accelerator, the problems of hardware scale adjustment and data format conversion are solved, and efficient convolution operation acceleration and memory access management are achieved.

CN120218148APending Publication Date: 2025-06-27HENAN XUNGU TECH CO LTD +1

Patent Information

Application Number
CN202510324937.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing convolutional computing accelerators have difficulty flexibly adjusting hardware scales according to different hardware platforms, and it is difficult to efficiently accelerate multiple types of convolutional computing. The data input format of the pulsating array requires additional modules to be converted, occupying additional hardware logic resources.

Method used

Design a flexible convolutional computing accelerator based on pulsating arrays, adopts a flexible cache module to realize data storage and format conversion functions, reduce hardware logic overhead, and support the configuration of accelerator scale and various convolution parameters.

Benefits of technology

It realizes efficient memory access efficiency, supports the acceleration of multiple convolution parameters, reduces the use of hardware resources, and improves the overall efficiency of convolutional operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218148A_ABST
    Figure CN120218148A_ABST
Patent Text Reader

Abstract

The invention provides a flexible convolution operation accelerator based on a systolic array. The flexible convolution operation accelerator is used for solving the technical problem that the hardware scale of an existing convolution operation accelerator is difficult to flexibly adjust according to different hardware platforms. The system comprises an input cache module, a pulsation matrix, an accumulation logic unit and an output cache module which are connected in sequence, the input cache module comprises a weight cache module and an image cache module, and the weight cache module and the image cache module are both connected with the pulsation matrix. The weight cache module, the image cache module, the pulsation matrix, the accumulation logic unit and the output cache module are all connected with the controller, and the controller and the output cache module are both connected with an external memory through a BUS. The systolic array and the flexible cache module are utilized, the functions of data storage and data format conversion can be achieved at the same time, hardware logic overhead is reduced, high memory access efficiency is achieved, and configuration of the accelerator scale and various convolution parameters is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hardware acceleration of algorithms, and particularly to a flexible convolution operation accelerator based on a systolic array. Background Art

[0002] Convolutional Neural Network (CNN) stands out in the field of image processing. With its unique structure and powerful learning ability, it has become a hot topic in computer vision research. In a convolutional neural network, the convolution operation is the most core operation. Its operation is essentially the accumulation after numerical multiplication, which is very suitable for using a dedicated hardware platform such as FPGA (Field-Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit) to perform hardware acceleration on the convolution operation, thereby effectively improving the overall operation efficiency of the convolutional neural network.

[0003] Among various hardware implementation schemes for convolution acceleration, the systolic array architecture is a widely adopted hardware acceleration architecture. This architecture consists of multiple processing elements (PEs), which are arranged according to certain rules to form an array. In this structure, the data flow between processing elements is like the flow of water in a water pipe, flowing regularly in a pre-set "flowing water" manner, thus achieving a high degree of parallel processing ability. The input data can be reused among the PE units in the systolic array structure, thereby reducing the number of data memory accesses, improving the utilization rate of the operation input data, and greatly reducing the pressure on the data bus bandwidth of the hardware computing unit.

[0004] When actually deploying a systolic array, the scale of the systolic array (the number of rows, columns, and dimensions of PEs) will affect the resources or area consumed by the hardware circuit. In the FPGA hardware implementation, different models of FPGA chips can provide different resources. When implementing in ASIC, the chip area will greatly affect the cost of chip manufacturing. Therefore, circuit design supporting reconfigurable hardware scale can make full use of FPGA hardware resources or balance the cost performance of chip products when implementing in ASIC. A convolutional neural network has a large number of convolutional layers, and in a convolutional neural network, the calculation parameters of each convolutional layer are usually inconsistent. These calculation parameters include but are not limited to the convolution kernel size, the number of convolution kernels, the convolution dimension, the size of the convolution input feature map, etc. A fully functional convolution accelerator must take these dynamic parameters as the input of the module and control the module to complete the convolution acceleration operation under the corresponding calculation parameters.

[0005] The invention patent with the application number 202411105645.1 discloses a configurable convolution operation acceleration device based on a systolic array, which includes the following modules: an instruction memory, a register bank, a scheduler, an arbiter, a weight cache, an input cache, a sliding window cache, and a systolic array; wherein, the instruction memory is used to store the instruction program applicable to the device; the register bank is used to store the working state of the device and instruction operands; the scheduler is used to read the instructions stored in the instruction memory and control the corresponding modules to work; the arbiter is used to select to send input data to the input cache or the weight cache under the control of the scheduler; the weight cache is used to store the weight data for convolution operation; the input cache is used to store the feature map data row by row and batch-send the sliding window data according to the size of the systolic array; the sliding window cache is used to receive the sliding window data output by the input cache and output it in the format applicable to the systolic array; the systolic array is used to broadcast in the row direction and pulsate in the column direction to complete the operation. The above invention proposes a systolic array with an improved operation mode, that is, an operation mode of broadcasting in the row direction and pulsating in the column direction, fixing the intermediate calculation results in the PE, without stopping the operation to update the weights, greatly improving the utilization rate of the PE, making the actual computing power close to the theoretical peak computing power, and because of adopting the broadcasting method in the row direction, the working states of each row of PEs are the same, so that the number of PEs that need to output the operation results in each cycle is the same, effectively ensuring the stability of the input-output data throughput rate and making full use of the data bandwidth of the output channels. However, this invention needs to design a dedicated "sliding window cache" to convert the data format of the data output by the input cache into the data format required by the systolic array, and the "sliding window cache" occupies additional hardware logic resources. Summary of the Invention

[0006] Aiming at the technical problems that existing convolution operation accelerators are difficult to flexibly adjust the hardware scale according to different hardware platforms, difficult to efficiently accelerate various types of convolution operations, and the data input format of the systolic array needs to be converted by an additional module, the present invention proposes a flexible convolution operation accelerator based on a systolic array, which is a new and efficient hardware accelerator supporting the acceleration of various neural network parameters. By using the systolic array and flexible cache modules, it can simultaneously realize the functions of data storage and data format conversion, reduce the hardware logic overhead, achieve high-efficiency memory access efficiency, support configuring the accelerator scale and various convolution parameters, and provide an efficient and practical acceleration method for the convolution calculation of various convolution neural networks with different configurations.

[0007] To achieve the above object, the technical solution of the present invention is implemented as follows: A flexible convolution operation accelerator based on a systolic array, including an input buffer module, a systolic matrix, an accumulation logic unit, and an output buffer module connected in sequence. The input buffer module includes a weight buffer module and an image buffer module. Both the weight buffer module and the image buffer module are connected to the systolic matrix. The weight buffer module, the image buffer module, the systolic matrix, the accumulation logic unit, and the output buffer module are all connected to a controller. The controller and the output buffer module are both connected to an external memory through a BUS bus.

[0008] Preferably, both the weight buffer module and the image buffer module are connected to a read arbiter. The read arbiter is connected to the external memory through a BUS bus; the read arbiter is connected to the controller.

[0009] Preferably, the controller includes an input logic controller, an operation logic controller, an accumulation logic controller, and an output logic controller. The input logic controller is respectively connected to the weight buffer module and the image buffer module. The operation logic controller is connected to the systolic matrix. The accumulation logic controller is connected to the accumulation logic unit. The output logic controller is connected to the output buffer module; the input logic controller is connected to the read arbiter. The output logic controller is connected to the external memory through a BUS bus; the input logic controller, the operation logic controller, the accumulation logic controller, and the output logic controller are generated by the cooperation of a state machine and a counter in digital logic design. The hardware implementation platform is an FPGA or an ASIC dedicated integrated circuit chip.

[0010] Preferably, the controller is connected to a configuration register. The user writes the configuration information required for convolution operation into the configuration register through the AXI-Lite interface configuration, controls the start of the convolution operation accelerator through the controller, and reads the running state of the convolution operation accelerator; the accumulation logic unit includes at least one convolution accumulator.

[0011] Preferably, the convolution operation acceleration method in a layer of convolutional neural network is as follows: After the configuration is completed and started, the input logic controller will generate read requests and data read instructions for the weight read channel and the image read channel. The read arbiter processes the read requests and authorizes the requests on the corresponding data channels. The authorized data channels start data transmission and store the read data into the input buffer module on the corresponding data channels; the authorized data channels have the access right to the data bus read channel, issue a read control signal, obtain the data at the corresponding address on the data bus, and the input logic controller issues a write enable signal to the weight buffer module or the image buffer module on the corresponding channel, and writes the data read from the data bus into the corresponding weight buffer module or image buffer module. After the weight cache module and the image cache module are ready with data, the input logic controller generates appropriate cache read control instructions according to the calculation requirements of the systolic array, retrieves weight data from the weight cache module, retrieves image data from the image cache module, and transfers them to the systolic array; under the control of the operation logic controller, the systolic array processes the input weight data and image data according to the specified convolution operation parameters and generates intermediate results of the convolution operation; Control the data flow direction of the systolic array, the data sources of each PE unit in the systolic array, and the output process of the calculation results of the systolic array; Subsequently, the intermediate results of the convolution operation are transferred to the convolution accumulator. Under the control of the accumulation logic controller, the convolution accumulator accumulates the intermediate results of the multi-dimensional convolution according to the dimension information of the convolution calculation, obtains the final result of the convolution calculation after accumulation, and transmits it to the output cache module; After detecting the final result of the convolution calculation obtained by the convolution accumulator, the output logic controller issues a cache control instruction to the output cache module, so that the final result of the convolution calculation is stored in the output cache module in the correct manner; After the output volume in the output cache module reaches the data volume required by the data bus, the output logic controller generates a cache control instruction to retrieve the corresponding data from the output cache module; At the same time, the output logic controller generates a bus write instruction to write the final result of the convolution calculation back to the external memory in the storage mode of an image, thus completing the acceleration process of the convolution operation in a convolutional neural network layer.

[0012] Preferably, a specified command is written to the REG_CONV_START register in the configuration register through the configuration bus. The configuration register generates a conv_start signal pulse representing the start instruction of the convolution calculation. After receiving the conv_start signal pulse, the controller starts the convolution acceleration calculation process; After completing the convolution operation of a convolutional neural network layer under the specified parameters, the output logic controller sends a conv_done pulse signal representing the completion of the convolution calculation to the configuration register. After receiving the conv_done pulse signal, the configuration register writes the value of the internal REG_CONV_DONE register as 1, and the convolution calculation of one layer is completely completed; Read the value of the REG_CONV_DONE register in the configuration register through the AXI-Lite configuration bus to obtain the running status of the convolution completion flag generated by the output logic controller.

[0013] Preferably, the systolic array is composed of basic computing units of m rows * n columns * k dimensions, and m, n, and k are all configurable static parameters; Each single-dimensional computing array in the systolic array is composed of basic computing units of m rows * n columns, and there is a shared path for data channels among the m * n basic computing units; There is no data sharing and no data reuse between the k-dimensional different systolic arrays in the longitudinal direction; The basic computing units in the last column and the last row of the systolic array directly obtain the input image data from the external memory or the input buffer module, while the image data of the basic computing units inside the systolic array is obtained from the adjacent basic computing units in the horizontal and vertical directions; in each clock cycle, each internal basic computing unit obtains the image data from one of the directions below in the vertical direction or to the right in the horizontal direction, and the control signals generated by the arithmetic logic controller of the controller control the path selection direction. In each clock cycle, a systolic array in each dimension inputs the value of a unique weight element, and the weight element is broadcast to all basic processing units in the one-dimensional systolic array. Different dimensions of the systolic array are used for parallel computing of different dimensions of the multi-dimensional convolution operation, and the input values of the weight elements are different; in each clock cycle, k different weight element data are input into the systolic arrays of different dimensions, where k is the number of dimensions of the systolic array. When the convolution operation does not wrap, the data in the systolic array pulsates in the horizontal direction; when the convolution operation wraps, the data in the systolic array pulsates in the vertical direction. The arithmetic logic controller generates the timing control commands for the product and accumulation operations required by the basic computing units. The timing control commands are related to the convolution operation parameters specified by the user and support variable convolution operation parameters. Each basic computing unit of the systolic array completes a multiplication and accumulation operation in each clock cycle under the drive of the clock signal; each basic computing unit has a horizontal image input port and a vertical image input port. In each clock cycle, when the convolution kernel moves to the right on the image, the data on the horizontal image input port is valid, and when the convolution kernel moves down on the image, the data on the vertical image input port is valid.

[0014] Preferably, the weight buffer module, the image buffer module, and the output buffer module are all flexible buffer modules.

[0015] Preferably, the convolution accumulator internally has an adder and a data buffer unit. After accumulating the intermediate result pixel units in multiple dimensions, they are cached in the data buffer unit. After all pixel units in the specified number of dimensions have been accumulated, the data is output to the downstream output buffer module; the final result of the convolution calculation is cached in the output buffer module. After the amount of data in the cache is sufficient, the output logic controller generates appropriate AXI bus commands and writes the final result of the convolution calculation back to the external memory in the storage organization form of an image; the output logic controller generates an address corresponding to each image data according to the position of the output image data segment in the image and writes the output image data segment back to the correct position in the external memory.

[0016] Preferably, the flexible cache module expands a one-dimensional storage unit into multiple parallel storage units. Each storage unit can store multiple independent data units. When inputting and writing, the written data is split into multiple independent data units and written into multiple parallel storage units. The storage units are indexed by "row" and "column". When performing read and write operations, the corresponding storage unit is accessed through "row index" and "column index". When reading, the data units are read out from multiple parallel storage units and spliced into a complete data unit. The way of reading data is the same as the way of writing data, so the read and write data directions are the same; or multiple consecutive data units are read out from one storage unit as the output data of the cache, and the read direction and the write direction are perpendicular. When writing data into the input cache module, the write bus data width is the "basic data unit width" multiplied by the "number of basic data units written". When reading data from the output cache module, the read bus data width is the "basic data unit width" multiplied by the "number of basic data units read". When reading out a data block, it supports reading out data in two modes: "consecutive in column direction" or "consecutive in row direction".

[0017] Compared with the prior art, the present invention has the following remarkable advantages and beneficial effects: The present invention provides an efficient convolution accelerator that supports circuit parameter reconstruction to adjust the hardware scale, supports accelerating various convolution parameter operations, and uses a systolic array as a hardware computing engine. It also includes a flexible and efficient cache module as the input and output cache in the accelerator, which is used to convert the data pattern of the bus interface and the data pattern of the accelerator architecture to improve the memory access efficiency of the convolution accelerator. The RTL code provided by the present invention adopts a fully parameterized design. By changing the global parameter configuration in the code, it supports flexible modification of the circuit scale to adapt to different hardware resource requirements, and supports specifying convolution calculation parameters at runtime to adapt to accelerating the convolution operations of various convolutional neural networks. The present invention can perform efficient image cache data management by proposing a reliable, efficient and flexible convolution accelerator and an efficient and novel flexible cache module. Combining the systolic array structure deployed by the present invention can effectively improve the memory access efficiency of the convolution accelerator. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1Schematic diagram of the internal structure of the convolution operation accelerator of the present invention.

[0020] Figure 2 Schematic diagram of the interaction between the accelerator of the present invention and the configuration register.

[0021] Figure 3 Schematic diagram of the basic structure of the systolic array of the present invention.

[0022] Figure 4 Schematic diagram of the operation logic of the basic operation unit of the systolic array of the present invention.

[0023] Among them, 1 is the weight cache module, 2 is the image cache module, 3 is the systolic array, 4 is the convolution accumulator, 5 is the output cache module, 6 is the controller, 7 is the input logic controller, 8 is the operation logic controller, 9 is the accumulation logic controller, 10 is the output logic controller, 11 is the read arbiter, 12 is the configuration register, 13 is the convolution accelerator, and 14 is the basic operation unit. Specific embodiments

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] Embodiment 1 As Figure 1As shown in the figure, a flexible convolution operation accelerator based on a systolic array, the architecture mainly includes: a configuration register 12 supporting AXI-Lite interface configuration, an input buffer module, a systolic array 3, a convolution accumulator 4, an output buffer module 5 and a controller 6. The input buffer module includes a weight buffer module 1 and an image buffer module 2, and both belong to the input logic unit. The weight buffer module 1 is used to temporarily store the convolution kernel weight data to be processed, obtain the weight data from the bus in advance and store it, and send it to the systolic array 3 when the acceleration engine of the systolic array 3 needs the weight data. The function of the image buffer module 2 is to temporarily store the input image data to be processed. Both the weight buffer module 1 and the image buffer module 2 are connected to the systolic array 3, the systolic array 3 is connected to the convolution accumulator 4, and the convolution accumulator 4 is connected to the output buffer module 5. The systolic array 3 is responsible for convolution operations and constitutes the operation logic unit. The convolution accumulator 4 is responsible for accumulating the results of multi-dimensional convolutions and constitutes the accumulation logic unit. The output buffer module 5 constitutes the output logic unit. The operation of each logic module is uniformly coordinated by the controller 6. Both the weight buffer module 1 and the image buffer module 2 are also connected to the read arbiter 11. Both the weight buffer module 1 and the image buffer module 2 have data reading requirements. However, the bus can only process one read request at a time. Therefore, an arbiter is needed to determine which cache module can obtain the access right of the bus currently. Both the read arbiter 11 and the output buffer module 5 are connected to the external memory through the BUS bus, and the data reading and writing operations of the external memory are realized through the BUS bus. The weight buffer module 1, the image buffer module 2, the systolic array 3, the convolution accumulator 4, and the output buffer module 5 are all connected to the controller 6. The controller 6 is connected to the BUS bus. The output logic controller is responsible for generating bus write commands, such as write addresses and write channel handshake signals, etc.; the output logic controller is responsible for generating bus read commands, such as read addresses and read channel handshake signals, etc. The controller 6 is connected to the read arbiter 11. The read arbiter 11 receives the weight read request and the image read request, and authorizes one of the weight read request initiator or the image read request initiator according to the bus situation, representing that the initiator can use the bus read channel.

[0026] Inside the controller 6, there are an arithmetic logic controller 8, an accumulation logic controller 9, an input logic controller 7, and an output logic controller 10. The input logic controller 7 is connected to the input logic unit composed of the weight cache module 1 and the image cache module 2. The arithmetic logic controller 8 is connected to the arithmetic logic unit composed of the systolic array 3. The accumulation logic controller 9 is connected to the accumulation logic unit composed of multiple convolution accumulators 4. The output logic controller 10 is connected to the output logic unit composed of the output cache module 5. Corresponding control signals are generated by these controllers to control the operation of the functional modules in the corresponding logic units. In terms of data reading, the access right of the data bus is coordinated through the read arbiter 11. The input logic controller 7, the arithmetic logic controller 8, the accumulation logic controller 9, and the output logic controller 10 are mainly generated through the cooperation of a state machine and a counter and other logics in digital logic design. The hardware implementation platform is an FPGA or an ASIC (Application Specific Integrated Circuit) chip.

[0027] The convolution operation accelerator provided by the present invention can be used as a coprocessor of a main control module such as a CPU. The user can write configuration information required for convolution operations (including but not limited to the convolution kernel size, the number of convolution kernels, the convolution dimension, the size of the convolution input feature map, etc.) into the configuration register through the AXI-Lite interface configuration. The configuration register supports controlling the start of the operation of the convolution operation accelerator and reading the running state of the convolution operation accelerator.

[0028] As Figure 1 shown, after the convolution operation accelerator is configured and started, the input logic controller 7 will generate read requests and data read instructions (including but not limited to the starting address, burst length, data volume, etc.) for the weight reading channel and the image reading channel. The read arbiter 11 processes the read requests and authorizes the requests on the corresponding data channels. The authorized data channels start data transmission and store the read data into the input cache module on the corresponding channels. The authorized data channels have the access right to the data bus read channel. A read control signal (such as a read address) is issued on this channel, and then the data at the corresponding address will be obtained on the data bus. The input logic controller 7 issues a write enable signal to the weight cache module 1 or the image cache module 2 on the corresponding channel to write the data read from the bus into the corresponding cache module. In terms of input cache management, the input logic controller 7 is responsible for generating appropriate cache control signals, and the cache control signals include but not limited to cache write enable, cache write row index, cache write column index, etc., so that the corresponding input data can be stored in the input cache module in the correct way.

[0029] After the corresponding data is ready in the weight cache module 1 and the image cache module 2, the input logic controller 7 generates appropriate cache read control instructions (including but not limited to cache read enable, cache read row index, cache read column index, etc.) according to the calculation requirements of the systolic array 3, retrieves the weights from the weight cache module 1 and the images from the image cache module 2, and transports both to the systolic array 3. Under the control of the operation logic controller 8, the systolic array 3 processes the input data and generates intermediate results of the convolution operation. The operation logic controller 8 controls the data flow direction of the systolic array, the data sources of each PE unit in the systolic array, and the output process of the calculation results of the systolic array according to the convolution operation parameters specified by the user (such as the convolution kernel size, etc.). Subsequently, the intermediate results of the convolution operation are transported to the convolution accumulator 4. Under the control of the accumulation logic controller 9, the convolution accumulator 4 accumulates the intermediate results of the multi-dimensional convolution according to the dimension information of the convolution calculation. The final result of the convolution calculation after the convolution accumulator 4 accumulates is transmitted to the output cache module 5. The accumulation logic controller 9 generates control signals for the convolution accumulator 4 according to the convolution operation parameters specified by the user (such as the convolution dimension), enabling the convolution accumulator to accumulate the calculation results of a specified number of one-dimensional convolution calculations to obtain the multi-dimensional convolution calculation result. The multi-dimensional convolution calculation is the accumulation of the one-dimensional convolution results at the same position in multiple dimensions. Two convolution accumulators can work alternately, so as to accumulate the intermediate results of the one-dimensional convolution sent from the systolic array 3 in a timely manner.

[0030] After detecting that the calculation result obtained by the convolution accumulator 4 is ready, the output logic controller 10 issues appropriate cache control instructions (including but not limited to cache write enable, cache write row index, cache write column index, etc.) to the output cache module 5, so that the corresponding calculation result data can be stored in the output cache module 5 in the correct manner. After detecting that the convolution accumulator 4 outputs a valid convolution result, the output logic controller 10 generates a cache write control signal (such as write index, write enable), and cooperates with the data provided by the convolution accumulator 4 to write the data to the output cache module 5. After the output volume in the output cache module 5 reaches the data volume required by the data bus, the output logic controller 10 generates appropriate cache control instructions (including but not limited to cache read enable, cache read row index, cache read column index, etc.) to retrieve the corresponding data from the output cache module 5. At the same time, the output logic controller 10 generates a bus write instruction (including but not limited to start address, burst length, data volume, etc.) to write the convolution calculation result back to the external memory in the storage mode of an image, thus completing the acceleration process of the convolution operation in a convolutional neural network layer.

[0031] The schematic top-level structure diagram of the present invention is as Figure 2As shown in the figure, the top layer of the accelerator consists of a configuration register 12 and a convolution accelerator 13. The configuration register 12 is directly connected to the controller 6 of the convolution accelerator 13 through signal wires. cfg_info, that is, configuration information, is the convolution operation description information configured by the user, including the convolution kernel size, the input feature map image size, etc. conv_state, that is, convolution acceleration state, is mainly the convolution calculation completion flag, used to notify the configuration register 12 that the current convolution acceleration operation has been completed, and the user can access the configuration register 12. The convolution accelerator 13 implements the main logical functions in the present invention, including calculation, control, and cache logic, and the relevant content has been described above. The role of the configuration register 12 in the present invention is mainly: 1) Provide runtime parameters for the convolution accelerator 13 (including but not limited to the convolution kernel size, convolution dimension, input feature map size, number of convolution kernels, etc.). The control logic inside the convolution accelerator 13 will generate appropriate control signals based on the runtime parameters specified by the user, and dynamically control the operation of the calculation unit and cache unit (including the weight cache module 1, the image cache module 2, the systolic array 3, and the output cache module 5), so that the present invention has the ability to perform convolution acceleration calculations on various convolutional neural networks. The runtime parameters specified by the user are used as the input signal of the controller 6, which is equivalent to a variable, and the internal logic of the controller generates control signals related to this variable based on this variable. The operation of all logics is controlled by the control signals sent by the controller 6.

[0032] 2) Generate the operation control signals required by the convolution accelerator 13. When the user writes a specified command to the REG_CONV_START register in the configuration register 12 through the configuration bus, the configuration register 12 will generate a conv_start signal pulse, representing the convolution calculation start instruction. After receiving the conv_start signal pulse, the controller 6 of the convolution accelerator 13 starts the convolution acceleration calculation process.

[0033] 3) Support reading and storing the operating status of the convolutional accelerator 13, including but not limited to the completion status of one-layer convolution calculation, the completion status of partial image output feature maps, etc. Taking the completion status of one-layer convolution calculation generated by the output logic controller 10 of the convolutional accelerator 13 as an example, after the convolutional accelerator 13 completes the convolution calculation of one-layer convolutional neural network under specified parameters, it sends a conv_done pulse signal to the configuration register 12, indicating that the convolution calculation is completed. After receiving the conv_done pulse signal, the configuration register 12 writes the value of the internal REG_CONV_DONE register as 1, indicating that the one-layer convolution calculation is completely completed. The user (the host with the AXI-Lite bus, such as the CPU) can read the value of the REG_CONV_DONE register in the configuration register 12 through the AXI-Lite configuration bus, and thus obtain the convolution completion flag generated by the output logic controller 10 of the convolutional accelerator 13. Other operating status query schemes are similar.

[0034] 4) Support querying the real-time operating parameters of the convolutional accelerator. The configuration register 12 stores the configuration information of the runtime convolution parameters specified by the user. These information are stored in each sub-register inside the configuration register 12. The values of these sub-registers support being read out through the AXI-Lite configuration bus, so that the user can obtain the configuration parameter information of the current convolutional accelerator 13 to confirm whether the specified user configuration parameters are configured correctly.

[0035] As Figure 3 shown in the basic structural schematic diagram of the arithmetic engine - systolic array of the present invention, the systolic array 3 is composed of basic computing units (Processing Element, PE) 14 with m rows * n columns * k dimensions. Here, m, n, and k are all configurable static parameters. By adjusting the values of m, n, and k, the circuit scale of the convolutional accelerator can be flexibly defined. It should be noted that m, n, and k are static parameters, and these values need to be determined before circuit synthesis. The synthesis tool (such as DesignCompiler of Synopsys or the FPGA development tool Vivado of Xilinx, etc., which is an EDA (Electronic Design Automation) tool that converts design code into a netlist) synthesizes the netlist under different circuit resource scales according to the values of m, n, and k specified in the code, and adjusts the circuit resource scale occupied by the accelerator according to the resources provided by the hardware to make full use of the circuit resources provided by various hardware platforms.

[0036] As Figure 3As shown, in the present invention, each one-dimensional computing array in the systolic array 3 is composed of basic computing units 14 of m rows * n columns. There is a shared path for data channels among these m * n basic computing units 14. Specifically, the basic computing units 14 at the rightmost (last column) and bottommost (last row) of the systolic array 3 directly obtain the input image data from the external memory or the input buffer module, while the image data of the basic computing units 14 inside the systolic array 3 is obtained from the adjacent processing units 14 in the horizontal and vertical directions. In each clock cycle, each internal basic computing unit 14 obtains the image data from one of the directions below it (vertical direction) or to its right (horizontal direction), and the specific path selection direction is controlled by the control signal generated by the operation logic controller 8 of the controller 6. There is no data sharing and no data reuse among the k-dimensional different systolic arrays 3 in the longitudinal direction. Since the internal basic computing units 14 do not need to obtain the image data from the external memory but reuse the data that has been taken out from the input buffer module before, the memory access bandwidth overhead of the present invention is greatly reduced.

[0037] In the present invention, in each dimension, the systolic array 3 inputs the value of a unique weight element in one clock cycle. This one weight value will be broadcast to all basic processing units 14 in the one-dimensional systolic array 3, and the different dimensions of the systolic array 3 are used for parallel computing of different dimensions of the multi-dimensional convolution operation, and the input weight values are not the same. For the multi-dimensional systolic array 3 of m rows * n columns * k dimensions, k different weight element data are input into the systolic arrays 3 of different dimensions in each clock cycle.

[0038] The data in the systolic array of the present invention will pulsate controllably in the row direction or the column direction, which is realized by the control signal generated by the operation logic controller 8 to control the data flow direction in the current systolic array. Specifically, when the convolution operation does not wrap lines, it pulsates in the horizontal direction; when the convolution operation wraps lines, it pulsates in the vertical direction. This vertical direction pulsation design enables no need to obtain a large amount of data from the external memory or cache when the convolution calculation wraps lines, further reducing the memory access bandwidth pressure.

[0039] The operation logic of the basic computing unit 14 is as Figure 4 shown. When performing the convolution operation, the basic control commands required by the basic computing unit 14 are provided by the operation logic controller. In each clock cycle, the basic computing unit 14 obtains the image data Image and the weight data Weight, multiplies the image data Image and the weight data Weight using the multiplier inside it, and then accumulates the product result through the adder, so as to obtain an updated intermediate result of the convolution calculation, and stores the accumulated result into the accumulation register inside the basic computing unit 14.

[0040] In Figure 3 Figure 3 , since the basic processing unit 14 performs a multiply-accumulate operation every clock cycle, and for each sub-array of the systolic array 3, only one convolutional kernel weight element is broadcast every clock cycle. Therefore, when performing one round of convolution operation, the number of cycles of its accumulation operation is directly related to the size of the convolutional kernel. For example, when the size of the convolutional kernel is 3 rows * 3 columns, a total of 3 * 3 = 9 accumulation operations need to be performed by the basic processing unit 14 of the systolic array. After 9 clock cycles, the accumulation calculation is completed, and the systolic array 3 outputs the calculation result of the one-dimensional convolution operation.

[0041] In Figure 4 Figure 4 , the timing control commands for the multiply-accumulate operations required by the basic processing unit 14 are provided by the operation logic controller 8. These control commands are related to the convolution operation parameters specified by the user, and supporting variable convolution operation parameters is also a feature of the present invention.

[0042] Each PE unit of the systolic array has independent computing capabilities. Driven by the clock signal, each PE unit of the systolic array completes a multiply-accumulate operation every clock cycle. For convolution operations, one weight data and one image data need to be multiplied every clock cycle, and the product is saved, waiting for the multiplication of new image data and weight data in the next clock cycle. This multiplication and accumulation process is repeated until all the weights on one convolutional kernel are used for multiplication operations, representing the end of one convolution operation process of the systolic array. In the present invention, each PE unit actually has two image data input ports (horizontal image input port and vertical image input port). At each clock cycle, only one of the two image ports is active (controlled by the control signal issued by the operation logic controller 8). Specifically, when the convolutional kernel moves to the right on the image, the data on the horizontal image input port is valid, and when the convolutional kernel moves down on the image, the data on the vertical image input port is valid.

[0043] After the user sends a start signal to the configuration register 12 through the configuration interface, the input logic controller 7 of the convolution operation accelerator uses internal digital circuits to generate corresponding control signals through logic such as state machines and counters, thereby automatically generating appropriate AXI bus commands (including memory access addresses, burst lengths, etc.), and moving part of the input feature map and part of the convolution kernel data required for the convolution operation to the input buffer module. When the data in the buffer is ready, the systolic array 3 obtains the convolution kernel weights and the input feature map from the input buffer module and performs convolution acceleration operations according to the input convolution configuration parameters. The operation process of the systolic array 3 is "pipelined", that is, while the calculation results of the systolic array are output, the basic calculation unit 14 of the systolic array 3 can already receive new input data for convolution operations. The data output by the systolic array 3 will be sent to the convolution accumulator 4 for the accumulation of multi-dimensional intermediate results. After the multi-dimensional convolution accumulation by the multi-dimensional accumulator, the calculated value obtained is the final convolution calculation result. In order to write the calculated convolution result back to the SDRAM memory in the form of image storage, it is necessary to first cache the convolution calculation result in the output buffer module 5. After the amount of data in the buffer is sufficient, the output logic controller 10 will generate appropriate AXI bus commands (including memory access addresses, burst lengths, etc.) and write the final convolution calculation result back to the external memory (such as the SDRAM memory) in the form of image storage organization. The convolution accumulator 4 has an adder and a data buffer unit inside. After accumulating the intermediate result pixel units in multiple dimensions, they are cached in the data buffer unit. After all the pixel units in the specified number of dimensions have been accumulated, the data is output to the downstream output buffer module 5. The AXI bus command is the output port of the output logic controller, and the control signal is generated through the circuit logic inside the output logic controller 10. The output logic controller 10 will generate the address corresponding to each specific image data according to the position of the output image data segment in the image, and write the output image data segment back to the correct position in the external memory.

[0044] In the present invention, there is no complex convolution calculation instruction scheduling process. The user only needs to send a "start" signal pulse after the convolution calculation information is configured. The internal controller 6 will automatically generate control signals corresponding to the configuration information according to the user configuration information, and the user then waits for the convolution accelerator to send back the "complete" flag.

[0045] In the present invention, the storage format of the data in the external memory conforms to the arrangement mode of the image data. The input data source of the accelerator does not need to be additionally converted, and the data obtained from the image signal source can be directly processed. The calculation results of the accelerator will also be written back to the memory address in the form of the arrangement of the image data and can be directly used as the data source for screen display and directly output to the display. Therefore, the present invention is also very suitable for use as a hardware accelerator for image processing.

[0046] The organization form of the acceleration engine for convolution calculation of the present invention is a systolic array. The scale of the computing units (Processing Element, PE) of the systolic array is represented by m rows * n columns * k groups, where m, n, and k are all configurable parameters in the design. The scale of the computing units of the one-dimensional systolic array is m rows * n columns. Among the m * n computing units, adjacent computing units are connected to each other to form a data path. Driven by the system clock, data flows between adjacent units inside the systolic array. Only the computing units at the edge of the systolic array need to receive new data from the external memory. By increasing the reuse times of data in the systolic array, the pressure of accessing the external memory caused by the accelerator calculation is reduced. By designing the global parameters in the code, the number of groups of the systolic array is adjusted, and multiple groups of systolic arrays are instantiated in the convolution accelerator. The present invention supports the deployment of multiple groups of systolic arrays, where each group of systolic arrays works independently and does not share data. The advantage of deploying multiple groups of systolic arrays is that different dimensions of convolution calculation can be calculated in parallel.

[0047] The present invention supports configuring parameters at runtime. These parameters include but are not limited to the convolution kernel size, convolution kernel dimension, number of convolution kernels, input feature map size, input feature map dimension, etc. The configuration information is saved by writing these values into the configuration register module. The configuration register 12 provides configuration parameters for the convolution accelerator 13. The configuration parameters are equivalent to variables, so that the controller 6 will generate corresponding control signals for each unit of the input logic, operation logic, accumulation logic, and output logic according to these parameters, thereby completing various convolution operations. The present invention supports convolution operations with various configurable convolution kernel sizes, numbers of convolution kernels, and input feature map sizes, has higher flexibility, and can support the operations of various convolutional neural networks. The configuration bus supported by the configuration register adopts the standard AXI-Lite protocol bus and supports register read operations to detect the operating status information of the accelerator.

[0048] In the present invention, the storage format of data in the external memory conforms to the arrangement mode of image data. The input data source of the convolution accelerator does not need to be additionally converted and can directly process the data from the image cache. The image cache module can directly process the data obtained from the image signal source. The calculation result of the convolution accelerator will also be written back to the memory address in the arrangement mode of image data, which can be conveniently output to the display device. In the present invention, the convolution acceleration calculation of multiple convolution kernels is scheduled by the controller 6. The controller will transfer the appropriate convolution kernel data to the basic computing units of the systolic array according to the number of convolution kernels configured by the user. The accumulation logic controller 9 inside the controller 6 will generate control signals to accumulate the data sent from the systolic array according to the specified number of convolution dimensions, thereby realizing the multi-dimensional convolution calculation process.

[0049] After the convolutional accelerator starts, the input logic controller 7 generates an AXI read command (address) to read out partial weight data and partial image data from the external memory, and stores the partial weights and partial images into the weight cache respectively. After a certain amount of data is stored in both caches, the operation logic controller issues a control instruction to control the systolic array, and the systolic array starts to perform convolutional operations. After a convolutional operation of the systolic array is completed, the calculation result is output to the convolution accumulator. The convolution accumulator accumulates the multi-dimensional convolution results under the control of the accumulation logic controller. After being processed by the convolution accumulator, the final convolutional calculation result is obtained. With the cooperation of the output logic controller 10, the convolutional result obtained by the convolution accumulator is written into the output cache 5 module. After the amount of data in the output cache is sufficient for the AXI bus to transmit, the output logic controller 10 generates an AXI write command (write address, etc.), reads the data from the output cache module 5, and sends it to the AXI bus, and writes the data to the external memory through the AXI bus. The overall effect achieved by the accelerator is the acceleration of convolutional operations on the hardware platform.

[0050] Embodiment 2 A flexible convolutional operation accelerator based on a systolic array. The weight cache module 1, the image cache module 2, and the output cache module 5 all use a new type of efficient and flexible cache module proposed by the present invention. The flexible cache module realizes two functions: efficiently converting bus data into accelerator data and data caching, without the need for an additional module for data conversion, and can efficiently perform input / output memory access behaviors. The read / write modes of the read / write cache module can be inconsistent. For example, when writing, it can be written by row, and when reading, it can be selected to read by row or by column. When the systolic array 3 is running, the required input image may be a data segment of a certain row of the image or a data segment of a certain column of the image. Since the cache in the present invention has the ability to read by row / by column in its design, there is no need for an additional data conversion process. The main features of the flexible cache module include but are not limited to: 1. Access the storage space in a manner similar to an image, and the provided access instruction interface includes a row index and a column index, rather than the single address form in traditional storage modules. Manage the storage space in an image manner, that is, the storage unit is indexed by "row" and "column", rather than the single address index in general memory. When performing read / write operations, the corresponding storage unit can be accessed through the "row index" and "column index". By expanding the one-dimensional storage unit in the traditional cache module into multiple parallel storage units, each storage unit can store multiple independent data units. When inputting and writing, the written data is split into multiple independent data units and written into multiple parallel storage units, so that the storage space can be accessed through a certain independent data unit in a certain storage unit in two dimensions.

[0051] 2. Support flexible data write position indexing, where users can specify the starting row index and column index of the data storage space.

[0052] 3. By externally passing in a control flag, which is equivalent to a switch, used to determine whether the current read is row-based or column-based. Support flexible data read position indexing and read mode control. Users can specify read and write, and provide the starting row index, starting column index, and read direction of the data storage space during read and write. The read and write indexes are directly specified by the pixel coordinates of the image. By expanding the one-dimensional storage units in the traditional cache module into multiple parallel storage units, each storage unit can store multiple independent data units. When inputting and writing, the written data is split into multiple independent data units and written into multiple parallel storage units, so that the storage space can be accessed through a certain independent data unit in a certain storage unit in two dimensions. The cache space size, row index, column index, and number of write units are adjusted through global parameters in the design code. When reading, data units can be read out from multiple parallel storage units and spliced into a complete data unit. At this time, the data reading method is the same as the data writing method, so the read and write data directions are the same; alternatively, multiple consecutive data units can be read out from one storage unit as the output data of the cache. At this time, the read direction and the write direction are perpendicular.

[0053] 4. Adjust through global parameters in the design code to support flexible, user-defined read and write interface bit widths. Users can specify the data bit width size of a single read and write operation to change the number of read and write data units.

[0054] The weight cache module 1, the image cache module 2, and the output cache module 5 are actually all flexible cache modules. Their interfaces and functions are the same, but the control signals and data signals they input are different. They play different roles in the entire convolutional accelerator. The flexible cache module is more general-purpose. One module can play different roles in the convolutional accelerator, effectively reducing duplicate development.

[0055] Users can flexibly specify the basic data unit bit width of read and write operations, the number of basic data units during writing, and the number of basic data units during reading. When writing data into the input cache module, the write bus data bit width is the "basic data unit bit width" multiplied by the "number of write basic data units". When reading data from the output cache module, the read bus data bit width is the "basic data unit bit width" multiplied by the "number of read basic data units".

[0056] When reading data blocks, it supports reading data in two modes: "continuous in column direction" or "continuous in row direction", that is, it supports the user to specify the splicing direction of the read data. For example, when the user selects to specify the column index as 3, the row index as 4, "continuous in row direction" as the read mode of the high-speed cache, and the number of basic data units read is 6, the data read on the data read bus will be the spliced version of the data on the six units with row index 4 and column indexes 3, 4, 5, 6, 7, and 8 respectively.

[0057] Through a new type of high-efficiency and flexible cache module proposed by the present invention, combined with the corresponding control logic, it can effectively manage the input data and calculation result output data required during convolution calculation.

[0058] The new flexible cache proposed by the present invention will be used as the input and output caches in the convolution operation accelerator. The input cache serves as a bridge between the DDR SDRAM memory and the input end of the accelerator. The data read from the SDRAM memory will first be written into the cache and wait until the accelerator needs data input, and then the data stored in the input cache module will be further sent to the accelerator. Through the cache bridge, it avoids the situation that when the accelerator needs data, the DDR SDRAM data bus is busy and cannot provide data to the accelerator in time. The output cache serves as a bridge between the output end of the accelerator and the DDR SDRAM memory. The calculation results of the accelerator will first be written into the output cache and wait until the AXI bus write transaction channel can handle the write-back transaction, and then the data in the output cache will be written back to the SDRAM. Through the cache bridge, it avoids the situation that when the accelerator calculates the result and needs to write data to the DDR SDRAM, the DDR SDRAM data bus is busy and cannot write the data back to the SDRAM in time.

[0059] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A flexible convolution operation accelerator based on a systolic array, characterized in that: The invention comprises an input buffer module, a pulsation matrix (3), an accumulation logic unit and an output buffer module (5) which are connected in sequence. The input buffer module comprises a weight buffer module (1) and an image buffer module (2). The weight buffer module (1) and the image buffer module (2) are both connected to the pulsation matrix (3). The weight buffer module (1), the image buffer module (2), the pulsation matrix (3), the accumulation logic unit and the output buffer module (5) are all connected to a controller (6). The controller (6) and the output buffer module (5) are both connected to an external memory via a BUS bus.

2. The flexible convolution operation accelerator based on a systolic array according to claim 1, characterized in that: The weight cache module (1) and the image cache module (2) are both connected to a read arbitrator (11), and the read arbitrator (11) is connected to an external memory via a BUS bus; the read arbitrator (11) is connected to a controller (6).

3. The flexible convolution operation accelerator based on systolic array according to claim 2, characterized in that: The controller (6) comprises an input logic controller (7), an operation logic controller (8), an accumulation logic controller (9) and an output logic controller (10); the input logic controller (7) is connected to the weight cache module (1) and the image cache module (2) respectively; the operation logic controller (8) is connected to the systolic matrix (3); the accumulation logic controller (9) is connected to the accumulation logic unit; and the output logic controller (10) is connected to the output cache module (5); the input logic controller (7) is connected to the read arbitrator (11); and the output logic controller (10) is connected to an external memory via a BUS bus; the input logic controller (7), the operation logic controller (8), the accumulation logic controller (9) and the output logic controller (10) are generated by the logic of the state machine in the digital logic design in conjunction with the counter, and the hardware implementation platform is an FPGA or ASIC special integrated circuit chip.

4. The flexible convolution operation accelerator based on a systolic array according to any one of claims 1 to 3, characterized in that: The controller (6) is connected to the configuration register (12), and the user writes the configuration information required for the convolution operation into the configuration register (12) through the AXI-Lite interface configuration, controls the start-up of the convolution operation acceleration operator through the controller (6), and reads the operating status of the convolution operation accelerator; the accumulation logic unit includes at least one convolution accumulation adder (4).

5. The flexible convolution operation accelerator based on systolic array according to claim 4, characterized in that: The convolution operation acceleration method in a layer of convolutional neural network is: After the configuration is completed and started, the input logic controller (7) will generate read requests and data read instructions for the weight read channel and the image read channel, the read arbitrator (11) processes the read request and authorizes the request on the corresponding data channel, the authorized data channel starts data transmission, and stores the read data in the input cache module on the corresponding data channel; the authorized data channel has the access right to the data bus read channel, sends a read control signal, obtains the data at the corresponding address on the data bus, and the input logic controller (7) sends a write enable signal to the weight cache module (1) or the image cache module (2) on the corresponding channel, and writes the data read from the data bus into the corresponding weight cache module (1) or the image cache module (2); When the data is ready in the weight cache module (1) and the image cache module (2), the input logic controller (7) generates a suitable cache read control instruction according to the calculation requirements of the systolic array (3), takes out the weight data from the weight cache module (1), takes out the image data from the image cache module (2), and transmits them to the systolic array (3); under the control of the operation logic controller (8), the systolic array (3) processes the input weight data and image data according to the specified convolution operation parameters and generates an intermediate result of the convolution operation; The data flow direction of the systolic array, the data source of each PE unit in the systolic array, and the output process of the systolic array calculation result are controlled; subsequently, the intermediate result of the convolution operation is transmitted to the convolution accumulator (4); under the control of the accumulation logic controller (9), the convolution accumulator (4) accumulates the intermediate results of the multi-dimensional convolution according to the dimension information of the convolution operation, obtains the accumulated final result of the convolution operation and transmits it to the output buffer module (5); after detecting the final result of the convolution operation obtained by the convolution accumulator (4), the output logic controller (10) issues a cache control instruction to the output buffer module (5), so that the final result of the convolution operation is stored in the output buffer module (5) in a correct manner; after the output amount in the output buffer module (5) reaches the amount of data required by the data bus, the output logic controller (10) generates a cache control instruction to retrieve the corresponding data from the output buffer module (5); at the same time, the output logic controller (10) generates a bus write instruction to write the final result of the convolution operation back to the external memory in the form of an image storage, thereby completing the convolution operation acceleration process in a layer of the convolution neural network.

6. The flexible convolution operation accelerator based on systolic array according to claim 5, characterized in that: A specified command is written to the REG_CONV_START register in the configuration register (12) through the configuration bus, and the configuration register (12) generates a conv_start signal pulse representing a convolution calculation start instruction. After the controller (6) receives the conv_start signal pulse, it starts the convolution acceleration calculation process; after completing the convolution operation of a layer of convolutional neural network under the specified parameters, the output logic controller (10) sends a conv_done pulse signal representing the completion of the convolution calculation to the configuration register (12). After the configuration register (12) receives the conv_done pulse signal, the value of the internal REG_CONV_DONE register is written to 1, and the convolution calculation of one layer is completed; the value of the REG_CONV_DONE register in the configuration register (12) is read through the AXI-Lite configuration bus to obtain the operating status of the convolution completion flag generated by the output logic controller (10).

7. The flexible convolution operation accelerator based on systolic array according to claim 5 or 6, characterized in that: The systolic array (3) is composed of basic computing units (14) of m rows*n columns*k dimensions, where m, n, and k are all configurable static parameters; Each single-dimensional computing array in the systolic array (3) is composed of m rows and n columns of basic computing units (14), and there is a shared path of data channels between the m*n basic computing units (14); there is no data sharing between the k-dimensional different systolic arrays (3) in the longitudinal direction, and no data reuse; The basic computing unit (14) in the last column and the last row of the systolic array (3) directly obtains input image data from an external memory or an input buffer module, while the image data of the basic computing unit (14) inside the systolic array (3) is obtained from adjacent basic computing units (14) in the horizontal direction and the vertical direction; in each clock cycle, each internal basic computing unit (14) obtains image data from one of the directions below in the vertical direction or to the right in the horizontal direction, and the control signal generated by the operation logic controller (8) of the controller (6) controls the path selection direction; The systolic array (3) in each dimension inputs a unique weight element value in one clock cycle, and the weight element is broadcast to all basic processing units (14) in the one-dimensional systolic array (3); Different dimensions of the systolic array (3) are used to parallelize different dimensions of the multi-dimensional convolution operation, and the values ​​of the input weight elements are different; data of k different weight elements are input into the systolic array (15) of different dimensions in each clock cycle, where k is the number of dimensions of the systolic array; When the convolution operation is not wrapped, the data in the systolic array (15) pulsates in the horizontal direction; when the convolution operation is wrapped, the data in the systolic array (15) pulsates in the vertical direction; The operation logic controller (8) generates a timing control command for the multiplication and accumulation operation required by the basic calculation unit (14), wherein the timing control command is related to the convolution operation parameter specified by the user and supports variable convolution operation parameters; Each basic computing unit (14) of the systolic array is driven by a clock signal to complete a multiplication-accumulation operation in each clock cycle; each basic computing unit (14) has a horizontal image input port and a vertical image input port. In each clock cycle, when the convolution kernel moves rightward on the image, the data on the horizontal image input port is valid, and when the convolution kernel moves downward on the image, the data on the vertical image input port is valid.

8. The flexible convolution operation accelerator based on a systolic array according to any one of claims 1-3, 5, and 6, characterized in that: The weight cache module (1), the image cache module (2) and the output cache module (5) are all flexible cache modules.

9. The flexible convolution operation accelerator based on systolic array according to claim 8, characterized in that: The convolution adder (4) has an adder and a data cache unit inside, which accumulates intermediate result pixel units in multiple dimensions and caches them in the data cache unit. After all pixel units in a specified number of dimensions have been accumulated, the data is output to a downstream output cache module (5); the final result of the convolution calculation is cached in the output cache module (5), and when the amount of data in the cache is sufficient, the output logic controller (10) generates a suitable AXI bus command to write the final result of the convolution calculation back to the external memory in the form of image storage organization; The output logic controller (10) generates an address of a corresponding position for each image data according to the position of the output image data segment in the image, and writes the output image data segment back to a correct position in the external memory.

10. The flexible convolution operation accelerator based on systolic array according to claim 9, characterized in that: The flexible cache module expands a one-dimensional storage unit into multiple parallel storage units, each storage unit can store multiple independent data units, and when inputting and writing, the written data is split into multiple independent data units and written into multiple parallel storage units; the storage unit is indexed by "row" and "column", and when performing read and write operations, the corresponding storage unit is accessed through the "row index" and "column index"; When reading, data units are read from multiple parallel storage units and spliced ​​into a complete data unit. The method of reading data is the same as the method of writing data, and the data directions of reading and writing are the same; or multiple consecutive data units are read from one storage unit as the output data of the cache, and the reading direction is perpendicular to the writing direction; When writing data to the input cache module, the write bus data bit width is "basic data unit bit width" multiplied by "write basic data unit number", and when reading data from the output cache module, the read bus data bit width is "basic data unit bit width" multiplied by "read basic data unit number"; When reading out data blocks, two modes of reading out data are supported: "continuous in the column direction" or "continuous in the row direction".

Citation Information

Patent Citations

  • Configurable convolution operation acceleration device and method based on systolic array

    CN118627565A

Cited By

  • SPAD imaging data compression system suitable for extremely low illumination

    CN120640145A

  • AI accelerator, processor, chip, board card and electronic equipment

    CN120950453A

  • Systolic array based on spatial domain

    CN121144258A

  • A spatial domain based systolic array

    CN121144258B