Mixed-base FFT hardware accelerator based on systolic array
By designing a hybrid-based FFT hardware accelerator based on pulsating arrays on the FPGA platform, the problem of inefficient calculations when the convolution kernel is large is solved, and rapid calculation and efficient resource utilization of FFTs of different points are realized.
Patent Information
- Application Number
- CN202510324938.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-06
AI Technical Summary
The calculation efficiency of existing convolutional neural networks is greatly reduced when the convolution kernel is large, making it difficult to efficiently calculate larger convolution kernels.
A hybrid-based FFT hardware accelerator based on pulsating arrays is designed, and it adopts fixed-point number operation and deploys on the FPGA hardware platform to accelerate FFT calculations through hybrid-based.
It realizes rapid calculation of FFTs of different points, supports 512 points, 1024 points, 2048 points, and 4096 points FFTs, improves computing efficiency, adapts to different performance needs, and reduces resource waste or insufficient resources.
Smart Images

Figure CN120104932A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hardware acceleration, and in particular to a mixed-radix FFT hardware accelerator based on a systolic array. Background Art
[0002] The main function of Fast Fourier Transform (FFT) is to realize the conversion of signals between time domain and frequency domain. It is an efficient and fast algorithm to reduce the number of discrete Fourier Transform (DFT) operations. FFT has many applications in various fields. For example, in the communication field, it can analyze the spectrum of the signal for modulation, demodulation and channel estimation. In the image field, it can be used for a variety of image algorithms, such as image enhancement, image filtering and image compression. In the radar field, the time-frequency domain characteristics of the echo signal can be analyzed to obtain information such as the speed and distance of the detected target.
[0003] In 2014, Google developed a tensor processing unit (TPU) with an ASIC (Application-Specific Integrated Circuit) architecture to accelerate the computational speed of convolutional neural networks (CNNs). Its core computing unit is a systolic array, which is used to accelerate convolution operations. The systolic array consists of multiple processing elements (PEs), which are arranged according to a certain rule to form an array consisting of rows, columns, and dimensions, and has strong parallel computing capabilities. The PE in the systolic array has a cache unit. After one calculation, data can be stored in the cache unit, and subsequent calculations can directly use this data, thereby reducing a large number of data access times, increasing data reuse, and greatly reducing the bandwidth pressure on the data bus. In the systolic array, data can flow regularly in the processing unit, which improves the computing speed. When deploying a systolic array, you can choose the appropriate size according to actual needs, such as changing the number of rows, columns, and dimensions of the systolic array. The larger the size of the systolic array, the greater the hardware overhead.
[0004] Convolutional neural networks have always been a hot research field, in which convolution operations account for more than 90% of the total amount of computation. When the convolution kernel is small, the use of convolution can meet the CNN calculation requirements. If a larger convolution kernel is encountered, the efficiency of convolution calculation will be greatly reduced. Using FFT to convert the convolution in the time domain into the product in the frequency domain can greatly reduce the amount of computation. Based on this, it is extremely important to design an efficient FFT algorithm architecture based on a systolic array.
[0005] FPGA (Field-Programmable Gate Array) not only has a large number of configurable logic devices, but also has the characteristics of high parallel computing capability, low latency, low power consumption and high throughput. Since FPGA uses digital chips, it is more adaptable to binary calculations. In binary multiplication, the circuit of fixed-point multiplication is simpler than that of floating-point multiplication. Although FPGA also supports floating-point operations, its implementation is relatively complex, and the latency and power consumption are higher. FFT uses the divide-and-conquer method to transform large-scale Fourier transform problems into multiple smaller transform units. Its core component is the butterfly unit, which is essentially a multiplication-addition calculation, and is very suitable for deployment on FPGA to achieve hardware acceleration of FFT.
[0006] The invention patent with application number 202410033848.8 discloses a lightweight neural network accelerator based on row input and its design method. The accelerator mainly includes: control module, input cache module, convolution pooling module, RELU module, output cache module, fully connected module, comparison module; the convolution pooling module is responsible for convolution operation on input data and weight data; the fully connected module is responsible for realizing the function of the fully connected layer; the comparison module is responsible for determining the recognition result of the convolution neural network by comparing the output size. The above invention aims to convert the convolution operation into the inner product of the vector, without the need to preprocess the input data, reduce the consumption of hardware resources, and at the same time, the accelerator combines the network level features for parallel optimization, realizes the parallel operation of multiple convolution kernels, improves the utilization of hardware resources, and reduces power consumption. However, this invention patent cannot efficiently calculate larger convolution kernels, because doing an inner product is equivalent to extracting the features of a point. If you want to extract the features of the entire sequence, you need to move the inner product on the signal, which only converts the convolution operation into the inner product of multiple vectors, which is essentially a convolution operation. Summary of the invention
[0007] In order to solve the technical problem that the computational efficiency of the convolution operation of the existing convolutional neural network is greatly reduced when the convolution kernel is large, the present invention proposes a mixed-radix FFT hardware accelerator based on a systolic array, which adopts fixed-point number operations and is deployed on an FPGA hardware platform to achieve better acceleration effect.
[0008] In order to achieve the above-mentioned purpose, the technical solution of the present invention is implemented as follows: a mixed-base FFT hardware accelerator based on a systolic array comprises an input logic unit, an operation logic unit, a data conversion unit and an output logic unit which are connected in sequence and arranged on an FPGA, wherein an operation unit is arranged in the operation logic unit, the operation unit is connected to the data conversion unit, a systolic array is arranged in the operation unit, and the systolic sequence uses a mixed basis to accelerate the calculation of FFT; the data conversion unit comprises a data truncation module and a data splicing unit which are connected in sequence, the operation unit is connected to the data truncation module, and the data splicing unit is connected to the output logic unit; the input logic unit, the operation unit, the data truncation module, the data splicing unit and the output logic unit are all connected to a controller, and the input logic unit, the data splicing unit, the data splicing unit and the output logic unit are all connected to a DDR memory through a data bus.
[0009] Preferably, the data splicing unit includes an FSM module and a RAM module, the data truncation module is connected to the FSM module, the FSM module is respectively connected to the RAM module and the output logic unit, and the RAM module is respectively connected to the input logic unit via a data bus.
[0010] Preferably, the input logic unit is connected to the first-layer hierarchical judgment module, and the first-layer hierarchical judgment module is respectively connected to the controller and the data bus; the output logic unit includes the last-layer hierarchical judgment module and the output cache module, the FSM module is connected to the last-layer hierarchical judgment module, the last-layer hierarchical judgment module is respectively connected to the output cache module and the input logic unit, and the output cache module is respectively connected to the controller and the data bus; the controller is connected to the configuration register module.
[0011] Preferably, the number of the operation units is two; the systolic array is an improved systolic sequence, each operation unit includes two 8*8 improved systolic arrays, and the improved systolic sequence supports the calculation of the radix-4 FFT butterfly unit and the radix-2 FFT butterfly unit.
[0012] Preferably, the controller includes an input logic control module, an operation logic control module, a storage logic control module and an output logic control module. The input logic unit includes a signal cache module and a weight cache module. The signal cache module, the weight cache module and the first-level hierarchical judgment module are all connected to the input logic control module, and the signal cache module and the weight cache module are all connected to the operation unit; the operation unit is connected to the operation logic control module, the data truncation module, the FSM module and the RAM module are all connected to the storage logic control module, and the output cache module is connected to the output logic control module.
[0013] Preferably, the signal and weight data generated in the DDR memory are sent to the first-layer hierarchical judgment module through the data bus. If it is judged to be the data of the first layer, the data is directly sent to the RAM module of the data conversion unit to cache the initial data; if it is other layers, the data is sent to the input logic unit through the data bus to start calculation; the systolic array is responsible for performing FFT butterfly unit calculation on the input weight data and signal data; the data truncation module is responsible for truncating the calculated data, and the FSM module uses the finite state machine method to process data with different points and different layers, including data position transposition, data inversion and data encoding; the RAM module is responsible for storing the processed data; the last-layer hierarchical judgment module of the output logic unit determines whether the data input by the data conversion unit is the last layer of data. If it is the last layer of data, the data will be placed in the output cache module, if not, the data will be sent to the input logic unit again; the output cache module transmits the calculation result to the DDR memory through the data bus; At the beginning of the calculation, the input logic control module automatically generates a control signal to cache the data of the weight matrix and the signal matrix. When the weight cache module and the signal cache module prepare valid data, they are sent to the operation unit for calculation. The operation logic control module generates a control signal to drive the operation unit to calculate the result; the storage logic control module generates a suitable control signal to store the data generated by the FSM module; when all the data of this layer are calculated, the last layer of the level judgment module performs a level judgment on the FFT of different numbers of points. If it is not the last layer, the storage logic control module generates a suitable read address, reads the data required for the next layer from the RAM module, and sends it to the operation unit for calculation again. If it is the last layer, the output logic control module generates a control signal to write the calculation result to the DDR memory through the AXI bus protocol; At the beginning of FPGA startup, the information of the weight matrix required for each point is written into the DDR memory through the PS end, the input logic control module generates a suitable address by itself, and stores the signal data from the DDR memory in the RAM module through the data bus for caching. When all signal information is cached, the input control logic module will generate a suitable address again, and store the weight data from the DDR memory in the weight cache module through the AXI bus for caching. When enough data is cached, the signal data and weight data are sent to the systolic array of the operation unit for calculation; when the systolic array generates valid data, the lower 14 bits of the valid data are truncated in the data truncation module, and the 16 bits after the 14th bit are retained, and all data are expanded by 2 14times, so that the next layer can perform FFT calculation on the data again; when the input calculation of this layer is completed, the storage logic control module generates a suitable address by itself, takes out the data required by the next layer from the RAM module, and sends it to the systolic array for calculation; when the calculation of all layers is completed, the final calculation result is written back to the DDR memory: the final result data generated is cached by the output cache module. When enough data is cached, the output logic control module generates a suitable data bus command and writes the FFT calculation result back to the DDR memory through the data bus.
[0014] Preferably, the method for improving the pulse sequence to perform radix-4 FFT butterfly unit is: There are multiple PE processing units in each systolic array. The weight matrix is input by the top row of PE processing units, and the signal matrix is input by the rightmost column of PE processing units. The butterfly unit of the radix-4FFT is composed of 4 complex input signals and four complex weight matrices. The butterfly unit of the radix-4FFT is decomposed into the form of multiplication of the signal matrix and the weight matrix, where the scale of the signal matrix is 1*8 and the scale of the weight matrix is 8*8. When the weight matrices are the same, the signal matrices are put together for calculation. The input data is divided into two directions, namely the vertical weight matrix and the horizontal signal matrix. Each column of PE processing units in the vertical direction shares the same weight data in the same clock. The PE processing units in each row in the horizontal direction do not share the same data in the same clock. Each clock in the horizontal direction sends the data to the next PE processing unit for calculation. The process of calculating the radix-4 FFT is: 1) 0~7 valid clocks: the weight matrix and signal matrix are input to the systolic array, and all the output data are invalid; 2) The 8th valid clock: All PE processing units in the first column from the right are valid; 3) 9th to 15th valid clocks: each clock makes the next column of PE processing units all valid; 4) 16~other valid clocks: The first column of PE processing units from the right starts to be valid again in sequence, entering a new cycle, and no invalid column data will be generated.
[0015] Preferably, the method for improving the pulse sequence to perform radix-2 FFT butterfly unit is: On the basis of not changing the scale and calculation process of the original systolic array, the 8*8 systolic array is divided into four 4*4 regions including the upper right region, the upper left region, the lower right region and the lower left region. Each region performs a different radix-2 FFT butterfly unit calculation, and changes the position of the PE processing unit where the systolic array outputs effective and the position of the weight matrix in the systolic array to correspond to the position of the calculation result; The radix-2FFT butterfly unit consists of two complex input signals and two complex rotation factors. The radix-2FFT butterfly unit is decomposed into the form of multiplication of the signal matrix and the weight matrix, where the scale of the signal matrix is 1*4 and the scale of the weight matrix is 4*4; the output order of the systolic array is changed to realize the calculation of the radix-2FFT butterfly unit. The calculation process is: a. 0~3 valid clocks: the weight matrix and signal matrix are input to the systolic array, and all output data are invalid; b. 4~7 valid clocks: PE processing units are valid in sequence starting from the first column on the right side of the upper right area of the systolic array. At this time, only four PE processing units are valid in each clock cycle; c. 8~11 valid clocks: PE processing units are valid in sequence starting from the first column on the right side of the upper right area and the lower left area. At this time, eight PE processing units are valid in each clock cycle; d. 12~15 valid clocks: The PE processing units start to become valid again in sequence starting from the first column on the right side of the lower left area. At this time, only four PE processing units are valid in each clock cycle. If data is input continuously, the valid PE processing units are the same as in step c above.
[0016] Preferably, the FSM module is provided with an initial coding unit and a base-4-to-base-2 coding, the initial coding unit is connected to the RAM module, and the base-4-to-base-2 coding is connected to the data truncation module; Different weight matrices are required for FFTs with different numbers of points, systolic arrays with different dimensions, and calculations with different numbers of layers. The weight matrix and signal matrix are stored in hexadecimal files in TXT format. The weight matrix and signal matrix are initially encoded. The encoded data is written from the PS end of the FPGA to the DDR memory at a predetermined address for storage. The non-last layer is calculated using a radix-4 FFT butterfly unit, and the radix-4 FFT of the last layer is converted into a two-layer radix-2 FFT through radix-4 to radix-2 encoding, and the radix-2 FFT butterfly unit is used for calculation. In the base-2FFT of the same layer, there is a certain symmetric relationship between the imaginary part and the real part of the rotation factor of two signals. The imaginary part and the real part of the rotation factor of one of the two signals with the symmetric relationship are matrix transformed, and the real part and the imaginary part of the signal are inversely transformed. The rotation factors of the two signals with the symmetric relationship are transformed into the same weight matrix, and the signal is multiplexed in the signal matrix.
[0017] Preferably, the initial encoding is a reverse encoding method; The calculated points, weights, base addresses of signals, and information for starting calculations are written into the configuration registers. The configuration registers use the standard AXI-Lite protocol bus and support register read and write operations. The number of rows, columns and dimensions of the PE processing units of the systolic array are modified before synthesis according to the actual operation speed and circuit scale.
[0018] Compared with the prior art, the present invention has the following significant advantages and beneficial effects: the present invention supports changing the scale of the circuit according to the requirements of hardware resources, operation speed and bandwidth before synthesis, which reflects the flexibility of design, can adapt to the requirements of different performances, and avoids waste or shortage of resources; supports specifying the number of FFT points at runtime, including 512-point, 1024-point, 2048-point and 4096-point FFT calculations, and FFTs with different numbers of points can meet the requirements of the accuracy, efficiency and boundary effects of the convolution results; the present invention also proposes a new mixed-base FFT architecture based on a systolic array to calculate FFTs with different numbers of points more quickly. The present invention also proposes a new systolic array architecture, which makes the output of the systolic array more regular without changing the original calculation time. The new systolic array architecture can support radix-4 and radix-2 FFT calculations, providing a calculation basis for the new mixed-base FFT architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 It is a schematic diagram of the structure of the internal module of the present invention.
[0021] Figure 2 A schematic diagram of the interaction between the accelerator and the configuration register of the present invention.
[0022] Figure 3 It is a schematic diagram of the structure of the pulsating array of the present invention.
[0023] Figure 4 It is a schematic diagram of the structure of the systolic array computing radix-2FFT of the present invention.
[0024] Figure 5 It is a schematic diagram of the mixed-basis FFT calculation process of the present invention.
[0025] Figure 6 This is a schematic diagram of encoding from one layer of base 4 to two layers of base 2 according to the present invention.
[0026] In the figure, 1 is a weight cache module, 2 is a signal cache module, 3 is an operation unit, 4 is a data truncation module, 5 is an FSM module, 6 is a RAM module, 7 is a last-layer level judgment module, 8 is an output cache module, and 9 is a controller, wherein: 10 is an input logic control module, 11 is an operation logic control module, 12 is a storage logic control module, 13 is an output logic control module, 14 is a first-layer level judgment module, 15 is a DDR memory, 16 is a configuration register module, 17 is an FFT accelerator, 18 is a weight matrix for vertical input, 19 is a signal matrix for horizontal input, 20 is a processing unit PE, 21 is a systolic array, 22, 23, 24 and 25 are radix-2 FFT butterfly units for calculating different regions of the systolic array, 26 is a weight data unit, 27 is an initial coding unit, 28 is a signal data unit, 29 is a non-last-layer FFT calculation unit, 30 is a radix-4 to radix-2 coding unit, and 31 is a last-layer FFT calculation unit. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] like Figure 1As shown, a mixed-base FFT hardware accelerator based on a systolic array supports 512-point, 1024-point, 2048-point, and 4096-point FFT algorithms. A mixed-base is used in the FFT algorithm to speed up the calculation process of FFT and is deployed in an FPGA. The accelerator architecture mainly includes: an input logic unit, an operation logic unit, a data conversion unit, an output logic unit, and a controller 9. The controller 9 includes an input logic control module 10, an operation logic control module 11, a storage logic control module 12, and an output logic control module 13. The input logic unit includes a signal cache module 2 and a weight cache module 1, and the signal cache module 2 and the weight cache module 1 are both connected to the input logic control module 10. The DDR memory 15 sends data to the first-layer level judgment module 14 through the data bus. If it is judged to be the data of the first layer, the data is directly sent to the RAM module 6 of the data conversion unit to cache the initial data. If it is other layers, the data is sent to the input logic unit to start calculation. The first-level hierarchical judgment module 14 is respectively connected to the signal cache module 2, the weight cache module 1, and the RAM module 6 of the data conversion module. The signal cache module 2 and the weight cache module 1 are both connected to the operation logic unit. The operation logic unit is provided with a systolic array calculation module, which is connected to the data conversion unit, and the data conversion unit is connected to the output logic unit. The weight cache module 1 is responsible for caching the weight data obtained from the data bus. The signal cache module 2 is responsible for caching the signal data obtained from the data bus. The systolic array calculation module is responsible for calculating the input weights and signals, and the data conversion unit is responsible for data processing of the calculation results to better calculate the next layer of FFT. Two operation units 3 are provided in the systolic array calculation module. The operation unit 3 is actually an improved systolic array responsible for the calculation of the FFT butterfly unit. Each operation unit 3 includes two 8*8 improved systolic arrays, so there are a total of four systolic arrays, namely 4 dimensions. Each systolic array has different inputs but the same architecture, and processes data in parallel, which is faster. The signal buffer module 2, the weight buffer module 1, the pulsation array calculation module, the data conversion unit and the output logic unit are all arranged on the FPGA and are all connected to the controller 9. The data conversion module includes a data truncation module 4 and a data splicing unit. The pulsation array calculation module is connected to the data truncation module 4, and the data truncation module 4 is connected to the data splicing unit. The data truncation module 4 is responsible for truncating the calculated data, and the data splicing unit is responsible for data processing of the truncated data. The data splicing unit includes an FSM module 5 and a RAM module 6. The FSM module 5 uses a finite state machine method to process data of different points and different layers. The processing includes position transposition, data inversion and data encoding of the data. The RAM module 6 is connected to the input logic unit after the last layer of hierarchical judgment, and the RAM module 6 is responsible for storing the processed data.The data conversion unit is connected to the output logic unit, and the output logic unit is connected to the input logic unit. The output logic unit is responsible for judging whether the data input by the data conversion unit is the last layer of data. If it is the last layer of data, the data will be placed in the output cache unit, and then written back to the DDR memory by the output cache unit. If not, the data will be sent to the input logic unit again. The output logic unit includes a last layer level judgment module 7 and an output cache module 8. The last layer level judgment module 7 is connected to the FSM module 5. The last layer level judgment module 7 is respectively connected to the output cache module 8 and the signal cache module 2 and the weight cache module 1 of the input logic unit. The last layer level judgment module 7 is responsible for judging the FFT level again, and the output cache module 8 is responsible for caching and outputting the data. The output cache module 8 is connected to the DDR memory 15 via the data bus. The DDR memory 15 is connected to the data bus, and the generated signal and weight data are stored in the DDR memory 15.
[0029] The operation of each module is coordinated by the controller 9, which includes an input logic control module 10, an operation logic control module 11, a storage logic control module 12, and an output logic control module 13. These controller modules can generate control signals to control the operation of the corresponding logic modules. The input logic control module 10 is connected to the input logic unit, the operation logic control module 11 is connected to the operation logic unit, the storage logic control module 12 is connected to the data conversion unit, and the output logic control module 13 is connected to the output logic unit.
[0030] At the beginning of the calculation, the input logic control module 10 will automatically generate corresponding control signals (including but not limited to cache read and write enables and addresses) to cache the data of the weight matrix and the signal matrix. When the weight cache module 1 and the signal cache module 2 are ready with valid data, the data of both will be sent to the operation logic unit for calculation. The operation logic control module 11 will generate control signals (including but not limited to data valid signals, start calculation signals, FFT mode selection signals, and output column valid signals) to drive the operation logic calculation results.
[0031] The data of the present invention are all input in 16 bits. Since the systolic array involves cumulative multiplication and accumulation operations, the bit width of the calculation result will be expanded to 32 bits. Therefore, it is necessary to restore the data to the original bit width in the data truncation module 4. Since the weights of FFT are all less than or equal to 1, in order to reduce errors, all weight data are expanded by 2. 14times, so in the data truncation module 4, the lower 14 bits are staged, and the last 16 bits of the 14th bit are retained so that the next layer can calculate the data again. In order to reduce hardware overhead, the truncated data needs to be processed and then stored. Since the data processed by each layer and each clock is different and the rotation factor symmetry of the base-2 FFT is used, the FSM (Finite-state machine) module 5 is used to process the data in different states, and the processed data is stored in the RAM module 6. The storage logic control module 12 will generate appropriate control signals (including but not limited to read and write enable signals, read and write addresses, selected RAM block signals and read intervals) to store the data generated by the FSM module 5. When all the data of this layer are calculated, the last layer level judgment module 7 will perform level judgment on the FFT of different number of points. If it is not the last layer, the storage logic control module 12 will generate a suitable read address, read the data required for the next layer from the RAM module 6, and send it to the operation logic unit for calculation again. If it is the last layer, the output logic control module 13 will generate a suitable signal (including but not limited to the cache read and write enable and address) to write the final calculated data into the DDR memory 15 through the AXI bus protocol.
[0032] At the beginning of FPGA startup, the information of the weight matrix required for each point needs to be written into the DDR memory through the PS (Processing System) end. The input logic control module 10 will automatically generate a suitable address, and store the signal data (including the real and imaginary parts of the signal required by the operation unit 3 of each dimension) from the DDR memory 15 through the data bus in the RAM module 6 for caching. The RAM module 6 is also a storage module for the generated intermediate data. When all signal information is cached, the input control logic module 10 will generate a suitable address again, and store the weight data (including the real and imaginary parts of the weight required by the PEA of each dimension) from the DDR memory 15 through the AXI bus in the weight cache module 1 for caching. When enough data is cached, the signal data and weight data will be sent to the systolic array of the operation unit 3 for calculation. The calculation of the systolic array adopts the form of "pipeline", that is, while calculating the output result, the systolic array will also receive new input data for calculation, and will continue to generate valid data, thereby improving the calculation efficiency. When the systolic array generates valid data, the data will be sent to the data conversion unit. The data is processed for the next layer of FFT calculation. When the input calculation of this layer is completed, the storage logic control module 12 will automatically generate a suitable address, take out the data required for the next layer from the RAM module 6, and send it to the systolic array for calculation. When the calculation of all layers is completed, the final calculation result will be written back to the DDR memory 15. First, the final result data generated needs to be cached through the output cache module 8. When enough data is cached, the output logic control module 13 will generate a suitable data bus command (including memory access address and burst length, etc.), and write the FFT calculation result back to the DDR memory 15 through the data bus.
[0033] Figure 2 The diagram is a schematic diagram of the interaction between the accelerator and the configuration register. It includes a configuration register module 16 and an FFT accelerator 17, wherein the configuration register module 16 is connected to the controller 9 of the FFT accelerator 17, wherein the FFT accelerator 17 is the controller of the present invention. Figure 1 As shown in the figure, the configuration register module 16 sends configuration information (including the accelerator start signal, the weight and the base address of the signal in DDR, the number of calculated FFT points, etc.) to the FFT accelerator 17. When the data calculation is completed, the output logic control module of the FFT accelerator 17 will also send the calculation completion status information to the configuration register 16.
[0034] Figure 3The figure shows a schematic diagram of the basic systolic array structure, which is only suitable for calculating the base-4 FFT operation, including a weight matrix 18, a signal matrix 19, a PE processing unit 20 and a systolic array 21. Each systolic array 21 is provided with a plurality of PE processing units 20. The weight matrix 18 is input by the top row of PEs, and the signal matrix 19 is input by the rightmost column of PEs. The scale of the systolic array (including the number of rows, columns and dimensions of PEs) can be modified before synthesis according to the actual operation speed and circuit scale. The present invention uses a four-dimensional 8*8 systolic array for calculation according to the circuit scale, operation speed and bandwidth requirements. Due to the special structure of FPGA, complex number operations cannot be performed directly, so a complex number needs to be decomposed into an imaginary part and a real part for separate calculation. For FFT operations, the signal and weight are both in complex form and need to be calculated separately. The butterfly unit of the base-4 FFT consists of 4 complex input signals and four complex weight matrices. The butterfly unit can be mathematically decomposed into the form of multiplication of the signal matrix and the weight matrix, where the signal matrix scale is 1*8 and the weight matrix scale is 8*8. Therefore Figure 3 The structure shown can well calculate the radix-4 FFT butterfly unit. When the weight matrices are the same, the signal matrices can be combined for calculation. For example, if there are two identical weight matrices, the original separate 1*8 signal matrices can be combined into a 2*8 signal matrix, thereby increasing data reuse and computational parallelism.
[0035] like Figure 3 As shown, in the present invention, the input data is divided into two directions, namely the vertical weight matrix 18 and the horizontal signal matrix 19. Each column of PE in the vertical direction shares the same weight data at the same clock, so there is no data transmission in the vertical direction. In the horizontal direction, each row of PE will not share the same data at the same clock, so each clock in the horizontal direction will send the data to the next PE for calculation. This architecture makes the output data more regular without changing the original computing efficiency. The architecture calculates the basis-4FFT process as follows: 1. 0~7 valid clocks: weight data and signal data are input to the systolic array, and all output data are invalid.
[0036] 2. The 8th valid clock: All PEs in the first column are valid (from the right).
[0037] 3. 9~15 valid clocks: Each clock in turn makes all the next column of PEs valid.
[0038] 4. 16~other valid clocks: The first column of PEs starts to be valid again in sequence, entering a new cycle, and no invalid column data will be generated.
[0039] Figure 3The calculation process of the systolic array is optimized, and the order of the output results of the systolic array is improved to make it more regular. The original systolic array calculation process is relatively complicated, and the effective PE regularity of each clock cycle is not obvious, which will increase the complexity of the code end and even affect the overall computing performance. In order to solve this problem, the present invention shares the same weight for the PEs of each column. At the beginning of the calculation, the weight data of each column of PE will generate a delay cycle in turn. For example, in the first clock cycle, only the PE of the rightmost column has weight data input, and the weight inputs of other columns are all 0, and the signal data does not need to be delayed and is directly input into the row of the systolic array. The signal data and weight data will enter the systolic array in turn according to the established order for calculation. Taking the 8*8 systolic array as an example, after seven effective clock cycles, the generated results begin to be effective. In the eighth clock cycle, all the PEs of the first column are effective, and all the PEs of the second column in the ninth cycle are effective. By analogy, each effective PE is all the PEs of the column. In this way, under the premise of not changing the calculation time, the effective PE output each time has a strong regularity, which reduces the complexity for the subsequent development of the code end.
[0040] Figure 4 The figure shows a schematic diagram of a systolic array calculating a radix-2 FFT. The present invention uses a systolic array to perform radix-2 FFT and radix-4 FFT operations. In order to be more suitable for butterfly unit calculation of FFT, an 8*8 systolic array is used. The systolic array of this size is very suitable for radix-4 FFT operations, but not suitable for radix-2 FFT operations. In order to solve this problem, the present invention Figure 3 Based on the systolic array shown in the figure, the optimization is continued. On the basis of not changing the original systolic array scale and calculation process, the 8*8 systolic array is divided into four 4*4 regions, and each region performs a radix-2 FFT butterfly unit calculation. In addition, it is necessary to change the effective PE position of the systolic array output and the position of the weight matrix in the systolic array, and its position corresponds to the position of the calculated result.
[0041] like Figure 4 As shown, the basic systolic array is divided into four parts, and four different radix-2FFT butterfly units are calculated respectively. For the radix-2FFT butterfly unit, there are two complex input signals and two complex rotation factors. The radix-2FFT butterfly unit can also be mathematically decomposed into the form of multiplication of the signal matrix and the weight matrix, where the scale of the signal matrix is 1*4 and the scale of the weight matrix is 4*4. In order not to change the scale and computing architecture of the deployed systolic array, it is necessary to change the output order of the systolic array to realize the calculation of the radix-2FFT butterfly unit. The specific calculation process is as follows: 1. 0~3 valid clocks: The weight matrix and signal matrix are input to the systolic array, and all output data are invalid.
[0042] 2. 4~7 valid clocks: Figure 4 The PEs starting from the first column (from the right) of the upper right area 23 of the middle systolic array are valid in sequence. At this time, only four PEs are valid in each clock cycle.
[0043] 3. 8~11 valid clocks: PEs in the first column of the upper right area 23 and the lower left area 25 in the figure are valid in sequence. At this time, eight PEs are valid in each clock cycle.
[0044] 4. 12~15 valid clocks: PEs start to become valid again in sequence from the first column of 25 in the lower left area of the figure. At this time, only four PEs are valid in each clock cycle. If data is input continuously, the valid PEs are the same as in process 3 above.
[0045] After the above process, the calculation of the radix-2FFT butterfly unit can be realized. If the data is always valid, except for the first four clock cycles and the last four clock cycles when four PEs are valid, the other clock cycles are all eight PEs valid. Therefore, the calculation speed is as efficient as that of the radix-4FFT. Although the complexity of the output rule is aggravated, the hardware resources are not increased and the calculation process is accelerated.
[0046] Figure 5 The figure shows a schematic diagram of the mixed-base FFT calculation process, including weight data 26, initial coding unit 27, signal data 28, non-last layer calculation 29, base 4 to base 2 coding 30 and last layer calculation 31.
[0047] like Figure 5 As shown, weight data 26 is generated by MATLAB. Different weight data are required to generate different FFTs of different numbers of points, pulsation arrays of different dimensions, and calculations of different numbers of layers. The weight data will be stored in a hexadecimal file in TXT format. Similarly, signal data 28 also needs to be generated by MATLAB. After obtaining both data, the data needs to be initially encoded 7. The initial encoding is a reverse encoding method. In order to simplify the code complexity of FPGA, the initial encoding 27 is completed by MATLAB. Only when the encoding is correct can the correct result be obtained. Afterwards, the encoded data is written to the DDR memory 15 of the predetermined address by the PS end of FPGA for storage.
[0048] The initial encoding of radix-2 and radix-4 FFT is different, and there is a re-encoding process when converting radix-4 to radix-2. Figure 5 The base 4 to base 2 encoding 30 is to solve this encoding problem, specifically in Figure 1 This is done in the data conversion unit.
[0049] Figure 6 This is a schematic diagram of converting one layer of base 4 to two layers of base 2 encoding. Figure 5In the specific implementation of the radix-4 to radix-2 encoding 30, one layer of radix-4 butterfly operation is equal to two layers of radix-2 butterfly operation. Among them, 32 is the initial code of the radix-4 butterfly unit, 33 is the output code after the radix-4 butterfly unit, 34 is the new code after the conversion to radix-2 operation, 35 is the code after the first layer of radix-2 operation, 36 is the initial code of the second layer of radix-2 butterfly, and 37 is the code after the second layer of radix-2 butterfly operation. During the conversion process, X0 and X2 of the initial code 32 form a group of radix-2 butterflies, and X1 and X3 of the initial code 30 form a group of radix-2 butterflies for operation. Y0 and Y2 of the output code 35 of the first layer form a group of radix-2 butterflies, and Y1 and Y3 of the output code 35 of the first layer form a group of radix-2 butterflies. After the above encoding process, the result can be calculated correctly.
[0050] The present invention adopts a mixed basis architecture to perform efficient FFT operations. Compared with the original traditional calculation method, using a systolic array to calculate FFT will produce some invalid calculation results, and the calculation results need to be processed to extract the valid data part. If the base-4FFT algorithm is used, although the algorithm will reduce some calculation scales compared to the base-2FFT algorithm, due to the limitation of the systolic array, the last layer does not have the same base-4FFT rotation factor, and the signal matrix cannot be pieced together. When the systolic array is used for calculation, it will take eight times more operation time than the first layer (the rotation factors of all butterflies in the first layer are the same), and the calculation efficiency is low. In order to speed up the FFT calculation process, a mixed basis of base-2FFT and base-4FFT is used for calculation, that is, the base-4FFT of the last layer is converted into a two-layer base-2FFT by base-4 to base-2 encoding. In the present invention, no matter which number of points of FFT is calculated, the last layer is calculated using base-2FFT, and the other layers are calculated using base-4FFT. In addition, the present invention performs matrix transformation on the rotation factor on the basis of using a mixed basis, further improving the calculation speed of FFT, and through the study of the rotation factor of base-2FFT, it is found that it has symmetry, that is, in the base-2FFT of the same layer, there is a certain symmetric relationship between the imaginary part and the real part of the rotation factor of two signals, and only the imaginary part and the real part of the rotation factor of one of the two signals with a symmetric relationship need to be transformed into a certain matrix, and the real part and the imaginary part of the signal are also transformed into an inverse matrix, so that the rotation factors of the two signals with a symmetric relationship can be transformed into the same weight matrix, thereby multiplexing the signal in the signal matrix. Through the above two optimization methods, the present invention can calculate FFT efficiently and quickly.
[0051] The present invention supports 512-point, 1024-point, 2048-point, and 4096-point FFTs. The accelerator can run and calculate on its own by simply writing the calculated points, weights, base addresses of signals, and startup calculation information into the configuration register. The configuration register adopts a standard AXI-Lite protocol bus and supports register read and write operations.
[0052] According to a preferred embodiment of the present invention, the optimized systolic array unit is used as a calculation module for FFT. The systolic array supports parameter configuration of the number of rows, columns and dimensions, that is, the scale of the systolic array can be changed, but the circuit needs to be re-synthesized. Systolic arrays of different dimensions work independently, do not affect each other, and do not share data (including control signals and input data of the systolic array). Deploying a multi-dimensional systolic array can calculate FFT in parallel, thereby having a faster calculation speed. The systolic array of the present invention can support radix-2FFT and radix-4FFT operations without changing the input data structure of the systolic array, the systolic array calculation process and the scale of the systolic array, and only changes the output order of the systolic array.
[0053] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A mixed-radix FFT hardware accelerator based on a systolic array, characterized in that: The invention comprises an input logic unit, an operation logic unit, a data conversion unit and an output logic unit which are connected in sequence and arranged on an FPGA. The operation logic unit is provided with an operation unit (3), the operation unit (3) is connected to the data conversion unit, a systolic array is provided in the operation unit (3), and the systolic sequence uses a mixed basis to accelerate the calculation of FFT; the data conversion unit comprises a data truncation module (4) and a data splicing unit which are connected in sequence, the operation unit (3) is connected to the data truncation module (4), and the data splicing unit is connected to the output logic unit; the input logic unit, the operation unit (3), the data truncation module (4), the data splicing unit and the output logic unit are all connected to a controller (9), and the input logic unit, the data splicing unit, the data splicing unit and the output logic unit are all connected to a DDR memory (15) via a data bus.
2. The mixed-radix FFT hardware accelerator based on a systolic array according to claim 1, characterized in that: The data splicing unit comprises an FSM module (5) and a RAM module (6); the data truncation module (4) is connected to the FSM module (5); the FSM module (5) is respectively connected to the RAM module (6) and the output logic unit; and the RAM module (6) is respectively connected to the input logic unit via a data bus.
3. The mixed-radix FFT hardware accelerator based on systolic array according to claim 2, characterized in that: The input logic unit is connected to the first layer level judgment module (14), and the first layer level judgment module (14) is respectively connected to the controller (9) and the data bus; the output logic unit includes the last layer level judgment module (7) and the output buffer module (8), the FSM module (5) is connected to the last layer level judgment module (7), the last layer level judgment module (7) is respectively connected to the output buffer module (8) and the input logic unit, and the output buffer module (8) is respectively connected to the controller (9) and the data bus; the controller (9) is connected to the configuration register module.
4. The mixed-radix FFT hardware accelerator based on systolic array according to claim 3, characterized in that: The number of the operation units (3) is set to two; the systolic array is an improved systolic sequence, each operation unit (3) includes two 8*8 improved systolic arrays, and the improved systolic sequence supports the calculation of a radix-4 FFT butterfly unit and a radix-2 FFT butterfly unit.
5. The mixed-radix FFT hardware accelerator based on a systolic array according to any one of claims 1 to 4, characterized in that: The controller (9) comprises an input logic control module (10), an operation logic control module (11), a storage logic control module (12) and an output logic control module (13); the input logic unit comprises a signal cache module (2) and a weight cache module (1); the signal cache module (2), the weight cache module (1) and the first-layer hierarchical judgment module (14) are all connected to the input logic control module (10); the signal cache module (2) and the weight cache module (1) are all connected to the operation unit (3); the operation unit (3) is connected to the operation logic control module (11); the data truncation module (4), the FSM module (5) and the RAM module (6) are all connected to the storage logic control module (12); and the output cache module (8) is connected to the output logic control module (13).
6. The mixed-radix FFT hardware accelerator based on systolic array according to claim 5, characterized in that: The signal and weight data generated in the DDR memory (15) are sent to the first layer level judgment module (14) through the data bus. If it is judged to be the data of the first layer, the data is directly sent to the RAM module (6) of the data conversion unit to cache the initial data; if it is other layers, the data is sent to the input logic unit through the data bus to start calculation; the systolic array is responsible for performing FFT butterfly unit calculation on the input weight data and signal data; the data truncation module (4) is responsible for truncation of the calculated data, and the FSM module (5) uses the finite state machine method to process data with different points and different layers, including data position transposition, data inversion and data encoding; the RAM module (6) is responsible for storing the processed data; the last layer level judgment module (7) of the output logic unit determines whether the data input by the data conversion unit is the last layer data. If it is the last layer data, the data will be placed in the output cache module (8); if not, the data will be sent to the input logic unit again; the output cache module (8) transmits the calculation result to the DDR memory (15) through the data bus; At the beginning of the calculation, the input logic control module (10) automatically generates a control signal to cache the data of the weight matrix and the signal matrix. When the weight cache module (1) and the signal cache module (2) prepare valid data, they are sent to the operation unit (3) for calculation. The operation logic control module (11) generates a control signal to drive the operation unit (3) to calculate the result. The storage logic control module (12) generates a suitable control signal to store the data generated by the FSM module (5). When all the data of this layer are calculated, the last layer level judgment module (7) performs level judgment on FFTs of different numbers of points. If it is not the last layer, the storage logic control module (12) generates a suitable read address to read the data required for the next layer from the RAM module (6) and send it to the operation unit (3) for calculation again. If it is the last layer, the output logic control module (13) generates a control signal to write the calculation result into the DDR memory (15) through the AXI bus protocol. At the beginning of FPGA startup, the information of the weight matrix required for each point is written into the DDR memory (15) through the PS terminal, and the input logic control module (10) generates a suitable address by itself, and stores the signal data from the DDR memory (15) in the RAM module (6) through the data bus for caching. When all signal information is cached, the input control logic module (10) will generate a suitable address again, and store the weight data from the DDR memory (15) in the weight cache module (1) through the AXI bus for caching. When enough data is cached, the signal data and weight data are sent to the systolic array of the operation unit (3) for calculation; when the systolic array generates valid data, the lower 14 bits of the valid data are truncated in the data truncation module (4), and the 16 bits after the 14th bit are retained, and all data are expanded by 2 14 times, so that the next layer can perform FFT calculation on the data again; when the input calculation of this layer is completed, the storage logic control module (12) generates a suitable address by itself, takes out the data required by the next layer from the RAM module (6), and sends it to the systolic array for calculation; when the calculation of all layers is completed, the final calculation result is written back to the DDR memory (15): the final result data generated is cached by the output cache module (8), and when enough data is cached, the output logic control module (13) generates a suitable data bus command, and writes the FFT calculation result back to the DDR memory (15) through the data bus.
7. The systolic array-based mixed-radix FFT hardware accelerator according to claim 4, characterized in that: The method for improving the pulse sequence to perform radix-4 FFT butterfly unit is: There are multiple PE processing units in each systolic array. The weight matrix is input by the top row of PE processing units, and the signal matrix is input by the rightmost column of PE processing units. The butterfly unit of the radix-4FFT is composed of 4 complex input signals and four complex weight matrices. The butterfly unit of the radix-4FFT is decomposed into the form of multiplication of the signal matrix and the weight matrix, where the scale of the signal matrix is 1*8 and the scale of the weight matrix is 8*8. When the weight matrices are the same, the signal matrices are put together for calculation. The input data is divided into two directions, namely the vertical weight matrix and the horizontal signal matrix. Each column of PE processing units in the vertical direction shares the same weight data in the same clock. The PE processing units in each row in the horizontal direction do not share the same data in the same clock. Each clock in the horizontal direction sends the data to the next PE processing unit for calculation. The process of calculating the radix-4 FFT is: 1) 0~7 valid clocks: the weight matrix and signal matrix are input to the systolic array, and all the output data are invalid; 2) The 8th valid clock: All PE processing units in the first column from the right are valid; 3) 9th to 15th valid clocks: each clock makes the next column of PE processing units all valid; 4) 16~other valid clocks: The first column of PE processing units from the right starts to be valid again in sequence, entering a new cycle, and no invalid column data will be generated.
8. The systolic array-based mixed-radix FFT hardware accelerator according to claim 7, characterized in that: The method for improving the pulse sequence to perform radix-2 FFT butterfly unit is: On the basis of not changing the scale and calculation process of the original systolic array, the 8*8 systolic array is divided into four 4*4 regions including the upper right region, the upper left region, the lower right region and the lower left region. Each region performs a different radix-2 FFT butterfly unit calculation, and changes the position of the PE processing unit where the systolic array outputs effective and the position of the weight matrix in the systolic array to correspond to the position of the calculation result; The radix-2FFT butterfly unit consists of two complex input signals and two complex rotation factors. The radix-2FFT butterfly unit is decomposed into the form of multiplication of the signal matrix and the weight matrix, where the scale of the signal matrix is 1*4 and the scale of the weight matrix is 4*4; the output order of the systolic array is changed to realize the calculation of the radix-2FFT butterfly unit. The calculation process is: a. 0~3 valid clocks: the weight matrix and signal matrix are input to the systolic array, and all output data are invalid; b. 4~7 valid clocks: PE processing units are valid in sequence starting from the first column on the right side of the upper right area of the systolic array. At this time, only four PE processing units are valid in each clock cycle; c. 8~11 valid clocks: PE processing units are valid in sequence starting from the first column on the right side of the upper right area and the lower left area. At this time, eight PE processing units are valid in each clock cycle; d. 12~15 valid clocks: The PE processing units start to become valid again in sequence starting from the first column on the right side of the lower left area. At this time, only four PE processing units are valid in each clock cycle. If data is input continuously, the valid PE processing units are the same as in step c above.
9. The mixed-radix FFT hardware accelerator based on a systolic array according to any one of claims 6 to 8, characterized in that: The FSM module (5) is provided with an initial coding unit and a base-4-to-base-2 coding, the initial coding unit is connected to the RAM module (6), and the base-4-to-base-2 coding is connected to the data truncation module (4); Different FFTs with different numbers of points, systolic arrays with different dimensions, and calculations with different numbers of layers all require the generation of different weight matrices. The weight matrix and signal matrix are stored in hexadecimal files in TXT format; Initially encode the weight matrix and the signal matrix; The encoded data is written from the PS side of the FPGA to the DDR memory (15) of the predetermined address for storage; the non-last layer uses the radix-4 FFT butterfly unit for calculation, and the radix-4 FFT of the last layer is converted into a two-layer radix-2 FFT through radix-4 to radix-2 encoding, and the radix-2 FFT butterfly unit is used for calculation; In the base-2FFT of the same layer, there is a certain symmetric relationship between the imaginary part and the real part of the rotation factor of two signals. The imaginary part and the real part of the rotation factor of one of the two signals with the symmetric relationship are matrix transformed, and the real part and the imaginary part of the signal are inversely transformed. The rotation factors of the two signals with the symmetric relationship are transformed into the same weight matrix, and the signal is multiplexed in the signal matrix.
10. The systolic array-based mixed-radix FFT hardware accelerator according to claim 9, characterized in that: The initial encoding is a reverse encoding method; The calculated points, weights, base addresses of signals, and information for starting calculations are written into the configuration registers. The configuration registers use the standard AXI-Lite protocol bus and support register read and write operations. The number of rows, columns and dimensions of the PE processing units of the systolic array are modified before synthesis according to the actual operation speed and circuit scale.
Citation Information
Patent Citations
Lightweight neural network accelerator based on row input and design method thereof
CN117852598A
Cited By
Systolic array based on spatial domain
CN121144258A
Computing device with zero detection-based clock gating and method for operating the same
KR102961283B1