Calculation unit based on coarse-grained reconfigurable architecture and FFT (Fast Fourier Transform) processor

By designing a computing unit based on coarse-grained reconfigurable architecture, one computing unit realizes that a base 2 butterfly operation is completed, solving the data transmission and Internet network complexity problems caused by the execution of butterfly operation by multiple computing units in the prior art, and achieving resource optimization and accuracy guarantee.

CN120045513AActive Publication Date: 2025-05-27CHINA SATELLITE NETWORK EXPLORATION CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510518376.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the existing FFT calculation scheme, performing a primary basis 2/base 4 butterfly operation requires multiple computing units, resulting in high data transmission and Internet network complexity between computing units.

Method used

A computing unit based on a coarse-grained reconfigurable architecture is designed, and a one-time basis 2 butterfly operation can be realized through one computing unit, reducing data transmission and interconnection resources between computing units. The calculation unit includes an input stage, an addition stage, a multiplication stage and an output stage, and a cutter is provided in the addition and multiplication stages to reduce accuracy loss.

Benefits of technology

The implementation of a base 2 butterfly operation requires only one computing unit, which reduces data transmission and interconnection resources, reduces the Internet complexity of the computing array, and minimizes accuracy losses while ensuring that the data flow bit width remains unchanged.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045513A_ABST
    Figure CN120045513A_ABST
Patent Text Reader

Abstract

The invention relates to the field of wireless communication, and provides a coarse-grained reconfigurable architecture-based computing unit and an FFT (Fast Fourier Transform) processor, the computing unit is used for executing one butterfly operation, and specifically comprises an input stage for providing computing data and twiddle factors required by the operation; the addition stage is used for carrying out first addition and subtraction operation on the calculation data; the multiplication stage is used for selecting a first addition and subtraction operation result to perform multiplication, second addition and subtraction operation or direct output with the twiddle factor according to the twiddle factor; and the output stage is used for outputting an operation result of the multiplication stage. According to the invention, one butterfly operation can be completed by only one calculation unit, and the precision loss is reduced as much as possible under the condition that the bit width of the data stream is not changed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of wireless communication, and particularly relates to a computing unit and an FFT processor based on a coarse-grained reconfigurable architecture. Background Art

[0002] In recent years, the coarse-grained reconfigurable architecture (CGRA) has become a research hotspot in the industry because it well balances the flexibility of GPP and the high performance of ASIC. In order to accelerate the fast Fourier transform (FFT), a core algorithm in the field of digital signal processing, a dedicated FFT processor needs to be designed. Since FFTs with different numbers of points are often used in the field of digital signal processing, and the coarse-grained reconfigurable architecture can well adapt to the processing of FFTs with different numbers of points and different radices, the coarse-grained reconfigurable architecture has become the mainstream implementation solution in the field of digital signal processing. Summary of the Invention

[0003] Embodiments of this application provide a computing unit and an FFT processor based on a coarse-grained reconfigurable architecture, which can effectively reduce the hardware overhead and the complexity of the interconnection network of the computing array.

[0004] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.

[0005] According to the first aspect of the embodiments of this application, a computing unit based on a coarse-grained reconfigurable architecture is provided. The computing unit is used to perform one butterfly operation, and specifically includes: an input stage for providing the calculation data and rotation factors required for the operation; an addition stage for performing the first addition and subtraction operation on the calculation data; a multiplication stage for selecting, according to the rotation factors, to multiply the result of the first addition and subtraction operation by the rotation factors, perform the second addition and subtraction operation, or directly output; and an output stage for outputting the operation result of the multiplication stage. Through the general PE unit applicable to butterfly operations proposed in this application, one butterfly operation only needs to be completed by using one computing unit, greatly reducing the data transmission and interconnection resources between computing units.

[0006] In one embodiment of the present application, the addition stage includes a complex number decomposition unit, a first data distribution unit, a first addition unit, a second addition unit, a first subtraction unit, and a second subtraction unit. Among them, the input end of the complex number decomposition unit is connected to the input stage, and the output ends of the complex number decomposition unit are respectively connected to the input end of the first data distribution unit and the multiplication stage. The complex number decomposition unit is used to decompose the rotation factor and the calculation data provided by the input stage into real part data and imaginary part data; wherein, the calculation data includes first input data and second input data; the output ends of the first data distribution unit are respectively connected to the input ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit. The first data distribution unit is used to respectively input the real part data of the first input data and the second input data into the first addition unit and the first subtraction unit, and input the imaginary part data of the first input data and the second input data into the second addition unit and the second subtraction unit; the output ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit are connected to the multiplication stage.

[0007] In one embodiment of the present application, the addition stage further includes a first data truncator, which is arranged between the output ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit and the multiplication stage, and is used to respectively intercept the required bit-width results from the operation results of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit.

[0008] In one embodiment of the present application, the first data truncator intercepts the operation results in a saturation truncation manner.

[0009] In one embodiment of the present application, the multiplication stage includes a second data distribution unit, a multiplier unit, a second data truncator, a third adder unit, a third subtractor unit, and a complex number combining unit. Among them, the input end of the second data distribution unit is respectively connected to the output end of the complex number decomposition unit, the output ends of the first adder unit, the second adder unit, the first subtractor unit, and the second subtractor unit. The output end of the second data distribution unit is respectively connected to the input ends of the multiplier unit and the complex number combining unit. The second data distribution unit is configured to select to input the real part data and the imaginary part data of the rotation factor, and the operation results of the first adder unit, the second adder unit, the first subtractor unit, and the second subtractor unit into the multiplier unit or the complex number combining unit according to the value of the rotation factor. The output end of the multiplier unit is connected to the input end of the second data truncator, the output end of the second data truncator is respectively connected to the input ends of the third adder unit and the third subtractor unit, and the output ends of the third adder unit and the third subtractor unit are respectively connected to the input end of the complex number combining unit. The second data truncator is used to truncate the integer part of the operation result of the multiplier unit. The complex number combining unit is used to combine the received data into a complex number and output it to the output stage through the output end.

[0010] In one embodiment of the present application, the multiplication stage further includes a third truncator, which is arranged between the third adder unit and the third subtractor unit and the complex number combining unit, and is used to truncate the operation results of the third adder unit and the third subtractor unit and then input them into the complex number combining unit.

[0011] In one embodiment of the present application, the input stage includes a rotation factor register and a local data register. The rotation factor register is used to provide a rotation factor, and the local data register is used to provide calculation data. The data in the local data register includes the calculation results of the current calculation unit and the data provided by the shared memory.

[0012] In one embodiment of the present application, the output stage includes a local data register shared with the input stage, which is used to store the operation results of the multiplication stage.

[0013] In one embodiment of the present application, the local data register includes a first sub-interval and a second sub-interval, which respectively provide first input data and second input data for the input stage, and respectively store first output data and second output data output by the output stage.

[0014] According to the second aspect of the embodiments of the present application, there is provided an FFT processor based on a coarse-grained reconfigurable architecture. The FFT processor includes: an array of computing units, including a plurality of homogeneous computing units as described in the first aspect, and is used to execute FFT calculation tasks in digital signal processing.

[0015] In one embodiment of the present application, the FFT processor includes: a shared memory for caching intermediate data and calculation results calculated by each calculation unit; each calculation unit in the calculation unit array is connected to an adjacent calculation unit for data transmission, and each non-adjacent calculation unit in the calculation unit array performs data transmission through the shared memory.

[0016] In one embodiment of the present application, the FFT processor includes: a main controller for generating configuration information according to an FFT calculation task; the configuration information is used to complete the reconstruction of the calculation unit array and the shared memory.

[0017] In one embodiment of the present application, the FFT processor further includes: a configuration memory for storing configuration information generated by the main controller, including configuration information of the calculation unit array, configuration information of the shared memory, and configuration information of the interconnection network.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic diagram of a calculation unit in an embodiment of the present application.

[0021] Figure 2 It is a schematic diagram of an FFT processor in an embodiment of the present application.

[0022] Figure 3 It is a schematic diagram of the network connection relationship between the calculation unit array and the shared memory of the present application.

[0023] Reference Numerals: 100 - input stage, 110 - rotation factor register, 120 - local data register; 200 - addition stage, 210 - complex number decomposition unit, 220 - first data distributor, 230 - first addition unit, 240 - second addition unit, 250 - first subtraction unit, 260 - second subtraction unit, 270 - first data truncator; 221 - first data distributor, 222 - second data distributor; 300 - Multiplication level, 310 - Second data distribution unit, 320 - Multiplier unit, 330 - Second data truncator, 340 - Third adder unit, 350 - Third subtractor unit, 360 - Complex number merging unit, 370 - Third data truncator, 321 - First multiplier, 322 - Second multiplier, 323 - Third multiplier, 324 - Fourth multiplier, 361 - First complex number merger, 362 - Second complex number merger; 400 - Output level; 500 - Shared memory. Detailed implementation

[0024] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts fall within the scope of protection of the present application. Without conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily combined with each other. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0025] The terms "first" and "second" in the specification and claims of the present application and the above - mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0026] In the current FFT calculation scheme, multiple computing units are required to perform a single radix - 2 / radix - 4 butterfly operation, and there is a large amount of overhead between the computing units. Based on this, the embodiments of the present application propose a computing unit based on a coarse - grained reconfigurable architecture. A single radix - 2 butterfly operation can be achieved through one computing unit, which can effectively reduce the interconnection between PEs and reduce the complexity of the interconnection network of the computing array.

[0027] Please refer to Figure 1, the computing unit mainly includes four parts: an input stage 100, an addition stage 200, a multiplication stage 300, and an output stage 400. Through these four parts, a computing unit can complete a radix-2 butterfly operation and multiply it by the corresponding rotation factor. By selecting multiple interconnected computing units to form a computing unit array according to requirements, FFT calculation can be realized according to the computing task.

[0028] Specifically, the input stage 100 of the computing unit serves as an entrance to provide the computing data and rotation factors required for the operation. It mainly includes a rotation factor register 110 and a local data register 120. Among them, the rotation factor register 110 is used to store the rotation factors required during the calculation process of the computing unit. During the configuration phase, all the rotation factors needed in all phases can be stored in the rotation factor register 110. In one embodiment, the depth of the rotation factor register 110 is designed to be 128. When the number of points of the FFT computing task executed by the computing unit does not exceed 1024 points, all the rotation factors can be stored in the rotation factor register before the execution of the computing task. When the number of points of the executed FFT computing task exceeds 1024 points, new rotation factors need to be stored in the rotation factor register 110 again at the end of a certain phase.

[0029] The local data register 120 in each computing unit only supports access by its own computing unit. The data sources of the local data register 120 include the data results calculated by the current computing unit and the data read from the shared memory 500. Among them, the shared memory 500 stores the intermediate data and computing results calculated by other computing units.

[0030] To avoid memory access conflicts of the computing unit to the local data register 120, in the embodiment of the present application, the local data register 120 of each computing unit is divided into a first sub-interval and a second sub-interval, which respectively provide the first input data and the second input data for the input stage 100, and store the first output data and the second output data output by the output stage 400 respectively. Specifically, the width of each sub-interval is 32 bits, which is used to store 32-bit complex numbers. Among them, the high 16 bits are the real part of the data, and the low 16 bits are the imaginary part of the data; the depth of each sub-interval is 64, which can store 64 data. The depth of the local data register 120 determines the size of the number of points for executing the FFT. In this specification, the depth of the local data register 120 is taken as 64 for illustration. In practical applications, it can be adjusted according to requirements.

[0031] The addition stage 200 of the computing unit, as the part for performing the first addition and subtraction operations on the data provided by the input stage, mainly includes a complex number decomposition unit 210, a first data distributor 220, a first adder 230, a second adder 240, a first subtractor 250, and a second subtractor 260. Among them, the input end of the complex number decomposition unit 210 is connected to the input stage, and the output ends of the complex number decomposition unit 210 are respectively connected to the input end of the first data distributor 220 and the multiplication stage 300; the output ends of the first data distributor 220 are respectively connected to the input ends of the first adder 230, the second adder 240, the first subtractor 250, and the second subtractor 260, and the output ends of the first adder 230, the second adder 240, the first subtractor 250, and the second subtractor 260 are connected to the multiplication stage 300.

[0032] Specifically, the complex number decomposition unit 210 is mainly used to divide the input 32-bit complex number into a 16-bit real part data and an imaginary part data. That is, through the complex number decomposition unit 210, the rotation factor, the first input data and the second input data included in the calculation data are respectively decomposed into two parts: a real part In1 and an imaginary part In2.

[0033] The first data distributor 220 includes two data distributors with the same structure, namely a first data distributor 221 and a second data distributor 222, which respectively process the real part and the imaginary part of the first input data and the second input data, that is, realize the real part data distribution and the imaginary part data distribution of the two input data. In one embodiment, the data distributor includes a first input end, a second input end, a first output end, a second output end, a third output end, and a fourth output end, wherein the first input end is respectively connected to the first output end and the third output end, and the second input end is respectively connected to the second output end and the fourth output end. When receiving data, the real parts of the two input data are respectively input to the first input end and the second input end of a data distributor, and the imaginary parts of the two input data are respectively input to the first input end and the second input end of another data distributor.

[0034] After the first data distribution unit 220, there are connected a first addition unit 230, a second addition unit 240, a first subtraction unit 250, and a second subtraction unit 260, which are respectively used to perform 16-bit signed addition or subtraction operations on the data output by the first data distribution unit 220. In the embodiment of the present application, the input end of the first addition unit 230 is respectively connected to the first output end and the second output end of the first data distributor 221 to complete the real part addition operation of the first input data and the second input data; the second addition unit 240 is respectively connected to the first output end and the second output end of the second data distributor 222 to complete the imaginary part addition operation of the first input data and the second input data; the first subtraction unit 250 is respectively connected to the third output end and the fourth output end of the first data distributor 221 to complete the real part subtraction operation of the first input data and the second input data; the second subtraction unit 260 is respectively connected to the third output end and the fourth output end of the second data distributor 222 to complete the imaginary part subtraction operation of the first input data and the second input data.

[0035] Since the addition or subtraction of two data may cause the bit width of the calculation result to overflow during the addition or subtraction process, in the embodiment of the present application, a first data truncator 270 is provided at the output end of the first addition unit 230, the second addition unit 240, the first subtraction unit 250, and the second subtraction unit 260, which is used to truncate one bit of the 17-bit input data and output 16-bit data. The data passing through the first data truncator 270 can be input to the multiplication stage 300 for subsequent operations. Specifically, the first data truncator 270 uses the method of saturation truncation to reduce the error caused by truncation. The saturation truncation process includes: if the calculation result exceeds the maximum value of the data that can be stored in the required data format, then the maximum value is used to represent this data; if the calculation result exceeds the minimum value of the data that can be stored in the required data format, then the minimum value is used to represent this data.

[0036] The multiplication stage 300 of the computing unit selects to multiply the result of the first addition and subtraction operation by the rotation factor, perform the second addition and subtraction operation, or directly output according to the rotation factor. Specifically, the multiplication stage 300 mainly includes a second data distribution unit 310, a multiplier unit 320, a second data truncator 330, a third addition unit 340, a third subtraction unit 350, and a complex number merging unit 360. Among them, the input end of the second data distribution unit 310 is respectively connected to the output end of the complex number decomposition unit 210, the output ends of the first addition unit 230, the second addition unit 240, the first subtraction unit 250, and the second subtraction unit 260 (after connecting to the first data truncator 270, it is connected to the output end of the first data truncator 270), and the output ends of the second data distribution unit 310 are respectively connected to the input ends of the multiplier unit 320 and the complex number merging unit 360; the output end of the multiplier unit 320 is connected to the input end of the second data truncator 330, the output end of the second data truncator 330 is respectively connected to the input ends of the third addition unit 340 and the third subtraction unit 350, and the output ends of the third addition unit 340 and the third subtraction unit 350 are respectively connected to the input ends of the complex number merging unit 360.

[0037] Specifically, the second data distribution unit 310 is used to directly output the data that does not need to be multiplied by the rotation factor to the complex number merging unit 360, and distribute other data to the multiplier unit 320 for calculation. It should be noted that the distribution path of the second data distribution unit 310 is controlled by the input rotation factor. When the real part data and the imaginary part data of the rotation factor are 0, the operation results of the first addition unit 230, the second addition unit 240, the first subtraction unit 250, and the second subtraction unit 260 are directly distributed to the complex number merging unit 360; when the real part data and the imaginary part data of the rotation factor are not 0, the real part data and the imaginary part data of the rotation factor, and the operation results of the first addition unit 230, the second addition unit 240, the first subtraction unit 250, and the second subtraction unit 260 are distributed to the multiplier unit 320.

[0038] Taking the butterfly operation with base 2 as an example, the second data distribution unit 310 includes input terminals i1, i2, i3, i4, i5, i6 and output terminals o1, o2, o3, o4, o5, o6, o7, o8, o9, o10, o11, o12. Among them, the input terminals i1 and i2 respectively receive the real part and the imaginary part obtained by the rotation factor decomposition, which serve as input data and control signals at the same time. The input terminals i3 and i4 respectively receive the real part data obtained by truncating the output of the first addition unit 230 and the first subtraction unit 250 by the first data truncator 270. The input terminals i5 and i6 respectively receive the imaginary part data obtained by truncating the output of the second addition unit 240 and the second subtraction unit 260 by the first data truncator 270. When performing data distribution, the data at the input terminals i3 and i4 are respectively distributed to the output terminals o1 and o2, and the outputs of the remaining ports are controlled by the input terminals i1 and i2. Specifically, when the input terminal i1 is 0 and the input i2 is 0, the data at the input terminal i5 is distributed to the output terminal o3, and the data at the input terminal i6 is distributed to the output terminal o4, and the remaining output terminals do not output data; when the input terminal i1 or the input terminal i2 is not 0, the data at the input terminal i1 is distributed to the output terminals o5 and o9, the data at the input terminal i2 is distributed to o7 and o11, the data at the input terminal i5 is distributed to o6 and o10, and the data at the input terminal i6 is distributed to o8 and o12. In the second data distribution unit 310, the data at the output terminals o1 to o4 are directly input to the complex number merging unit 360, and the data at the output terminals o5 to o12 are output to the multiplier unit 320 to perform a 16-bit signed multiplication operation.

[0039] In the embodiment of the present application, the multiplier unit 320 includes a first multiplier 321, a second multiplier 322, a third multiplier 323 and a fourth multiplier 324. Among them, the first multiplier 321 receives the data at the output terminals o5 and o6 of the second data distribution unit 310, the second multiplier 322 receives the data at the output terminals o7 and o8 of the second data distribution unit 310, the third multiplier 323 receives the data at the output terminals o9 and o10 of the second data distribution unit 310, and the fourth multiplier 324 receives the data at the output terminals o11 and o12 of the second data distribution unit 310. Since in the FFT calculation process, one of the operands of the multiplication operation is a rotation factor, by the definition of the rotation factor:

[0040] It can be seen that both the real part and the imaginary part of the rotation factor are less than 1. Therefore, during the multiplication operation, data overflow does not occur in the integer part of the calculation result, and the integer part of the multiplication calculation result does not require a truncation operation. Multiplying two data with 7 decimal places will cause the decimal width of the calculation result to reach a maximum of 14 bits, and a maximum of 7 bits of the decimal part need to be truncated. Therefore, in the embodiment of the present application, a second data truncator 330 is added after the multiplier unit 320. The second data truncator 330 uses the method of rounding truncation to truncate the overflowing decimal places to reduce the error caused by truncation. The rounding truncation process is as follows: if the highest bit of the decimal places to be truncated is 1, directly discard the decimal places to be truncated, and add 1 to the result after truncation; if the highest bit of the decimal places to be truncated is 0, directly discard the decimal places to be truncated, and add 0 to the result after truncation.

[0041] The data truncated by the second data truncator 330 still needs to be input into the third addition unit 340 and the third subtraction unit 350. Among them, the third subtraction unit 350 receives the data truncated by the second data truncator 330 output by the first multiplier 321 and the second multiplier 322, and completes the subtraction operation; the third addition unit 340 receives the data truncated by the second data truncator 330 output by the third multiplier 323 and the fourth multiplier 324, and completes the addition operation.

[0042] In the multiplication stage of the embodiment of the present application, a third data truncator 370 is also provided, which is used to truncate the output data of the third addition unit 340 and the third subtraction unit 350. Similar to the addition stage 200, the third data truncator 370 uses the method of saturation truncation to reduce the error caused by truncation. Through the first data truncator 270, the second data truncator 330, and the third data truncator 370, the precision loss is minimized as much as possible while ensuring that the data stream width remains unchanged.

[0043] The complex number merging unit 360 mainly merges the input data into a complex number for output. The calculation unit implements the butterfly operation of two complex numbers, and the final output is also two complex numbers. Therefore, in the embodiment of the present application, the complex number merging unit 360 includes a first complex number merger 361 and a second complex number merger 362. The first complex number merger 361 receives the data at the output terminals o1 and o2 of the second data distribution unit 310, and forms the first output data as the high 16 bits and the low 16 bits of the merger result respectively; the second complex number merger 362 receives the data at the output terminals o3 and o4 of the second data distribution unit 310 or the output data of the third data truncator 370 according to the value of the rotation factor, and merges to form the second output data, which is finally sent to the output stage 400.

[0044] The output stage 400 of the calculation unit is used to store the calculation result of the multiplication stage 300 in the local data register 120.

[0045] The general computing unit applicable to butterfly operations proposed in this application enables a radix-2 butterfly operation to be completed using only one computing unit at a time, greatly reducing the data transmission and interconnection resources when using an array of computing units. An addition / subtraction truncation unit and a multiplication truncation unit are provided inside the computing unit to minimize precision loss as much as possible while keeping the data stream bit width unchanged.

[0046] Please refer to Figure 2 , an FFT processor is also proposed in an embodiment of this application. The processor includes an array of computing units and a shared memory.

[0047] Specifically, the array of computing units includes several of the aforementioned homogeneous computing units for implementing digital signal processing computing tasks. Each computing unit is reconfigurable and applicable to the radix-2 or radix-4 fast Fourier transform algorithm; the shared memory is used to cache the intermediate data and computing results calculated by each computing unit. The interconnection network of the array of computing units is divided into two levels of interconnection. Each computing unit in the array of computing units is connected to an adjacent computing unit for data transmission, and each non-adjacent computing unit in the array of computing units performs data transmission through the shared memory. In this application, at the level of the array of computing units, a two-level interconnection network is used to implement data transmission in the FFT algorithm process, and data transmission between non-adjacent units is realized through the shared memory, reducing the complexity of the interconnection network while satisfying the data dependence relationship of the FFT algorithm.

[0048] In an embodiment of this application, the array of computing units includes M rows and M columns, a total of M M computing units; each row in the array of computing units is divided into two groups of computing units, denoted as the p-th row group 1 and the p-th row group 2, and the grouping method of taking even terms as one group and odd terms as one group is adopted. Each column in the array of computing units is divided into two groups of computing units, denoted as the q-th column group 1 and the q-th column group 2, and the grouping method of taking even terms as one group and odd terms as one group is adopted, where p ≤ M, q ≤ M, and M ≥ 2. The shared memory includes 2M sub-memories, namely sub-memory b00 to sub-memory b2M-1. Each sub-memory in the shared memory is respectively interconnected with the computing units of a preset row, a preset group, a preset column, and a preset group. Specifically, the connection relationship between each sub-memory and each computing unit in the array of computing units is as follows: Sub-memory b00 is interconnected with row group 1 of the first row and column group 1 of the first column in the array of computing units; Sub-memory b01 is interconnected with row group 2 of the first row and column group 1 of the second column in the array of computing units; Sub-memory b02 is interconnected with row group 1 of the second row and column group 2 of the first column in the array of computing units; Sub-memory b03 is interconnected with row group 2 of the second row and column group 2 of the second column in the array of computing units; Until The sub-memory b2M-4 is interconnected with row group 1 of the M-1th row and column group 1 of the M-1th column in the computing unit array; The sub-memory b2M-3 is interconnected with row group 2 of the M-1th row and column group 1 of the Mth column in the computing unit array; The sub-memory b2M-2 is interconnected with row group 1 of the Mth row and column group 2 of the M-1th column in the computing unit array; The sub-memory b2M-1 is interconnected with row group 2 of the Mth row and column group 2 of the Mth column in the computing unit array.

[0049] Please refer to Figure 3 , taking the computing unit array with 64 computing units as an example for illustration below. The first-level interconnection of the computing unit array is that each computing unit is directly interconnected with its adjacent PE unit, and the second-level interconnection is the interconnection between the PE and the shared memory to achieve data transmission between non-adjacent computing units. At this time, the shared memory consists of 16 sub-memories, namely b00~b15. The four PEs in the odd-numbered columns of each row are classified into row group 1, and the four PEs in the even-numbered columns are classified into row group 2; the four PEs in the odd-numbered rows of each column are classified into column group 1, and the four PEs in the even-numbered rows are classified into column group 2. The situations of the sub-memories b00~b15 being interconnected with specific PEs are as follows: b00 is interconnected with row group 1 of the 1st row and column group 1 of the 1st column; b01 is interconnected with row group 2 of the 1st row and column group 1 of the 2nd column; b02 is interconnected with row group 1 of the 2nd row and column group 2 of the 1st column; b03 is interconnected with row group 2 of the 2nd row and column group 2 of the 2nd column; b04 is interconnected with row group 1 of the 3rd row and column group 1 of the 3rd column; b05 is interconnected with row group 2 of the 3rd row and column group 1 of the 4th column; b06 is interconnected with row group 1 of the 4th row and column group 2 of the 3rd column; b07 is interconnected with row group 2 of the 4th row and column group 2 of the 4th column; b08 is interconnected with row group 1 of the 5th row and column group 1 of the 5th column; b09 is interconnected with row group 2 of the 5th row and column group 1 of the 6th column; b10 is interconnected with row group 1 of the 6th row and column group 2 of the 5th column; b11 is interconnected with row group 2 of the 6th row and column group 2 of the 6th column; b12 is interconnected with row group 1 of the 7th row and column group 1 of the 7th column; b13 is interconnected with row group 2 of the 7th row and column group 1 of the 8th column; b14 is interconnected with row group 1 of the 8th row and column group 2 of the 7th column; b15 is interconnected with row group 2 of the 8th row and column group 2 of the 8th column.

[0050] For a computing unit array containing 64 computing units, in one stage, one computing unit receives two input data to complete a two-point butterfly operation. Then, the computing unit array can process 128 data in one stage and store 128 intermediate computing results in the local data register. The depth of the local data register being 64 determines that the FFT processor can compute an FFT algorithm with a maximum number of points of 8K. In the embodiments of the present application, the data is cached in each storage unit in the fixed-point format, and the bit width of all storage units is 32 bits. Among them, the upper 16 bits represent the real part of the data, the highest bit is the sign bit of the real part, the 2nd to 9th bits are the integer part of the real part, and the 10th to 16th bits are the fractional part of the real part; the lower 16 bits represent the imaginary part of the data, the 17th bit is the sign bit of the imaginary part, the 18th to 24th bits are the integer part of the imaginary part, and the 25th to 32nd bits are the fractional part of the imaginary part.

[0051] For the shared memory interconnected with the computing unit array, the width of the sub-memory is designed to be 32 bits and is used to store 32-bit complex numbers, where the upper 16 bits are the real part of the data and the lower 16 bits are the imaginary part of the data. Due to the memory access conflict problem, multiple computing units cannot access the same address of the same shared memory at the same time. In the execution of the fast Fourier transform algorithm, there will be a situation where four computing units pass through the same shared memory for data transmission in a certain stage. In the embodiments of the present application, the depth of each sub-memory of the shared memory is designed to be 4, and different computing units can read or write to four addresses in the sub-memory at the same time. However, it should be noted that it is impossible to read and write to the same address of the sub-memory at the same time.

[0052] It should be added that in each computing stage, the computing results will only exist in the local data register. After one stage of computing is completed and before entering the next stage, it is necessary to refresh the data in the shared memory and the local memory, that is, to obtain the data in the local data register of other computing units through the shared memory. Further, the FFT processing also includes a main controller and a configuration memory. Among them, the main controller is responsible for starting the processor, compiling the FFT computing task to generate corresponding configuration information, and storing it in the configuration memory; at the same time, controlling the configuration memory to transmit the configuration information to the computing unit array and the shared memory, and completing the reconstruction of the PEA and the shared memory through the configuration information. The configuration memory is used to store the configuration information generated by the main controller, including the configuration information of the computing unit, the configuration information of the shared memory, and the configuration information of the interconnection network.

[0053] The FFT processor proposed in the embodiments of the present application uses a two-level interconnection network to implement data transmission in the FFT algorithm process; the data transmission between non-adjacent units is realized through a shared memory, which reduces the complexity of the interconnection network while satisfying the data dependence relationship of the FFT algorithm. Based on the flexible configuration method of the reconfigurable coarse-grained architecture and the reconfigurable performance of the arithmetic unit array, the FFT processing tasks with different numbers of points and radices can be dealt with more flexibly and richly.

[0054] To better verify the FFT processor proposed in the present application, taking the radix-2 frequency division 1024-point fast Fourier transform algorithm (FFT) as an example for further illustration, the FFT processor calculates the radix-2 frequency division 1024-point FFT in a total of 10 stages, and the distance difference between the input data of the butterfly operation in the nth stage is 1024 / 2 n , and the rotation factor used is , where m is from 0 to 1024 / 2 n . For a computing unit array containing 64 computing units (PEs), the operand of each PE is 2, and 128 data can be input simultaneously for calculation in one computing cycle. Each stage needs to be divided into eight inputs to calculate all the data. The data input positions of the 10 stages are shown in Table 1 below: Table 1 Input data table for each stage

[0055] Since the input data sources of the computing units are different in each stage, the FFT processor has cached the 1024-point initial data in the local data registers of 64 PEs in a frequency division manner during the configuration stage. Each PE stores 16 data. Since each PE has two sub-intervals, each sub-interval stores 8 data. The data in the corresponding addresses of the two sub-intervals are input into one PE for butterfly operation. The input data sources of the PEs in the 10 stages are shown in Table 2 below: Table 2 Input data source table for each stage

[0056] It should be noted that the data sources of non-adjacent PEs shown in Table 2 realize data transmission through a shared memory.

[0057] Furthermore, when implementing FFT algorithms with other numbers of points, the interconnection resources will not increase additionally. When performing FFT with a number of points greater than 1K, the number of input data times in each stage will increase. Taking the 8K-point FFT as an example, the computing array can calculate 128 data points at a time. Therefore, the input data needs to be divided into 64 groups in each stage, and each group is sequentially input into the computing array. The complete computing process includes 13 stages. In the first 7 stages, the PE only accesses the intermediate data in the local data register. The data sources in the last 6 stages are the same as those in the 1024-point FFT. Therefore, no additional interconnection resource overhead will be generated when calculating FFT with other numbers of points.

[0058] Furthermore, the FFT processor proposed in the embodiment of the present application is not only applicable to the radix-2 frequency division 1024-point fast Fourier transform algorithm, but also applicable to the radix-4 frequency division 1024-point fast Fourier transform algorithm. When the computing task is the radix-4 frequency division 1024-point fast Fourier transform algorithm, during the configuration stage, generating a configuration to change the data distribution method of the second data distribution unit inside the computing unit can meet the data stream of the radix-4 butterfly operation.

[0059] At this time, the second data distribution unit performs data distribution operations in two times, that is, performs two data distribution controls according to the values of the twiddle factors. The first twiddle factor controls the data output of the input ends i3 and i4, and the second twiddle factor controls the data output of the input ends i5 and i6. The data input from the input ends i3 and i4 is no longer directly distributed to the output ends o1 and o2 respectively, but the distribution situation is determined by using the twiddle factor as a control signal. Specifically: During the first distribution, the twiddle factors i11 and i12 are used as input data and control signals, and the distribution process is as follows: When the input end i11 is 0 and the input end i12 is 0, the data of the input end i3 is directly distributed to the output end o1, and the data of the input end i6 is directly distributed to the output end o2. At the same time, the rest of the output ends do not output data.

[0060] When the input end i11 is not 0 or the input end i12 is not 0, the data of the input end i11 is distributed to the output ends o5 and o9, the data of the input end i3 is distributed to the output ends o6 and o10, the data of the input end i12 is distributed to the output ends o7 and o11, and the data of the input end i6 is distributed to the output ends o8 and o12.

[0061] During the second distribution, the twiddle factors i21 and i22 are used as input data and control signals, and the distribution process is as follows: When the input end i21 is 0 and the input end i22 is 0, the data of the input end i5 is directly distributed to the output end o3, and the data of the input end i6 is directly distributed to the output end o4. At the same time, the rest of the output ends do not output data.

[0062] When the input terminal i21 is not 0 or the input terminal i22 is 0, the data of the input terminal i21 is distributed to the output terminals o5 and o9, the data of the input terminal i5 is distributed to the output terminals o6 and o10, the data of the input terminal i22 is distributed to the output terminals o7 and o11, and the data of the input terminal i6 is distributed to the output terminals o8 and o12.

[0063] At this time, the multiplication units in the computing unit are multiplexed. After the four multipliers complete the multiplication of the first rotation factor and the intermediate result, they immediately accept the second rotation factor to perform the multiplication operation with another intermediate result, improving the utilization rate of the computing unit and reducing the consumption of hardware resources by sacrificing the computing cycle.

[0064] After the FFT processor completes the above reconstruction of the computing unit, the process of executing the radix-4 frequency division 1024-point fast Fourier transform algorithm is as follows: The first stage: First, calculate sub-stage 1. First, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the a-th and the (a + 512)-th, the (a + 256)-th and the (a + 768)-th respectively, where a = 0~255. Then calculate sub-stage 2. Similarly, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the a-th and the (a + 256)-th, the (a + 512)-th and the (a + 768)-th respectively, where a = 0~255.

[0065] The second stage: First, calculate sub-stage 1. First, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the b-th and the (b + 128)-th, the (b + 64)-th and the (b + 192)-th respectively, where b = 256 k~256 k + 63, k = 0~3. Then calculate sub-stage 2. Similarly, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the b-th and the (b + 64)-th, the (b + 128)-th and the (b + 192)-th respectively, where b = 0~255.

[0066] The third stage: First, calculate sub-stage 1. First, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the c-th and the (c + 32)-th, the (c + 16)-th and the (c + 48)-th respectively, where c = 128 k~128 k + 79, where k = 0 to 7. Then calculate sub - stage 2. Similarly, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the c - th and the (c + 16)-th, the (c + 32)-th and the (c + 48)-th respectively, where c = 128 k to 128 k + 79, where k = 0 to 7.

[0067] Fourth stage: First, calculate sub - stage 1. First, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the d - th and the (d + 8)-th, the (d + 4)-th and the (d + 12)-th respectively, where d = 16 k to 16 k + 3, where k = 0 to 63. Then calculate sub - stage 2. Similarly, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the d - th and the (d + 4)-th, the (d + 8)-th and the (d + 12)-th respectively, where d = 16 k to 16 k + 3, where k = 0 to 63.

[0068] Fifth stage: First, calculate sub - stage 1. First, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the e - th and the (e + 2)-th, the (e + 1)-th and the (e + 3)-th respectively, where e = 4 k, where k = 0 to 255. Then calculate sub - stage 2. Similarly, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the e - th and the (e + 1)-th, the (e + 2)-th and the (e + 3)-th respectively, where e = 4 k, where k = 0 to 255.

[0069] Through the above five stages, the calculation task of the radix - 4 frequency - division 1024 - point fast Fourier transform algorithm can be completed.

[0070] In this application, by designing a general - purpose computing unit suitable for butterfly operations, one radix - 2 butterfly operation only needs to be completed by using one PE, which greatly reduces the data transmission and interconnection resources between PEs. And an addition - subtraction truncation unit and a multiplication truncation unit are set inside the computing unit, so as to ensure that the precision loss is minimized as much as possible while the data stream bit - width remains unchanged. In the FFT processor applying the computing unit, a two - level interconnection network is used to realize the data transmission in the FFT algorithm process. The data transmission between non - adjacent units is realized through the shared memory, which reduces the complexity of the interconnection network while satisfying the data - dependence relationship of the FFT algorithm.

[0071] The features, structures, or characteristics described in this application may be combined in any suitable manner in one or more embodiments. Unless otherwise specified, the terms "connected", "coupled", and "connected to" are used to specify an electrical connection between circuit elements that may be direct or may be via one or more other elements. In contrast, when an element is said to be "directly connected to" or "directly coupled to" another element, there are no intervening elements. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood in specific circumstances; the accompanying drawings in the embodiments are used to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0072] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A computing unit based on a coarse-grained reconfigurable architecture, characterized in that: The computing unit is used to perform a butterfly operation, specifically including: An input stage, used to provide calculation data and rotation factors required for operation; An addition stage, used for performing a first addition and subtraction operation on the calculation data; a multiplication stage, for selecting, according to the rotation factor, a first addition and subtraction operation result and the rotation factor for multiplication, a second addition and subtraction operation, or direct output; and The output stage is used to output the multiplication stage operation result.

2. The calculation unit according to claim 1, characterized in that The addition stage comprises a complex number decomposition unit, a first data allocation unit, a first addition unit, a second addition unit, a first subtraction unit and a second subtraction unit, wherein: The input end of the complex decomposition unit is connected to the input stage, and the output end of the complex decomposition unit is respectively connected to the input end of the first data allocation unit and the multiplication stage, and the complex decomposition unit is used to decompose the rotation factor and the calculation data provided by the input stage into real data and imaginary data; wherein the calculation data includes the first input data and the second input data; The output end of the first data distribution unit is connected to the input ends of the first addition unit, the second addition unit, the first subtraction unit and the second subtraction unit respectively, and the first data distribution unit is used to input the real part data of the first input data and the second input data to the first addition unit and the first subtraction unit respectively, and to input the imaginary part data of the first input data and the second input data to the second addition unit and the second subtraction unit; Output ends of the first adding unit, the second adding unit, the first subtracting unit and the second subtracting unit are connected to the multiplication stage.

3. The calculation unit according to claim 2, characterized in that The addition stage also includes a first data truncation device, which is arranged between the output ends of the first addition unit, the second addition unit, the first subtraction unit and the second subtraction unit and the multiplication stage, and is used to respectively truncate the required bit width results from the operation results of the first addition unit, the second addition unit, the first subtraction unit and the second subtraction unit.

4. The calculation unit according to claim 3, characterized in that The first data truncation device truncates the operation result by using a saturation truncation method.

5. The calculation unit according to claim 1 or 2, characterized in that: The multiplication stage includes a second data allocation unit, a multiplier unit, a second data truncation unit, a third addition unit, a third subtraction unit and a complex merging unit, wherein: The input end of the second data allocation unit is respectively connected to the output end of the complex decomposition unit, the first addition unit, the second addition unit, the first subtraction unit and the output end of the second subtraction unit, the output end of the second data allocation unit is respectively connected to the input end of the multiplier unit and the complex merging unit, and the second data allocation unit is used to select the real part data and the imaginary part data of the rotation factor, the operation results of the first addition unit, the second addition unit, the first subtraction unit and the second subtraction unit to be input into the multiplier unit or the complex merging unit according to the value of the rotation factor; The output end of the multiplier unit is connected to the input end of the second data truncation device, the output end of the second data truncation device is respectively connected to the input end of the third addition unit and the third subtraction unit, and the output end of the third addition unit and the third subtraction unit is respectively connected to the input end of the complex merging unit; the second truncation device data is used to truncate the integer part of the operation result of the multiplier unit; the complex merging unit is used to merge the received data into a complex number and output it to the output stage through the output end.

6. The calculation unit according to claim 5, characterized in that The multiplier stage further includes a third data truncation device, which is arranged between the third addition unit, the third subtraction unit and the complex merging unit, and is used for truncating the operation results of the third addition unit and the third subtraction unit and then inputting them into the complex merging unit.

7. The calculation unit according to claim 1, characterized in that The input stage includes a rotation factor register and a local data register, the rotation factor register is used to provide a rotation factor, and the local data register is used to provide calculation data; the data in the local data register includes a calculation result of a current calculation unit and data provided by a shared memory.

8. The calculation unit according to claim 7, characterized in that The output stage includes a local data register shared with the input stage and used for storing the operation result of the multiplication stage.

9. The calculation unit according to claim 8, characterized in that The local data register includes a first sub-interval and a second sub-interval, which respectively provide first input data and second input data to the input stage, and respectively store first output data and second output data output by the output stage.

10. An FFT processor based on a coarse-grained reconfigurable architecture, characterized in that: The FFT processor comprises: A computing unit array, comprising a plurality of isomorphic computing units as claimed in any one of claims 1 to 9, for executing FFT computing tasks in digital signal processing.

11. The FFT processor according to claim 10, characterized in that: The FFT processor comprises: Shared memory, used to cache intermediate data and calculation results calculated by each computing unit; Each computing unit in the computing unit array is connected to an adjacent computing unit and performs data transmission, and each non-adjacent computing unit in the computing unit array performs data transmission through a shared memory.

12. The FFT processor according to claim 10, characterized in that: The FFT processor comprises: The main controller is used to generate configuration information according to the FFT calculation task; the configuration information is used to complete the reconstruction of the calculation unit array and the shared memory.

13. The FFT processor according to claim 10, characterized in that: The FFT processor also includes: The configuration memory is used to store configuration information generated by the main controller, including configuration information of the computing unit array, configuration information of the shared memory and configuration information of the interconnection network.

Citation Information

Patent Citations

  • Configuration method for pipeline-architecture fixed-point FFT word length

    CN103761074A

  • Point-changeable floating point FFT (fast Fourier transform) processor

    CN104268122A

  • Fast Fourier transform hardware design method based on base 2-2 algorithm

    CN109522674A

  • An FFT operation device and method for a power line carrier communication chip

    CN109948112A

  • Radix-2 fast Fourier transform hardware design method based on an FPGA

    CN110765709A