A computing unit and an FFT processor based on a coarse-grained reconfigurable architecture
Through the computing unit and two-level Internet network based on coarse-grained reconfigurable architecture, the efficient completion of the first-order 2 butterfly operation in FFT calculation is achieved, and the problem of excessive interconnection resources between computing units in the prior art is solved, reducing the complexity of the Internet network and reducing accuracy loss.
Patent Information
- Application Number
- CN202510518376.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In the existing FFT calculation scheme, performing a primary basis 2/base 4 butterfly operation requires multiple computing units, resulting in excessive data transmission and interconnection resources between computing units and high Internet complexity.
A computing unit based on a coarse-grained reconfigurable architecture is adopted, and a one-level basis 2 butterfly operation is realized through a computing unit to reduce data transmission and interconnection resources between calculation units, and add, subtract and multiplication cut-off units are set inside the calculation unit to reduce accuracy loss. At the same time, data transmission between non-adjacent units is achieved using two-level Internet networks.
It reduces the Internet complexity of the computing array, reduces data transmission and interconnected resources, improves computing efficiency, and minimizes accuracy losses while ensuring that the data flow bit width remains unchanged.
Smart Images

Figure CN120045513B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of wireless communication, and particularly relates to a computing unit and an FFT processor based on a coarse-grained reconfigurable architecture. Background Art
[0002] In recent years, the coarse-grained reconfigurable architecture (CGRA) has become a research hotspot in the industry because it well compromises the flexibility of GPP and the high performance of ASIC. In order to accelerate the fast Fourier transform (FFT), a core algorithm in the field of digital signal processing, a dedicated FFT processor needs to be designed. Since different-point FFTs are often used in the field of digital signal processing, and the coarse-grained reconfigurable architecture can well adapt to the processing of FFTs with different points and different radices, the coarse-grained reconfigurable architecture has become the mainstream implementation solution in the field of digital signal processing. Summary of the Invention
[0003] The embodiments of this application provide a computing unit and an FFT processor based on a coarse-grained reconfigurable architecture, which can effectively reduce the hardware overhead and the complexity of the interconnection network of the computing array.
[0004] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0005] According to the first aspect of the embodiments of this application, a computing unit based on a coarse-grained reconfigurable architecture is provided. The computing unit is used to perform one butterfly operation, and specifically includes: an input stage for providing the calculation data and rotation factors required for the operation; an addition stage for performing the first addition and subtraction operations on the calculation data; a multiplication stage for selecting, according to the rotation factor, to multiply the result of the first addition and subtraction operation by the rotation factor, perform the second addition and subtraction operation, or directly output; and an output stage for outputting the operation result of the multiplication stage. Through the general PE unit applicable to the butterfly operation proposed in this application, one butterfly operation only needs to be completed by using one computing unit, greatly reducing the data transmission and interconnection resources between computing units.
[0006] In an embodiment of the present application, the addition stage includes a complex number decomposition unit, a first data distribution unit, a first addition unit, a second addition unit, a first subtraction unit, and a second subtraction unit. Among them, the input end of the complex number decomposition unit is connected to the input stage, and the output ends of the complex number decomposition unit are respectively connected to the input end of the first data distribution unit and the multiplication stage. The complex number decomposition unit is configured to decompose the rotation factor and the calculation data provided by the input stage into real part data and imaginary part data; wherein, the calculation data includes first input data and second input data; the output ends of the first data distribution unit are respectively connected to the input ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit. The first data distribution unit is configured to respectively input the real part data of the first input data and the second input data into the first addition unit and the first subtraction unit, and input the imaginary part data of the first input data and the second input data into the second addition unit and the second subtraction unit; the output ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit are connected to the multiplication stage.
[0007] In an embodiment of the present application, the addition stage further includes a first data truncator, which is arranged between the output ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit and the multiplication stage, and is used to respectively intercept the required bit-width results from the operation results of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit.
[0008] In an embodiment of the present application, the first data truncator intercepts the operation results in a saturation truncation manner.
[0009] In one embodiment of the present application, the multiplication stage includes a second data distribution unit, a multiplier unit, a second data truncator, a third addition unit, a third subtraction unit, and a complex number merging unit. Among them, the input end of the second data distribution unit is respectively connected to the output end of the complex number decomposition unit, the output ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit. The output end of the second data distribution unit is respectively connected to the input ends of the multiplier unit and the complex number merging unit. The second data distribution unit is configured to select and input the real part data and imaginary part data of the rotation factor, and the operation results of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit into the multiplier unit or the complex number merging unit according to the value of the rotation factor. The output end of the multiplier unit is connected to the input end of the second data truncator, and the output end of the second data truncator is respectively connected to the input ends of the third addition unit and the third subtraction unit. The output ends of the third addition unit and the third subtraction unit are respectively connected to the input end of the complex number merging unit. The second truncator data is used to truncate the integer part of the operation result of the multiplier unit. The complex number merging unit is configured to merge the received data into a complex number and output it to the output stage through the output end.
[0010] In one embodiment of the present application, the multiplication stage further includes a third truncator, which is arranged between the third addition unit and the third subtraction unit and the complex number merging unit, and is used to truncate the operation results of the third addition unit and the third subtraction unit and then input them into the complex number merging unit.
[0011] In one embodiment of the present application, the input stage includes a rotation factor register and a local data register. The rotation factor register is used to provide a rotation factor, and the local data register is used to provide calculation data. The data in the local data register includes the calculation result of the current calculation unit and the data provided by the shared memory.
[0012] In one embodiment of the present application, the output stage includes a local data register shared with the input stage, which is used to store the operation results of the multiplication stage.
[0013] In one embodiment of the present application, the local data register includes a first sub-interval and a second sub-interval, which respectively provide a first input data and a second input data for the input stage, and respectively store a first output data and a second output data output by the output stage.
[0014] According to the second aspect of the embodiments of the present application, there is provided an FFT processor based on a coarse-grained reconfigurable architecture. The FFT processor includes: an array of computing units, including a plurality of homogeneous computing units as described in the first aspect, for performing FFT calculation tasks in digital signal processing.
[0015] In one embodiment of the present application, the FFT processor includes: a shared memory for caching intermediate data and calculation results calculated by each calculation unit; each calculation unit in the calculation unit array is connected to an adjacent calculation unit for data transmission, and each non-adjacent calculation unit in the calculation unit array performs data transmission through the shared memory.
[0016] In one embodiment of the present application, the FFT processor includes: a main controller for generating configuration information according to the FFT calculation task; the configuration information is used to complete the reconstruction of the calculation unit array and the shared memory.
[0017] In one embodiment of the present application, the FFT processor further includes: a configuration memory for storing the configuration information generated by the main controller, including the configuration information of the calculation unit array, the configuration information of the shared memory, and the configuration information of the interconnection network.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic diagram of a calculation unit in an embodiment of the present application.
[0021] Figure 2 It is a schematic diagram of an FFT processor in an embodiment of the present application.
[0022] Figure 3 It is a schematic diagram of the network connection relationship between the calculation unit array and the shared memory of the present application.
[0023] Reference Numerals:
[0024] 100 - input stage, 110 - rotation factor register, 120 - local data register;
[0025] 200 - addition stage, 210 - complex number decomposition unit, 220 - first data distributor, 230 - first addition unit, 240 - second addition unit, 250 - first subtraction unit, 260 - second subtraction unit, 270 - first data truncator; 221 - first data distributor, 222 - second data distributor;
[0026] 300 - Multiplication level, 310 - Second data distribution unit, 320 - Multiplier unit, 330 - Second data truncator, 340 - Third adder unit, 350 - Third subtractor unit, 360 - Complex number merging unit, 370 - Third data truncator, 321 - First multiplier, 322 - Second multiplier, 323 - Third multiplier, 324 - Fourth multiplier, 361 - First complex number merger, 362 - Second complex number merger;
[0027] 400 - Output level;
[0028] 500 - Shared memory. Detailed implementation
[0029] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0030] The terms "first" and "second" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0031] In the current FFT calculation scheme, multiple calculation units are required to perform a single radix - 2 / radix - 4 butterfly operation, and there is a large amount of overhead between the individual calculation units. Based on this, the embodiments of the present application propose a calculation unit based on a coarse - grained reconfigurable architecture. With one calculation unit, a single radix - 2 butterfly operation can be achieved, which can effectively reduce the interconnection between PEs and reduce the complexity of the interconnection network of the calculation array.
[0032] Please refer to Figure 1, the computing unit mainly includes four parts: an input stage 100, an addition stage 200, a multiplication stage 300, and an output stage 400. Through these four parts, a computing unit can complete a radix-2 butterfly operation and multiply it by the corresponding rotation factor. By selecting multiple interconnected computing units to form a computing unit array according to requirements, FFT calculations can be realized according to the computing tasks.
[0033] Specifically, the input stage 100 of the computing unit serves as an entry for providing the computing data and rotation factors required for operations. It mainly includes a rotation factor register 110 and a local data register 120. Among them, the rotation factor register 110 is used to store the rotation factors required during the computing process of the computing unit. During the configuration phase, all the rotation factors needed in all phases can be stored in the rotation factor register 110. In one embodiment, the depth of the rotation factor register 110 is designed to be 128. When the number of points of the FFT computing task executed by the computing unit does not exceed 1024 points, all the rotation factors can be stored in the rotation factor register before the execution of the computing task. When the number of points of the executed FFT computing task exceeds 1024 points, new rotation factors need to be stored in the rotation factor register 110 again at the end of a certain phase.
[0034] The local data register 120 in each computing unit only supports access by its own computing unit. The data sources of the local data register 120 include the computing data results of the current computing unit and the data read from the shared memory 500. Among them, the shared memory 500 stores the intermediate data and computing results of other computing units.
[0035] To avoid memory access conflicts of the computing unit to the local data register 120, in the embodiments of the present application, the local data register 120 of each computing unit is divided into a first sub-interval and a second sub-interval, which respectively provide the first input data and the second input data for the input stage 100, and store the first output data and the second output data output by the output stage 400 respectively. Specifically, the width of each sub-interval is 32 bits, which is used to store 32-bit complex numbers. Among them, the upper 16 bits are the real part of the data, and the lower 16 bits are the imaginary part of the data; the depth of each sub-interval is 64, which can store 64 data. The depth of the local data register 120 determines the size of the number of points for executing the FFT. In this specification, the depth of the local data register 120 is taken as 64 for illustration. In practical applications, it can be adjusted according to requirements.
[0036] The addition stage 200 of the computing unit, which is the part for performing the first addition and subtraction operations on the data provided by the input stage, mainly includes a complex number decomposition unit 210, a first data distributor 220, a first adder 230, a second adder 240, a first subtractor 250, and a second subtractor 260. Among them, the input end of the complex number decomposition unit 210 is connected to the input stage, and the output ends of the complex number decomposition unit 210 are respectively connected to the input end of the first data distributor 220 and the multiplication stage 300; the output ends of the first data distributor 220 are respectively connected to the input ends of the first adder 230, the second adder 240, the first subtractor 250, and the second subtractor 260, and the output ends of the first adder 230, the second adder 240, the first subtractor 250, and the second subtractor 260 are connected to the multiplication stage 300.
[0037] Specifically, the complex number decomposition unit 210 is mainly used to divide the input 32-bit complex number into a 16-bit real part data and an imaginary part data. That is, through the complex number decomposition unit 210, the rotation factor, the first input data and the second input data included in the calculation data are respectively decomposed into two parts: a real part In1 and an imaginary part In2.
[0038] The first data distributor 220 includes two data distributors with the same structure, namely a first data distributor 221 and a second data distributor 222, which respectively process the real part and the imaginary part of the first input data and the second input data, that is, realize the real part data distribution and the imaginary part data distribution of the two input data. In one embodiment, the data distributor includes a first input end, a second input end, a first output end, a second output end, a third output end, and a fourth output end, where the first input end is respectively connected to the first output end and the third output end, and the second input end is respectively connected to the second output end and the fourth output end. When receiving data, the real parts of the two input data are respectively input to the first input end and the second input end of a data distributor, and the imaginary parts of the two input data are respectively input to the first input end and the second input end of another data distributor.
[0039] After the first data distribution unit 220, there are a first addition unit 230, a second addition unit 240, a first subtraction unit 250, and a second subtraction unit 260, which are respectively used to perform 16-bit signed addition or subtraction operations on the data output by the first data distribution unit 220. In the embodiment of the present application, the input ends of the first addition unit 230 are respectively connected to the first output end and the second output end of the first data distributor 221 to complete the real part addition operation of the first input data and the second input data; the second addition unit 240 is respectively connected to the first output end and the second output end of the second data distributor 222 to complete the imaginary part addition operation of the first input data and the second input data; the first subtraction unit 250 is respectively connected to the third output end and the fourth output end of the first data distributor 221 to complete the real part subtraction operation of the first input data and the second input data; the second subtraction unit 260 is respectively connected to the third output end and the fourth output end of the second data distributor 222 to complete the imaginary part subtraction operation of the first input data and the second input data.
[0040] Since the addition or subtraction of two data may cause the bit width of the calculation result to overflow during the addition or subtraction process, in the embodiment of the present application, a first data truncator 270 is provided at the output ends of the first addition unit 230, the second addition unit 240, the first subtraction unit 250, and the second subtraction unit 260, which is used to truncate one bit of the 17-bit input data and output 16-bit data. The data passing through the first data truncator 270 can be input to the multiplication stage 300 for subsequent operations. Specifically, the first data truncator 270 uses the method of saturation truncation to reduce the error caused by truncation. The saturation truncation process includes: if the calculation result exceeds the maximum value of the data that can be stored in the required data format, then the maximum value is used to represent this data; if the calculation result exceeds the minimum value of the data that can be stored in the required data format, then the minimum value is used to represent this data.
[0041] The multiplication stage 300 of the computing unit selects to multiply the result of the first addition and subtraction operation by the rotation factor, perform the second addition and subtraction operation, or directly output according to the rotation factor. Specifically, the multiplication stage 300 mainly includes a second data distribution unit 310, a multiplier unit 320, a second data truncator 330, a third addition unit 340, a third subtraction unit 350, and a complex number merging unit 360. Among them, the input end of the second data distribution unit 310 is respectively connected to the output end of the complex number decomposition unit 210, the output ends of the first addition unit 230, the second addition unit 240, the first subtraction unit 250, and the second subtraction unit 260 (after connecting to the first data truncator 270, it is connected to the output end of the first data truncator 270). The output end of the second data distribution unit 310 is respectively connected to the input ends of the multiplier unit 320 and the complex number merging unit 360. The output end of the multiplier unit 320 is connected to the input end of the second data truncator 330. The output end of the second data truncator 330 is respectively connected to the input ends of the third addition unit 340 and the third subtraction unit 350. The output ends of the third addition unit 340 and the third subtraction unit 350 are respectively connected to the input ends of the complex number merging unit 360.
[0042] Specifically, the second data distribution unit 310 is used to directly output the data that does not need to be multiplied by the rotation factor to the complex number merging unit 360, and distribute other data to the multiplier unit 320 for calculation. It should be noted that the distribution path of the second data distribution unit 310 is controlled by the input rotation factor. When the real part data and the imaginary part data of the rotation factor are 0, the operation results of the first addition unit 230, the second addition unit 240, the first subtraction unit 250, and the second subtraction unit 260 are directly distributed to the complex number merging unit 360. When the real part data and the imaginary part data of the rotation factor are not 0, the real part data and the imaginary part data of the rotation factor, and the operation results of the first addition unit 230, the second addition unit 240, the first subtraction unit 250, and the second subtraction unit 260 are distributed to the multiplier unit 320.
[0043] Taking the butterfly operation in base 2 as an example, the second data distribution unit 310 includes input terminals i1, i2, i3, i4, i5, i6 and output terminals o1, o2, o3, o4, o5, o6, o7, o8, o9, o10, o11, o12. Among them, the input terminals i1 and i2 respectively receive the real part and the imaginary part obtained by the factorization of the rotation factor, which serve as both input data and control signals at the same time. The input terminals i3 and i4 respectively receive the real part data obtained by truncating the output of the first adder unit 230 and the first subtractor unit 250 by the first data truncator 270. The input terminals i5 and i6 respectively receive the imaginary part data obtained by truncating the output of the second adder unit 240 and the second subtractor unit 260 by the first data truncator 270. When performing data distribution, the data of the input terminals i3 and i4 are respectively distributed to the output terminals o1 and o2, and the outputs of the remaining ports are controlled by the input terminals i1 and i2. Specifically, when the input terminal i1 is 0 and the input i2 is 0, the data of the input terminal i5 is distributed to the output terminal o3, and the data of the input terminal i6 is distributed to the output terminal o4, and the remaining output terminals do not output data. When the input terminal i1 or the input terminal i2 is not 0, the data of the input terminal i1 is distributed to the output terminals o5 and o9, the data of the input terminal i2 is distributed to o7 and o11, the data of the input terminal i5 is distributed to o6 and o10, and the data of the input terminal i6 is distributed to o8 and o12. In the second data distribution unit 310, the data of the output terminals o1 to o4 are directly input to the complex number merging unit 360, and the data of the output terminals o5 to o12 are output to the multiplier unit 320 to perform a 16-bit signed multiplication operation.
[0044] In the embodiment of the present application, the multiplier unit 320 includes a first multiplier 321, a second multiplier 322, a third multiplier 323 and a fourth multiplier 324. Among them, the first multiplier 321 receives the data of the output terminals o5 and o6 of the second data distribution unit 310, the second multiplier 322 receives the data of the output terminals o7 and o8 of the second data distribution unit 310, the third multiplier 323 receives the data of the output terminals o9 and o10 of the second data distribution unit 310, and the fourth multiplier 324 receives the data of the output terminals o11 and o12 of the second data distribution unit 310. Since in the FFT calculation process, one of the operands of the multiplication operation is the rotation factor, by the definition of the rotation factor:
[0045]
[0046] It can be seen that both the real part and the imaginary part of the rotation factor are less than 1. Therefore, during the multiplication operation, no data overflow occurs in the integer part of the calculation result, and no truncation operation is required for the integer part of the multiplication calculation result. The multiplication of two data with 7 decimal places will cause the decimal width of the calculation result to reach a maximum of 14 bits, and a maximum of 7 bits of the decimal part need to be truncated. Therefore, in the embodiment of the present application, a second data truncator 330 is added after the multiplier unit 320. The second data truncator 330 uses the method of rounding truncation to truncate the overflowing decimal places to reduce the error caused by truncation. The rounding truncation process is as follows: If the highest bit of the decimal place to be truncated is 1, directly discard the decimal place to be truncated, and add 1 to the result after truncation; if the highest bit of the decimal place to be truncated is 0, directly discard the decimal place to be truncated, and add 0 to the result after truncation.
[0047] The data truncated by the second data truncator 330 still needs to be input into the third addition unit 340 and the third subtraction unit 350. Among them, the third subtraction unit 350 receives the data truncated by the second data truncator 330 output by the first multiplier 321 and the second multiplier 322, and completes the subtraction operation; the third addition unit 340 receives the data truncated by the second data truncator 330 output by the third multiplier 323 and the fourth multiplier 324, and completes the addition operation.
[0048] In the multiplication stage of the embodiment of the present application, a third data truncator 370 is also provided, which is used to truncate the output data of the third addition unit 340 and the third subtraction unit 350. Similar to the addition stage 200, the third data truncator 370 uses the method of saturation truncation to reduce the error caused by truncation. Through the first data truncator 270, the second data truncator 330 and the third data truncator 370, the precision loss is minimized as much as possible while ensuring that the data stream width remains unchanged.
[0049] The complex number merging unit 360 mainly merges the input data into a complex number for output. The calculation unit implements the butterfly operation of two complex numbers, and the final output is also two complex numbers. Therefore, in the embodiment of the present application, the complex number merging unit 360 includes a first complex number merger 361 and a second complex number merger 362. The first complex number merger 361 receives the data at the output terminals o1 and o2 of the second data distribution unit 310, and forms the first output data as the high 16 bits and the low 16 bits of the merger result respectively; the second complex number merger 362 receives the data at the output terminals o3 and o4 of the second data distribution unit 310 or the output data of the third data truncator 370 according to the value of the rotation factor, and merges them to form the second output data, which is finally sent to the output stage 400.
[0050] The output stage 400 of the calculation unit is used to store the calculation result of the multiplication stage 300 in the local data register 120.
[0051] Based on the general computing unit applicable to butterfly operations proposed in this application, a radix-2 butterfly operation only needs to use one computing unit to complete, greatly reducing the data transmission and interconnection resources when using an array of computing units. And an addition / subtraction truncation unit and a multiplication truncation unit are arranged inside the computing unit, minimizing the precision loss as much as possible while ensuring the unchanged bit width of the data stream.
[0052] Please refer to Figure 2 , an FFT processor is also proposed in an embodiment of this application. The processor includes an array of computing units and a shared memory.
[0053] Specifically, the array of computing units includes several of the aforementioned homogeneous computing units, which are used to implement digital signal processing computing tasks. Each computing unit is reconfigurable and applicable to the radix-2 or radix-4 fast Fourier transform algorithm; the shared memory is used to cache the intermediate data and computing results of each computing unit. The interconnection network of the array of computing units is divided into two levels of interconnection. Each computing unit in the array of computing units is connected to adjacent computing units for data transmission, and each non-adjacent computing unit in the array of computing units performs data transmission through the shared memory. In this application, at the level of the array of computing units, a two-level interconnection network is used to implement data transmission in the FFT algorithm process, and data transmission between non-adjacent units is realized through the shared memory, reducing the complexity of the interconnection network while satisfying the data dependence relationship of the FFT algorithm.
[0054] In an embodiment of this application, the array of computing units includes M rows and M columns, a total of M M computing units; each row in the array of computing units is divided into two groups of computing units, denoted as the p-th row group 1 and the p-th row group 2, using the grouping method of taking even terms as one group and odd terms as one group. Each column in the array of computing units is divided into two groups of computing units, denoted as the q-th column group 1 and the q-th column group 2, using the grouping method of taking even terms as one group and odd terms as one group, where p ≤ M, q ≤ M, and M ≥ 2. The shared memory includes 2M sub-memories, namely sub-memory b00 to sub-memory b2M-1. Each sub-memory in the shared memory completes the interconnection with the computing units of a preset row, preset group, and preset column, preset group. Specifically, the connection relationship between each sub-memory and each computing unit in the array of computing units is as follows:
[0055] Sub-memory b00 is interconnected with the row group 1 of the first row and the column group 1 of the first column in the array of computing units;
[0056] Sub-memory b01 is interconnected with the row group 2 of the first row and the column group 1 of the second column in the array of computing units;
[0057] Sub-memory b02 is interconnected with the row group 1 of the second row and the column group 2 of the first column in the array of computing units;
[0058] The sub-memory b03 is interconnected with row group 2 of the second row and column group 2 of the second column in the computing unit array;
[0059] until
[0060] The sub-memory b2M - 4 is interconnected with row group 1 of the (M - 1)th row and column group 1 of the (M - 1)th column in the computing unit array;
[0061] The sub-memory b2M - 3 is interconnected with row group 2 of the (M - 1)th row and column group 1 of the Mth column in the computing unit array;
[0062] The sub-memory b2M - 2 is interconnected with row group 1 of the Mth row and column group 2 of the (M - 1)th column in the computing unit array;
[0063] The sub-memory b2M - 1 is interconnected with row group 2 of the Mth row and column group 2 of the Mth column in the computing unit array.
[0064] Please refer to Figure 3 , taking the computing unit array with 64 computing units as an example for illustration below. The first-level interconnection of the computing unit array is that each computing unit is directly interconnected with its adjacent PE unit, and the second-level interconnection is the interconnection between the PE and the shared memory to achieve data transmission between non-adjacent computing units. At this time, the shared memory consists of 16 sub-memories, namely b00~b15. The four PEs in the odd-numbered columns of each row are classified into row group 1, and the four PEs in the even-numbered columns are classified into row group 2; the four PEs in the odd-numbered rows of each column are classified into column group 1, and the four PEs in the even-numbered rows are classified into column group 2. The situations of the sub-memories b00~b15 being interconnected with specific PEs are as follows: b00 is interconnected with row group 1 of the first row and column group 1 of the first column; b01 is interconnected with row group 2 of the first row and column group 1 of the second column; b02 is interconnected with row group 1 of the second row and column group 2 of the first column; b03 is interconnected with row group 2 of the second row and column group 2 of the second column; b04 is interconnected with row group 1 of the third row and column group 1 of the third column; b05 is interconnected with row group 2 of the third row and column group 1 of the fourth column; b06 is interconnected with row group 1 of the fourth row and column group 2 of the third column; b07 is interconnected with row group 2 of the fourth row and column group 2 of the fourth column; b08 is interconnected with row group 1 of the fifth row and column group 1 of the fifth column; b09 is interconnected with row group 2 of the fifth row and column group 1 of the sixth column; b10 is interconnected with row group 1 of the sixth row and column group 2 of the fifth column; b11 is interconnected with row group 2 of the sixth row and column group 2 of the sixth column; b12 is interconnected with row group 1 of the seventh row and column group 1 of the seventh column; b13 is interconnected with row group 2 of the seventh row and column group 1 of the eighth column; b14 is interconnected with row group 1 of the eighth row and column group 2 of the seventh column; b15 is interconnected with row group 2 of the eighth row and column group 2 of the eighth column.
[0065] For a computing unit array containing 64 computing units, in one stage, one computing unit receives two input data to complete a two-point butterfly operation. Then, the computing unit array can process 128 data in one stage and store 128 intermediate computing results in the local data register. The depth of the local data register being 64 determines that the FFT processor can compute an FFT algorithm with a maximum number of points of 8K. In the embodiments of the present application, the data is cached in each storage unit in the fixed-point format, and the bit width of all storage units is 32 bits. Among them, the upper 16 bits represent the real part of the data, the highest bit is the sign bit of the real part, the 2nd to 9th bits are the integer part of the real part, and the 10th to 16th bits are the fractional part of the real part; the lower 16 bits represent the imaginary part of the data, the 17th bit is the sign bit of the imaginary part, the 18th to 24th bits are the integer part of the imaginary part, and the 25th to 32nd bits are the fractional part of the imaginary part.
[0066] For the shared memory interconnected with the computing unit array, the width of the sub-memory is designed to be 32 bits and is used to store 32-bit complex numbers, where the upper 16 bits are the real part of the data and the lower 16 bits are the imaginary part of the data. Due to the memory access conflict problem, multiple computing units cannot access the same address of the same shared memory simultaneously at the same time. In the execution of the fast Fourier transform algorithm, there will be a situation where four computing units pass through the same shared memory for data transmission in a certain stage. In the embodiments of the present application, the depth of each sub-memory of the shared memory is designed to be 4, and different computing units can read or write to four addresses in the sub-memory at the same time. However, it should be noted that it is not possible to read and write to the same address of the sub-memory at the same time.
[0067] It should be added that in each computing stage, the computing results will only exist in the local data register. After one stage of computing is completed and before entering the next stage, it is necessary to refresh the data in the shared memory and the local memory, that is, to obtain the data in the local data register of other computing units through the shared memory. Further, the FFT processing also includes a main controller and a configuration memory. Among them, the main controller is responsible for starting the processor, compiling the FFT computing task to generate corresponding configuration information, and storing it in the configuration memory; at the same time, controlling the configuration memory to transmit the configuration information to the computing unit array and the shared memory, and completing the reconstruction of the PEA and the shared memory through the configuration information. The configuration memory is used to store the configuration information generated by the main controller, including the configuration information of the computing units, the configuration information of the shared memory, and the configuration information of the interconnection network.
[0068] The FFT processor proposed in the embodiments of the present application uses a two-level interconnection network to implement data transmission in the FFT algorithm process. The data transmission between non-adjacent units is realized through a shared memory, which reduces the complexity of the interconnection network while satisfying the data dependence relationship of the FFT algorithm. Based on the flexible configuration method of the reconfigurable coarse-grained architecture and the reconfigurable performance of the arithmetic unit array, the FFT processing tasks with different numbers of points and radices can be handled more flexibly and richly.
[0069] To better verify the FFT processor proposed in the present application, taking the radix-2 frequency-division 1024-point fast Fourier transform algorithm (FFT) as an example for further illustration, the FFT processor calculates the radix-2 frequency-division 1024-point FFT in a total of 10 stages. The distance difference between the input data of the butterfly operation in the nth stage is 1024 / 2 n , and the rotation factor used is , where m is from 0 to 1024 / 2 n . For a computing unit array containing 64 computing units (PEs), the operand of each PE is 2, and 128 data can be input simultaneously for calculation in one computing cycle. Each stage needs to be divided into eight inputs to calculate all the data. The data input positions of the 10 stages are shown in Table 1 below:
[0070] Table 1 Data input table for each stage
[0071]
[0072] Since the source of the input data of the computing unit is different in each stage, the FFT processor has cached the 1024-point initial data in the local data registers of 64 PEs in a frequency-division manner during the configuration stage. Each PE stores 16 data. Since each PE has two sub-intervals, each sub-interval stores 8 data. The data in the corresponding addresses of the two sub-intervals are input into one PE for butterfly operation. The source of the input data of the 10-stage PEs is shown in Table 2 below:
[0073] Table 2 Source table of input data for each stage
[0074]
[0075] It should be noted that the data transmission of the non-adjacent PEs shown in Table 2 is realized through a shared memory.
[0076] Further, when implementing FFT algorithms with other numbers of points, the interconnection resources will not increase additionally. When performing FFT with a number of points greater than 1K, the number of input data times in each stage will increase. Taking the 8K-point FFT as an example, the calculation array can calculate 128 data points at a time. Therefore, the input data needs to be divided into 64 groups in each stage, and each group is sequentially input into the calculation array. The complete calculation process includes 13 stages. In the first 7 stages, the PE only accesses the intermediate data in the local data register. The data source in the last 6 stages is the same as that in the 1024-point FFT. Therefore, no additional interconnection resource overhead will be generated when calculating FFT with other numbers of points.
[0077] Further, the FFT processor proposed in the embodiment of the present application is not only applicable to the radix-2 frequency division 1024-point fast Fourier transform algorithm, but also applicable to the radix-4 frequency division 1024-point fast Fourier transform algorithm. When the calculation task is the radix-4 frequency division 1024-point fast Fourier transform algorithm, during the configuration stage, generating a configuration to change the data distribution method of the second data distribution unit inside the calculation unit can meet the data stream of the radix-4 butterfly operation.
[0078] At this time, the second data distribution unit performs data distribution operations in two times, that is, two data distribution controls are performed according to the values of the twiddle factors. The first twiddle factor controls the data output of the input ends i3 and i4, and the second twiddle factor controls the data output of the input ends i5 and i6. The data input from the input ends i3 and i4 is no longer directly distributed to the output ends o1 and o2 respectively, but the distribution situation is determined by the twiddle factor as a control signal. Specifically:
[0079] During the first distribution, the twiddle factors i11 and i12 are used as input data and control signals, and the distribution process is as follows:
[0080] When the input end i11 is 0 and the input end i12 is 0, the data of the input end i3 is directly distributed to the output end o1, and the data of the input end i6 is directly distributed to the output end o2. At the same time, no data is output from the other output ends.
[0081] When the input end i11 is not 0 or the input end i12 is not 0, the data of the input end i11 is distributed to the output ends o5 and o9, the data of the input end i3 is distributed to the output ends o6 and o10, the data of the input end i12 is distributed to the output ends o7 and o11, and the data of the input end i6 is distributed to the output ends o8 and o12.
[0082] During the second distribution, the twiddle factors i21 and i22 are used as input data and control signals, and the distribution process is as follows:
[0083] When the input terminal i21 is 0 and the input terminal i22 is 0, the data of the input terminal i5 is directly allocated to the output terminal o3, and the data of the input terminal i6 is directly allocated to the output terminal o4. At the same time, no data is output from the remaining output terminals.
[0084] When the input terminal i21 is not 0 or the input terminal i22 is 0, the data of the input terminal i21 is allocated to the output terminals o5 and o9, the data of the input terminal i5 is allocated to the output terminals o6 and o10, the data of the input terminal i22 is allocated to the output terminals o7 and o11, and the data of the input terminal i6 is allocated to the output terminals o8 and o12.
[0085] At this time, the multiplication units in the computing unit are multiplexed. After the four multipliers complete the multiplication of the first rotation factor and the intermediate result, they immediately accept the second rotation factor to perform the multiplication operation with another intermediate result, sacrificing the computing cycle to improve the utilization rate of the computing unit and reduce the consumption of hardware resources.
[0086] After the FFT processor completes the above reconstruction of the computing unit, the process of executing the radix-4 frequency division 1024-point fast Fourier transform algorithm is as follows:
[0087] The first stage: First, calculate sub-stage 1. First, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the a-th and the (a + 512)-th, the (a + 256)-th and the (a + 768)-th, where a = 0~255. Then calculate sub-stage 2. Similarly, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the a-th and the (a + 256)-th, the (a + 512)-th and the (a + 768)-th, where a = 0~255.
[0088] The second stage: First, calculate sub-stage 1. First, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the b-th and the (b + 128)-th, the (b + 64)-th and the (b + 192)-th, where b = 256 k~256 k + 63, k = 0~3. Then calculate sub-stage 2. Similarly, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the b-th and the (b + 64)-th, the (b + 128)-th and the (b + 192)-th, where b = 0~255.
[0089] The third stage: First, calculate sub-stage 1. First, divide the 1024-point data into eight inputs, with 128 data inputs each time. The data input to two adjacent PE units are the c-th and the (c + 32)-th, the (c + 16)-th and the (c + 48)-th, where c = 128 k ~ 128 k + 79, where k = 0 to 7. Then calculate sub - stage 2. Similarly, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the c - th and the (c + 16)-th, the (c + 32)-th and the (c + 48)-th respectively, where c = 128 k ~ 128 k + 79, k = 0 to 7.
[0090] Fourth stage: First, calculate sub - stage 1. First, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the d - th and the (d + 8)-th, the (d + 4)-th and the (d + 12)-th respectively, where d = 16 k ~ 16 k + 3, where k = 0 to 63. Then calculate sub - stage 2. Similarly, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the d - th and the (d + 4)-th, the (d + 8)-th and the (d + 12)-th respectively, where d = 16 k ~ 16 k + 3, k = 0 to 63.
[0091] Fifth stage: First, calculate sub - stage 1. First, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the e - th and the (e + 2)-th, the (e + 1)-th and the (e + 3)-th respectively, where e = 4 k, where k = 0 to 255. Then calculate sub - stage 2. Similarly, divide the 1024 - point data into eight inputs, with 128 data points input each time. The data input to two adjacent PE units are the e - th and the (e + 1)-th, the (e + 2)-th and the (e + 3)-th respectively, where e = 4 k, k = 0 to 255.
[0092] Through the above five stages, the calculation task of the radix - 4 frequency - division 1024 - point fast Fourier transform algorithm can be completed.
[0093] This application designs a general computing unit suitable for butterfly operations, enabling a single - PE implementation of a single - radix - 2 butterfly operation, significantly reducing data transfer and interconnection resources between PEs. An addition / subtraction truncation unit and a multiplication truncation unit are set inside the computing unit, thus ensuring minimal precision loss while keeping the data - flow bit - width unchanged. In the FFT processor applying the computing unit, a two - level interconnection network is used to implement data transfer during the FFT algorithm. Data transfer between non - adjacent units is achieved through a shared memory, reducing the complexity of the interconnection network while satisfying the data - dependency relationship of the FFT algorithm.
[0094] The features, structures, or characteristics described in this application can be combined in any suitable manner in one or more embodiments. Unless otherwise specified, the terms "connected", "coupled", and "connected to" are used to specify an electrical connection between circuit elements that can be direct or can be via one or more other elements. In contrast, when an element is said to be "directly connected to" or "directly coupled to" another element, there is no intermediate element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances; the accompanying drawings in the embodiments are used to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in a variety of different configurations.
[0095] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A computing unit based on a coarse-grained reconfigurable architecture, characterized in that, The computing unit is used to perform a butterfly operation, specifically including: An input stage for providing the calculation data and rotation factors required for the operation; the calculation data includes a first input data and a second input data; the input stage includes a rotation factor register and a local data register, the rotation factor register is used to provide the rotation factor, and the local data register is used to provide the calculation data; the data in the local data register includes the calculation result of the current computing unit and the data provided by the shared memory; An addition stage for performing a first addition and subtraction operation on the calculation data; A multiplication stage, at least including a second data distribution unit, the distribution path of the second data distribution unit is controlled by the input rotation factor. When the real part data and the imaginary part data of the rotation factor are 0, the result of the first addition and subtraction operation is directly output; when the real part data or the imaginary part data of the rotation factor is not 0, the result of the first addition and subtraction operation is multiplied by the rotation factor and then a second addition and subtraction operation is performed; the second data distribution unit completes the conversion of the radix-4 FFT calculation task or the radix-2 FFT calculation task by reconstructing the data distribution path; and An output stage for outputting the operation result of the multiplication stage.
2. The computing unit according to claim 1, wherein The addition stage includes a complex number decomposition unit, a first data distribution unit, a first addition unit, a second addition unit, a first subtraction unit, and a second subtraction unit, where The input end of the complex number decomposition unit is connected to the input stage, and the output ends of the complex number decomposition unit are respectively connected to the input end of the first data distribution unit and the multiplication stage. The complex number decomposition unit is used to decompose the rotation factor and the calculation data provided by the input stage into real part data and imaginary part data; The output ends of the first data distribution unit are respectively connected to the input ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit. The first data distribution unit is used to input the real part data of the first input data and the second input data into the first addition unit and the first subtraction unit respectively, and input the imaginary part data of the first input data and the second input data into the second addition unit and the second subtraction unit; The output ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit are connected to the multiplication stage.
3. The computing unit according to claim 2, characterized in that The addition stage further includes a first data truncator, which is arranged between the output ends of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit and the multiplication stage, and is used to respectively intercept the required bit-width results from the operation results of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit.
4. The computing unit according to claim 3, wherein The first data truncator intercepts the operation result in a saturation truncation manner.
5. The computing unit according to claim 1 or 2, characterized in that, The multiplication stage includes a second data distribution unit, a multiplier unit, a second data truncator, a third addition unit, a third subtraction unit, and a complex number merging unit, where The input ends of the second data distribution unit are respectively connected to the output ends of the complex number decomposition unit, the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit. The output ends of the second data distribution unit are respectively connected to the input ends of the multiplier unit and the complex number merging unit. The second data distribution unit is configured to select, according to the value of the rotation factor, to input the real part data and the imaginary part data of the rotation factor, and the operation results of the first addition unit, the second addition unit, the first subtraction unit, and the second subtraction unit into the multiplier unit or the complex number merging unit; The output end of the multiplier unit is connected to the input end of the second data truncator. The output end of the second data truncator is respectively connected to the input ends of the third addition unit and the third subtraction unit. The output ends of the third addition unit and the third subtraction unit are respectively connected to the input end of the complex number merging unit. The second data truncator is used to truncate the integer part of the operation result of the multiplier unit. The complex number merging unit is configured to merge the received data into a complex number and output it to the output stage through the output end.
6. The computing unit according to claim 5, wherein The multiplier stage further includes a third data truncator, which is arranged between the third addition unit and the third subtraction unit and the complex number merging unit, and is configured to truncate the operation results of the third addition unit and the third subtraction unit and then input them into the complex number merging unit.
7. The computing unit according to claim 1, wherein The output stage includes a local data register shared with the input stage, which is used to store the operation results of the multiplier stage.
8. The computing unit according to claim 7, characterized in that, The local data register includes a first sub-interval and a second sub-interval, which respectively provide the first input data and the second input data for the input stage, and respectively store the first output data and the second output data output by the output stage.
9. An FFT processor based on a coarse-grained reconfigurable architecture, characterized in that, The FFT processor includes: An array of computing units, including a plurality of homogeneous computing units as described in any one of claims 1 to 8, which is used to execute the FFT computing task in digital signal processing; A shared memory, which is used to cache the intermediate data and computing results calculated by each computing unit; Each computing unit in the array of computing units is connected to an adjacent computing unit for data transmission, and each non-adjacent computing unit in the array of computing units performs data transmission through the shared memory; The computing unit array includes M computing units arranged in M rows and M columns; each row in the computing unit array is divided into two groups of computing units, with even-numbered terms in one group and odd-numbered terms in the other group, denoted as row group 1 and row group 2 of the p-th row; each column in the computing unit array is divided into two groups of computing units, with even-numbered terms in one group and odd-numbered terms in the other group, denoted as column group 1 and column group 2 of the q-th column, where p ≤ M, q ≤ M, and M ≥ 2; the shared memory includes 2M sub-memories, namely sub-memory b00 to sub-memory b2M-1; The computing unit array includes M computing units arranged in M rows and M columns; each row in the computing unit array is divided into two groups of computing units, with even-numbered terms in one group and odd-numbered terms in the other group, denoted as row group 1 and row group 2 of the p-th row; each column in the computing unit array is divided into two groups of computing units, with even-numbered terms in one group and odd-numbered terms in the other group, denoted as column group 1 and column group 2 of the q-th column, where p ≤ M, q ≤ M, and M ≥ 2; the shared memory includes 2M sub-memories, namely sub-memory b00 to sub-memory b2M-1; The connection relationship between each sub-memory and each computing unit in the array of computing units is as follows: The sub-memory b00 is interconnected with the row group 1 of the first row and the column group 1 of the first column in the array of computing units; The sub-memory b01 is interconnected with the row group 2 of the first row and the column group 1 of the second column in the array of computing units; The sub-memory b02 is interconnected with the row group 1 of the second row and the column group 2 of the first column in the array of computing units; The sub-memory b03 is interconnected with the row group 2 of the second row and the column group 2 of the second column in the array of computing units; Until, The sub-memory b2M-4 is interconnected with the row group 1 of the (M-1)th row and the column group 1 of the (M-1)th column in the array of computing units; The sub-memory b2M-3 is interconnected with the row group 2 of the (M-1)th row and the column group 1 of the Mth column in the array of computing units; The sub-memory b2M-2 is interconnected with the row group 1 of the Mth row and the column group 2 of the (M-1)th column in the array of computing units; The sub-memory b2M-1 is interconnected with row group 2 of the M-th row and column group 2 of the M-th column in the computing unit array.
10. The FFT processor according to claim 9, wherein The FFT processor includes: A main controller for generating configuration information according to the FFT calculation task; the configuration information is used to complete the reconstruction of the computing unit array and the shared memory.
11. The FFT processor according to claim 9, characterized in that, The FFT processor further includes: A configuration memory for storing the configuration information generated by the main controller, including the configuration information of the computing unit array, the configuration information of the shared memory, and the configuration information of the interconnection network.
Citation Information
Patent Citations
Configuration method for pipeline-architecture fixed-point FFT word length
CN103761074A
Processor based on coarse-grained reconfigurable architecture
CN111581148A
FFT processor
CN112231626A
Butterfly computing unit, butterfly computing unit array, reconfigurable array and chip
CN119829005A