Universal conflict-free memory access method used in storage-based parallel FFT (Fast Fourier Transform) computing structure
By designing a general conflict-free memory access method in a parallel FFT computing structure based on a single N-depth memory, the problems of large storage resource consumption and low computing efficiency in the prior art are solved, and efficient conflict-free memory access and flexible parallel read and write configuration are realized.
Patent Information
- Application Number
- CN202510231268.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, in the storage-based parallel FFT computing structure, it is difficult to implement a general conflict-free memory access method, resulting in large consumption of storage resources and limited computing efficiency.
A general conflict-free memory access method using a parallel FFT computing structure based on a single N-depth memory is proposed. By analyzing the arrangement of input data flows, a general conflict-free parallel memory access method is designed, which supports the variable cardinality of each level of butterfly calculation, and realizes conflict-free memory access when only one memory with depth N is used.
This method effectively saves storage resource consumption, enhances the application scope of parallel FFT structures, and can flexibly configure parallel reading and writing of FFT computing, which is suitable for the requirements of different hardware to realize resources and computing efficiency.
Smart Images

Figure CN120104931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of FFT calculation, and in particular to a general non-conflicting memory access method used in a storage-based parallel FFT calculation structure. Background Art
[0002] Fast Fourier Transform (FFT) is one of the most commonly used algorithms in the field of digital signal processing and is widely used in digital communications, radar, electronic measurement and other fields. Compared with software-implemented FFT calculations, hardware-implemented FFT calculations have higher computational efficiency and are therefore widely used. Among the hardware-implemented FFT calculation structures, the storage-based parallel FFT calculation structure is a commonly used structure. Whether this structure can effectively read and store the operands of the butterfly calculation is the key to affecting the overall computing efficiency.
[0003] Conflict-free memory access is an effective way to maintain parallel read and write efficiency. n Point FFT structure design, most current methods use 2 or 3 memories with a depth of N to separate the read data operation and the storage (write) data operation in space to achieve conflict-free memory access. In order to reduce the consumption of storage resources, there are also methods that use only one memory with a depth of N to achieve conflict-free memory access, but it is only applicable to butterfly calculations of a specific base and is not very universal. Summary of the invention
[0004] In order to solve the above problems existing in the prior art, the present invention further proposes a general non-conflicting memory access method used in a storage-based parallel FFT calculation structure.
[0005] The technical solution adopted by the present invention to solve the above problems is:
[0006] A general non-conflicting memory access method used in a storage-based parallel FFT computing structure according to the present invention comprises the following steps:
[0007] Step 1: Analysis of the arrangement of input data streams in P-parallel FFT calculation; Analysis N = 2 n Data flow arrangement of the front and back butterfly calculations in the point P parallel FFT calculation and the use of butterfly calculation cardinality 2 r The relationship between the two is analyzed, and the data stream with data reverse order correction is analyzed, and the data stream arrangement is effectively expressed using symbols.
[0008] Step 2: Propose a general conflict-free parallel memory access method for P-parallel FFT calculation structures based on a single N-depth memory; the conflict-free memory access method is used to implement the input data stream arrangement in Step 1, which will be implemented based on reading from and writing to a memory with a single N-depth and P read / write ports, and the read and write operations are located before and after each stage of butterfly calculation; the radix supported by each stage of butterfly calculation is 2 r (r can take any value from 1, 2,..., p_bits; P = 2 p _bits ).
[0009] Step 3: N = 2 n Number of points decomposition.
[0010] Furthermore, in Step 1, the analysis of the arrangement method of the input data stream in the P-parallel FFT calculation is as follows:[[]]
[0011] For the N-point P-parallel FFT calculation by frequency extraction, the input data arrangement to the butterfly calculation unit: The N-point data is divided into P paths. The data from the 0th to the N / P - 1st form the 0th data stream, the data from the N / P to the 2N / P - 1st form the 1st data stream, and so on to form P paths of parallel input data; each column of the data stream is sequentially input as the input data of the butterfly calculation unit for butterfly calculation.
[0012] Furthermore, in Step 2, the design of the general conflict-free memory access method specifically includes the following 4 stages:
[0013] Calculation stage 1, when remainderBits > p_bits and (remainderBits - r thisStage ) > p_bits:
[0014] Perform read and write operations according to Equation (6), where S now->nexts is the transition arrangement method between the write operation of writing to the memory after the completion of the current stage of butterfly calculation and the read operation of reading data from the memory at the start of the next stage of butterfly calculation, corresponding to the actual storage arrangement of the data in the memory, (where the symbol rotl(a, b) represents circularly shifting the data a to the left by b bits);
[0015]
[0016] Calculation stage 2, when remainderBits > p_bits and (remainderBits - r thisStage ) < p_bits:
[0017] This stage is essentially the process stage from calculation stage 1 to calculation stage 3 and is only executed once. If in calculation stage 1 (remainderBits - rthisStage ) is equal to p_bits, then this stage is not executed, but directly proceeds from stage 1 to stage 3; the permutation read / write expression for this stage is shown in Equation (7); compared to calculation stage 1, the shift amount of the circular left shift operation rotl is not r thisStage , but n - r sum -p_bits(=remainderBits - p_bits), where n represents the number of bits of N-point data (N = 2 n ), r sum represents that the radix-2 rsum butterfly calculation has been completed;
[0018]
[0019] Calculation stage 3, when remainderBits <= p_bits and (remainderBits - r thisStage ) < p_bits:
[0020] The data stream permutation situation does not change until the entire FFT calculation is completed, that is, when remainderBits == 0; the permutation read / write expression for this stage is shown in Equation (8),
[0021]
[0022] Calculation stage 4, complete the reverse order correction after the frequency extraction FFT calculation, that is, complete the bit reversal and output the FFT calculation result in sequential order; in the mixed butterfly calculation, 2 r butterfly calculations are used. From the perspective of the overall FFT calculation, the sequence numbers of the output data x(n) and the output FFT result X(k) are in a reverse order relationship by 1 bit, as shown in Equation (9);
[0023]
[0024] The read / write permutation relationship for the reverse order result to be output in sequential order is expressed using Equation (10),
[0025]
[0026] Furthermore, in step 3, N = 2 n The point decomposition suggestions are as follows:
[0027] For N = 2 n points P = 2 p_bits parallel data stream inputs, the butterfly calculation at this level can use a radix-2 r butterfly calculation (r can take values 1, 2,..., p_bits). For N = 2 nIn order to reduce the overall FFT calculation time, it is recommended to prioritize the base 2 p_bits Butterfly calculation. As calculated in formula (11), it is necessary to perform S-base 2 p_bits Butterfly calculation, the last level performs radix 2 rlastStage Butterfly computing;
[0028] n / p_bits=S···r lastStage (11)
[0029] For FFT calculations with different numbers of points, the final r lastStage is different, and may be a number among 0, 1, 2, ..., p_bits-1, so the adapted butterfly computing unit is required to support a base of 2 r Butterfly calculation, r can be 1, 2, …, p_bits.
[0030] The beneficial effects of the present invention are:
[0031] 1. This invention proposes a general non-conflicting memory access method for a parallel FFT computing structure based on a single N-depth memory, which is applicable to N=2 n Point P = 2 p_bits Parallel FFT calculation, and the cardinality of each butterfly calculation is variable (cardinality is 2 r , r<=p_bits).
[0032] 2. The present invention enhances the application scope of the storage-based parallel FFT structure. Under the requirements of different hardware implementation resources and computing efficiency, the parallel reading of FFT calculation can be flexibly configured, rather than being limited to butterfly calculations of a specific cardinality. The structure only uses a single memory with a depth of N, which effectively saves the consumption of storage resources. In scenarios with strict resource consumption requirements, this method plays an important role.
[0033] 3. The conflict-free memory access method proposed in the present invention supports a parallelism of P = 2 when only one memory with a depth of N is used. p_bits , the cardinality of each butterfly calculation is 2 r N = 2 n Point FFT calculation, where r can take values such as 1, 2..., p_bits according to the user's setting, without affecting the efficiency of P parallel reading and writing, reflecting a more "universal" feature. At the same time, the present invention also designs a method for the problem of reverse order correction that is rarely considered in most similar methods. The present invention is a universal and conflict-free memory access method used in a storage-based parallel FFT calculation structure with strong versatility and low resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1It is a 16-point radix-2 butterfly data flow graph of the present invention;
[0035] Figure 2 is the arrangement corresponding to the expression of formula (3) of the present invention;
[0036] Figure 3 It is a P parallel FFT calculation structure diagram based on a single N-depth memory of the present invention;
[0037] Figure 4 It is a data flow diagram of conflict-free memory access read and write operations of 32-point 4-parallel FFT calculation of the present invention;
[0038] Figure 5 It is a timeline diagram of the FFT calculation process of the present invention;
[0039] Figure 6 is the state machine of the single FFT controller of the present invention;
[0040] Figure 7 This is the data routing situation of the write operation when P=8 of the present invention;
[0041] Figure 8 It is a data selector with a label of the present invention;
[0042] Fig. 9 It is a block diagram of wrSel_x and wr_mem_addr_x generation during CTRL_trans_0 of the present invention;
[0043] Fig.10 It is a block diagram of wrSel_x and wr_mem_addr_x generation during CTRL_trans_x of the present invention;
[0044] Fig.11 It is a block diagram of wrSel_x and wr_mem_addr_x generation during CTRL_trans_last of the present invention;
[0045] Fig.12 It is a block diagram of the data routing of the read controller crossBarB of the present invention;
[0046] Fig.13 It is a block diagram of rdSel_x and rd_mem_addr_selx generation during CTRL_trans_x of the present invention;
[0047] Fig.14 It is a block diagram of rdSel_x and rd_mem_addr_selx generation during CTRL_trans_last of the present invention;
[0048] Fig.15This is a block diagram of rdSel_x and rd_mem_addr_selx generation during CTRL_fftout_trans of the present invention. DETAILED DESCRIPTION
[0049] The general non-conflicting memory access method used in a storage-based parallel FFT computing structure described in this embodiment includes the following steps:
[0050] Step 1: Step 1 is the analysis of the FFT calculation rules, which serves as the design basis for the conflict-free memory access method in step 2. In step 1, the arrangement of the input data stream in the P-parallel FFT calculation is analyzed. For the N-point P-parallel FFT calculation by frequency decimation (Decimation In Frequency, DIF), the input data input to the butterfly calculation unit is arranged as shown in Table 1: N-point data is divided into P paths, the 0th to N / P-1th data form the 0th data stream, and the N / Pth to 2N / P-1th data form the 1st data stream, thus forming P-path parallel input data. Each column of the data stream is sequentially input as the input data of the butterfly calculation unit for butterfly calculation.
[0051] Table 1 Arrangement rules of initial data
[0052] Butterfly input data stream No. 0 0 1 2 … N / P-1 Butterfly input data flow 1 N / P … 2N / P-1 Second butterfly input data stream 2N / P … 3N / P-1 … … … … … … The P-1 butterfly input data stream (P-1)N / P (P-1)N / P+1 (P-1)N / P+2 … N-1
[0053] Table 1 The N data to be calculated are in natural order 0, 1, ..., N-1 using binary correspondence b n-1 b n-2 …b 1 b 0 The initialization rule in the above table can be expressed by formula (1), where the first half of the vertical line “|” is b n-p_bits-1 b n-p_bits-2 …b 1 b 0 Indicates the order in a corresponding data stream; the second half b n-1 b n-2 …b n-p_bits Indicates which data stream it is in, with a total of p_bits bits.
[0054] S 1 =b n-p_bits-1 b n-p_bits-2 ...b 1 b 0 |b n-1 b n-2 ...b n-p_bits (1)
[0055] In the 16-point radix-2 butterfly calculation, each level performs radix-2 butterfly calculation. Figure 1As shown in the figure, it is a 16-point FFT data flow diagram, in which 8 radix-2 butterfly operations are performed at each level. The first level is 1 group of 8 radix-2 butterflies; the second level is 2 groups, 1 group of 4 radix-2 butterflies; the third level is 4 groups, 1 group of 2 radix-2 butterflies; the fourth level is 8 groups, 1 group of 1 radix-2 butterfly. A better understanding is that the entire operation is a 1-time radix-16 butterfly calculation; after the first-level radix-2 butterfly calculation, 2 radix-8 butterfly calculations are required, and the 0th to 7th data are the input of the first radix-8 butterfly; the 8th to 15th data are the input of the second radix-8 butterfly. However, the 8 parallel data inputs only perform radix-2 butterfly calculations, so the next level performs 4 radix-4 butterfly calculations.
[0056] The above process can be expressed by formula (2). The "group" part indicates how many groups the data at this level is divided into, and the "reminder" part indicates the butterfly operation that needs to be completed next. " / " indicates "empty". The logical form of formula (2) is a "shift" logic, which reflects the rule of FFT data sequence number. The shift logic only shifts 1 bit each time, because the base 2 is always used. 1 The butterfly operation corresponds to the highest 1 bit in the remainder.
[0057]
[0058] Using a more general expression such as formula (3), in this case, the base 2 rthisStage After the butterfly calculation, shift r thisStage bits.
[0059]
[0060] The number of bits on the group side, groupBits, is the sum of the previous several times with a base of 2. rs After the butterfly unit operation, the current stage can be expressed as formula (4), where remainderBits is the number of bits on the remainder side, indicating that there is a cardinality of 2 remainderBits Butterfly calculations to be completed:
[0061]
[0062] For the current parallelism, P = 2 p_bits The memory MEM, MEM_0 to MEM_P-1 sub-storage units at the same address store P data, occupying the high p_bits bits of the remainder part in equation (3). Figure 2 As shown, the spliced bit data is in order {b n-1 b n-2 …b n-rsum ,bn-rsum-p_bits-1 b n-rsum-p_bits-2 …b 1 b 0} is written to the on-chip address of MEM_x.
[0063] like Figure 2 As shown, the expression of formula (1) is used to obtain formula (5):
[0064]
[0065] Step 2: Propose a general conflict-free parallel memory access method using a P-parallel FFT calculation structure based on a single N-depth memory, as follows:
[0066] (1) P-parallel FFT calculation structure based on a single N-depth memory
[0067] The P-parallel FFT calculation structure based on a single N-depth memory is as follows Figure 3 As shown, it includes a data storage unit MEM, a storage control unit (write controller crossBarA and read controller crossBarB), a single FFT controller fft_once_ctrl and a butterfly computing unit. The general non-conflicting memory access method of this invention is embodied by this computing structure.
[0068] Compared with the structure of using two data memories with a depth of N, this method only uses one memory MEM with a depth of N to realize the reading and storage of the front and back data of the FFT butterfly calculation, reducing the storage resource consumption by 50%. The memory MEM is composed of P sub-storage units MEM_0~MEM_P-1, forming P parallel read and write ports.
[0069] The general conflict-free memory access method is implemented by the write controller crossBarA, the read controller crossBarB and the controller once_fft_ctrl in the single FFT calculation process. In order to achieve efficient parallel transmission, conflict-free transmission is necessary, which requires the three controllers to work together.
[0070] The butterfly computing unit supports a variety of butterfly operations of power r with a base of 2 (r=1, 2, ..., p_bits). The data to be calculated comes from the data stream dataflowBuIn output by the read controller crossBarB, and then the data stream dataflowBuOut after the butterfly calculation is completed is output.
[0071] The conflict-free memory access method proposed in this paper is described based on Figure 3 The structure implements a control method that complies with the data flow arrangement in step 1, without paying attention to the specific implementation of the butterfly computing unit.
[0072] Combination Figure 3 , the following gives N = 2 5 Point 4: Conflict-free memory access during parallel FFT calculations, such as Figure 4 As shown. The decomposition method is as suggested in step 3, which is divided into three levels of butterfly operations (32 = 4*4*2), and the butterfly operation bases are 2 r1 =2 r2 =2 2 , 2 r3 =2 1 , satisfying r 1 +r 2 +r 3 =5, and finally the data bits in the output stage are output in reverse order.
[0073] (2) FFT calculation process control implementation
[0074] like Figure 3 As shown, the input data stream dataflowInit to be calculated by FFT is first written into the memory MEM. The writing order and the first reading of these data to the butterfly calculation unit are all in the order of reading data in sequence. The write operation is completed by the write data controller crossBarA, which can be divided into two states: the last write operation is performed to correct the output data in reverse order, corresponding to the calculation stage 4 of step 2; except for the last write operation, the remaining write operations are all write operations performed in the FFT calculation to achieve the next level of conflict-free parallel read operations, corresponding to the calculation stages 1 to 3 of step 2. The read operation is completed by the read controller crossBarB. After each write operation is completed, the read controller generates a conflict-free address for reading different sub-memory units MEM_x, so as to continuously read out in parallel and re-array the data stream dataflowBuIn and input it into the butterfly unit. Except for calculation stage 4, the remaining stages are butterfly calculation processes, and each level of butterfly calculation process needs to perform read operation read_x and write operation write_x first. The single butterfly calculation process is collectively called trans_x, which is completed by the read controller crossBarB, the write controller crossBarA and the single FFT controller once_fft_ctrl respectively. Calculation stage 4 corresponds to the write operation write_S-1 of the last level of butterfly calculation and the read operation fft_out_read of the last read out of the reverse order correction result.
[0075] The overall process is controlled by Figure 3The single FFT controller once_fft_ctrl in is completed. Among them, trans_0 has no read operation, but directly receives the initial data stream dataflowInit to be calculated. In the subsequent transmission, except for the read request signal of the fftout_ctrl stage which comes from the subsequent module, the start signals of the remaining stages from trans_1 to trans_S-1 all come from the previous stage single write completion signal write_once_done. During the transmission process, the delay between the data read stage and the data write stage comes from the delay of the butterfly calculation unit. After each transmission stage is completed, remainderBits is subtracted from the current stage butterfly calculation base 2 rthisStage The index r thisStage As for the current stage trans_x is located in the calculation stage 1 to 3 in step 2, the single FFT controller once_fft_ctrl determines the value of remainderBits and then sends the control signal to the read-write controller.
[0076] The state machine transition of the single FFT controller once_fft_ctrl is as follows Figure 6 As shown, when the external controller needs to start an FFT calculation and the single FFT controller is in the idle state CTRL_IDLE, the external controller sends a signal once_fft_start to make the single FFT controller enter the CTRL_trans_0 state; in the CTRL_trans_0 state, the input data stream dataflowInit to be calculated by the FFT is received and sequentially written into the memory MEM. After writing all the data, the write controller crossBarA sets the single write completion signal write_once_done to 1 to indicate that the write operation is completed, and passes it to the single FFT controller once_fft_ctrl. Then the single FFT controller sends a start signal to the read controller, and the read controller reads data from the memory to form the butterfly calculation unit input data stream dataflowInit, which forms the data stream dataflowBuOut after passing through the butterfly calculation unit. The data stream dataflowBuOut will be received and routed by the write controller crossBarA according to the proposed conflict-free parallel memory access rules to be written into the MEM. When the writing is completed, a single write completion signal write_once_done will be formed and enter the subsequent transmission state CTRL_trans_x. In the CTRL_trans_x state, after each read and write operation forms write_once_done, the remainderBits parameter is updated and the state jump judgment state CTRL_update_param is entered. In the state, if (remainderBits-r thisStage) is less than p_bits, the next level still performs radix 2 p_bits Butterfly calculation and update parameter remainderBits = remainderBits-r thisStage , the next state is CTRL_trans_x; if remainderBits-r thisStage Between 0 and p_bits, indicating that base 2 is required next time remainderBits-rthisStage The butterfly calculation and update parameter remainderBits = remainderBits-r thisStage , the next state is CTRL_trans_last, corresponding to trans_S-1 in the figure; if remainderBits-r thisStage If it is equal to 0, the next state is the state of waiting for transmission CTRL_fftout_trans. In the state CTRL_trans_last, the completed operation is consistent with the state CTRL_trans_x, but it is used as an identifier to control the read / write controller. The state CTRL_fftout_trans is a waiting state waiting for the DMA controller to read. When the DMA controller sends a read request, the read controller crossBarB completes the read operation and generates the read completion signal read_once_done. The single FFT controller once_fft_ctrl returns the value CTRL_IDLE to wait for the next FFT calculation.
[0077] (3) Write controller crossBarA to implement
[0078] Write controller crossBarA to output P of butterfly computing unit = 2 p_bits The data flow dataflowBuOut is written into the MEM of the storage unit according to the conflict-free parallel memory access rule. There are P possible data routing situations. For example, when P = 8, there are 8 data routing situations as follows: Figure 7 As shown. p_bits Butterfly calculation, only in the last stage, the radix 8, radix 4 or radix 2 butterfly calculation is performed according to the different number of FFT calculation points. The last stage belongs to the calculation stage 3, which requires bit reversal operation. Therefore, the data routing method used is the same as k in formula (10). p_bits-1 …k 1 k 0 Correspondingly, in the process of writing each level of data flow of FFT calculation into the storage unit MEM, all data routing situations will be covered.
[0079] In order to realize data routing in various situations, the write controller crossBarA will use a data multiplexer with data labels to realize data routing. Figure 8 As shown, the output data dataflowWrite_x is the "x"th data port, and the data selection signal muxSel_x is equal to "x". The input n-channel data is indicated by the label signals wrSel_0~wrSel_P-1. During a data write-back process, wrSel_0~wrSel_P-1 is constantly changing and is generated by the write controller crossBarA.
[0080] Data routing is controlled by label signals wrSel_0~wrSel_P-1, generated by the write controller crossBarA, and is related to the current corresponding write MEM address. According to the conflict-free parallel memory access rule in Section 2.5.1, the counter write_cnt does not correspond to the address of the sub-memory block MEM_x in MEM (all sub-memory blocks MEM_x in the write controller have the same on-chip address wr_mem_addr_x=wr_mem_addr), and a certain conversion mapping is required.
[0081] In the CTRL_trans_0 state (remainderBits == n): Since the initial data stream is input in sequence, the write MEM_x address wr_mem_addr_x is equal to the sequential count write counter write_cnt. However, wrSel_0~wrSel_P-1 need to prepare for the next level of parallel readout. According to formula (6), the high r of write_cnt is intercepted. thisStage bit and add it to the sequence number x of wrSel_x to get the value of wrSel_x. Fig. 9 shown.
[0082] In the CTRL_trans_x stage (remainderBits>=p_bits): This stage is implemented as follows Fig.10 As shown, the read controller performs a jump address read, so the low (n-remainderBits) bit write_cnt_lowBits and the high (remainderBits-p_bits) bit write_cnt_highBits of the write counter write_cnt need to be exchanged to form the write MEM_x address wr_mem_addr_x. The rule for generating the values of wrSel_0 to wrSel_P-1 is also based on formula (6), cyclically shifting the sequence number x of wrSel_x to the left by r thisStageThe bits are added to the high p_bits bits in write_cnt_highBits to obtain the value of wrSel_x. Specifically, when remainderBits > p_bits and remainderBits < 2p_bits, it means that the next-level butterfly calculation is the last-level butterfly calculation of the FFT. Then, according to Equation (7), the sequence number x of wrSel_x is circularly shifted left by remainderBits - r thisStage bit positions, and added to the high remainderBits - r thisStage bit positions of write_cnt_highBits to obtain the value of wrSel_x.
[0083] In the CTRL_trans_last stage (remainderBits < p_bits): In this stage, the read controller completes the reading in the read_S-1 stage. Its jump address reading interval is 1, that is, data is read from the MEM sequentially. Therefore, the write address wr_mem_addr_x of MEM_x is equal to write_cnt. And in this stage, the write controller needs to generate the wrSel_x signal for the operation of bit reversal and sequential reading of the FFT calculation result. According to Equation (10), the high p_bits bits of write_cnt are intercepted, reversed by p_bits bit positions, and added to the sequence number x of wrSel_x to obtain the value of wrSel_x. The above implementation is as Fig.11 shown.
[0084] Through the control of the above write controller crossBarA, the data stream dataflowBuOut is routed according to wrSel_x and wr_mem_addr_x to form the data stream dataflowWrite for writing to the MEM, thus completing the specific implementation of the write operation in the non-conflicting parallel memory access rule.
[0085] (4) Implementation of the read controller crossBarB
[0086] The read controller crossBarB reads data from the MEM to form the data stream dataflowBuIn for performing butterfly operations. Compared with the write controller crossBarA, the read addresses read_mem_addr_0 to read_mem_addr_P-1 generated by the read controller for MEM_0 to MEM_P-1 are all different. The data routing control of the read controller is as Fig.12 shown, and the read tags rdSel_x and the corresponding addresses read_mem_addr_x generated by it are passed through as Figure 8The selector distributes the read data address read_mem_addr_x to each sub-memory block and outputs the read label. For example, when rdSel_2 is equal to 0, rd_mem0_sel is assigned a value of 2, and the corresponding read address of MEM_0 is read_mem_addr_sel2. The data read_memx_data returned by MEM_x is aligned with the label rd_memx_sel_rr, indicating that the data read_memx_data should be routed to the rd_memx_sel_rr data path. Because MEM_x is delayed by 2clk from the read address writing to the read data output, rd_memx_sel_rr is a signal delayed by 2clk from rd_memx_sel. The data read_memx_data read from MEM_x and the synchronization alignment label rd_memx_sel_rr are used by the multiplexer to output P data streams. These data streams are output as data streams dataflowRuIn entering the butterfly calculation unit in the CTRL_trans_0~CTRL_trans_S-1 stages, and are output as FFT calculation result data streams dataflowResult after reverse correction in the fftout_ctrl stage. During the read operation, the read counter read_cnt is used to count, and the count is counted from 0 to 2 n-p _ bits -1, when counting the last column of data, the signal read_once_done is set to "1", and its bit width is n-p_bits.
[0087] exist Fig.12 In the process, the generation of label rdSel_x and corresponding read address read_mem_addr_selx is the key step.
[0088] In the CTRL_trans_0 state (remainderBits==n): there is only write operation but no read operation.
[0089] In the CTRL_trans_x state (remainderBits>=p_bits): the read controller performs jump address reading, referring to equations (6) and (7). Fig.13 As shown, the read counter read_cnt is divided into a high (remainderBits-3) bit read_cnt_highBits and a low (n-remainderBits) bit read_cnt_lowBits, where the low bit of read_cnt_lowBits is lastStageThe bit read_cnt_lowBits_rlastStage is the corresponding bit when the previous butterfly calculation was completed. Adding the sequence number x to read_cnt_lowBits_rlastStage gives the value of rdSel_x. And by taking the lower r lastStage bits of the sequence number x and replacing the lower r lastStage bits of read_cnt_lowBits, here the lower r lastStage bits of the sequence number x are added to the result of read_cnt_lowBits minus read_cnt_lowBits_rlastStage, thus mapping to the corresponding read_cnt_lowBits of the previous-stage butterfly calculation. The result after replacement is swapped with read_cnt_highBits by shifting the result after replacement left by (remainderBits - p_bits) bits and adding it to read_cnt_highBits, obtaining the value of read_mem_addr_selx.
[0090] In the CTRL_trans_last state (remainderBits < p_bits): In this stage, referring to Equation (8), the read data is in column order and there is no need for jump address reading. As Fig.14 shown, the previous stage performs a radix-2 rlastStage butterfly calculation. Taking the lower r lastStage bits of the read counter read_cnt, read_cnt_lowBits_lastStage, and adding the sequence number x gives the value of the tag signal rdSel_x. Taking the lower r lastStage bits of the sequence number x and replacing the lower r lastStage bits of read_cnt_lowBits_lastStage in read_cnt gives the value of the address signal read_mem_addr_selx.
[0091] In the CTRL_fftout_trans stage: In this stage, the read controller crossBarB completes the sequential output of the FFT calculation result, referring to Equation (10). As Fig.15As shown, the intercepted high (p_bits) bit position of read_cnt is bit-reversed and added to the sequence number x to obtain the value of rdSel_x. Since the high (p_bits) of write_cnt is bit-reversed in the last write stage write_S-1, the bit-reversed result x_reverse of the sequence number x in this stage should be replaced by x_reverse for subsequent calculations when calculating read_mem_addr_selx. The bit-reversed result of the low (n-2*p_bits) bit position of read_cnt is added to the result of x_reverse shifted left by (n-2*p_bits) bits to obtain the value of read_mem_addr_selx.
[0092] Step 3: N = 2 n Point decomposition; for N = 2 n Point P = 2 p_bits For the parallel FFT calculation, a decomposition method is proposed, which can be used as a supplementary suggestion for step 2.
[0093] In this method, for N=2 n Point P = 2 p_bits Parallel data stream input, this level of butterfly calculation can use a base of 2 r Butterfly calculation (r can be 1, 2, ..., p_bits). For N = 2 n In order to reduce the overall FFT calculation time, it is recommended to prioritize the base 2 p_bits Butterfly calculation. As calculated in formula (11), it is necessary to perform S-base 2 p_bits Butterfly calculation, the last level performs radix 2 rlastStage Butterfly computing.
[0094] n / p_bits=S···r lastStage (11)
[0095] For FFT calculations with different numbers of points, the final r lastStage is different, and may be a number among 0, 1, 2, ..., p_bits-1, so the adapted butterfly computing unit is required to support a base of 2 r The butterfly calculation (r can be 1, 2, ..., p_bits) is performed. This method focuses on the conflict-free memory access data flow control in the parallel FFT calculation structure based on a single N-depth memory, so no specific design is made for the specific implementation of the butterfly calculation unit.
[0096] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this profession can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the protection scope of the technical solution of the present invention.
Claims
1. A general conflict-free memory access method used in a storage-based parallel FFT computing structure, characterized in that: The method comprises the following steps: Step 1: Analysis of the arrangement of input data streams in P-parallel FFT calculation; Analysis N = 2 n Data flow arrangement of the front and back butterfly calculations in the point P parallel FFT calculation and the use of butterfly calculation cardinality 2 r The relationship between the two is analyzed, and the reverse-corrected data stream is analyzed, and the data stream arrangement is effectively expressed using symbols; Step 2: Propose a general conflict-free parallel memory access method for a P-parallel FFT calculation structure based on a single N-depth memory; the conflict-free memory access method is used to implement the input data stream arrangement in step 1, which will be implemented based on reading and writing a single N-depth memory with P read and write ports. The read and write operations are located before and after each level of butterfly calculation; each level can support a cardinality of 2 r Butterfly calculation (where r can be any value in 1, 2, ..., p_bits, P = 2 p_bits ); Step 3: N = 2 n Point breakdown.
2. A general conflict-free memory access method for use in a storage-based parallel FFT computing structure according to claim 1, characterized in that: In step 1, the arrangement of the input data stream in the P-parallel FFT calculation is analyzed as follows: For the N-point P-parallel FFT calculation by frequency extraction, the input data input to the butterfly calculation unit is arranged as follows: the N-point data is divided into P paths, the 0th to N / P-1th data form the 0th data stream, and the N / Pth to 2N / P-1th data form the 1st data stream, thus forming P-path parallel input data; Each column of the data stream is sequentially input as input data of the butterfly computing unit for butterfly computing.
3. The general conflict-free memory access method used in a storage-based parallel FFT computing structure according to claim 1, characterized in that: Step 2, design of a general conflict-free memory access method, specifically including the following 4 stages: Calculation phase 1, when remainderBits>p_bits and (remainderBits-r thisStage )>p_bits: Read and write operations are performed according to formula (6), where S now->nexts It is the transition arrangement between the write operation to the memory after the current level butterfly calculation is completed and the read operation from the memory when the next level butterfly calculation starts, corresponding to the actual storage arrangement of the data in the memory (where the symbol rotl(a,b) means cyclically shifting data a to the left by b bits); Calculation stage 2, when remainderBits > p_bits and (remainderBits - r thisStage ) < p_bits: This stage is essentially the process from calculation stage 1 to calculation stage 3, and is executed only once; if in calculation stage 1 (remainderBits-r thisStage ) is equal to p_bits, then this stage is not executed, but directly goes from stage 1 to stage 3; the arrangement read and write expression of this stage is shown in formula (7); Compared to calculation stage 1, the number of shifts in the circular left shift operation rotl is not r thisStage , but nr sum -p_bits(=remainderBits-p_bits), where n represents the number of bits of data at N points (N=2 n ), r sum Indicates that the base 2 has been completed rsum Butterfly calculation of Calculation stage 3, when remainderBits <= p_bits and (remainderBits - r thisStage ) < p_bits: The data stream arrangement does not change until the entire FFT calculation is completed, that is, when remainderBits==0; the arrangement read and write expressions at this stage are shown in formula (8), Calculation stage 4, complete the reverse order correction after frequency extraction FFT calculation, that is, complete the bit reversal and output the FFT calculation results in sequential order; in the hybrid butterfly calculation, 2 r Butterfly calculation, from the perspective of the overall FFT calculation, the sequence numbers of the output data x(n) and the output FFT result X(k) are in reverse order by 1 bit, as shown in formula (9); The read-write arrangement relationship of the reverse result order output is expressed using formula (10):
4. The general conflict-free memory access method used in a storage-based parallel FFT computing structure according to claim 1, characterized in that: In step 3, N = 2 n The recommended point breakdown is as follows: For N = 2 n Point P = 2 p_bits Parallel data stream input, this level of butterfly calculation can use a base of 2 r Butterfly calculation (r can be 1, 2, ..., p_bits); for N = 2 n In order to reduce the overall FFT calculation time, it is recommended to prioritize the base 2 p _bits Butterfly calculation; as in formula (11), it is necessary to perform S-base 2 p_bits Butterfly calculation, the last level performs radix 2 rlastStage Butterfly computing; n / p_bits=S···r lastStage (11) For FFT calculations with different numbers of points, the final r lastStage is different, and may be a number among 0, 1, 2, ..., p_bits-1, so the adapted butterfly computing unit is required to support a base of 2 r Butterfly calculation, r can be 1, 2, …, p_bits.