Discrete Fourier transform (DFT) processing method and device and electronic equipment
By decomposing DFT into multi-stage hybrid butterfly operation and determining the parallelism based on processing delay and resource number, the problem of insufficient flexibility in processing parallelism is solved, and the processing performance of DFT is improved.
Patent Information
- Application Number
- CN202411068128.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the parallelism degree of DFT processing is relatively flexible, which affects the processing performance of parallel DFT.
The DFT is decomposed into a multi-level mixed base butterfly operation, and the parallelism degree of each base multi-level butterfly operation is determined based on at least one of the processing delay of the DFT, the number of arithmetic logic unit ALU resources required for each base butterfly operation, the number of slicing resources required for the DFT, and the number of series of each base butterfly operation.
It realizes flexible setting of parallelism of different DFT points, improving the processing performance of parallel DFTs.
Smart Images

Figure CN120372140A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of communications and image processing, and particularly to a method, apparatus, electronic device, storage medium, and chip for processing a Discrete Fourier Transform (DFT). Background Art
[0002] Currently, the Discrete Fourier Transform (DFT) has been widely applied in the fields of signal analysis and processing, communications, image processing, radar, etc. Parallel DFT has the advantages of parallel processing and high processing efficiency, and is an important development direction in DFT technology. However, in related technologies, the flexibility of the parallelism of DFT processing is relatively low, which affects the processing performance of parallel DFT. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and chip for processing a Discrete Fourier Transform (DFT), so as to at least solve the problem that the flexibility of the parallelism of DFT processing in related technologies is relatively low, which affects the processing performance of parallel DFT. The technical solutions of the present disclosure are as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a method for processing a Discrete Fourier Transform (DFT), including: decomposing the DFT into multi-stage mixed-radix butterfly operations based on the number of DFT points; determining the parallelism of each multi-stage butterfly operation based on at least one of the processing delay of the DFT, the number of Arithmetic Logic Unit (ALU) resources required for each butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each butterfly operation, where the parallelism of any multi-stage butterfly operation is the number of data processed in parallel in the any multi-stage butterfly operation.
[0005] According to a second aspect of an embodiment of the present disclosure, there is provided a device for processing a Discrete Fourier Transform (DFT), including: a decomposition module configured to perform decomposing the DFT into multi-stage mixed-radix butterfly operations based on the number of DFT points; a determination module configured to perform determining the parallelism of each multi-stage butterfly operation based on at least one of the processing delay of the DFT, the number of Arithmetic Logic Unit (ALU) resources required for each butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each butterfly operation, where the parallelism of any multi-stage butterfly operation is the number of data processed in parallel in the any multi-stage butterfly operation.
[0006] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the steps of the method described in the first aspect of the embodiments of the present disclosure.
[0007] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method described in the first aspect of the embodiments of the present disclosure are implemented.
[0008] According to a fifth aspect of the embodiments of the present disclosure, there is provided a chip, including one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal and send the signal to the processor, and the signal includes computer instructions stored in a memory. When the processor executes the computer instructions, the chip executes the steps of the method described in the first aspect of the embodiments of the present disclosure.
[0009] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: The parallelism of the per-radix multi-stage butterfly operation can be determined by considering at least one of the processing delay of the DFT, the number of ALU resources required for each radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each radix butterfly operation. The flexible setting of the parallelism for different DFT point numbers can be realized, and the processing performance of the parallel DFT is improved.
[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0012] Figure 1 It is a flowchart of a method for processing a discrete Fourier transform (DFT) shown according to an exemplary embodiment.
[0013] Figure 2 It is a flowchart of a method for processing a discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0014] Figure 3 It is a flowchart of a method for processing a discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0015] Figure 4 It is a flowchart of a method for processing a discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0016] Figure 5 It is a flowchart of a method for processing a discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0017] Figure 6It is a flowchart of a processing method for discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0018] Figure 7 It is a flowchart of a processing method for discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0019] Figure 8 It is a flowchart of a processing method for discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0020] Figure 9 It is a flowchart of a processing method for discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0021] Figure 10 It is a schematic diagram of a processing method for discrete Fourier transform (DFT) shown according to an exemplary embodiment.
[0022] Figure 11 It is a schematic diagram of a processing method for discrete Fourier transform (DFT) shown according to another exemplary embodiment.
[0023] Figure 12 It is a block diagram of a processing device for discrete Fourier transform (DFT) shown according to an exemplary embodiment.
[0024] Figure 13 It is a block diagram of an electronic device shown according to an exemplary embodiment.
[0025] Figure 14 It is a block diagram of a chip shown according to an exemplary embodiment. Detailed implementation
[0026] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0028] In the technical solution of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the provisions of relevant laws and regulations.
[0029] Figure 1 It is a flowchart of a processing method for discrete Fourier transform (DFT) shown according to an exemplary embodiment, as Figure 1 shown, the processing method for discrete Fourier transform (DFT) in the embodiment of the present disclosure includes the following steps.
[0030] S101, based on the number of DFT points, decompose the DFT into multiple levels of mixed-radix butterfly operations.
[0031] It should be noted that the execution subject of the processing method for DFT (Discrete Fourier Transform) in the embodiment of the present disclosure is an electronic device, such as a mobile phone, a notebook, a desktop computer, a vehicle-mounted terminal, a smart home appliance, a wearable device, etc. Among them, the wearable device may include a wrist-worn device (such as a smart watch, a smart bracelet), a head-worn device, a foot-worn device, etc. The processing method for discrete Fourier transform (DFT) in the embodiment of the present disclosure can be executed by the processing device for discrete Fourier transform (DFT) in the embodiment of the present disclosure, and the processing device for discrete Fourier transform (DFT) in the embodiment of the present disclosure can be configured in any electronic device to execute the processing method for discrete Fourier transform (DFT) in the embodiment of the present disclosure.
[0032] It should be noted that the number of DFT points refers to the number of discrete data for DFT processing (also called the number of sampling points), and is also the number of discrete data in the time domain or spatial domain considered during frequency domain transformation, also called the transform length of the DFT. There is no excessive limitation on the number of DFT points. For example, it can be any number of DFT points in communication systems such as NR (new radio) and LTE (Long Term Evolution).
[0033] It should be noted that based on the number of DFT points, decomposing the DFT into multiple levels of mixed-radix butterfly operations can be implemented by using any DFT decomposition method in related technologies, and there is no excessive limitation here. Among them, the DFT decomposition methods include PFA (Prime-factor FFT algorithm) and FFT (Fast Fourier Transform).
[0034] It should be noted that there is no excessive limitation on the radix of the mixed radix for DFT decomposition. For example, it can include radix 2, radix 3, radix 5, etc.
[0035] In one embodiment, based on the number of DFT points, the DFT is decomposed into multi - level mixed - radix butterfly operations, including decomposing the DFT into α2 - level radix - 2 butterfly operations, α3 - level radix - 3 butterfly operations, and α5 - level radix - 5 butterfly operations based on the number of DFT points, where α2 and α3 are positive integers greater than or equal to 1, and α5 is a non - negative integer.
[0036] It can be understood that α2 is the number of levels of the radix - 2 butterfly operation, α3 is the number of levels of the radix - 3 butterfly operation, and α5 is the number of levels of the radix - 5 butterfly operation. If α5 = 0, at this time, the DFT is only decomposed into α2 - level radix - 2 butterfly operations and α3 - level radix - 3 butterfly operations. If α5≥1, at this time, the DFT is decomposed into α2 - level radix - 2 butterfly operations, α3 - level radix - 3 butterfly operations, and α5 - level radix - 5 butterfly operations. Among them, the symbol "0" is used to represent the number zero.
[0037] For example, based on the number of DFT points, decomposing the DFT into multi - level mixed - radix butterfly operations can be achieved through the following formula:
[0038]
[0039] where M sc is the number of DFT points.
[0040] It can be understood that if the DFT is decomposed into α2 - level radix - 2 butterfly operations, α3 - level radix - 3 butterfly operations, and α5 - level radix - 5 butterfly operations, then first perform times of α2 - level radix - 2 butterfly operations, then perform times of α3 - level radix - 3 butterfly operations. If α5≥1, then perform times of α5 - level radix - 5 butterfly operations. The α2 - level radix - 2 butterfly operation is a DFT with a point number of , the α3 - level radix - 3 butterfly operation is a DFT with a point number of , and the α5 - level radix - 5 butterfly operation is a DFT with a point number of .
[0041] If α5 = 0, the input sequence of the α2 - level radix - 2 butterfly operation is the input sequence of the DFT, the output sequence of the α2 - level radix - 2 butterfly operation is the input sequence of the α3 - level radix - 3 butterfly operation, and the output sequence of the α3 - level radix - 3 butterfly operation is the output sequence of the DFT.
[0042] If α5≥1, the input sequence of the α2 - level radix - 2 butterfly operation is the input sequence of the DFT, the output sequence of the α2 - level radix - 2 butterfly operation is the input sequence of the α3 - level radix - 3 butterfly operation, the output sequence of the α3 - level radix - 3 butterfly operation is the input sequence of the α5 - level radix - 5 butterfly operation, and the output sequence of the α5 - level radix - 5 butterfly operation is the output sequence of the DFT.
[0043] It should be noted that the multi-level butterfly operation per stage can be implemented by any butterfly operation method in the related technologies, and no excessive limitation is made here. Among them, the butterfly operation method may include CFA (common factor algorithm).
[0044] S102, determine the parallelism of the multi-level butterfly operation per stage based on at least one of the processing delay of the DFT, the number of arithmetic logic unit (ALU) resources required for the butterfly operation per stage, the number of partitions of the storage resources required for the DFT, and the number of levels of the butterfly operation per stage, where the parallelism of any stage of the multi-level butterfly operation is the number of data processed in parallel in any stage of the multi-level butterfly operation.
[0045] It should be noted that the parallelism of different levels of the butterfly operation in any stage may be different, and the parallelism of the multi-level butterfly operations of different stages may be different. For example, the parallelism of any stage of the multi-level butterfly operation is the number of data read in parallel (input), calculated in parallel, and stored in parallel (output) in any stage of the multi-level butterfly operation.
[0046] It should be noted that no excessive limitation is made on the ALU (Arithmetic Logic Unit) resources. For example, it may include adders, multipliers, etc. The number of ALU resources required for the butterfly operations of different stages may be different.
[0047] For example, the number of adders required for the radix-2 butterfly operation is 6, and the number of multipliers is 4.
[0048] For example, the number of adders required for the radix-3 butterfly operation is 16, and the number of multipliers is 12.
[0049] For example, the number of adders required for the radix-5 butterfly operation is 42, and the number of multipliers is 26.
[0050] It should be noted that no excessive limitation is made on the storage resources required for the DFT. For example, it may include SPRAM (Single-Port Random-access memory). If the SPRAM adopts a dual-buffer strategy, such as Ping-pong SPRAM, the number of partitions of the storage resources required for the DFT at this time is twice the original number of partitions.
[0051] In one implementation, the processing delay of the DFT is negatively correlated with the parallelism of at least one stage of the multi-level butterfly operation.
[0052] In one embodiment, the number of partitions of the storage resources required for DFT is positively correlated with the parallelism of at least one radix-multistage butterfly operation, or the number of partitions of the storage resources required for DFT is positively correlated with the first maximum value of the parallelism of each radix-multistage butterfly operation. It should be noted that the first maximum value refers to the maximum value of the parallelism of each radix-multistage butterfly operation. For example, if the parallelism of the radix-2 multistage butterfly operation is 6, the parallelism of the radix-3 multistage butterfly operation is 9, and the parallelism of the radix-5 multistage butterfly operation is 10, then the first maximum value is 10.
[0053] In one embodiment, based on at least one of the processing delay of DFT, the number of arithmetic logic unit (ALU) resources required for each radix butterfly operation, the number of partitions of the storage resources required for DFT, and the number of stages of each radix butterfly operation, determine the parallelism of each radix-multistage butterfly operation, including obtaining the total constraint condition of the parallelism of each radix-multistage butterfly operation based on at least one of the processing delay of DFT, the number of ALU resources required for each radix butterfly operation, the number of partitions of the storage resources required for DFT, and the number of stages of each radix butterfly operation, and determining the parallelism of each radix-multistage butterfly operation based on the total constraint condition.
[0054] The processing method of discrete Fourier transform (DFT) provided by the embodiments of the present disclosure decomposes DFT into multistage mixed-radix butterfly operations based on the number of DFT points, and determines the parallelism of each radix-multistage butterfly operation based on at least one of the processing delay of DFT, the number of ALU resources required for each radix butterfly operation, the number of partitions of the storage resources required for DFT, and the number of stages of each radix butterfly operation. Thus, considering at least one of the processing delay of DFT, the number of ALU resources required for each radix butterfly operation, the number of partitions of the storage resources required for DFT, and the number of stages of each radix butterfly operation, the parallelism of each radix-multistage butterfly operation can be determined, and the flexible setting of the parallelism for different DFT points can be realized, improving the processing performance of parallel DFT.
[0055] Figure 2 is a flowchart of a processing method of discrete Fourier transform (DFT) shown according to another exemplary embodiment. As Figure 2 shown, the processing method of discrete Fourier transform (DFT) of the embodiments of the present disclosure includes the following steps.
[0056] S201, decompose DFT into multistage mixed-radix butterfly operations based on the number of DFT points.
[0057] For the relevant content of step S201, reference can be made to the above embodiments, which will not be elaborated here.
[0058] S202, obtain the first correlation relationship between the processing delay of DFT and the parallelism of each radix-multistage butterfly operation.
[0059] It should be noted that the first correlation relationship is not overly restricted. For example, it may include positive correlation, negative correlation, linear relationship, non-linear relationship, polynomial curve, functional relationship represented by a function expression, corresponding relationship between value ranges, etc. To obtain the first correlation relationship, any method for obtaining a correlation relationship in related technologies can be used, and no further restrictions are imposed here.
[0060] For example, taking the function relationship as an example of the first correlation relationship, the function relationship is not overly restricted. For example, it may include linear function relationship, inverse proportional function relationship, quadratic function relationship, exponential function relationship, logarithmic function relationship, power function relationship, etc. For example, the first correlation relationship includes a function relationship with the processing delay of DFT as the dependent variable and the parallelism of the multi-level butterfly operation per stage as the independent variable.
[0061] For example, to obtain the first correlation relationship, any function construction method in related technologies can be used, and no further restrictions are imposed here. Among them, the function construction method may include linear regression, symbolic regression, polynomial regression, random forest regression, decision tree regression, etc. The symbolic regression method may include symbolic regression methods based on genetic algorithms, evolutionary strategies, particle swarm optimization, etc.
[0062] In one implementation, obtaining the first correlation relationship between the processing delay of DFT and the parallelism of the multi-level butterfly operation per stage includes obtaining a sample data set. Among them, the z-th sample data in the sample data set includes the sample processing delay of DFT and the sample parallelism of the multi-level butterfly operation per stage, where z is a positive integer, and constructing a third function between the processing delay of DFT and the parallelism of the multi-level butterfly operation per stage based on the sample data set as the first correlation relationship.
[0063] In one implementation, obtaining the first correlation relationship between the processing delay of DFT and the parallelism of the multi-level butterfly operation per stage includes constructing a first function between the data fetching delay of any stage of the multi-level butterfly operation and the parallelism of any stage of the multi-level butterfly operation based on the number of DFT points, constructing a second function between the data fetching delay of DFT and the parallelism of the multi-level butterfly operation per stage based on the first function corresponding to each stage of the multi-level butterfly operation, and constructing a third function between the processing delay of DFT and the parallelism of the multi-level butterfly operation per stage based on the second function as the first correlation relationship.
[0064] It should be noted that the data fetching delay of any stage of the multi-level butterfly operation refers to the duration of reading data from the storage space for butterfly operation during the process of any stage of the multi-level butterfly operation.
[0065] For example, based on the number of DFT points, construct a first function corresponding to any radix multi - stage butterfly operation, including determining the coefficients of the first function corresponding to any radix multi - stage butterfly operation based on the number of DFT points and the number of stages of any radix butterfly operation, taking the parallelism of any radix multi - stage butterfly operation as the independent variable of the first function corresponding to any radix multi - stage butterfly operation, and taking the data fetch latency of any radix multi - stage butterfly operation as the dependent variable of the first function corresponding to any radix multi - stage butterfly operation.
[0066] For example, when the radices of the mixed radix in DFT decomposition include radix 2, radix 3, and radix 5, the first function corresponding to the radix - 2 multi - stage butterfly operation is as follows:
[0067]
[0068] The first function corresponding to the radix - 3 multi - stage butterfly operation is as follows:
[0069]
[0070] The first function corresponding to the radix - 5 multi - stage butterfly operation is as follows:
[0071]
[0072] Among them, latency_r2 is the data fetch latency of the radix - 2 multi - stage butterfly operation, latency_r3 is the data fetch latency of the radix - 3 multi - stage butterfly operation, latency_r5 is the data fetch latency of the radix - 5 multi - stage butterfly operation, N is the number of DFT points, r2_p is the parallelism of the radix - 2 multi - stage butterfly operation, r3_p is the parallelism of the radix - 3 multi - stage butterfly operation, and r5_p is the parallelism of the radix - 5 multi - stage butterfly operation.
[0073] For example, the second function is as follows:
[0074] latency_1 = latency_r2 + latency_r3 + latency_r5
[0075] Among them, latency_1 is the data fetch latency of the DFT.
[0076] For example, the third function is as follows:
[0077] latency = latency_1 + latency_2
[0078] Among them, latency is the processing latency of the DFT, and latency_2 is the calculation latency of the DFT.
[0079] It should be noted that the calculation delay of the DFT refers to the duration of performing the butterfly operation, and there is no excessive limitation on the calculation delay of the DFT. For example, the calculation delay of the DFT is less than or equal to 200 cycles. For example, the unit of various delays in the embodiments of the present disclosure is cycle, and cycle is a clock cycle, which is a basic time unit.
[0080] S203. Take the processing delay of the DFT being less than or equal to the first set threshold as the first constraint condition.
[0081] S204. Based on the first correlation and the first constraint condition, determine the parallelism of each radix multi - stage butterfly operation.
[0082] It should be noted that there is no excessive limitation on the first set threshold. For example, it may include 2500 cycles.
[0083] For example, with the DFT point number N = M sc = 3072 as an example, based on the DFT point number, the DFT is decomposed into multi - stage mixed - radix butterfly operations, which can be implemented through the following formula:
[0084] M sc = 2 10 ·3 1 = 3072
[0085] If the parallelism of the radix - 2 multi - stage butterfly operation is 16 and the parallelism of the radix - 3 multi - stage butterfly operation is 12, the processing delay of the DFT can be obtained through the following formula:
[0086]
[0087] latency = latency_1 + latency_2 = 2176 + 200 = 2376
[0088] If the first set threshold is 2500 cycles, and at this time the processing delay of the DFT is 2376 cycles, it can be known that the processing delay of the DFT meets the first constraint condition. The final parallelism of the radix - 2 multi - stage butterfly operation can be determined to be 16, and the final parallelism of the radix - 3 multi - stage butterfly operation can be determined to be 12.
[0089] In one implementation manner, based on the first correlation and the first constraint condition, determining the parallelism of each radix multi - stage butterfly operation includes obtaining the parallelism of each radix multi - stage butterfly operation when the processing delay of the DFT meets the first constraint condition based on the first correlation, as the final parallelism of each radix multi - stage butterfly operation.
[0090] For example, based on the parallelism of the per - radix multi - stage butterfly operation and the first correlation relationship, the processing delay of the DFT is obtained. If the processing delay of the DFT does not meet the first constraint condition, the parallelism of at least one per - radix multi - stage butterfly operation is optimized, and based on the optimized parallelism of the per - radix multi - stage butterfly operation and the first correlation relationship, the processing delay of the DFT is obtained again. If the processing delay of the DFT obtained again meets the first constraint condition, the parallelism of the per - radix multi - stage butterfly operation obtained last time is used as the final parallelism of the per - radix multi - stage butterfly operation.
[0091] In one implementation, based on the first correlation relationship and the first constraint condition, to determine the parallelism of the per - radix multi - stage butterfly operation, it includes: based on the first correlation relationship, obtaining the value range of the parallelism of the per - radix multi - stage butterfly operation when the processing delay of the DFT meets the first constraint condition, constructing a search space based on the value range of the parallelism of the per - radix multi - stage butterfly operation, and performing parameter search within the search space to obtain the parallelism of the per - radix multi - stage butterfly operation.
[0092] The processing method of the discrete Fourier transform (DFT) provided by the embodiments of the present disclosure obtains the first correlation relationship between the processing delay of the DFT and the parallelism of the per - radix multi - stage butterfly operation, takes that the processing delay of the DFT is less than or equal to the first set threshold as the first constraint condition, and determines the parallelism of the per - radix multi - stage butterfly operation based on the first correlation relationship and the first constraint condition. Thus, taking that the processing delay of the DFT is less than or equal to the first set threshold as the first constraint condition makes the determined parallelism meet the constraint of a smaller processing delay of the DFT, which helps to reduce the processing delay of the parallel DFT, and can comprehensively consider the first correlation relationship and the first constraint condition to determine the parallelism of the per - radix multi - stage butterfly operation.
[0093] Figure 3 is a flowchart of a processing method of a discrete Fourier transform (DFT) shown according to another exemplary embodiment. As Figure 3 shown, the processing method of the discrete Fourier transform (DFT) of the embodiments of the present disclosure includes the following steps.
[0094] S301, decompose the DFT into multi - stage mixed - radix butterfly operations based on the number of DFT points.
[0095] For the relevant content of step S301, reference can be made to the above - mentioned embodiments and will not be elaborated here.
[0096] S302, based on the number of ALU resources required for any one - radix butterfly operation, obtain the second correlation relationship between the total number of ALU resources required for any one - radix multi - stage butterfly operation and the parallelism of any one - radix multi - stage butterfly operation.
[0097] It should be noted that for the relevant content of the second correlation relationship, reference may be made to the relevant content of the first correlation relationship in the above embodiments, which will not be elaborated here. For example, the second correlation relationship corresponding to any radix-multistage butterfly operation includes a functional relationship with the total number of ALU resources required for any radix-multistage butterfly operation as the dependent variable and the parallelism of any radix-multistage butterfly operation as the independent variable.
[0098] In one implementation, based on the number of ALU resources required for any radix butterfly operation, obtaining the second correlation relationship between the total number of ALU resources required for any radix-multistage butterfly operation and the parallelism of any radix-multistage butterfly operation includes constructing a fourth function between the parallel number of any radix-multistage butterfly operation and the parallelism of any radix-multistage butterfly operation based on the radix of any radix, where the parallel number of any radix-multistage butterfly operation is the number of parallel butterfly operations of any radix-multistage butterfly operation, and constructing a fifth function between the total number of ALU resources required for any radix-multistage butterfly operation and the parallelism of any radix-multistage butterfly operation based on the number of ALU resources required for any radix butterfly operation and the fourth function, as the second correlation relationship.
[0099] It should be noted that the parallel number of any radix-multistage butterfly operation refers to the number of parallel radix butterfly operations during any radix-multistage butterfly operation.
[0100] For example, constructing a fourth function between the parallel number of any radix-multistage butterfly operation and the parallelism of any radix-multistage butterfly operation based on the radix of any radix includes determining the coefficients of the fourth function of any radix-multistage butterfly operation based on the radix of any radix, taking the parallelism of any radix-multistage butterfly operation as the independent variable of the fourth function corresponding to any radix-multistage butterfly operation, and taking the parallel number of any radix-multistage butterfly operation as the dependent variable of the fourth function corresponding to any radix-multistage butterfly operation.
[0101] For example, when the radixes of the mixed radix in DFT decomposition include radix 2, radix 3, and radix 5, the fourth function corresponding to the radix-2 multistage butterfly operation is as follows:
[0102]
[0103] The fourth function corresponding to the radix-3 multistage butterfly operation is as follows:
[0104]
[0105] The fourth function corresponding to the radix-5 multistage butterfly operation is as follows:
[0106]
[0107] Among them, p_r2 is the parallel quantity of the radix-2 multi-stage butterfly operation, p_r3 is the parallel quantity of the radix-3 multi-stage butterfly operation, and p_r5 is the parallel quantity of the radix-5 multi-stage butterfly operation.
[0108] For example, the fifth function corresponding to the radix-2 multi-stage butterfly operation is as follows:
[0109]
[0110] The fifth function corresponding to the radix-3 multi-stage butterfly operation is as follows:
[0111]
[0112] The fifth function corresponding to the radix-5 multi-stage butterfly operation is as follows:
[0113]
[0114]
[0115] Among them, qadd_r2 is the total number of adders required for the radix-2 multi-stage butterfly operation, qmul_r2 is the total number of multipliers required for the radix-2 multi-stage butterfly operation, add_r2 is the number of adders required for the radix-2 butterfly operation, and mul_r2 is the number of multipliers required for the radix-2 butterfly operation.
[0116] qadd_r3 is the total number of adders required for the radix-3 multi-stage butterfly operation, qmul_r3 is the total number of multipliers required for the radix-3 multi-stage butterfly operation, add_r3 is the number of adders required for the radix-3 butterfly operation, and mul_r3 is the number of multipliers required for the radix-3 butterfly operation.
[0117] qadd_r5 is the total number of adders required for the radix-5 multi-stage butterfly operation, qmul_r5 is the total number of multipliers required for the radix-5 multi-stage butterfly operation, add_r5 is the number of adders required for the radix-5 butterfly operation, and mul_r5 is the number of multipliers required for the radix-5 butterfly operation.
[0118] For example, if the parallelism of the radix-2 multi-stage butterfly operation is 16, 16 / 2 = 8, and 8 is the number of parallel radix-2 butterfly operations during the radix-2 multi-stage butterfly operation. If the number of adders required for the radix-2 butterfly operation is 6 and the number of multipliers is 4, then the total number of adders required for the radix-2 multi-stage butterfly operation is 48, and the number of multipliers is 32.
[0119] For example, if the parallelism of the radix-3 multi-stage butterfly operation is 12, 12 / 3 = 4, and 4 is the number of parallel radix-3 butterfly operations during the radix-3 multi-stage butterfly operation. If the number of adders required for the radix-3 butterfly operation is 16 and the number of multipliers is 12, then the total number of adders required for the radix-3 multi-stage butterfly operation is 64, and the number of multipliers is 48.
[0120] For example, if the parallelism of the radix-5 multi-stage butterfly operation is 10, 10 / 5 = 2, and 2 is the number of parallel radix-5 butterfly operations during the radix-5 multi-stage butterfly operation. If the number of adders required for the radix-5 butterfly operation is 42 and the number of multipliers is 26, then the total number of adders required for the radix-5 multi-stage butterfly operation is 84, and the number of multipliers is 52.
[0121] S303, minimize the difference between the total quantities corresponding to different radix multi-stage butterfly operations as the optimization goal.
[0122] S304, based on the second correlation relationship corresponding to each radix multi-stage butterfly operation and the optimization goal, determine the parallelism of each radix multi-stage butterfly operation.
[0123] In one implementation, based on the second correlation relationship corresponding to each radix multi-stage butterfly operation and the optimization goal, determining the parallelism of each radix multi-stage butterfly operation includes constructing an objective function between the difference between the total quantities corresponding to different radix multi-stage butterfly operations and the parallelism of each radix multi-stage butterfly operation based on the second correlation relationship corresponding to each radix multi-stage butterfly operation, performing optimization processing on the objective function, and obtaining the parallelism of each radix multi-stage butterfly operation when the difference between the total quantities corresponding to different radix multi-stage butterfly operations is the minimum value as the final parallelism of each radix multi-stage butterfly operation.
[0124] It should be noted that the optimization processing of the objective function can be implemented by any function optimization method in related technologies, and no excessive limitation is made here. Among them, the function optimization methods can include gradient descent, Newton's method, quasi-Newton method, conjugate gradient method, trust region method, etc.
[0125] In one implementation, based on the second correlation relationship corresponding to each radix multi-stage butterfly operation and the optimization goal, determining the parallelism of each radix multi-stage butterfly operation includes obtaining the parallelism of each radix multi-stage butterfly operation when the difference between the total quantities corresponding to different radix multi-stage butterfly operations is the minimum value based on the second correlation relationship as the final parallelism of each radix multi-stage butterfly operation.
[0126] For example, based on the parallelism of the multi - stage butterfly operation per radix and the second correlation relationship, the difference between the total quantities corresponding to different - radix multi - stage butterfly operations is obtained. If the above - mentioned difference is not the minimum value, the parallelism of at least one - radix multi - stage butterfly operation is optimized, and based on the optimized parallelism of the multi - stage butterfly operation per radix and the second correlation relationship, the difference between the total quantities corresponding to different - radix multi - stage butterfly operations is obtained again. If the re - obtained difference is the minimum value, the parallelism of the multi - stage butterfly operation per radix obtained last time is used as the final parallelism of the multi - stage butterfly operation per radix.
[0127] The processing method of the discrete Fourier transform (DFT) provided by the embodiments of the present disclosure obtains the second correlation relationship between the total quantity of ALU resources required for any - radix multi - stage butterfly operation and the parallelism of any - radix multi - stage butterfly operation based on the quantity of ALU resources required for any - radix butterfly operation. Taking the minimum difference between the total quantities corresponding to different - radix multi - stage butterfly operations as the optimization goal, based on the second correlation relationship corresponding to each - radix multi - stage butterfly operation and the optimization goal, the parallelism of each - radix multi - stage butterfly operation is determined. Thus, taking the minimum difference between the total quantities corresponding to different - radix multi - stage butterfly operations as the optimization goal to determine the parallelism of each - radix multi - stage butterfly operation makes the determined parallelism satisfy the constraint that the difference in the total quantity of ALU resources corresponding to different - radix multi - stage butterfly operations is small, that is, the total quantity of ALU resources required for different - radix multi - stage butterfly operations is as uniform as possible, and the ALU resources can be time - shared and reused between different - radix multi - stage butterfly operations, improving the time - sharing reuse efficiency of ALU resources. The DFT requires less ALU resources, which can save ALU resources, and can comprehensively consider the second correlation relationship and the optimization goal to determine the parallelism of each - radix multi - stage butterfly operation.
[0128] Figure 4 It is a flowchart of a processing method of a discrete Fourier transform (DFT) shown according to another exemplary embodiment. As Figure 4 shown, the processing method of the discrete Fourier transform (DFT) of the embodiments of the present disclosure includes the following steps.
[0129] S401, based on the number of DFT points, decompose the DFT into multi - stage mixed - radix butterfly operations.
[0130] For the relevant content of step S401, reference can be made to the above - mentioned embodiments and will not be elaborated here.
[0131] S402, obtain the third correlation relationship between the number of partitions of the storage resources required for the DFT and the first maximum value of the parallelism of the multi - stage butterfly operation per radix, where the number of partitions of the storage resources required for the DFT is greater than or equal to the first maximum value.
[0132] It should be noted that for the relevant content of the third correlation relationship, reference can be made to the relevant content of the first correlation relationship in the above embodiments, which will not be elaborated here. For example, the number of partitions of the storage resources required for DFT in the third correlation relationship is greater than or equal to the first maximum value.
[0133] For example, if the first maximum value is 18, the number of partitions of the storage resources required for DFT can be 20.
[0134] S403. Use the condition that the number of partitions of the storage resources required for DFT is less than or equal to a second set threshold as the second constraint condition.
[0135] S404. Based on the third correlation relationship and the second constraint condition, determine the parallelism of each-radix multi-stage butterfly operation.
[0136] It should be noted that no excessive limitation is imposed on the second set threshold. For example, it can be 25. It can be understood that the larger the number of partitions of the storage resources required for DFT, the larger the area occupied by each storage space (such as one bit), resulting in a lower storage efficiency of the storage resources. Therefore, to ensure the storage efficiency of the storage resources, the number of partitions of the storage resources required for DFT is less than or equal to the second set threshold.
[0137] In one implementation manner, determining the parallelism of each-radix multi-stage butterfly operation based on the third correlation relationship and the second constraint condition includes, based on the third correlation relationship, obtaining the parallelism of each-radix multi-stage butterfly operation when the number of partitions of the storage resources required for DFT meets the second constraint condition as the final parallelism of each-radix multi-stage butterfly operation.
[0138] For example, obtain the maximum value in the parallelism of each-radix multi-stage butterfly operation as the first maximum value. Based on the first maximum value and the third correlation relationship, obtain the number of partitions of the storage resources required for DFT. If the number of partitions of the storage resources required for DFT does not meet the second constraint condition, optimize the parallelism of at least one-radix multi-stage butterfly operation, and based on the optimized parallelism of each-radix multi-stage butterfly operation and the third correlation relationship, re-obtain the number of partitions of the storage resources required for DFT. If the re-obtained number of partitions of the storage resources required for DFT meets the second constraint condition, use the parallelism of each-radix multi-stage butterfly operation obtained last time as the final parallelism of each-radix multi-stage butterfly operation.
[0139] In one implementation manner, determining the parallelism of each-radix multi-stage butterfly operation based on the third correlation relationship and the second constraint condition includes, based on the third correlation relationship, obtaining the value range of the parallelism of each-radix multi-stage butterfly operation when the number of partitions of the storage resources required for DFT meets the second constraint condition, constructing a search space based on the value range of the parallelism of each-radix multi-stage butterfly operation, and performing parameter search within the search space to obtain the parallelism of each-radix multi-stage butterfly operation.
[0140] The processing method of discrete Fourier transform (DFT) provided by the embodiments of the present disclosure obtains a third correlation relationship between the number of partitions of the storage resources required for DFT and the first maximum value of the parallelism of each radix multi-stage butterfly operation, where the number of partitions of the storage resources required for DFT is greater than or equal to the first maximum value. Taking the number of partitions of the storage resources required for DFT being less than or equal to a second set threshold as a second constraint condition, based on the third correlation relationship and the second constraint condition, the parallelism of each radix multi-stage butterfly operation is determined. Thus, taking the number of partitions of the storage resources required for DFT being less than or equal to the second set threshold as the second constraint condition, the parallelism of each radix multi-stage butterfly operation is determined, so that the determined parallelism satisfies the constraint of a smaller number of partitions of the storage resources, which can ensure a higher storage efficiency of the storage resources, and can comprehensively consider the third correlation relationship and the second constraint condition to determine the parallelism of each radix multi-stage butterfly operation.
[0141] Figure 5 is a flowchart of a processing method of discrete Fourier transform (DFT) shown according to another exemplary embodiment, as Figure 5 shown, the processing method of discrete Fourier transform (DFT) of the embodiments of the present disclosure includes the following steps.
[0142] S501, based on the number of DFT points, decompose the DFT into multi-stage mixed-radix butterfly operations.
[0143] For the relevant content of step S501, reference can be made to the above embodiments, which will not be elaborated here.
[0144] S502, obtain the maximum value of the parallelism of any radix multi-stage butterfly operation.
[0145] It should be noted that for the relevant content of step S502, reference can be made to the relevant content of step S603 in the following embodiments, which will not be elaborated here.
[0146] S503, determine that the parallelism of the first i - stage butterfly operation of any radix is the maximum value of the parallelism of any radix multi-stage butterfly operation.
[0147] S504, determine that the parallelism of the last j - stage butterfly operation of any radix is less than or equal to the maximum value of the parallelism of any radix multi-stage butterfly operation, where both i and j are non - negative integers, and the sum of i and j is the number of stages of any radix butterfly operation.
[0148] It can be understood that the parallelism of a certain stage of butterfly operation of any radix is less than or equal to the maximum value of the parallelism of any radix multi-stage butterfly operation. It should be noted that no excessive restrictions are imposed on both i and j. For example, if the number of stages of any radix butterfly operation is 4, then i = 3 and j = 1.
[0149] In one embodiment, j = 1. At this time, the first i levels refer to non - final levels. That is, at this time, the parallelism of any base non - final - level butterfly operation is determined to be the maximum value of the parallelism of any base multi - level butterfly operation, and the parallelism of any base final - level butterfly operation is determined to be less than or equal to the maximum value of the parallelism of any base multi - level butterfly operation.
[0150] In one embodiment, when the radices of the mixed radix in the DFT decomposition include radix 2, radix 3, and radix 5, the method further includes obtaining a second maximum value of the parallelism of the radix - 3 multi - level butterfly operation, determining the parallelism of the radix - 3 non - final - level butterfly operation to be the second maximum value, and determining the parallelism of the radix - 3 final - level butterfly operation to be less than the second maximum value.
[0151] It should be noted that the second maximum value refers to the maximum value of the parallelism of the radix - 3 multi - level butterfly operation, and the parallelism of any level of the radix - 3 butterfly operation is less than or equal to the second maximum value.
[0152] It can be understood that during the radix - 5 multi - level butterfly operation process, it is necessary to re - order the output sequence of the radix - 3 multi - level butterfly operation. If the parallelism of the radix - 3 final - level butterfly operation is larger, the re - order complexity of the output sequence of the radix - 3 multi - level butterfly operation is greater. Therefore, in order to reduce the re - order complexity, the parallelism of the radix - 3 final - level butterfly operation can be less than or equal to the second maximum value.
[0153] For example, when the radices of the mixed radix in the DFT decomposition include radix 2, radix 3, and radix 5, if the number of levels of the radix - 3 multi - level butterfly operation is 4, and the second maximum value of the parallelism of the radix - 3 multi - level butterfly operation is 18, it can be determined that the parallelism of the first 3 levels of the radix - 3 butterfly operation is 18, and the parallelism of the radix - 3 final - level butterfly operation is 9.
[0154] The DFT processing method provided by the embodiments of the present disclosure obtains the maximum value of the parallelism of any base multi - level butterfly operation, determines the parallelism of the first i levels of any base butterfly operation to be the maximum value of the parallelism of any base multi - level butterfly operation, and determines the parallelism of the last j levels of any base butterfly operation to be less than or equal to the maximum value of the parallelism of any base multi - level butterfly operation, where i and j are both non - negative integers, and the sum of i and j is the number of levels of any base butterfly operation. Thus, the parallelism of the first several levels of a certain base butterfly operation is the maximum value of the parallelism, and the parallelism of the last several levels of a certain base butterfly operation is less than the maximum value of the parallelism, which can reduce the re - order complexity of the next - base multi - level butterfly operation of a certain base multi - level butterfly operation.
[0155] Figure 6 is a flowchart of a DFT processing method shown according to another exemplary embodiment, as Figure 6 shown, the DFT processing method of the embodiments of the present disclosure includes the following steps.
[0156] S601. Decompose the DFT into multi - stage mixed - radix butterfly operations based on the number of DFT points.
[0157] For the relevant content of step S601, reference can be made to the above - mentioned embodiments and will not be elaborated here.
[0158] S602. Determine the value range of the parallelism of each - radix multi - stage butterfly operation based on at least one of the processing delay of the DFT, the number of ALU resources required for each - radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each - radix butterfly operation.
[0159] In one implementation, determining the value range of the parallelism of each - radix multi - stage butterfly operation based on at least one of the processing delay of the DFT, the number of ALU resources required for each - radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each - radix butterfly operation includes determining the total constraint conditions of the parallelism of each - radix multi - stage butterfly operation based on at least one of the processing delay of the DFT, the number of ALU resources required for each - radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each - radix butterfly operation, and determining the value range of the parallelism of each - radix multi - stage butterfly operation based on the total constraint conditions of the parallelism of each - radix multi - stage butterfly operation.
[0160] In some examples, determining the value range of the parallelism of each - radix multi - stage butterfly operation based on the total constraint conditions of the parallelism of each - radix multi - stage butterfly operation includes performing parameter search within the search space according to the total constraint conditions of the parallelism of each - radix multi - stage butterfly operation to obtain the value range of the parallelism of each - radix multi - stage butterfly operation.
[0161] It should be noted that the value ranges of the parallelisms of different - radix multi - stage butterfly operations may be different, or there may be overlapping parallelisms.
[0162] In one implementation, when the radix of the mixed - radix in the DFT decomposition includes radix 2, the value range of the parallelism of the radix - 2 multi - stage butterfly operation includes at least one of 6, 8, 12, and 16.
[0163] In one implementation, when the radix of the mixed - radix in the DFT decomposition includes radix 3, the value range of the parallelism of the radix - 3 multi - stage butterfly operation includes at least one of 3, 6, 9, 12, and 18.
[0164] In one implementation, when the radix of the mixed - radix in the DFT decomposition includes radix 5, the value range of the parallelism of the radix - 5 multi - stage butterfly operation includes 10 and / or 15.
[0165] S603. Determine the parallelism of any - radix multi - stage butterfly operation from the value range of the parallelism of any - radix multi - stage butterfly operation.
[0166] In one embodiment, determining the parallelism of any radix multi - stage butterfly operation from the value range of the parallelism of any radix multi - stage butterfly operation includes randomly screening out the parallelism of any radix multi - stage butterfly operation from the value range of the parallelism of any radix multi - stage butterfly operation.
[0167] In one embodiment, determining the parallelism of any radix multi - stage butterfly operation from the value range of the parallelism of any radix multi - stage butterfly operation includes determining the parallelism of any radix multi - stage butterfly operation from the value range of the parallelism of any radix multi - stage butterfly operation based on at least one of the number of DFT points, the number of stages of at least one radix butterfly operation, and the parallelism of the remaining radix multi - stage butterfly operations other than any radix. Thus, considering at least one of the number of DFT points, the number of stages of at least one radix butterfly operation, and the parallelism of the remaining radix multi - stage butterfly operations other than any radix, determining the parallelism of any radix multi - stage butterfly operation from the value range of the parallelism of any radix multi - stage butterfly operation can achieve flexible setting of the parallelism for different numbers of DFT points and improve the processing performance of parallel DFT.
[0168] In the embodiments of the present disclosure, determining the parallelism of any radix multi - stage butterfly operation from the value range of the parallelism of any radix multi - stage butterfly operation based on at least one of the number of DFT points, the number of stages of at least one radix butterfly operation, and the parallelism of the remaining radix multi - stage butterfly operations other than any radix includes the following possible embodiments:
[0169] Method 1: Taking the number of stages of the radix - 2 butterfly operation being 1 as the first condition, the number of stages of the radix - 2 butterfly operation being greater than 3 as the second condition, and the number of DFT points being 3240 as the third condition.
[0170] If the first condition is satisfied, determine the parallelism of the radix - 2 multi - stage butterfly operation to be 6.
[0171] If the second condition is satisfied, determine the third maximum value of the parallelism of the radix - 2 multi - stage butterfly operation to be 16.
[0172] If the third condition is satisfied, determine the third maximum value to be 8.
[0173] If the first condition, the second condition, and the third condition are not satisfied, determine the third maximum value to be 12.
[0174] It should be noted that in this embodiment, the radix of the mixed radix of DFT decomposition includes radix - 2, and the value range of the parallelism of the radix - 2 multi - stage butterfly operation includes at least one of 6, 8, 12, and 16.
[0175] It should be noted that the third maximum value refers to the maximum value of the parallelism of the radix - 2 multi - stage butterfly operation, and the parallelism of any stage of the radix - 2 butterfly operation is less than or equal to the third maximum value.
[0176] It should be noted that if the second condition is satisfied, the third maximum value for determining the parallelism of the radix-2 multi-stage butterfly operation is 16, which means that if the second condition is satisfied, it is determined that the parallelism of any stage of the radix-2 butterfly operation is less than or equal to 16, that is, if the second condition is satisfied, it is determined that the parallelism of any stage of the radix-2 butterfly operation is any one of 6, 8, 12, and 16.
[0177] It should be noted that if the third condition is satisfied, the third maximum value is determined to be 8, which means that if the third condition is satisfied, it is determined that the parallelism of any stage of the radix-2 butterfly operation is less than or equal to 8, that is, if the third condition is satisfied, it is determined that the parallelism of any stage of the radix-2 butterfly operation is any one of 6 and 8.
[0178] It should be noted that if the first to third conditions are not satisfied, the third maximum value is determined to be 12, which means that if the first to third conditions are not satisfied, it is determined that the parallelism of any stage of the radix-2 butterfly operation is less than or equal to 12, that is, if the first to third conditions are not satisfied, it is determined that the parallelism of any stage of the radix-2 butterfly operation is any one of 6, 8, and 12.
[0179] Method 2: Take any one of 600, 750, and 2400 for the DFT points as the fourth condition, take 1 for both the number of stages of the radix-2 butterfly operation and the number of stages of the radix-3 butterfly operation as the fifth condition, take 1200 or 3000 for the DFT points as the sixth condition, take 1 for the number of stages of the radix-3 butterfly operation as the seventh condition, take 2916 or 3240 for the DFT points as the eighth condition, and take the number of stages of the radix-3 butterfly operation greater than 1 as the ninth condition.
[0180] If the fourth condition is satisfied, the parallelism of the radix-3 multi-stage butterfly operation is determined to be 3.
[0181] If the fifth condition is satisfied and the DFT points are not 750, the second maximum value for determining the parallelism of the radix-3 multi-stage butterfly operation is determined to be 6.
[0182] If the sixth condition is satisfied, the second maximum value is determined to be 6.
[0183] If the seventh condition is satisfied and the fourth, fifth, and sixth conditions are not satisfied, the second maximum value is determined to be 12.
[0184] If the eighth condition is satisfied, the second maximum value is determined to be 18.
[0185] If the ninth condition is satisfied and the eighth condition is not satisfied, the second maximum value is determined to be 9.
[0186] It should be noted that in this embodiment, the radix of the mixed radix of the DFT decomposition includes radix-3, and the value range of the parallelism of the radix-3 multi-stage butterfly operation includes at least one of 3, 6, 9, 12, and 18. For the relevant content of Method 2, reference can be made to the relevant content of Method 1, which will not be elaborated here.
[0187] Method 3: Taking the number of DFT points as 2916 or 3240 as the eighth condition, and taking the number of levels of the radix-3 butterfly operation being greater than 1 as the ninth condition.
[0188] If the ninth condition is satisfied and the eighth condition is not satisfied, determine that the parallelism of the radix-5 multi-stage butterfly operation is 10.
[0189] If the ninth condition is not satisfied, or if the eighth condition is satisfied, determine that the fourth maximum value of the parallelism of the radix-5 multi-stage butterfly operation is 15.
[0190] It should be noted that in this embodiment, the radix of the mixed radix of the DFT decomposition includes radix-5, and the value range of the parallelism of the radix-5 multi-stage butterfly operation includes 10 and / or 15.
[0191] It should be noted that the fourth maximum value refers to the maximum value of the parallelism of the radix-5 multi-stage butterfly operation, and the parallelism of any stage of the radix-5 butterfly operation is less than or equal to the fourth maximum value. For the relevant content of Method 3, reference can be made to the relevant content of Method 1, which will not be elaborated here.
[0192] Method 4: Taking the second maximum value of the parallelism of the radix-3 multi-stage butterfly operation being 9 as the tenth condition.
[0193] If the tenth condition is satisfied, determine that the parallelism of the radix-5 multi-stage butterfly operation is 10.
[0194] If the tenth condition is not satisfied, determine that the fourth maximum value of the parallelism of the radix-5 multi-stage butterfly operation is 15.
[0195] It should be noted that in this embodiment, the radix of the mixed radix of the DFT decomposition includes radix-3 and radix-5, and the value range of the parallelism of the radix-5 multi-stage butterfly operation includes 10 and / or 15. For the relevant content of Method 4, reference can be made to the relevant content of Method 1, which will not be elaborated here.
[0196] The processing method of discrete Fourier transform (DFT) provided by the embodiments of the present disclosure determines the value range of the parallelism of each base multi-stage butterfly operation based on at least one of the processing delay of DFT, the number of ALU resources required for each base butterfly operation, the number of divisions of the storage resources required for DFT, and the number of stages of each base butterfly operation, and determines the parallelism of any base multi-stage butterfly operation from the value range of the parallelism of any base multi-stage butterfly operation. Thus, considering at least one of the processing delay of DFT, the number of ALU resources required for each base butterfly operation, the number of divisions of the storage resources required for DFT, and the number of stages of each base butterfly operation, the value range of the parallelism of each base multi-stage butterfly operation can be determined, and the parallelism of any base multi-stage butterfly operation can be determined from the value range of the parallelism of any base multi-stage butterfly operation, so as to flexibly set the parallelism of different DFT points and improve the processing performance of parallel DFT.
[0197] Figure 7 is a flowchart of a processing method of discrete Fourier transform (DFT) shown according to another exemplary embodiment, as Figure 7 shown, the processing method of discrete Fourier transform (DFT) of the embodiments of the present disclosure includes the following steps.
[0198] S701, decompose the DFT into multi-stage mixed base butterfly operations based on the DFT points.
[0199] For the relevant content of step S701, reference can be made to the above embodiments and will not be elaborated here.
[0200] S702, obtain multiple reordering processing rules for any base multi-stage butterfly operation based on the DFT points.
[0201] It should be noted that the reordering processing rule is to reorder the data in the original sequence to obtain a new sequence. The new sequence and the original sequence are composed of the same data, and the sorting of the data in the new sequence and the original sequence is different. There are no excessive limitations on the multiple reordering processing rules. For example, the multiple reordering processing rules include at least two of the input sequence rearrangement rule (Scramble), the input sequence reshaping rule (Reshape_in), the input sequence flipping rule (Flip), and the output sequence reshaping rule (Reshape_out). Among them, the input sequence flipping rule (Flip) may include the DIT (Decimation In Time) algorithm.
[0202] In one embodiment, when the bases of the mixed bases in the DFT decomposition include base 2, base 3, and base 5, the multiple reordering processing rules for the base-2 multi-stage butterfly operations include an input sequence rearrangement rule (Scramble), an input sequence reshaping rule (Reshape_in), and an input sequence flipping rule (Flip). The multiple reordering processing rules for the base-3 and base-5 multi-stage butterfly operations both include an output sequence reshaping rule (Reshape_out), an input sequence reshaping rule (Reshape_in), and an input sequence flipping rule (Flip).
[0203] For example, taking the original sequence as DataIp[i] and the new sequence as DataOp[j i , where the number of DFT points is N, both the original sequence and the new sequence include N data. DataIp[i] is the i-th data of the original sequence, and DataOp[j i is the j i -th data of the new sequence, that is, i is the original sorting of the data, and j i is the sorting of the data after reordering.
[0204] The input sequence rearrangement rule (Scramble) can be implemented by the following formula:
[0205] DataOp[j i = DataIp[i] i = 0, 1,..., N - 1
[0206] j0 = 0
[0207] j i = (j i-1 + Δ) mod N i > 0
[0208]
[0209] Wherein,
[0210] The input sequence reshaping rule (Reshape_in) can be implemented by the following formula:
[0211] DataOp[j i = DataIp[i] j i = 0, 1,..., N - 1
[0212] i = (M2·M3·n2 + M1·n2') mod N
[0213] n2 = j i mod M1
[0214]
[0215] The output sequence reshaping rule (Reshape_out) can be implemented by the following formula:
[0216] DataOp[j i = DataIp[i] i = 0, 1, …, N - 1
[0217] j i = (M2·M3·n2 + M1·n2') mod N
[0218] n2 = i mod M1
[0219]
[0220] S703. Based on multiple transposition processing rules of any - base multi - stage butterfly operation, obtain the combined transposition processing rule of any - base multi - stage butterfly operation.
[0221] It should be noted that the combined transposition processing rule of any - base multi - stage butterfly operation refers to the total transposition processing rule after combining (integrating) multiple transposition processing rules of any - base multi - stage butterfly operation.
[0222] In one implementation, the combined transposition processing rule of any - base multi - stage butterfly operation includes at least one of the number of DFT points, the number of points of any - base multi - stage butterfly operation, the sorting difference between two adjacent data in the same column and the same group of the target transposition sequence of any - base multi - stage butterfly operation, the sorting difference between two adjacent data in different groups of the same column of the target transposition sequence, and the sorting of the first - row data of the target transposition sequence. Among them, r adjacent data in the same column of the target transposition sequence are in the same group, and there is no overlapping data between different groups of the target transposition sequence, and r is the base number of any base.
[0223] It can be understood that the target transposition sequence of any - base multi - stage butterfly operation is a matrix of A rows and B columns, where A is the number of points of any - base multi - stage butterfly operation, B is the number of times of any - base multi - stage butterfly operation, and the product of A and B is the number of DFT points.
[0224] For example, taking M sc = 2 2 ·3 1 = 12 as an example, the process of obtaining the combined transposition processing rule of the radix - 2 multi - stage butterfly operation based on multiple transposition processing rules of the radix - 2 multi - stage butterfly operation is as follows:
[0225] The input sequence of the radix - 2 multi - stage butterfly operation is as follows:
[0226] 0 1 2 3 4 5 6 7 8 9 10 11
[0227] Based on the input sequence rearrangement rule of the radix-2 multi-stage butterfly operation, the input sequence of the radix-2 multi-stage butterfly operation can be reordered once to obtain the first reordered sequence of the radix-2 multi-stage butterfly operation. The first reordered sequence of the radix-2 multi-stage butterfly operation is as follows:
[0228] 0 7 2 9 4 11 6 1 8 3 10 5
[0229] Based on the input sequence reshaping rule of the radix-2 multi-stage butterfly operation, the first reordered sequence of the radix-2 multi-stage butterfly operation is reordered once to obtain the second reordered sequence of the radix-2 multi-stage butterfly operation. The second reordered sequence of the radix-2 multi-stage butterfly operation is as follows:
[0230] 0 4 8 9 1 5 6 10 2 3 7 11
[0231] Based on the input sequence flipping rule of the radix-2 multi-stage butterfly operation, the second reordered sequence of the radix-2 multi-stage butterfly operation is reordered once to obtain the target reordered sequence of the radix-2 multi-stage butterfly operation. The target reordered sequence of the radix-2 multi-stage butterfly operation is as follows:
[0232] 0 4 8 6 10 2 9 1 5 3 7 11
[0233] It should be noted that 0 to 11 in the above sequences are all used to represent the sorting of data in the input sequence of the radix-2 multi-stage butterfly operation, not the specific values of the data.
[0234] The target reordered sequence of the radix-2 multi-stage butterfly operation satisfies the following rules:
[0235] 1. The starting sorting of each column is r2_len·x, where r2_len is the number of points of the radix-2 multi-stage butterfly operation, x is the column number, and x is a natural number.
[0236] 2. Two adjacent data in the same column of the target reordered sequence are in the same group, and there is no overlapping data between different groups of the target reordered sequence. The sorting difference between two adjacent data in the same group in the same column of the target reordered sequence is 6, and the sorting difference between two adjacent data in different groups in the same column of the target reordered sequence is 3. When the sorting exceeds N (12), it means overflow and N needs to be subtracted.
[0237] 3. The target reordered sequence is a matrix of A rows and B columns, A = r2_len, B is the number of times of the radix-2 multi-stage butterfly operation, and the product of A and B is the DFT number of points.
[0238] Therefore, when the DFT number of points is 12, the combined reordering processing rule of the radix-2 multi-stage butterfly operation includes N = 12, r2_len, the sorting difference between two adjacent data in the same group in the same column of the target reordered sequence is 6, and the sorting difference between two adjacent data in different groups in the same column of the target reordered sequence is 3.
[0239] For example, taking M sc = 2 1 · 3 2 = 18 as an example, based on multiple reordering processing rules of the radix-3 multi-stage butterfly operation, the process of obtaining the combined reordering processing rule of the radix-3 multi-stage butterfly operation is as follows:
[0240] The input sequence of the radix-3 multi-stage butterfly operation is as follows:
[0241] 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17
[0242] Based on the output sequence reshaping rule of the radix-3 multi-stage butterfly operation, perform a reordering process on the input sequence of the radix-3 multi-stage butterfly operation to obtain the first reordering sequence of the radix-3 multi-stage butterfly operation. The first reordering sequence of the radix-3 multi-stage butterfly operation is as follows:
[0243] 0 11 2 13 4 15 6 17 8 1 10 3 12 5 14 7 16 9
[0244] Based on the input sequence reshaping rule of the radix-3 multi-stage butterfly operation, perform a reordering process on the first reordering sequence of the radix-3 multi-stage butterfly operation to obtain the second reordering sequence of the radix-3 multi-stage butterfly operation. The second reordering sequence of the radix-3 multi-stage butterfly operation is as follows:
[0245] 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17
[0246] Based on the input sequence flipping rule of the radix-3 multi-stage butterfly operation, perform a reordering process on the second reordering sequence of the radix-3 multi-stage butterfly operation to obtain the target reordering sequence of the radix-3 multi-stage butterfly operation. The target reordering sequence of the radix-3 multi-stage butterfly operation is as follows:
[0247] 0 1 6 7 12 13 2 3 8 9 14 15 4 5 10 11 16 17
[0248] It should be noted that 0 to 17 in the above sequences are all used to represent the sorting of data in the input sequence of the radix-3 multi-stage butterfly operation, not the specific values of the data.
[0249] The target reordering sequence of the radix-3 multi-stage butterfly operation satisfies the following rules:
[0250] 1. The starting sorting of each column is x, where x is the column number and x is a natural number.
[0251] 2. The adjacent 3 data in the same column of the target reordering sequence are in the same group, and there is no overlapping data between different groups of the target reordering sequence. The sorting difference between adjacent two data in the same group in the same column of the target reordering sequence is 6. The starting sorting of each group in the same column of the target reordering sequence is where y is the row number and y is a natural number.
[0252] 3. The target permutation sequence is a matrix of A rows and B columns, where A = r3_len, B is the number of times of the radix-3 multi-stage butterfly operation, and the product of A and B is the DFT point number. Here, r3_len is the number of points of the radix-3 multi-stage butterfly operation.
[0253] Therefore, when the DFT point number is 18, the merging permutation processing rules of the radix-3 multi-stage butterfly operation include N = 18, r3_len, the sorting difference between two adjacent data in the same group of the same column of the target permutation sequence is 6, and the starting sorting of each group in the same column of the target permutation sequence is
[0254] For example, taking M sc = 2 1 · 3 1 · 5 1 = 30 as an example, the process of obtaining the merging permutation processing rules of the radix-5 multi-stage butterfly operation based on multiple permutation processing rules of the radix-5 multi-stage butterfly operation is as follows:
[0255] The input sequence of the radix-5 multi-stage butterfly operation is as follows:
[0256] 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
[0257] 20 21 22 23 24 25 26 27 28 29
[0258] It should be noted that the data with sorting 20 is the next data of the data with sorting 19 in the input sequence of the radix-5 multi-stage butterfly operation.
[0259] Based on the reshaping rule of the output sequence of the radix-5 multi-stage butterfly operation, the input sequence of the radix-5 multi-stage butterfly operation is permuted once to obtain the first permutation sequence of the radix-5 multi-stage butterfly operation. The first permutation sequence of the radix-5 multi-stage butterfly operation is as follows:
[0260] 0 22 14 3 25 17 6 28 20 9 1 23 12 4 26 15 7 29 18 10
[0261] 2 21 13 5 24 16 8 27 19 11
[0262] It should be noted that the data with sorting 2 is the next data of the data with sorting 10 in the first permutation sequence of the radix-5 multi-stage butterfly operation.
[0263] Based on the reshaping rule of the input sequence of the radix-5 multi-stage butterfly operation, the first permutation sequence of the radix-5 multi-stage butterfly operation is permuted once to obtain the second permutation sequence of the radix-5 multi-stage butterfly operation. The second permutation sequence of the radix-5 multi-stage butterfly operation is as follows:
[0264] 0 17 1 15 2 16 6 23 7 21 8 22 12 29 13 27 14 28 18 5 19 3 20 4 24 11 25 9 26 10
[0265] Based on the input sequence reversal rule of the radix-5 multi-stage butterfly operation, perform a reordering process on the second reordered sequence of the radix-5 multi-stage butterfly operation to obtain the target reordered sequence of the radix-5 multi-stage butterfly operation. The target reordered sequence of the radix-5 multi-stage butterfly operation is as follows:
[0266] 0 17 1 15 2 16 6 23 7 21 8 22 12 29 13 27 14 28 18 5 19 3 20 4 24 11 25 9 26 10
[0267] It should be noted that 0 to 29 in the above sequence are used to represent the sorting of data in the input sequence of the radix-5 multi-stage butterfly operation, not the specific values of the data.
[0268] The target reordered sequence of the radix-5 multi-stage butterfly operation satisfies the following rules:
[0269] 1. The sorting of the data in the first row of the target reordered sequence is (0, 17, 1, 15, 2, 16).
[0270] 2. The same column in the target reordered sequence is in the same group. The sorting difference between two adjacent data in the same column of the target reordered sequence is 6. When the sorting exceeds N (30), it means overflow and N needs to be subtracted.
[0271] 3. The target reordered sequence is a matrix with A rows and B columns. A = r5_len, B is the number of times of the radix-5 multi-stage butterfly operation, and the product of A and B is the DFT point number. Where r5_len is the number of points of the radix-5 multi-stage butterfly operation.
[0272] Therefore, when the DFT point number is 30, the combined reordering processing rule of the radix-5 multi-stage butterfly operation includes N = 30, r5_len, the sorting difference between two adjacent data in the same column of the target reordered sequence is 6, and the sorting of the data in the first row of the target reordered sequence is (0, 17, 1, 15, 2, 16).
[0273] It should be noted that for the process of obtaining the combined reordering processing rule of any radix multi-stage butterfly operation under other DFT point numbers, reference can be made to the above embodiments and will not be elaborated here.
[0274] S704, store the combined reordering processing rule of each radix multi-stage butterfly operation.
[0275] In one embodiment, the method further includes obtaining an input sequence of any radix multi - stage butterfly operation, performing a re - ordering process on the input sequence of any radix multi - stage butterfly operation according to the merging and re - ordering processing rule of any radix multi - stage butterfly operation to obtain a target re - ordered sequence of any radix multi - stage butterfly operation, and performing any radix multi - stage butterfly operation on the target re - ordered sequence of any radix multi - stage butterfly operation to obtain an output sequence of any radix multi - stage butterfly operation. Thus, only by performing a re - ordering process on the input sequence of a certain radix multi - stage butterfly operation according to the merging and re - ordering processing rule of the certain radix multi - stage butterfly operation, the target re - ordered sequence of the certain radix multi - stage butterfly operation can be obtained. Compared with the related art where multiple re - ordering processes need to be performed on the input sequence, in this solution, only one re - ordering process needs to be performed on the input sequence, which simplifies the re - ordering steps, can save the re - ordering processing time, improve the re - ordering processing efficiency, and further improve the DFT processing efficiency.
[0276] In some examples, when the radices of the mixed radix in the DFT decomposition include radix 2, radix 3, and radix 5, the input sequence of the DFT can be obtained as the input sequence of the radix - 2 multi - stage butterfly operation. According to the merging and re - ordering processing rule of the radix - 2 multi - stage butterfly operation, perform a re - ordering process on the input sequence of the radix - 2 multi - stage butterfly operation to obtain a target re - ordered sequence of the radix - 2 multi - stage butterfly operation, and perform the radix - 2 multi - stage butterfly operation on the target re - ordered sequence of the radix - 2 multi - stage butterfly operation to obtain an output sequence of the radix - 2 multi - stage butterfly operation.
[0277] Take the output sequence of the radix - 2 multi - stage butterfly operation as the input sequence of the radix - 3 multi - stage butterfly operation. According to the merging and re - ordering processing rule of the radix - 3 multi - stage butterfly operation, perform a re - ordering process on the input sequence of the radix - 3 multi - stage butterfly operation to obtain a target re - ordered sequence of the radix - 3 multi - stage butterfly operation, and perform the radix - 3 multi - stage butterfly operation on the target re - ordered sequence of the radix - 3 multi - stage butterfly operation to obtain an output sequence of the radix - 3 multi - stage butterfly operation.
[0278] If α5 = 0, take the output sequence of the radix - 3 multi - stage butterfly operation as the output sequence of the DFT.
[0279] If α5≥1, take the output sequence of the radix - 3 multi - stage butterfly operation as the input sequence of the radix - 5 multi - stage butterfly operation. According to the merging and re - ordering processing rule of the radix - 5 multi - stage butterfly operation, perform a re - ordering process on the input sequence of the radix - 5 multi - stage butterfly operation to obtain a target re - ordered sequence of the radix - 5 multi - stage butterfly operation, and perform the radix - 5 multi - stage butterfly operation on the target re - ordered sequence of the radix - 5 multi - stage butterfly operation to obtain an output sequence of the radix - 5 multi - stage butterfly operation, and take the output sequence of the radix - 5 multi - stage butterfly operation as the output sequence of the DFT.
[0280] The processing method of discrete Fourier transform (DFT) provided by the embodiments of the present disclosure obtains multiple reordering processing rules for any radix-multistage butterfly operation based on the number of DFT points, obtains the combined reordering processing rule for any radix-multistage butterfly operation based on the multiple reordering processing rules for any radix-multistage butterfly operation, and stores the combined reordering processing rule for each radix-multistage butterfly operation. Thus, multiple reordering processing rules for a certain radix-multistage butterfly operation can be integrated to obtain the combined reordering processing rule for the certain radix-multistage butterfly operation, and the combined reordering processing rule is stored. Compared with the problem in the related art that storing the entire reordering sequence consumes a large amount of storage resources, in this solution, only the combined reordering processing rule needs to be stored, and the reordering sequence does not need to be stored, so the consumed storage resources are less, and the storage resources can be saved.
[0281] Figure 8 is a flowchart of a processing method of discrete Fourier transform (DFT) shown according to another exemplary embodiment, as Figure 8 shown, the processing method of discrete Fourier transform (DFT) in the embodiments of the present disclosure includes the following steps.
[0282] S801, based on the number of DFT points, decompose the DFT into multiple radix-mixed butterfly operations.
[0283] For the relevant content of step S801, reference can be made to the above embodiments, and details are not described herein again.
[0284] S802, obtain the output sequence of the first-radix multistage butterfly operation, where the output sequence of the first-radix multistage butterfly operation includes sub-output sequences at p moments, the c-th sub-output sequence includes q output data at the c-th moment, q is the parallelism of the first-radix multistage butterfly operation, both p and q are positive integers, the product of p and q is the number of DFT points, and c is a positive integer not greater than p.
[0285] S803, obtain the target reordering sequence of the second-radix multistage butterfly operation, where the target reordering sequence of the second-radix multistage butterfly operation includes sub-reordering sequences at s moments, the d-th sub-reordering sequence includes t output data at the d-th moment, t is the parallelism of the second-radix multistage butterfly operation, both s and t are positive integers, the product of s and t is the number of DFT points, and d is a positive integer not greater than s.
[0286] It should be noted that in this embodiment, the radixes of the mixed radix of the DFT decomposition include the first radix and the second radix. The next-radix multistage butterfly operation of the first-radix multistage butterfly operation is the second-radix multistage butterfly operation, and the output sequence of the first-radix multistage butterfly operation and the target reordering sequence of the second-radix multistage butterfly operation are composed of the same data. For the relevant content of the target reordering sequence, reference can be made to the above embodiments, and details are not described herein again.
[0287] It should be noted that the cth moment refers to the storage moment, and the dth moment refers to the reading (data acquisition) moment.
[0288] For example, if the cardinality of the first basis is 2 and the cardinality of the second basis is 3, the number of DFT points is 48, and M sc =2 4 ·3 1 =48, the parallelism of radix-2 multi-stage butterfly operation is 16, and the parallelism of radix-3 multi-stage butterfly operation is 12, that is, p=3, q=16, s=4, t=12.
[0289] The output sequence r2_dout of the radix-2 multi-stage butterfly operation is shown in Table 1.
[0290] Table 1 Output sequence r2_dout of radix 2 multi-stage butterfly operation
[0291]
[0292] Among them, dout0 to dout15 are used to represent 16 output data at one moment, and cycle=0, 1, and 2 are used to represent one moment respectively. The output sequence of the radix-2 multi-stage butterfly operation includes sub-output sequences at three moments, the first sub-output sequence includes 16 output data at cycle=0 (i.e., the first moment), the second sub-output sequence includes 16 output data at cycle=1 (i.e., the second moment), and the third sub-output sequence includes 16 output data at cycle=2 (i.e., the third moment). The first column of the output sequence of the radix-2 multi-stage butterfly operation is used to represent the first sub-output sequence, the second column of the output sequence of the radix-2 multi-stage butterfly operation is used to represent the second sub-output sequence, and the third column of the output sequence of the radix-2 multi-stage butterfly operation is used to represent the third sub-output sequence.
[0293] The target reordering sequence r3_premap of the radix-3 multi-stage butterfly operation is shown in Table 2.
[0294] Table 2 Target reordering sequence r3_premap for radix 3 multi-stage butterfly operation
[0295]
[0296] Among them, din0 to din11 are used to represent 12 output data at one moment. The target permutation sequence of the radix-3 multi-stage butterfly operation includes sub-permutation sequences at 4 moments, the first sub-permutation sequence includes 12 output data of the first cycle, the second sub-permutation sequence includes 12 output data of the second cycle, the third sub-permutation sequence includes 12 output data of the third cycle, and the fourth sub-permutation sequence includes 12 output data of the fourth cycle.
[0297] Among them, the first column of the target reordering sequence of the radix-3 multi-stage butterfly operation is used to represent the first sub-reordering sequence, the second column of the target reordering sequence of the radix-3 multi-stage butterfly operation is used to represent the second sub-reordering sequence, the third column of the target reordering sequence of the radix-3 multi-stage butterfly operation is used to represent the third sub-reordering sequence, and the fourth column of the target reordering sequence of the radix-3 multi-stage butterfly operation is used to represent the fourth sub-reordering sequence.
[0298] It should be noted that 0 to 47 in the above sequence are all used to represent the sorting of data in the input sequence of the radix-2 multi-stage butterfly operation, not the specific values of the data.
[0299] S804, store the q output data at the same moment in any sub-output sequence into different real storage units respectively, and store the t output data at the same moment in any sub-reordering sequence into different real storage units respectively as the storage rule.
[0300] S805, store the output sequence of the first radix multi-stage butterfly operation according to the storage rule.
[0301] It should be noted that there are no excessive restrictions on the real storage unit. For example, it may include SPRAM.
[0302] It should be noted that storing the q output data at the same moment in any sub-output sequence into different real storage units respectively means storing 1 output data at the same moment in any sub-output sequence into 1 real storage unit, and storing any two output data at the same moment in any sub-output sequence into different real storage units. That is, to achieve the storage of any sub-output sequence, q real storage units are required.
[0303] It should be noted that storing the t output data at the same moment in any sub-reordering sequence into different real storage units respectively means storing 1 output data at the same moment in any sub-reordering sequence into 1 real storage unit, and storing any two output data at the same moment in any sub-reordering sequence into different real storage units. That is, to achieve the storage of any sub-reordering sequence, t real storage units are required.
[0304] For example, continue to take Tables 1 and 2 above as examples.
[0305] Storing the q output data at the same moment in any sub-output sequence into different real storage units respectively means storing the q output data in any column of the output sequence of the radix-2 multi-stage butterfly operation into different real storage units respectively. For example, taking the first column of the output sequence of the radix-2 multi-stage butterfly operation as an example, the output data sorted as 0, 8, 1, 9, 2, 10, 3, 11, 4, 12, 5, 13, 6, 14, 7, 15 are stored into different real storage units respectively.
[0306] At the same moment, the t output data in any sub-permutation sequence are respectively stored in different real storage units, which means that the t output data in any column of the target permutation sequence of the radix-3 multi-stage butterfly operation are respectively stored in different real storage units. For example, taking the first column of the target permutation sequence of the radix-3 multi-stage butterfly operation as an example, the output data sorted as 0, 16, 32, 1, 17, 33, 2, 18, 34, 3, 19, 35 are respectively stored in different real storage units.
[0307] In one embodiment, after storing the output sequence of the first-radix multi-stage butterfly operation, it further includes reading one data from each of the t real storage units at the d-th moment to read the d-th sub-permutation sequence, that is, the t output data in any sub-permutation sequence can be taken out simultaneously, reducing the data fetching delay of the DFT and improving the processing performance of the parallel DFT.
[0308] The processing method of the discrete Fourier transform (DFT) provided by the embodiments of the present disclosure obtains the output sequence of the first-radix multi-stage butterfly operation, obtains the target permutation sequence of the second-radix multi-stage butterfly operation, stores the q output data at the same moment in any sub-output sequence in different real storage units, and stores the t output data at the same moment in any sub-permutation sequence in different real storage units as the storage rule. According to the storage rule, the output sequence of the first-radix multi-stage butterfly operation is stored. Thus, taking storing the q output data at the same moment in any sub-output sequence in different real storage units as the storage rule can achieve the simultaneous storage of the q output data at the same moment in any sub-output sequence, and taking storing the t output data at the same moment in any sub-permutation sequence in different real storage units as the storage rule, so as to take out the t output data in any sub-permutation sequence simultaneously subsequently, reducing the data fetching delay of the DFT and improving the processing performance of the parallel DFT.
[0309] Figure 9 It is a flowchart of a processing method of a discrete Fourier transform (DFT) shown according to another exemplary embodiment, as Figure 9 shown, the processing method of the discrete Fourier transform (DFT) in the embodiments of the present disclosure includes the following steps.
[0310] S901, decompose the DFT into a multi-stage mixed-radix butterfly operation based on the DFT points.
[0311] S902, obtain the output sequence of the first-radix multi-stage butterfly operation.
[0312] S903, obtain the target permutation sequence of the second-radix multi-stage butterfly operation.
[0313] S904. As a storage rule, for any q output data at the same moment in any sub-output sequence, store them in different real storage units respectively, and for any t output data at the same moment in any sub-permutation sequence, store them in different real storage units respectively.
[0314] For the relevant content of steps S901 - S904, reference can be made to the above embodiments and will not be elaborated here.
[0315] S905. According to the storage rule, divide q real storage units into t virtual storage units, where a virtual storage unit is jointly composed of one storage space of each of s real storage units.
[0316] For example, continuing with Tables 1 and 2 above, according to the storage rule, 16 real storage units can be divided into 12 virtual storage units, where a virtual storage unit is jointly composed of one storage space of each of 4 real storage units.
[0317] For example, the division result of 12 virtual storage units is shown in Table 3.
[0318] Table 3 Division result of 12 virtual storage units
[0319]
[0320] Among them, ram_vir0 to 11 are respectively used to represent a virtual storage unit, and addr = 0, 1, 2, 3 are respectively used to represent a storage space.
[0321] Taking ram_vir0 as an example, the storage space with address 0 of ram_vir0 is used to store the data of sorting 0, the storage space with address 1 of ram_vir0 is used to store the data of sorting 4, the storage space with address 2 of ram_vir0 is used to store the data of sorting 8, and the storage space with address 3 of ram_vir0 is used to store the data of sorting 12.
[0322] In one implementation, according to the storage rule, divide q real storage units into t virtual storage units, including based on the storage rule, storing any sub-output sequence in the storage spaces of multiple virtual storage units, which are jointly composed of one storage space of each of q real storage units, as a division rule, and according to the division rule, divide q real storage units into t virtual storage units.
[0323] It can be understood that storing any sub-output sequence in the storage spaces of multiple virtual storage units, which are jointly composed of one storage space of each of q real storage units, can achieve storing any q output data at the same moment in any sub-output sequence in different real storage units respectively.
[0324] In some examples, the method further includes, based on a storage rule, forming the storage space of the d-th address of a virtual storage unit from the storage space of the d-th address of a real storage unit as a partitioning rule. Thus, the storage space of the d-th address of the virtual storage unit corresponds one-to-one with the storage space of the d-th address of the real storage unit, which helps reduce the partitioning complexity.
[0325] For example, continuing with the above Tables 1 to 3 as an example. The relationship between 12 virtual storage units and 16 real storage units is as Figure 10 shown, and the sorting of the data stored in each storage space of the 16 real storage units is as Figure 11 shown.
[0326] As Figure 10 shown, ram0 to 15 are respectively used to represent a real storage unit.
[0327] Taking ram_vir0 as an example, ram_vir0 is jointly composed of the storage space of address 0 of ram0, the storage space of address 1 of ram4, the storage space of address 2 of ram8, and the storage space of address 3 of ram12. Among them, the storage space of address 0 of ram_vir0 is composed of the storage space of address 0 of ram0, the storage space of address 1 of ram_vir0 is composed of the storage space of address 1 of ram4, the storage space of address 2 of ram_vir0 is composed of the storage space of address 2 of ram8, and the storage space of address 3 of ram_vir0 is composed of the storage space of address 3 of ram12.
[0328] Taking ram_vir3 as an example, ram_vir3 is jointly composed of the storage space of address 0 of ram1, the storage space of address 1 of ram5, the storage space of address 2 of ram9, and the storage space of address 3 of ram13. Among them, the storage space of address 0 of ram_vir3 is composed of the storage space of address 0 of ram1, the storage space of address 1 of ram_vir3 is composed of the storage space of address 1 of ram5, the storage space of address 2 of ram_vir3 is composed of the storage space of address 2 of ram9, and the storage space of address 3 of ram_vir3 is composed of the storage space of address 3 of ram13.
[0329] As Figure 11 shown, taking ram0 as an example, the storage space of address 0 of ram0 is used to store the data of sorting 0, the storage space of address 1 of ram0 has no data stored, the storage space of address 2 of ram0 is used to store the data of sorting 40, and the storage space of address 3 of ram0 is used to store the data of sorting 28.
[0330] S906. For any sub-sequence rearrangement, store the t output data at the same moment in any sub-sequence rearrangement into different virtual storage units respectively.
[0331] For example, continuing with Table 3 as an example, the 12 output data at the same moment in any sub-sequence rearrangement can be stored into different virtual storage units respectively.
[0332] In one implementation, for any sub-sequence rearrangement, storing the t output data at the same moment in any sub-sequence rearrangement into different virtual storage units respectively includes storing the t output data at the d-th moment into the storage spaces at the d-th address of the t virtual storage units to store the d-th sub-sequence rearrangement.
[0333] For example, continuing with Table 3 as an example, the 12 output data at the 1st cycle can be stored into the storage spaces at address 0 of the 12 virtual storage units to store the 1st sub-sequence rearrangement. The 12 output data at the 2nd cycle can be stored into the storage spaces at address 1 of the 12 virtual storage units to store the 2nd sub-sequence rearrangement.
[0334] In some examples, the method further includes reading one data respectively from the storage spaces at the d-th address of the t virtual storage units at the d-th moment to read the d-th sub-sequence rearrangement, that is, the t output data in any sub-sequence rearrangement can be taken out simultaneously.
[0335] For example, continuing with Table 3 as an example, one data can be read respectively from the storage spaces at address 0 of the 12 virtual storage units at the 1st cycle to read the 1st sub-sequence rearrangement. One data can be read respectively from the storage spaces at address 1 of the 12 virtual storage units at the 2nd cycle to read the 2nd sub-sequence rearrangement.
[0336] In one implementation, for any sub-sequence rearrangement, storing the t output data at the same moment in any sub-sequence rearrangement into different virtual storage units respectively includes storing the t output data at the d-th moment into the storage spaces at the d-th address of the t real storage units to store the d-th sub-sequence rearrangement.
[0337] For example, continuing with Figure 11 as an example, the 12 output data at the 1st cycle can be stored into the storage spaces at address 0 of ram0 to 11 to store the 1st sub-sequence rearrangement. The 12 output data at the 2nd cycle can be stored into the storage spaces at address 1 of ram4 to 15 to store the 2nd sub-sequence rearrangement.
[0338] In some examples, the method further includes reading a data from the storage spaces of the d-th address of the t true storage units at the d-th moment respectively to read the d-th sub-permutation sequence, that is, the t output data in any sub-permutation sequence can be taken out simultaneously.
[0339] For example, continuing with Figure 11 as an example, one data can be read from the storage spaces of address 0 of ram0 to 11 respectively at the 1st cycle to read the 1st sub-permutation sequence. One data can be read from the storage spaces of address 1 of ram4 to 15 respectively at the 2nd cycle to read the 2nd sub-permutation sequence.
[0340] The processing method of discrete Fourier transform (DFT) provided by the embodiments of the present disclosure divides q true storage units into t virtual storage units according to the storage rule, where the virtual storage unit is jointly composed of one storage space of each of s true storage units. For any sub-permutation sequence, the t output data at the same moment in any sub-permutation sequence are stored in different virtual storage units respectively. Thus, q true storage units can be divided into t virtual storage units according to the storage rule to realize the storage of the output sequence of the first-stage multi-radix butterfly operation.
[0341] Figure 12 It is a block diagram of a processing device for discrete Fourier transform (DFT) shown according to an exemplary embodiment. Referring to Figure 12 , the processing device 100 for discrete Fourier transform (DFT) according to the embodiments of the present disclosure includes: a decomposition module 110 and a determination module 120.
[0342] The decomposition module 110 is configured to perform decomposing the DFT into multi-radix butterfly operations based on the number of DFT points.
[0343] The determination module 120 is configured to determine the parallelism of each multi-radix butterfly operation based on at least one of the processing delay of the DFT, the number of arithmetic logic unit (ALU) resources required for each multi-radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each multi-radix butterfly operation, where the parallelism of any multi-radix butterfly operation is the number of data processed in parallel in the any multi-radix butterfly operation.
[0344] In an embodiment of the present disclosure, the determination module 120 is further configured to perform: obtaining a first correlation relationship between the processing delay of the DFT and the parallelism of each multi-radix butterfly operation; taking that the processing delay of the DFT is less than or equal to a first set threshold as a first constraint condition; and determining the parallelism of each multi-radix butterfly operation based on the first correlation relationship and the first constraint condition.
[0345] In one embodiment of the present disclosure, the determining module 120 is further configured to perform: constructing a first function between the fetch latency of any radix-multistage butterfly operation and the parallelism of the any radix-multistage butterfly operation based on the number of DFT points; constructing a second function between the fetch latency of the DFT and the parallelism of each radix-multistage butterfly operation based on the first function corresponding to each radix-multistage butterfly operation; constructing a third function between the processing latency of the DFT and the parallelism of each radix-multistage butterfly operation based on the second function, as the first correlation relationship.
[0346] In one embodiment of the present disclosure, the determining module 120 is further configured to perform: obtaining a second correlation relationship between the total number of ALU resources required for any radix-multistage butterfly operation and the parallelism of the any radix-multistage butterfly operation based on the number of ALU resources required for any radix butterfly operation; taking the minimum difference between the total quantities corresponding to different radix-multistage butterfly operations as the optimization target; determining the parallelism of each radix-multistage butterfly operation based on the second correlation relationship corresponding to each radix-multistage butterfly operation and the optimization target.
[0347] In one embodiment of the present disclosure, the determining module 120 is further configured to perform: constructing a fourth function between the parallel number of any radix-multistage butterfly operation and the parallelism of the any radix-multistage butterfly operation based on the radix of any radix, where the parallel number of any radix-multistage butterfly operation is the number of parallel butterfly operations of the any radix-multistage butterfly operation; constructing a fifth function between the total number of ALU resources required for any radix-multistage butterfly operation and the parallelism of the any radix-multistage butterfly operation based on the number of ALU resources required for any radix butterfly operation and the fourth function, as the second correlation relationship.
[0348] In one embodiment of the present disclosure, the determining module 120 is further configured to perform: obtaining a third correlation relationship between the number of partitions of the storage resources required for the DFT and the first maximum value of the parallelism of each radix-multistage butterfly operation, where the number of partitions of the storage resources required for the DFT is greater than or equal to the first maximum value; taking the number of partitions of the storage resources required for the DFT being less than or equal to a second set threshold as a second constraint condition; determining the parallelism of each radix-multistage butterfly operation based on the third correlation relationship and the second constraint condition.
[0349] In one embodiment of the present disclosure, the determining module 120 is further configured to perform: obtaining the maximum value of the parallelism of any radix multi - stage butterfly operation; determining that the parallelism of the first i - stage butterfly operation of any radix is the maximum value of the parallelism of any radix multi - stage butterfly operation; determining that the parallelism of the last j - stage butterfly operation of any radix is less than or equal to the maximum value of the parallelism of any radix multi - stage butterfly operation, where i and j are both non - negative integers, and the sum of i and j is the number of stages of any radix butterfly operation.
[0350] In one embodiment of the present disclosure, when the radices of the mixed radix in DFT decomposition include radix 2, radix 3, and radix 5, the determining module 120 is further configured to perform: obtaining the second maximum value of the parallelism of radix 3 multi - stage butterfly operation; determining that the parallelism of the non - last - stage butterfly operation of radix 3 is the second maximum value; determining that the parallelism of the last - stage butterfly operation of radix 3 is less than the second maximum value.
[0351] In one embodiment of the present disclosure, the determining module 120 is further configured to perform: determining the value range of the parallelism of each radix multi - stage butterfly operation based on at least one of the processing delay of the DFT, the number of ALU resources required for each radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each radix butterfly operation; determining the parallelism of any radix multi - stage butterfly operation from the value range of the parallelism of any radix multi - stage butterfly operation.
[0352] In one embodiment of the present disclosure, when the radix of the mixed radix in DFT decomposition includes radix 2, the value range of the parallelism of radix 2 multi - stage butterfly operation includes at least one of 6, 8, 12, and 16; when the radix of the mixed radix in DFT decomposition includes radix 3, the value range of the parallelism of radix 3 multi - stage butterfly operation includes at least one of 3, 6, 9, 12, and 18; when the radix of the mixed radix in DFT decomposition includes radix 5, the value range of the parallelism of radix 5 multi - stage butterfly operation includes 10 and / or 15.
[0353] In one embodiment of the present disclosure, the determining module 120 is further configured to perform: determining the parallelism of any radix multi - stage butterfly operation from the value range of the parallelism of any radix multi - stage butterfly operation based on at least one of the DFT points, the number of stages of at least one radix butterfly operation, and the parallelism of the multi - stage butterfly operations of the other radices except any radix.
[0354] In one embodiment of the present disclosure, when the radix of the mixed radix in DFT decomposition includes radix 2, and the value range of the parallelism of radix 2 multi - stage butterfly operation includes at least one of 6, 8, 12, and 16;
[0355] the determining module 120 is further configured to perform:
[0356] Take the number of stages of the radix-2 butterfly operation being 1 as the first condition;
[0357] Take the number of stages of the radix-2 butterfly operation being greater than 3 as the second condition;
[0358] Take the number of DFT points being 3240 as the third condition;
[0359] If the first condition is satisfied, determine that the parallelism of the radix-2 multi-stage butterfly operation is 6;
[0360] If the second condition is satisfied, determine that the third maximum value of the parallelism of the radix-2 multi-stage butterfly operation is 16;
[0361] If the third condition is satisfied, determine that the third maximum value is 8;
[0362] If the first condition, the second condition, and the third condition are not satisfied, determine that the third maximum value is 12.
[0363] In an embodiment of the present disclosure, when the radix of the mixed radix in the DFT decomposition includes radix-3, and the value range of the parallelism of the radix-3 multi-stage butterfly operation includes at least one of 3, 6, 9, 12, 18;
[0364] The determining module 120 is further configured to execute:
[0365] Take the number of DFT points being any one of 600, 750, 2400 as the fourth condition;
[0366] Take the number of stages of both the radix-2 butterfly operation and the radix-3 butterfly operation being 1 as the fifth condition;
[0367] Take the number of DFT points being 1200 or 3000 as the sixth condition;
[0368] Take the number of stages of the radix-3 butterfly operation being 1 as the seventh condition;
[0369] Take the number of DFT points being 2916 or 3240 as the eighth condition;
[0370] Take the number of stages of the radix-3 butterfly operation being greater than 1 as the ninth condition;
[0371] If the fourth condition is satisfied, determine that the parallelism of the radix-3 multi-stage butterfly operation is 3;
[0372] If the fifth condition is satisfied and the number of DFT points is not 750, determine that the second maximum value of the parallelism of the radix-3 multi-stage butterfly operation is 6;
[0373] If the sixth condition is satisfied, determine that the second maximum value is 6;
[0374] If the seventh condition is satisfied and the fourth, fifth, and sixth conditions are not satisfied, determine that the second maximum value is 12;
[0375] If the eighth condition is satisfied, determine that the second maximum value is 18;
[0376] If the ninth condition is satisfied and the eighth condition is not satisfied, determine that the second maximum value is 9.
[0377] In an embodiment of the present disclosure, when the radix of the mixed radix in the DFT decomposition includes radix 5 and the value range of the parallelism of the radix 5 multi - stage butterfly operation includes 10 and / or 15;
[0378] The determining module 120 is further configured to execute:
[0379] Take the DFT point number of 2916 or 3240 as the eighth condition, and take the number of stages of the radix 3 butterfly operation being greater than 1 as the ninth condition;
[0380] If the ninth condition is satisfied and the eighth condition is not satisfied, determine that the parallelism of the radix 5 multi - stage butterfly operation is 10;
[0381] If the ninth condition is not satisfied, or if the eighth condition is satisfied, determine that the fourth maximum value of the parallelism of the radix 5 multi - stage butterfly operation is 15.
[0382] In an embodiment of the present disclosure, when the radix of the mixed radix in the DFT decomposition includes radix 3 and radix 5, and the value range of the parallelism of the radix 5 multi - stage butterfly operation includes 10 and / or 15;
[0383] The determining module 120 is further configured to execute:
[0384] Take the first maximum value of the parallelism of the radix 3 multi - stage butterfly operation being 9 as the tenth condition;
[0385] If the tenth condition is satisfied, determine that the parallelism of the radix 5 multi - stage butterfly operation is 10;
[0386] If the tenth condition is not satisfied, determine that the fourth maximum value of the parallelism of the radix 5 multi - stage butterfly operation is 15.
[0387] In an embodiment of the present disclosure, the apparatus 100 further includes: a re - ordering module, configured to execute: based on the DFT point number, obtain multiple re - ordering processing rules for any radix multi - stage butterfly operation; based on the multiple re - ordering processing rules of the any radix multi - stage butterfly operation, obtain the combined re - ordering processing rule of the any radix multi - stage butterfly operation; store the combined re - ordering processing rule for each radix multi - stage butterfly operation.
[0388] In an embodiment of the present disclosure, the merging and transposition processing rule of any - radix multi - stage butterfly operation includes at least one of the DFT points, the number of points of the any - radix multi - stage butterfly operation, the sorting difference between two adjacent data in the same column and the same group of the target transposition sequence of the any - radix multi - stage butterfly operation, the sorting difference between two adjacent data in different groups of the same column of the target transposition sequence, and the sorting of the data in the first row of the target transposition sequence; wherein, r adjacent data in the same column of the target transposition sequence are in the same group, and there is no overlapping data between different groups of the target transposition sequence, and r is the radix of the any - radix.
[0389] In an embodiment of the present disclosure, the transposition module is further configured to perform: obtaining an input sequence of any - radix multi - stage butterfly operation; performing a transposition process on the input sequence of the any - radix multi - stage butterfly operation according to the merging and transposition processing rule of the any - radix multi - stage butterfly operation to obtain a target transposition sequence of the any - radix multi - stage butterfly operation; performing the any - radix multi - stage butterfly operation on the target transposition sequence of the any - radix multi - stage butterfly operation to obtain an output sequence of the any - radix multi - stage butterfly operation.
[0390] In an embodiment of the present disclosure, the multiple transposition processing rules include at least two of an input sequence rearrangement rule, an input sequence reshaping rule, an input sequence flipping rule, and an output sequence reshaping rule.
[0391] In an embodiment of the present disclosure, when the radix of the mixed - radix in the DFT decomposition includes a first radix and a second radix, and the next - radix multi - stage butterfly operation of the first - radix multi - stage butterfly operation is the second - radix multi - stage butterfly operation, the output sequence of the first - radix multi - stage butterfly operation and the target transposition sequence of the second - radix multi - stage butterfly operation are composed of the same data;
[0392] The apparatus 100 further includes: a storage module, configured to perform:
[0393] Obtaining the output sequence of the first - radix multi - stage butterfly operation, wherein the output sequence of the first - radix multi - stage butterfly operation includes sub - output sequences at p moments, the c - th sub - output sequence includes q output data at the c - th moment, q is the parallelism of the first - radix multi - stage butterfly operation, p and q are both positive integers, the product of p and q is the DFT points, and c is a positive integer not greater than p;
[0394] Obtaining the target transposition sequence of the second - radix multi - stage butterfly operation, wherein the target transposition sequence of the second - radix multi - stage butterfly operation includes sub - transposition sequences at s moments, the d - th sub - transposition sequence includes t output data at the d - th moment, t is the parallelism of the second - radix multi - stage butterfly operation, s and t are both positive integers, the product of s and t is the DFT points, and d is a positive integer not greater than s;
[0395] Store the q output data at the same moment in any sub-output sequence into different real storage units respectively, and store the t output data at the same moment in any sub-permutation sequence into different real storage units respectively as the storage rule;
[0396] Store the output sequence of the first base multi-stage butterfly operation according to the storage rule.
[0397] In an embodiment of the present disclosure, the apparatus 100 further includes: a reading module configured to perform: reading one data from each of the t real storage units at the d-th moment to read the d-th sub-permutation sequence.
[0398] In an embodiment of the present disclosure, the storage module is further configured to perform: dividing the q real storage units into t virtual storage units according to the storage rule, where each virtual storage unit is jointly composed of one storage space of each of the s real storage units; for any sub-permutation sequence, store the t output data at the same moment in the any sub-permutation sequence into different virtual storage units respectively.
[0399] In an embodiment of the present disclosure, the storage module is further configured to perform: based on the storage rule, store any sub-output sequence into the storage spaces of multiple virtual storage units, which are jointly composed of one storage space of each of the q real storage units, as the division rule; divide the q real storage units into t virtual storage units according to the division rule.
[0400] In an embodiment of the present disclosure, the storage module is further configured to perform: based on the storage rule, the storage space of the d-th address of a virtual storage unit is composed of the storage space of the d-th address of a real storage unit as the division rule.
[0401] In an embodiment of the present disclosure, the storage module is further configured to perform: store the t output data at the d-th moment into the storage spaces of the d-th addresses of the t virtual storage units respectively to store the d-th sub-permutation sequence.
[0402] In an embodiment of the present disclosure, the apparatus 100 further includes: a reading module configured to perform: reading one data from the storage spaces of the d-th addresses of the t virtual storage units at the d-th moment to read the d-th sub-permutation sequence.
[0403] In an embodiment of the present disclosure, the storage module is further configured to perform: store the t output data at the d-th moment into the storage spaces of the d-th addresses of the t real storage units respectively to store the d-th sub-permutation sequence.
[0404] In one embodiment of the present disclosure, the apparatus 100 further includes: a reading module configured to perform: reading a data from the storage space of the d-th address of t true storage units at the d-th moment respectively, so as to read the d-th sub-permutation sequence.
[0405] Regarding the apparatus in the above embodiment, the specific manners in which each module performs operations have been described in detail in the embodiment related to the method, and will not be elaborated herein.
[0406] The processing apparatus for discrete Fourier transform (DFT) provided by the embodiment of the present disclosure decomposes the DFT into multi-level mixed-radix butterfly operations based on the number of DFT points, and determines the parallelism of each multi-level butterfly operation based on at least one of the processing delay of the DFT, the number of ALU resources required for each radix butterfly operation, the number of divisions of the storage resources required for the DFT, and the number of levels of each radix butterfly operation. Thus, the parallelism of each multi-level butterfly operation can be determined in consideration of at least one of the processing delay of the DFT, the number of ALU resources required for each radix butterfly operation, the number of divisions of the storage resources required for the DFT, and the number of levels of each radix butterfly operation, and the flexible setting of the parallelism of different DFT points can be realized, improving the processing performance of the parallel DFT.
[0407] Figure 13 It is a block diagram of an electronic device shown according to an exemplary embodiment.
[0408] As Figure 13 shown, the above-mentioned electronic device 200 includes:
[0409] A memory 210 and a processor 220, a bus 230 connecting different components (including the memory 210 and the processor 220), and the memory 210 stores a computer program, and when the processor 220 executes the program, the processing method for discrete Fourier transform (DFT) described in the embodiment of the present disclosure is implemented.
[0410] The bus 230 represents one or more of several types of bus architectures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus architecture in a variety of bus architectures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0411] The electronic device 200 typically includes a variety of electronically readable media. These media can be any available media that can be accessed by the electronic device 200, including volatile and non-volatile media, removable and non-removable media.
[0412] The memory 210 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 240 and / or cache memory 250. The electronic device 200 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 260 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 13 not shown, commonly referred to as a "hard disk drive"). Although Figure 13 not shown in, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus 230 through one or more data media interfaces. The memory 210 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present disclosure.
[0413] A program / utility 280 having a set (at least one) of program modules 270 can be stored, for example, in the memory 210. Such program modules 270 include - but are not limited to - an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 270 generally perform the functions and / or methods in the embodiments described in the present disclosure.
[0414] The electronic device 200 can also communicate with one or more external devices 290 (such as a keyboard, a pointing device, a display 291, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 200, and / or communicate with any device that enables the electronic device 200 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 292. And, the electronic device 200 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through a network adapter 293. As Figure 13As shown, network adapter 293 communicates with other modules of electronic device 200 via bus 230. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with electronic device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0415] Processor 220 performs various functional applications and data processing by running programs stored in memory 210.
[0416] It should be noted that for the implementation process and technical principle of the electronic device in this embodiment, refer to the foregoing explanation of the processing method of the discrete Fourier transform DFT in the embodiments of the present disclosure, which will not be elaborated here.
[0417] The electronic device provided in the embodiments of the present disclosure can execute the processing method of the discrete Fourier transform DFT as described above. Based on the number of DFT points, the DFT is decomposed into multi-stage mixed-radix butterfly operations. Based on at least one of the processing delay of the DFT, the number of ALU resources required for each radix butterfly operation, the number of divisions of the storage resources required for the DFT, and the number of stages of each radix butterfly operation, the parallelism of each radix multi-stage butterfly operation is determined. Thus, considering at least one of the processing delay of the DFT, the number of ALU resources required for each radix butterfly operation, the number of divisions of the storage resources required for the DFT, and the number of stages of each radix butterfly operation, the parallelism of each radix multi-stage butterfly operation can be determined, and the flexible setting of the parallelism of different DFT points can be realized, improving the processing performance of parallel DFT.
[0418] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the processing method of the discrete Fourier transform DFT provided by the present disclosure are implemented.
[0419] Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0420] Figure 14 is a block diagram of a chip shown according to an exemplary embodiment.
[0421] As Figure 14 shown, the above chip 300 includes one or more interface circuits 302 and one or more processors 301; the interface circuit 302 is used to receive signals and send signals to the processor 301. The signals include computer instructions stored in the memory. When the processor 301 executes the computer instructions, the chip 300 executes the steps of the processing method of the discrete Fourier transform DFT provided by the present disclosure.
[0422] The processor 301 and the interface circuit 302 can be interconnected by a line.
[0423] The chip 300 further includes a memory 303, and all or part of the memory 303 may be outside the chip 300.
[0424] The interface circuit 302 is connected to the memory 303. The interface circuit 302 can be used to receive signals from the memory 303 or other devices, and the interface circuit 302 can be used to send signals to the processor 301 or other devices. For example, the interface circuit 302 can read the instructions stored in the memory 303 and send the instructions to the processor 301.
[0425] The interface circuit 302 can obtain data, program instructions, and / or information, etc. in the internal storage area of the chip 300; it can also obtain data, program instructions, and / or information, etc. outside the chip 300.
[0426] It should be noted that terms such as interface circuit, interface, transceiver pin, transceiver, etc. can be replaced with each other.
[0427] To implement the above embodiments, the present disclosure also provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor of an electronic device, it implements the processing method of the discrete Fourier transform DFT as described above.
[0428] Those skilled in the art will readily think of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure aims to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0429] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A processing method for discrete Fourier transform (DFT), characterized in that, Including: Decompose the DFT into multi-level mixed-radix butterfly operations based on the number of DFT points. Determine the parallelism of each-level multi-radix butterfly operation based on at least one of the processing delay of the DFT, the number of arithmetic logic unit (ALU) resources required for each-radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of levels of each-radix butterfly operation, where the parallelism of any one-level multi-radix butterfly operation is the number of data processed in parallel in the any one-level multi-radix butterfly operation.
2. The method according to claim 1, wherein The determining the parallelism of each-level multi-radix butterfly operation based on at least one of the processing delay of the DFT, the number of arithmetic logic unit (ALU) resources required for each-radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of levels of each-radix butterfly operation includes: Obtain a first correlation relationship between the processing delay of the DFT and the parallelism of each-level multi-radix butterfly operation. Take the processing delay of the DFT being less than or equal to a first set threshold as a first constraint condition. Based on the first correlation relationship and the first constraint condition, determine the parallelism of each-level multi-radix butterfly operation.
3. The method according to claim 2, characterized in that, The obtaining the first correlation relationship between the processing delay of the DFT and the parallelism of each-level multi-radix butterfly operation includes: Based on the number of DFT points, construct a first function between the fetch delay of any one-level multi-radix butterfly operation and the parallelism of the any one-level multi-radix butterfly operation. Based on the first function corresponding to each-level multi-radix butterfly operation, construct a second function between the fetch delay of the DFT and the parallelism of each-level multi-radix butterfly operation. Based on the second function, construct a third function between the processing delay of the DFT and the parallelism of each-level multi-radix butterfly operation as the first correlation relationship.
4. The method according to claim 1, wherein The determining the parallelism of each-level multi-radix butterfly operation based on at least one of the processing delay of the DFT, the number of arithmetic logic unit (ALU) resources required for each-radix butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of levels of each-radix butterfly operation includes: Based on the number of ALU resources required for any one-radix butterfly operation, obtain a second correlation relationship between the total number of ALU resources required for the any one-level multi-radix butterfly operation and the parallelism of the any one-level multi-radix butterfly operation. Take the minimum difference between the total quantities corresponding to different-level multi-radix butterfly operations as an optimization objective. Based on the second correlation relationship corresponding to each-level multi-radix butterfly operation and the optimization objective, determine the parallelism of each-level multi-radix butterfly operation.
5. The method according to claim 4, characterized in that, The obtaining the second correlation relationship between the total number of ALU resources required for any one-level multi-radix butterfly operation and the parallelism of the any one-level multi-radix butterfly operation based on the number of ALU resources required for any one-radix butterfly operation includes: Based on the radix of any one level, construct a fourth function between the number of parallel operations of the any one-level multi-radix butterfly operation and the parallelism of the any one-level multi-radix butterfly operation, where the number of parallel operations of the any one-level multi-radix butterfly operation is the number of parallel butterfly operations of the any one-level multi-radix butterfly operation. Construct a fifth function between the total number of ALU resources required for any base multi - stage butterfly operation and the parallelism of any base multi - stage butterfly operation based on the number of ALU resources required for any base butterfly operation and the fourth function, as the second correlation.
6. The method according to claim 1, wherein Determining the parallelism of each base multi - stage butterfly operation based on at least one of the DFT - based processing delay, the number of arithmetic logic unit (ALU) resources required for each base butterfly operation, the number of partitions of the storage resources required for DFT, and the number of stages of each base butterfly operation includes: Obtain a third correlation between the number of partitions of the storage resources required for DFT and the first maximum value of the parallelism of each base multi - stage butterfly operation, where the number of partitions of the storage resources required for DFT is greater than or equal to the first maximum value; Take the number of partitions of the storage resources required for DFT being less than or equal to a second set threshold as the second constraint condition; Determine the parallelism of each base multi - stage butterfly operation based on the third correlation and the second constraint condition.
7. The method according to claim 1, characterized in that Determining the parallelism of each base multi - stage butterfly operation based on at least one of the DFT - based processing delay, the number of arithmetic logic unit (ALU) resources required for each base butterfly operation, the number of partitions of the storage resources required for DFT, and the number of stages of each base butterfly operation includes: Obtain the maximum value of the parallelism of any base multi - stage butterfly operation; Determine the parallelism of the first i - stage butterfly operation of any base to be the maximum value of the parallelism of any base multi - stage butterfly operation; Determine that the parallelism of the last j - stage butterfly operation of any base is less than or equal to the maximum value of the parallelism of any base multi - stage butterfly operation, where i and j are both non - negative integers, and the sum of i and j is the number of stages of any base butterfly operation.
8. The method according to claim 7, wherein When the radix of the mixed base in DFT decomposition includes base 2, base 3, and base 5, the method further includes: Obtain the second maximum value of the parallelism of the base 3 multi - stage butterfly operation; Determine the parallelism of the non - last - stage butterfly operation of base 3 to be the second maximum value; Determine that the parallelism of the last - stage butterfly operation of base 3 is less than the second maximum value.
9. The method according to claim 1, wherein Determining the parallelism of each base multi - stage butterfly operation based on at least one of the DFT - based processing delay, the number of arithmetic logic unit (ALU) resources required for each base butterfly operation, the number of partitions of the storage resources required for DFT, and the number of stages of each base butterfly operation includes: Determine the value range of the parallelism of each base multi - stage butterfly operation based on at least one of the DFT - based processing delay, the number of ALU resources required for each base butterfly operation, the number of partitions of the storage resources required for DFT, and the number of stages of each base butterfly operation; Determine the parallelism of any base multi - stage butterfly operation from the value range of the parallelism of any base multi - stage butterfly operation.
10. The method according to claim 9, characterized in that, When the radix of the mixed base in DFT decomposition includes base 2, the value range of the parallelism of the base 2 multi - stage butterfly operation includes at least one of 6, 8, 12, and 16; When the radix of the mixed base in DFT decomposition includes base 3, the value range of the parallelism of the base 3 multi - stage butterfly operation includes at least one of 3, 6, 9, 12, and 18; When the radix of the mixed radix in the DFT decomposition includes radix 5, the value range of the parallelism of the radix-5 multi-stage butterfly operation includes 10 and / or 15.
11. The method according to claim 9, wherein Determining the parallelism of any radix multi-stage butterfly operation from the value range of the parallelism of any radix multi-stage butterfly operation includes: Based on at least one of the DFT points, the number of stages of at least one radix butterfly operation, and the parallelism of the multi-stage butterfly operations of the other radixes except any radix, determining the parallelism of the multi-stage butterfly operation of any radix from the value range of the parallelism of the multi-stage butterfly operation of any radix.
12. The method according to claim 11, wherein When the radix of the mixed radix in the DFT decomposition includes radix 2, and the value range of the parallelism of the radix-2 multi-stage butterfly operation includes at least one of 6, 8, 12, and 16; The method further includes: Taking the number of stages of the radix-2 butterfly operation being 1 as the first condition; Taking the number of stages of the radix-2 butterfly operation being greater than 3 as the second condition; Taking the DFT points being 3240 as the third condition; The determining the parallelism of the multi-stage butterfly operation of any radix from the value range of the parallelism of the multi-stage butterfly operation of any radix based on at least one of the DFT points, the number of stages of at least one radix butterfly operation, and the parallelism of the multi-stage butterfly operations of the other radixes except any radix includes at least one of the following: If the first condition is satisfied, determining the parallelism of the radix-2 multi-stage butterfly operation to be 6; If the second condition is satisfied, determining the third maximum value of the parallelism of the radix-2 multi-stage butterfly operation to be 16; If the third condition is satisfied, determining the third maximum value to be 8; If the first condition, the second condition, and the third condition are not satisfied, determining the third maximum value to be 12.
13. The method according to claim 11, wherein When the radix of the mixed radix in the DFT decomposition includes radix 3, and the value range of the parallelism of the radix-3 multi-stage butterfly operation includes at least one of 3, 6, 9, 12, and 18; The method further includes: Taking any one of the DFT points being 600, 750, and 2400 as the fourth condition; Taking the number of stages of both the radix-2 butterfly operation and the radix-3 butterfly operation being 1 as the fifth condition; Taking the DFT points being 1200 or 3000 as the sixth condition; Taking the number of stages of the radix-3 butterfly operation being 1 as the seventh condition; Taking the DFT points being 2916 or 3240 as the eighth condition; Taking the number of stages of the radix-3 butterfly operation being greater than 1 as the ninth condition; The determining the parallelism of the multi-stage butterfly operation of any radix from the value range of the parallelism of the multi-stage butterfly operation of any radix based on at least one of the DFT points, the number of stages of at least one radix butterfly operation, and the parallelism of the multi-stage butterfly operations of the other radixes except any radix includes at least one of the following: If the fourth condition is satisfied, determining the parallelism of the radix-3 multi-stage butterfly operation to be 3; If the fifth condition is satisfied and the DFT points are not 750, determining the second maximum value of the parallelism of the radix-3 multi-stage butterfly operation to be 6; If the sixth condition is satisfied, determining the second maximum value to be 6; If the seventh condition is satisfied and the fourth condition, the fifth condition, and the sixth condition are not satisfied, determining the second maximum value to be 12; If the eighth condition is satisfied, determining the second maximum value to be 18; If the ninth condition is satisfied and the eighth condition is not satisfied, determine that the second maximum value of the second is 9.
14. The method according to claim 11, wherein In the case where the radix of the mixed radix of the DFT decomposition includes radix 5, and the value range of the parallelism of the radix 5 multi-stage butterfly operation includes 10 and / or 15; The method further includes: Taking the DFT number of points being 2916 or 3240 as the eighth condition; Taking the number of stages of the radix 3 butterfly operation being greater than 1 as the ninth condition; Based on at least one of the DFT number of points, the number of stages of at least one radix butterfly operation, and the parallelism of the multi-stage butterfly operation of the remaining radices other than any one radix, determining the parallelism of the multi-stage butterfly operation of any one radix from the value range of the parallelism of the multi-stage butterfly operation of any one radix includes at least one of the following: If the ninth condition is satisfied and the eighth condition is not satisfied, determine that the parallelism of the radix 5 multi-stage butterfly operation is 10; If the ninth condition is not satisfied, or if the eighth condition is satisfied, determine that the fourth maximum value of the parallelism of the radix 5 multi-stage butterfly operation is 15.
15. The method according to claim 11, wherein In the case where the radix of the mixed radix of the DFT decomposition includes radix 3 and radix 5, and the value range of the parallelism of the radix 5 multi-stage butterfly operation includes 10 and / or 15; The method further includes: Taking the second maximum value of the parallelism of the radix 3 multi-stage butterfly operation being 9 as the tenth condition; Based on at least one of the DFT number of points, the number of stages of at least one radix butterfly operation, and the parallelism of the multi-stage butterfly operation of the remaining radices other than any one radix, determining the parallelism of the multi-stage butterfly operation of any one radix from the value range of the parallelism of the multi-stage butterfly operation of any one radix includes at least one of the following: If the tenth condition is satisfied, determine that the parallelism of the radix 5 multi-stage butterfly operation is 10; If the tenth condition is not satisfied, determine that the fourth maximum value of the parallelism of the radix 5 multi-stage butterfly operation is 15.
16. The method according to any one of claims 1 to 15, characterized in that, The method further includes: Based on the DFT number of points, obtaining multiple reordering processing rules for the multi-stage butterfly operation of any one radix; Based on the multiple reordering processing rules for the multi-stage butterfly operation of any one radix, obtaining the combined reordering processing rule for the multi-stage butterfly operation of any one radix; Storing the combined reordering processing rule for the multi-stage butterfly operation of each radix.
17. The method according to claim 16, characterized in that, The combined reordering processing rule for the multi-stage butterfly operation of any one radix includes at least one of the DFT number of points, the number of points of the multi-stage butterfly operation of any one radix, the sorting difference between two adjacent data in the same column and the same group of the target reordering sequence of the multi-stage butterfly operation of any one radix, the sorting difference between two adjacent data in different groups in the same column of the target reordering sequence, and the sorting of the data in the first row of the target reordering sequence; where The adjacent r data in the same column of the target reordering sequence are in the same group, and there is no overlapping data between different groups of the target reordering sequence, and r is the radix of any one radix.
18. The method according to claim 16, wherein The method further includes: Obtaining the input sequence of the multi-stage butterfly operation of any one radix; According to the combined reordering processing rule for the multi-stage butterfly operation of any one radix, performing a single reordering process on the input sequence of the multi-stage butterfly operation of any one radix to obtain the target reordering sequence of the multi-stage butterfly operation of any one radix; Performing any one of the base multi - stage butterfly operations on the target re - ordering sequence of any one of the base multi - stage butterfly operations to obtain the output sequence of any one of the base multi - stage butterfly operations.
19. The method according to claim 16, wherein The multiple re - ordering processing rules include at least two of an input sequence rearrangement rule, an input sequence reshaping rule, an input sequence flipping rule, and an output sequence reshaping rule.
20. The method according to any one of claims 1-15, characterized in that, In the case where the bases of the mixed - base in the DFT decomposition include a first base and a second base, and the next - stage multi - stage butterfly operation of the first - base multi - stage butterfly operation is the second - base multi - stage butterfly operation, the output sequence of the first - base multi - stage butterfly operation and the target re - ordering sequence of the second - base multi - stage butterfly operation are composed of the same data; The method further includes: Obtaining the output sequence of the first - base multi - stage butterfly operation, where the output sequence of the first - base multi - stage butterfly operation includes sub - output sequences at p time instants, the c - th sub - output sequence includes q output data at the c - th time instant, q is the parallelism of the first - base multi - stage butterfly operation, both p and q are positive integers, the product of p and q is the DFT point number, and c is a positive integer not greater than p; Obtaining the target re - ordering sequence of the second - base multi - stage butterfly operation, where the target re - ordering sequence of the second - base multi - stage butterfly operation includes sub - re - ordering sequences at s time instants, the d - th sub - re - ordering sequence includes t output data at the d - th time instant, t is the parallelism of the second - base multi - stage butterfly operation, both s and t are positive integers, the product of s and t is the DFT point number, and d is a positive integer not greater than s; Storing the q output data at the same time instant in any one sub - output sequence into different real storage units respectively, and storing the t output data at the same time instant in any one sub - re - ordering sequence into different real storage units respectively as the storage rule; Storing the output sequence of the first - base multi - stage butterfly operation according to the storage rule.
21. The method according to claim 20, wherein After storing the output sequence of the first - base multi - stage butterfly operation, it further includes: Reading one data from t real storage units respectively at the d - th time instant to read the d - th sub - re - ordering sequence.
22. The method according to claim 20, wherein The storing the output sequence of the first - base multi - stage butterfly operation according to the storage rule includes: Dividing q real storage units into t virtual storage units according to the storage rule, where the virtual storage unit is jointly composed of one storage space of each of s real storage units; For any one sub - re - ordering sequence, storing the t output data at the same time instant in the any one sub - re - ordering sequence into different virtual storage units respectively.
23. The method according to claim 22, wherein The dividing q real storage units into t virtual storage units according to the storage rule includes: Based on the storage rule, storing any one sub - output sequence into the storage spaces of multiple virtual storage units, which are jointly composed of one storage space of each of q real storage units, as the division rule; Dividing q real storage units into t virtual storage units according to the division rule.
24. The method according to claim 23, wherein The method further includes: Based on the storage rule, the storage space of the d - th address of a virtual storage unit is composed of the storage space of the d - th address of a real storage unit as the division rule.
25. A processing device for discrete Fourier transform (DFT), characterized in that, Including: A decomposition module, configured to perform a decomposition of a DFT into a multi-stage mixed-radix butterfly operation based on the number of DFT points; A determination module, configured to determine the parallelism of each multi-stage butterfly operation based on at least one of the processing delay of the DFT, the number of arithmetic logic unit (ALU) resources required for each butterfly operation, the number of partitions of the storage resources required for the DFT, and the number of stages of each butterfly operation, wherein the parallelism of any multi-stage butterfly operation is the number of data processed in parallel in the any multi-stage butterfly operation.
26. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: Implement the steps of the method according to any one of claims 1-24.
27. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the method according to any one of claims 1-24 are implemented.
28. A chip, characterized in that, Comprising one or more interface circuits and one or more processors; the interface circuit is used to receive a signal and send the signal to the processor, the signal includes computer instructions stored in the memory, and when the processor executes the computer instructions, the chip is caused to execute the steps of the method according to any one of claims 1-24.