FFT calculation method and device based on intelligent processor

Through the FFT calculation method based on the intelligent processor, the convolution operator and weight matrix table are used to perform convolution calculation of the signal sequence, which solves the problem of low efficiency of the traditional FFT algorithm on FPGA and DSP processors and realizes efficient and flexible FFT calculation.

CN120744290APending Publication Date: 2025-10-03TIANJIN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510809336.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional FFT algorithms are difficult to implement efficiently in vector or matrix calculations on FPGA and DSP processors, and multi-level iterations lead to error accumulation and low computational efficiency.

Method used

The FFT calculation method based on intelligent processor is adopted to perform convolution calculation of signal sequence through preset weight matrix table and convolution operator, which is simplified to simple convolution operation and utilizes the parallel computing capability of modern processors.

Benefits of technology

It significantly improves the FFT calculation efficiency, reduces the calculation complexity and resource consumption, adapts to signal sequences of different lengths and types, and is easy to integrate into existing signal processing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744290A_ABST
    Figure CN120744290A_ABST
Patent Text Reader

Abstract

The invention relates to an FFT (Fast Fourier Transform) calculation method and device based on an intelligent processor. The method comprises the following steps: inputting a signal sequence; searching and obtaining a weight matrix corresponding to the signal sequence according to the sequence length of the signal sequence through a preset weight matrix table; calling a convolution operator and setting convolution parameters; and performing convolution calculation on the signal sequence and the weight matrix through the convolution operator to obtain a calculation result. According to the method, the calculation efficiency is remarkably improved, and the calculation complexity and the implementation difficulty are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent computing technology, and in particular to an FFT computing method and device based on an intelligent processor. Background Art

[0002] FFT (Fast Fourier Transform) is a key algorithm in the field of digital signal processing and is widely used in scenarios such as radar signal processing, communication signal processing, and data analysis. The performance of the FFT algorithm is a key factor in many application scenarios. Traditional FFT algorithms, limited by the characteristics of processors such as FPGAs and DSPs, have difficulty implementing vector or matrix calculations efficiently. They often rely on scalar calculations, parallel multi-stage butterfly operations, or multi-stage pipeline butterfly operations to implement large-scale FFT calculations. This multi-stage operation method, on the one hand, will continuously accumulate errors, ultimately causing FFT noise to increase with the number of FFT points. On the other hand, the multi-stage iterative FFT calculation method will significantly reduce FFT computational efficiency.

[0003] Therefore, there is a need in the art for an FFT algorithm that improves computational efficiency. Summary of the Invention

[0004] The present invention provides an FFT calculation method and device based on an intelligent processor, which are used to solve the defects of the prior art.

[0005] The present invention provides an FFT calculation method based on an intelligent processor, comprising: S1: input signal sequence; S2: Obtaining a weight matrix corresponding to the signal sequence according to the sequence length of the signal sequence through a preset weight matrix table; S3: Call the convolution operator and set the convolution parameters; S4: performing convolution calculation on the signal sequence and the weight matrix by the convolution operator to obtain a calculation result.

[0006] According to an FFT calculation method based on an intelligent processor provided by the present invention, the expression of the weight matrix in step S2 is:

[0007] in, is the sequence length of the signal sequence, The length of the input sequence is The signal sequence, Is an imaginary unit.

[0008] According to an FFT calculation method based on an intelligent processor provided by the present invention, step S2 further includes: S21: Receive the signal sequence, and write the signal sequence into a first memory; S22: Obtaining the sequence length of the signal sequence; S23: according to the sequence length, searching a weight matrix table to obtain a weight matrix corresponding to the signal sequence; S24: Writing the weight matrix into the second memory.

[0009] According to an intelligent processor-based FFT calculation method provided by the present invention, the first memory is a non-volatile random access memory, and the second memory is a write random access memory.

[0010] According to an FFT calculation method based on an intelligent processor provided by the present invention, the convolution parameters in step S3 include: the number of convolution channels, the convolution kernel size, the convolution horizontal stride, and the convolution vertical stride.

[0011] According to an FFT calculation method based on an intelligent processor provided by the present invention, the number of convolution channels corresponds to the dimensions of the signal sequence and the weight matrix, and the size of the convolution kernel corresponds to the size of the weight matrix.

[0012] According to the FFT calculation method based on an intelligent processor provided by the present invention, the calculation result output in step S4 is stored in the first memory.

[0013] A second aspect of the present invention provides an FFT calculation device based on an intelligent processor, comprising: Input module: used to receive input signal sequence; Search module: used to search and obtain the weight matrix corresponding to the signal sequence according to the sequence length of the signal sequence through a preset weight matrix table; Call module: used to call the convolution operator and set the convolution parameters; An intelligent processor, wherein the intelligent processor includes a convolution operator, and when the convolution operator is called by the calling module, is used to perform a convolution calculation on the signal sequence and the weight matrix to obtain a calculation result; Storage module: used to store the signal sequence and the weight matrix so that the intelligent processor can obtain the signal sequence and the weight matrix for convolution calculation, and also used to store the calculation results output by the intelligent processor.

[0014] A third aspect of the present invention provides an FFT calculation device based on an intelligent processor, comprising: a memory and at least one processor, wherein instructions are stored in the memory; At least one of the processors calls the instructions in the memory to enable the FFT calculation device based on the intelligent processor to execute any one of the above FFT calculation methods based on the intelligent processor.

[0015] A fourth aspect of the present invention provides a computer-readable storage medium having instructions stored thereon, wherein the instructions, when executed by a processor, implement the FFT calculation method based on an intelligent processor as described in any one of the above.

[0016] The present invention provides an FFT calculation method, device, equipment and storage medium based on an intelligent processor. By utilizing the convolution operator in the intelligent processor to perform FFT calculation, the parallel computing capability and optimization algorithm of modern processors can be fully utilized. Compared with the traditional FFT calculation method, the computing efficiency can be significantly improved, especially when processing large-scale signal sequences. Traditional FFT calculation involves complex mathematical transformations and iterative processes. The present invention simplifies the FFT calculation process into a simple convolution operation through a preset weight matrix table and convolution calculation, thereby reducing the complexity of the calculation and the difficulty of implementation. In the process, the present invention also flexibly adapts to signal sequences of different lengths and types by adjusting the convolution parameters, and has strong scalability and adaptability. Moreover, since the convolution calculation has been widely optimized in modern processors, the use of the convolution operator for FFT calculation can reduce the consumption of computing resources. Combined with the hardware of the present invention, the method of the present invention can be easily integrated into the existing signal processing system without the need for large-scale modification of the system, reducing the cost of deployment and integration, and being easy to implement. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A schematic diagram of the flow of an FFT calculation method based on an intelligent processor provided by the present invention; Figure 2 A schematic diagram of the structure of an FFT calculation device based on an intelligent processor provided by the present invention; Figure 3 A schematic diagram of the performance of the FFT calculation method based on the intelligent processor provided by the present invention.

[0019] Reference numerals: 100, input module; 200, search module; 300, call module; 400, intelligent processor; 500, storage module. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the embodiments described are part of the embodiments of the present invention, not all of the embodiments, and they should not be understood as limitations on the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0021] The following describes embodiments of the present invention with reference to the accompanying drawings.

[0022] like Figure 1 As shown, the present invention provides an FFT calculation method based on an intelligent processor, comprising: S1: Input signal sequence.

[0023] In step S1 , a signal sequence is first received. The sequence may be an audio signal, image data or other forms of digital information, ie, input data for subsequent Fourier transformation.

[0024] S2: Obtain a weight matrix corresponding to the signal sequence by searching according to the sequence length of the signal sequence through a preset weight matrix table.

[0025] The weight matrix in step S2 is expressed as follows:

[0026] in, is the sequence length of the signal sequence, The length of the input sequence is The signal sequence, Is an imaginary unit.

[0027] by For example, Time ,Right now The weight matrix expression is:

[0028] in,

[0029] Wherein, step S2 further includes: S21: Receive the signal sequence, and write the signal sequence into a first memory.

[0030] S22: Acquire the sequence length of the signal sequence.

[0031] S23: According to the sequence length, a weight matrix corresponding to the signal sequence is obtained by searching a weight matrix table.

[0032] S24: Writing the weight matrix into the second memory.

[0033] In steps S21 to S24, the signal sequence from S1 is first received, and then the length of the signal sequence is determined. The length information is the key to finding the corresponding weight matrix. Subsequently, the sequence length is used as an index to find and obtain the corresponding weight matrix in the preset weight matrix table. Finally, the found weight matrix is ​​stored in the WRAM for subsequent calls.

[0034] The first memory is a non-volatile random access memory, and the second memory is a write random access memory.

[0035] Furthermore, non-volatile random access memory (NVRAM) is non-volatile, meaning that stored data will not be lost even if power is lost. It is primarily used to store data or configuration information that needs to be preserved for a long time. In the matrix multiplication scenario of the present invention, NVRAM is selected to store the transposed matrix so that it can be quickly accessed when needed while ensuring data persistence. Random access memory (RAM) is a volatile memory, meaning that data will be lost after power failure, but it has advantages such as fast read and write speeds and large capacity, and is therefore used to store temporary data being processed. In the scenario of the present invention, the second matrix is ​​stored in RAM so that it can be quickly accessed and modified during the calculation process.

[0036] S3: Call the convolution operator and set the convolution parameters.

[0037] Among them, the convolution parameters in step S3 include: the number of convolution channels, the convolution kernel size, the convolution horizontal stride, and the convolution vertical stride.

[0038] The number of convolution channels corresponds to the dimensions of the signal sequence and the weight matrix, and the size of the convolution kernel corresponds to the size of the weight matrix.

[0039] The number of convolution channels usually corresponds to the dimensions of the input signal sequence and the weight matrix. In multidimensional signal processing (such as color image processing), each dimension can be regarded as an independent signal channel and convolution operations are performed separately. Therefore, setting the number of convolution channels can ensure that the convolution operation can process all relevant dimensions of the input signal, thereby extracting complete feature information.

[0040] The convolution kernel size corresponds to the size of the weight matrix. The convolution kernel is a small matrix used to perform a sliding window operation on the input signal to extract local features. The size of the convolution kernel determines the size of the input signal area covered by each operation. By adjusting the convolution kernel size, you can control the local sensitivity of the convolution operation to the input signal. Larger convolution kernels can capture broader contextual information, while smaller convolution kernels focus more on detailed features.

[0041] The horizontal stride and vertical stride of convolution respectively define the step size of the convolution kernel moving horizontally and vertically on the input signal. These two parameters determine the spatial resolution of the output signal after the convolution operation. By adjusting the stride parameters, the size of the output signal and the granularity of feature extraction can be controlled. A larger stride will result in a smaller output signal size, which may lose some detail information, but can improve computational efficiency. A smaller stride can retain more detail information, but may increase the amount of computation.

[0042] S4: performing convolution calculation on the signal sequence and the weight matrix by the convolution operator to obtain a calculation result.

[0043] The calculation result output in step S4 is stored in the first memory.

[0044] Convolution is a mathematical operation commonly used in fields such as signal processing, image processing, and machine learning. In signal processing, the convolution operation involves point-by-point multiplication and summation of a signal (or sequence) with another signal called a convolution kernel or filter. This process can be regarded as a form of weighted summation of the input signal to extract features or perform transformations.

[0045] In the FFT calculation method of the present invention, a convolution operator is used to perform convolution calculation on an input signal sequence and a weight matrix, and the convolution operator can efficiently perform convolution operations on an intelligent processor.

[0046] The convolution calculation process is to multiply the signal sequence and the weight matrix point by point and sum them. Specifically, each row (or column) of the weight matrix is ​​convolved with the signal sequence to generate a transformed output sequence. After the convolution calculation is completed, the output is a complex sequence with the same length as the input signal sequence. The sequence contains the representation of the original signal in the frequency domain, that is, the result of the FFT transform. This result can be used for subsequent spectrum analysis, filtering or other signal processing tasks.

[0047] Generally speaking, the process of the present invention is: first pre-calculate the weight matrices of different sizes, and construct a lookup table of the weight matrix ;Write the input signal sequence into the NRAM of the intelligent processor;According to the length of the input sequence , from the lookup table Read out the corresponding weight matrix , and store it in WRAM; call the CONV convolution operator of the intelligent processor to convert the data in WRAM and NRAM and signal sequence as its input; reasonably set the parameters such as the number of channels k, convolution kernel height m, convolution kernel width l, convolution horizontal stride p, and convolution vertical stride w of CONV convolution calculation to make them consistent with the input of CONV operator. Match the dimension of the input sequence; store the output of the CONV operator in NRAM and use it as The FFT calculation results of .

[0048] like Figure 2 As shown, the present invention also provides an FFT calculation device based on an intelligent processor, comprising: Input module 100: used to receive an input signal sequence; Search module 200: configured to search for a weight matrix corresponding to the signal sequence according to the sequence length of the signal sequence through a preset weight matrix table; Calling module 300: used to call the convolution operator and set convolution parameters; An intelligent processor 400, wherein the intelligent processor 400 includes a convolution operator. When the convolution operator is called by the calling module 300, the convolution operator is used to perform a convolution calculation on the signal sequence and the weight matrix to obtain a calculation result; Storage module 500: used to store the signal sequence and the weight matrix so that the calling module 300 can obtain the signal sequence and the weight matrix for convolution calculation, and also used to store the calculation results output by the intelligent processor 400.

[0049] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0050] The present invention also provides an FFT calculation device based on an intelligent processor, comprising: a memory and at least one processor, wherein instructions are stored in the memory; At least one of the processors calls the instructions in the memory to enable the FFT calculation device based on the intelligent processor to execute any one of the above FFT calculation methods based on the intelligent processor.

[0051] The present invention also provides a computer-readable storage medium having instructions stored thereon, and when the instructions are executed by a processor, the FFT calculation method based on an intelligent processor as described in any one of the above items is implemented.

[0052] Furthermore, the FFT computing device based on an intelligent processor provided by the present invention may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU), for example, one or more processors and memories, one or more storage media for storing applications or data, such as one or more massive storage devices, wherein the memories and storage media may be short-term storage or persistent storage, and the program stored in the storage medium may include one or more modules, each module may include a series of instruction operations in the FFT computing device based on the intelligent processor, and further, the processor may be configured to communicate with the storage medium and execute a series of instruction operations in the storage medium on the FFT computing device based on the intelligent processor.

[0053] It may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input and output interfaces, and one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that the intelligent processor-based FFT calculation structure provided in the present invention does not limit the device, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0054] like Figure 3As shown, in a specific embodiment, the present invention divides 2048-point FFT into 11 levels of operations, and the performance of implementing FFT on the MLU220 processor is compared with that of DSP2146. The running time can be improved. At the same time, when the FFT scale is 4096*4096, the running speed is 0.095ms, when the FFT scale is 2048*2048, the running speed is 0.018ms, when the FFT scale is 1024*1024, the running speed is 0.005ms, and when the FFT scale is 512*512, the running speed is 0.002ms. In comparison, FFT calculations of different scales in the intelligent processor are all implemented by only one level of calculation. Therefore, the FFT calculation accuracy does not change with the increase of the FFT scale. In addition, since there is only one level of operation, there is no error accumulation.

[0055] In another specific embodiment, the present invention runs the FFT calculation method based on the intelligent processor on the MLU220 intelligent processor, and the results are shown in Table 1.

[0056] Table 1 Comparison of the results of the algorithm of the present invention when performing FFT calculations on an MLU processor and an FPGA

[0057] As can be seen from Table 1, the FFT scale calculated in this embodiment is 2048 points, and the data bit number is 8 bits. That is, when performing the FFT calculation, 2048 data points will be processed, each data point occupies 8 bits. The speed of running the FFT calculation method on the MLU220 intelligent processor. In this example, the MLU running speed is 5.5us, while the speed of running the same scale FFT calculation on the FPGA platform is 27.43us. In comparison, the FFT calculation method based on the intelligent processor of the present invention has a calculation speed increased by 4 times when completing the same scale FFT, which is more efficient.

[0058] The present invention provides an FFT calculation method, device, equipment and storage medium based on an intelligent processor. By using the convolution operator of the intelligent processor to perform FFT calculation, the complex iteration and mathematical transformation process in the traditional FFT algorithm is simplified to a convolution operation. The method can make full use of the optimization of convolution operation by modern processors and significantly improve the calculation efficiency of FFT, especially when processing large-scale signal sequences. The present invention can dynamically search and apply the corresponding weight matrix according to the length of the input signal sequence through a preset weight matrix table, so that the present invention can flexibly adapt to signal sequences of different lengths and has strong scalability. At the same time, by adjusting the convolution parameters (such as the number of convolution channels, the size of the convolution kernel, etc.), the calculation performance can be further optimized to meet the needs of different application scenarios. The present invention uses a non-volatile random access memory (N The present invention uses a write random access memory (VRAM) to store the input signal sequence, ensuring that data will not be lost in the event of a power outage. At the same time, the write random access memory (WRAM) is used to store the weight matrix and convolution calculation results to improve the reading and writing speed and computing efficiency, which not only ensures the persistence of data but also improves the overall performance of the system. The traditional FFT algorithm is relatively complex to implement and requires processing a large number of mathematical transformations and iterative processes. However, the present invention implements FFT through convolution operations, reducing the implementation complexity and making the FFT algorithm easier to implement and optimize on hardware platforms such as intelligent processors. Since FFT has a wide range of applications in signal processing, spectrum analysis, filtering and other fields, the present invention's FFT calculation method based on an intelligent processor can be applied to multiple fields such as audio processing, image processing, and communication system design, providing an efficient and accurate FFT calculation tool.

[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An FFT calculation method based on an intelligent processor, characterized in that: include: S1: input signal sequence; S2: Obtaining a weight matrix corresponding to the signal sequence according to the sequence length of the signal sequence through a preset weight matrix table; S3: Call the convolution operator and set the convolution parameters; S4: performing convolution calculation on the signal sequence and the weight matrix by the convolution operator to obtain a calculation result.

2. The FFT calculation method based on an intelligent processor according to claim 1, characterized in that: The expression of the weight matrix in step S2 is: in, is the sequence length of the signal sequence, The length of the input sequence is The signal sequence, Is an imaginary unit.

3. The FFT calculation method based on an intelligent processor according to claim 1, characterized in that: Step S2 further comprises: S21: Receive the signal sequence, and write the signal sequence into a first memory; S22: Obtaining the sequence length of the signal sequence; S23: according to the sequence length, searching a weight matrix table to obtain a weight matrix corresponding to the signal sequence; S24: Writing the weight matrix into the second memory.

4. The FFT calculation method based on an intelligent processor according to claim 1, characterized in that: The first memory is a non-volatile random access memory, and the second memory is a write random access memory.

5. The FFT calculation method based on an intelligent processor according to claim 1, characterized in that: The convolution parameters in step S3 include: the number of convolution channels, the convolution kernel size, the convolution horizontal stride, and the convolution vertical stride.

6. The FFT calculation method based on an intelligent processor according to claim 5, characterized in that: The number of convolution channels corresponds to the dimensions of the signal sequence and the weight matrix, and the size of the convolution kernel corresponds to the size of the weight matrix.

7. The FFT calculation method based on an intelligent processor according to claim 1, characterized in that: The calculation result output in step S4 is stored in the first memory.

8. An FFT calculation device based on an intelligent processor, characterized in that: include: Input module: used to receive input signal sequence; Search module: used to search and obtain the weight matrix corresponding to the signal sequence according to the sequence length of the signal sequence through a preset weight matrix table; Call module: used to call the convolution operator and set the convolution parameters; An intelligent processor, wherein the intelligent processor includes a convolution operator, and when the convolution operator is called by the calling module, is used to perform a convolution calculation on the signal sequence and the weight matrix to obtain a calculation result; Storage module: used to store the signal sequence and the weight matrix so that the intelligent processor can obtain the signal sequence and the weight matrix for convolution calculation, and also used to store the calculation results output by the intelligent processor.

9. An FFT computing device based on an intelligent processor, characterized in that: include: a memory and at least one processor, wherein instructions are stored in the memory; At least one of the processors calls the instructions in the memory to enable the FFT calculation device based on the intelligent processor to execute the FFT calculation method based on the intelligent processor according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed by the processor, the FFT calculation method based on the intelligent processor according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Winograd convolution operation acceleration method and acceleration module

    CN113283587A

  • Convolution calculation method and device based on photon calculation chip

    CN114723019A

  • Hardware implementation of discrete Fourier correlation transforms

    CN115994565A

  • Voice signal processing method and device and readable storage medium

    CN116741202A