Multi-core DSP-FFT processing method applied to radar signal processing
Through the multi-core DSP-FFT processing method, the FFT operation process is optimized, the problem of long data processing delay is solved, real-time processing under high sampling rate conditions is achieved, and the computing performance and stability of DSP are improved.
Patent Information
- Application Number
- CN202510833954.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-17
AI Technical Summary
The existing technology has a long data processing delay in FFT operations and cannot meet the real-time processing requirements under high sampling rate conditions. In addition, the module processing capacity becomes a limiting factor when processing large-scale data, leading to performance bottlenecks.
Adopting multi-core DSP-FFT processing mode, short data or long data FFT operation is performed by judging the data length, and multi-core CPU is used to allocate and cache resources, and the calculation pipeline is optimized, including the specific step design of short data FFT operation and long data FFT operation.
It improves the pipeline performance of data processing, fully utilizes the multi-core resources of DSP, reduces the control complexity, ensures the stability and parallelism of the calculation process, and meets the real-time processing requirements.
Smart Images

Figure CN120804482A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing technology, and in particular to a multi-core DSP-FFT processing method applied to radar signal processing. Background Art
[0002] DSP chips are widely used in fields requiring digital signal processing, such as aviation, aerospace, and radar. Among them, FFT calculation is the most time-consuming processing method in the overall data processing.
[0003] In FFT operations, due to the continuous input data stream, the DSP core typically copies the input data into memory before performing the FFT calculation. The FFT calculation is performed after completing one data accumulation cycle, and the next FFT and other signal processing calculations are performed after the second data accumulation cycle.
[0004] like Figure 1 The processing method shown, primarily a pipeline structure, allows the DSP's processor core to operate under load. However, due to the large amount of data to be processed and the long processing latency, it cannot meet the real-time processing requirements of current high sampling rates. The complex multiplication and addition operations in the FFT calculation will scale logarithmically with the length, significantly increasing the DSP's cycle and time for processing data in the slow clock domain. Beyond a certain threshold, it will no longer be able to complete processing within the current accumulation cycle, severely impacting radar detection accuracy and response time.
[0005] Patent document CN10395544A discloses an FFT accelerator based on a DSP chip. The FFT accelerator uses an FFT operation control module to determine whether the operation scale N is greater than a threshold N1, and selects to perform a small-scale FFT operation or a large-scale FFT operation, thereby improving execution performance and hardware resource utilization.
[0006] However, the FFT accelerator has the following problems:
[0007] ① Using a single data access module to process data read and output can lead to performance bottlenecks. Both data read and address calculation take up module processing time and resources. Especially when processing large amounts of data, module processing power can become a limiting factor. Summary of the Invention
[0008] To solve the above problems, the technical solutions of the present invention are as follows:
[0009] A multi-core DSP-FFT processing method for radar signal processing includes the following steps:
[0010] Acquire the data to be processed, the data to be processed is M*2 N Length data matrix, wherein M, N are natural numbers;
[0011] Determine the length of the data to be processed, if N < 12, perform short data FFT operation on the data; if N ≥ 12, perform long data FFT operation on the data;
[0012] According to the determination result, select the processing mode of the data to be calculated; the processing mode includes short data FFT operation and long data FFT operation;
[0013] The short data FFT operation is: first parameter configuration, then M / 8 times of loop, 8 CPUs perform 2 N point FFT calculation each time;
[0014] The long data FFT operation is: first parameter configuration, then M times of loop, in each loop: 8 CPUs jointly perform split FFT operation, then respectively perform complex multiplication of split FFT operation result and rotation factor, then 1 CPU performs configuration distribution, 1 CPU performs calculation and writing, and the remaining CPUs perform 2 N-3 point 2
[0015] The processing steps of short data FFT operation include:
[0016] Step 1: configure the parameters required for data operation of CPU0~7 through CPU0, the parameters include data start length and address;
[0017] Step 2: enable the calculation process of CPU0~7, initiate the command of moving the related length data from the DDR original data address through the EDMA enhanced DMA controller, complete the data preparation work of 8 cores of CPU0~7;
[0018] Step 3: start the calculation of DSP-CPU0~7, complete the FFT operation of 8 groups of 0~7, after completing the operation of 8 groups, the process returns to step 2, starts to move the next round of 8 groups of data, until the M groups of FFT calculation are completed.
[0019] The processing steps of long data FFT operation include:
[0020] Step 1: configure the parameters required for data operation of CPU0~7 through CPU0, the parameters include data start length and address;
[0021] Step 2: enable the calculation process of CPU0~7, and complete the data preparation work for 8 cores of CPU0~7 by initiating the command of moving the relevant length data from the DDR original data address to the EDMA enhanced DMA controller; wherein 2 N-3 points of FFT calculation are completed by each CPU selecting 2 N-3 points of FFT calculation from the original data address, and then performing point FFT calculation on the data length;
[0022] Step 3: since the 2 N point FFT operation process is split, after each CPU completes the split FFT calculation of 2 N-3 points, 8 groups of rotation factors corresponding to the coordinates need to be read from the cache, and each group of rotation factors is 2 N-3 complex numbers, and after multiplication, the rotated results are written into the L2 configured as SRAM;
[0023] Step 4: through the second configuration of CPU0 for CPU0~7, all cores of the DSP will start the second 8-point FFT calculation, CPU0 is responsible for the configuration and distribution of the data input of all CPUs1~6, and CPU7 is responsible for the calculation and writing of the new address index of the cache after the data end calculation; after step 4 is completed, it will jump back to step 2 to start the next round of split operation of points.
[0024] In the process of step 4, at the last 10 calculation times, CPU0 starts to read the next cluster of data (M-1)*2 N of the data matrix from the DDR again through the EDMA enhanced DMA controller; the data will be moved from the DDR memory to the MSMC-L3 memory of the DSP by the EDMA enhanced DMA controller.
[0025] The present application has the following beneficial effects:
[0026] 1) The multi-core distribution data calculation is adopted, the DSP calculation resources and the cache resources of the multi-core are fully utilized, and the data processing pipeline is ensured;
[0027] 2) The principle of FFT operation is combined with the split DSP calculation process of large point number FFT, the multi-core distribution, calculation and cache resources of the DSP are fully utilized, and the overall performance of the algorithm design pipeline is improved;
[0028] 3) The task scheduling, calculation and cache resources of the multi-core in the DSP are reasonably designed, the calculation parallel degree is improved while the control complexity is maintained to a certain extent, and the stability of the calculation process program is ensured. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 Figure of prior art DSP data cache architecture diagram;
[0030] Figure 2 Figure of DSP multi-core FFT data organization structure provided by the embodiment of the application;
[0031] Figure 3 Figure of processing flow provided by the embodiment of the application. DETAILED DESCRIPTION
[0032] The specific implementation of the application will be described in detail below with reference to the accompanying drawings.
[0033] As shown in the accompanying Figure 2 , 3 , a multi-core DSP-FFT processing mode applied to radar signal processing includes the following steps:
[0034] S100: obtaining data to be processed, the data to be processed being an M*2 N length data matrix, wherein M and N are natural numbers;
[0035] Two embodiments are provided in the application: in the embodiment 1, the input data matrix is 64*512 length; in the embodiment 2, the input data is 16 groups of 16384 length operation.
[0036] S110: judging the length of the data to be processed, if N<12, performing S200 short data FFT operation on the data; if N≥12, performing S300 long data FFT operation on the data;
[0037] Through the judgment of the data length of the embodiment 1, it is determined that the 64-point 512 length calculation is completed in the slow clock dimension, that is, S200 short data FFT operation is performed on the data.
[0038] For the embodiment 2, if the calculation mode of short data FFT operation is adopted, although the multi-core can be better called for operation, the calculation ability of the single-core CPU is limited due to the significant increase of the single-time FFT calculation length.
[0039] S200 short data FFT operation is: first parameter configuration, then M / 8 times of loop, and 2 N point FFT calculation is performed by 8 CPUs respectively each time.
[0040] Further, the processing steps of S200 short data FFT operation include:
[0041] S210: parameter configuration by CPU0 for data operation of CPU0-7, the parameters including data start length and address;
[0042] For example 1, in this step, the data operation parameter configuration of CPU0~7 is performed by CPU0, the data start length, address and other parameters are configured;
[0043] S220: enable the calculation process of CPU0~7, initiate the command of moving the related length data from the DDR original data address to the EDMA enhanced DMA controller, and complete the data preparation work of 8 cores of CPU0~7;
[0044] For example 1, in this step, the data preparation work of 8 cores of CPU0~7 is completed;
[0045] S230: start the calculation of DSP-CPU0~7, complete the FFT operation of 8 groups in total, and after completing the operation of 8 groups, the process returns to S220 to start the movement of the next 8 groups of data, until the M group FFT calculation is completed.
[0046] In example 1, after completing the operation of 8 groups, the process returns to S220 to start the movement of the next 8 groups of data, and then step 4 is executed again to complete a total of 64 group FFT calculation.
[0047] S300 long data FFT operation: first, parameter configuration is performed, then M times of loop is performed, in each loop: N 8 CPUs jointly perform 2 N-3 point split FFT operation, then the split FFT operation result is multiplied by the rotation factor, then 1 CPU performs configuration distribution, 1 CPU performs calculation and writing, and the remaining CPUs perform 2
[0048] For example 2, the long data FFT operation can split the DSP calculation process of large point number FFT by combining the principle of FFT operation, fully utilize the multi-core distribution, calculation and cache resources of DSP, and improve the overall performance of algorithm design pipeline.
[0049] Further, the processing steps of S300 long data FFT operation include:
[0050] S310: parameter configuration for data operation of CPU0~7 is performed by CPU0, and the parameters include data start length and address;
[0051] For example 2, in this step, the data operation parameter configuration of CPU0~7 is performed by CPU0, the data start length, address and other parameters are configured.
[0052] S320: enable the calculation process of CPU0~7, initiate the command of moving the relevant length data from the DDR original data address to the EDMA enhanced DMA controller, complete the data preparation work of 8 cores of CPU0~7; wherein each CPU selects 2048 data lengths from the original data address to perform 2048-point FFT calculation, and 8 CPUs can jointly complete the first calculation of 8*2048=16384-point FFT. N-3 N-3 N
[0053] For example 2, the data preparation work of 8 cores of CPU0~7 is completed in this step. Each CPU selects 2048 data lengths from the original data address to perform 2048-point FFT calculation, and 8 CPUs can jointly complete the first calculation of 8*2048=16384-point FFT.
[0054] S330: Since the 2 N point FFT operation process is split, each CPU completes the split FFT calculation of 2 N-3 points, and then reads 8 groups of rotation factors corresponding to the coordinates from the cache, each group of rotation factors has 2 N-3 complex numbers, and after multiplication, the rotated results are written into L2 configured as SRAM.
[0055] For example 2, the 16384-point FFT operation process is split in this step, each CPU completes the FFT calculation of 2048 points, and then reads 8 groups of rotation factors corresponding to the coordinates from the cache, each group of rotation factors has 2048 complex numbers, and after multiplication, the rotated results are written into L2 configured as SRAM.
[0056] S340: CPU0 performs the second configuration for CPU0~7, all cores of the DSP will start 2 N-3 times of 8-point FFT calculation, CPU0 is responsible for configuring and distributing the data input of all CPUs 1~6, and CPU7 is responsible for calculating and writing the new address index of the cache after the data end calculation; after step 4, it will jump back to S320 to start the next round of 2 N point split operation.
[0057] For example 2, in this step, CPU0 performs the second configuration for CPU0~7, all cores of the DSP start 2048 times of 8-point FFT calculation, CPU0 configures and distributes the data input of all CPUs 1~6, and CPU7 is responsible for calculating and writing the new address index of the cache after the data end calculation.
[0058] For example 2, in this step, 6 cores of 1-6 will complete 2048 times of 8-point FFT operation by the configuration of S340. Due to the reduction of single data amount, the cache of DSP can complete the cache of single data to improve the operation speed of the whole process.
[0059] In the process of S340, at the last 10 times of calculation time, CPU0 starts to read the remaining (M-1)*2 data of the next slow clock FFT calculation from DDR again through the EDMA enhanced DMA controller. N The next cluster of 2 data matrices N The data will be moved from the DDR memory to the MSMC-L3 memory of the DSP through the EDMA enhanced DMA controller.
[0060] For example 2, in the process of S340, at the last 10 times of calculation time, CPU0 starts to read the remaining 15*16384 data of the next slow clock FFT calculation from DDR again through the EDMA enhanced DMA. The next cluster of 16384 data of 16384 data matrices will be moved from the DDR memory to the MSMC-L3 memory of the DSP through the EDMA enhanced DMA to improve the data rate of the whole data processing algorithm. The MSMC controller is configured as MSMC-L3. The storage space of the MSMC is 4MB, and the single 16384-point data occupies a space of 64kB, with a total of 4 bytes for real and imaginary parts, which fully meets the exchange of at least 2 times of 16384-point data. After completing S340, it will jump back to S320 to start the second round of 16384-point splitting operation.
[0061] In the whole process of 16*16384 data matrix processing, the special storage data in MSMC-SRAM includes 2 times of 16384-point calculation, 16384-point rotation factor, and 16384*2-point exchange space, so the total calculation requires 320KB of cache space, which meets the exchange cache required by the second calculation method. The calculation of the rotation factor is through the setting of 128KB L2 as its 8-core complex multiplication calculation data cache space. At this time, 256KB in the L2 of the DSP needs to be configured as a non-cache addressable memory access space (L2 SRAM).
[0062] The application makes full use of multi-core distribution data calculation, DSP calculation resources and multi-core cache resources to ensure data processing pipeline; combines the principle of FFT operation to split the DSP calculation process of large point number FFT, fully utilizes the multi-core distribution, calculation and cache resources of DSP to improve the overall performance of algorithm design pipeline; reasonably designs the task scheduling, calculation and cache resources of multi-core in DSP, improves the calculation parallel degree while maintaining a low control complexity to a certain extent, and ensures the stability of the calculation process program.
[0063] The above disclosed are only several specific embodiments of the present application, but the present application is not limited to this. Any changes that can be thought of by those skilled in the art shall fall within the protection scope of the present application.
Claims
1. A multi-core DSP-FFT processing method for radar signal processing, characterized by: The following steps are involved: Get the data to be processed, the data to be processed is M*2 N Length data matrix, where M and N are natural numbers; Determine the length of the data to be processed, if N < 12, perform a short data FFT operation on the data; if N ≥ 12, perform a long data FFT operation on the data; According to the judgment result, a processing method of the data to be calculated is selected; the processing method includes short data FFT operation and long data FFT operation; The short data FFT operation is as follows: first configure the parameters, then perform M / 8 cycles, with 8 CPUs performing 2 cycles each time. N Point FFT calculation; The long data FFT operation is as follows: first configure the parameters, then perform M cycles, in each cycle: 8 CPUs perform 2 N The split FFT operation of the point, and then the complex multiplication of the split FFT operation results and the rotation factors, then 1 CPU performs configuration distribution, 1 CPU performs calculation and writing, and the remaining CPUs perform 2 N-3 8-point FFT calculations are performed.
2. The multi-core DSP-FFT processing method for radar signal processing according to claim 1 is characterized in that: The processing steps of the short data FFT operation include: Step 1: CPU0 configures the parameters required for data operation on CPU0-7, including the data start length and address; Step 2: Enable the computing process of CPUs 0 to 7 and complete data preparation for the eight cores of CPUs 0 to 7 by initiating a command to the EDMA enhanced DMA controller to move data of a certain length from the DDR original data address. Step 3: Start the calculation of DSP-CPU0~7 and complete 8 groups of FFT calculations from 0 to 7. After completing these 8 groups of calculations, the process returns to step 2 and starts moving the next round of 8 groups of data until M groups of FFT calculations are completed.
3. The multi-core DSP-FFT processing method for radar signal processing according to claim 1 is characterized in that: The processing steps of the long data FFT operation include: Step 1: CPU0 configures the parameters required for data operation on CPU0-7, including the data start length and address; Step 2: Enable the computing process of CPU0~7, and complete the data preparation work for the 8 cores of CPU0~7 by issuing a command to the EDMA enhanced DMA controller to move the relevant length data from the DDR original data address; each CPU selects 2 from the original data address. N-3 The data length is 2 N-3 8 CPUs can complete 2 point FFT calculations together. N Calculation of point FFT; Step 3: Since for 2 N The FFT operation process is split, and each CPU completes 2 N-3 After the split FFT calculation of the point, it is necessary to read 8 groups of rotation factors of the corresponding coordinates from the cache, each group of rotation factors 2 N-3 After multiplying the complex numbers, the rotated results are written to L2 configured as SRAM; Step 4: Configure CPU0 to CPU7 for the second time through CPU0, and all cores of DSP will start 2 N-3 In the 8-point FFT calculation, CPU0 is responsible for configuring and distributing the data input of all CPUs 1 to 6, and CPU7 is responsible for calculating and writing the new address index of the cache after the data calculation is completed; after completing step 4, it will jump back to step 2 and start the next round of 2 N Point splitting operation.
4. The multi-core DSP-FFT processing method for radar signal processing according to claim 3 is characterized in that: During the process of step 4, during the last 10 calculation times, CPU0 starts to read the remaining (M-1)*2 of the next slow clock FFT calculation from DDR again through the EDMA enhanced DMA controller. N Next cluster of data matrix 2 N Data; the data will be moved from the DDR memory to the DSP's MSMC-L3 memory through the EDMA enhanced DMA controller.