A parallel coherent accumulation method for frequency-hopping echoes based on Chebyshev NUFFT

By adopting a parallel coherent accumulation method for frequency-hopping echoes based on Chebyshev NUFFT, the problems of large computational load, memory access bottleneck and low detection probability in frequency-hopping radar echo processing are solved. This method achieves efficient target detection and data processing, reduces false alarm rate, and improves the real-time performance and detection accuracy of radar system.

CN122085237APending Publication Date: 2026-05-26XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-01-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies suffer from enormous computational demands when processing frequency-hopping radar echoes. The rotation factor dynamically changes with the carrier frequency, making parallelization difficult. Non-uniform sampling processing results in discontinuous memory access, leading to memory bottlenecks. The computation time is too long, the detection probability is too low, and the sidelobe leakage is high, resulting in an even lower detection probability.

Method used

A parallel coherent accumulation method based on Chebyshev NUFFT for frequency hopping echoes is adopted. Utilizing a multi-core parallel architecture of a DSP, the main core dynamically acquires the dimension information of the input RD matrix, allocates task blocks and descriptors, and each slave core performs NUFFT algorithm processing. Non-uniform sampling is mapped to uniform grid points, and adaptive one-dimensional coarse screening and adaptive variable window topology convergence are performed using precise phase decoupling and cascade detection algorithms to confirm the target center coordinates and amplitude.

Benefits of technology

It achieved an 85% increase in processing efficiency, a 3dB~6dB increase in coherent gain, a reduction of more than 30% in false alarm rate, improved utilization of computing resources, and met real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122085237A_ABST
    Figure CN122085237A_ABST
Patent Text Reader

Abstract

This invention discloses a parallel coherent accumulation method for frequency-hopping echoes based on Chebyshev NUFFT, applied to a multi-core parallel architecture of DSP based on task space slicing and double buffering. By introducing the NUFFT algorithm combined with the multi-core parallel architecture of DSP, this invention significantly increases the scale of data processing, reducing the total processing time by 85% compared to traditional serial schemes, with a measured processing time of approximately 70ms~80ms, demonstrating a substantial improvement in processing efficiency. Through Gaussian kernel latticeization and precise phase decoupling, the echo consistency loss caused by frequency hopping is recovered to the maximum extent. Compared to traditional methods, the coherent gain is improved by 3dB~6dB. Utilizing adaptive variable window topology convergence intelligently suppresses false targets caused by residual sidelobes from frequency hopping, reducing the false alarm rate of output traces by more than 30% compared to fixed window algorithms, greatly reducing the huge computational load and mistracking risk brought by the backend track association and tracking algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar data processing technology, specifically relating to a parallel coherent accumulation method for frequency-hopping echoes based on Chebyshev NUFFT. Background Technology

[0002] Frequency-hopping radar employs frequency-agile waveforms to counter jamming, but inter-pulse carrier frequency variations cause phase discontinuities. Highly maneuverable targets move across multiple range cells within the Coherent Processing Interval (CPI), requiring Range Cell Migration Correction (RCMC) and phase alignment to achieve energy accumulation. Non-uniform sampling spectral analysis (e.g., NUFFT) has become crucial, and low-rank approximation processing VT-AF (Velocity Filtering) is essential. Time Amplitude Frequency, speed time Amplitude-frequency coupling. Utilizing the TI C6678 octa-core architecture and KeyStone hardware support, real-time RD (Range-Doppler) imaging can be achieved.

[0003] Existing solutions are mostly based on CZT (Chirp Z-Transform) to process FAR (Frequency-Agile Radar) signals, and use the Bluestein algorithm to achieve convolution through three FFT / IFFT (Fast Fourier Transform / Inverse Fast Fourier Transform). CZT requires processing each distance unit separately, resulting in a huge computational load and excessively high overall complexity. The rotation factor changes dynamically with the carrier frequency, making it impossible to pre-compute a unified convolution kernel, which disrupts the DSP (Digital Signal Processor) SIMD (Digital Signal Processor, Single Instruction Multiple Data) pipeline and makes parallelization difficult. In non-uniform sampling processing, memory access is discontinuous, causing severe invalidation of the C6678 L1 / L2 cache, leading to DDR (Double Data Rate memory) waiting. Actual measurements show that CZT takes far longer than the real-time requirements, resulting in low DSP resource utilization, bottlenecks in storage and memory access, excessively long computation time, high sidelobe leakage, and low detection probability under non-uniform sampling. Summary of the Invention

[0004] To address the aforementioned problems in the existing technology, this invention provides a parallel coherent accumulation method for frequency-hopping echoes based on Chebyshev NUFFT. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a parallel coherent accumulation method for frequency-hopping echoes based on Chebyshev NUFFT, applied to a multi-core parallel architecture of DSP based on task space slicing-double buffering. The method includes: In a DSP multi-core parallel architecture, the master core dynamically acquires the dimension information of the input RD matrix and outputs the task block and task descriptor corresponding to each slave core according to the number of slave cores in the DSP multi-core parallel architecture. In the DSP multi-core parallel architecture, each slave core processes data through a double-buffered parallel pipeline. Based on the task block and task descriptor, the NUFFT algorithm is used to perform non-uniform sampling to uniform grid mapping through Gaussian kernel convolution. Precise phase decoupling is then used to obtain the phase-compensated matrix. A cascaded detection algorithm is then used to perform adaptive one-dimensional coarse screening and adaptive variable window topology convergence on the phase-compensated matrix to obtain the search window. Based on local extremum constraints and gradient convergence constraints, the target center coordinates and magnitude are confirmed according to the search window. The master core outputs the final target list based on the target center coordinates and magnitudes obtained from the parallel computation of all slave cores.

[0005] In one embodiment of the present invention, the master core in the DSP multi-core parallel architecture acts as a management node, and divides the pulses in the dimension information of the dynamically acquired input RD matrix into task blocks of the same size according to the number of slave cores, with each task block associated with a task descriptor.

[0006] In one embodiment of the present invention, the task descriptor includes: The physical base address, step size, and frequency hopping sequence index required for coherent accumulation of the task block in DDR3.

[0007] In one embodiment of the present invention, the step of using the NUFFT algorithm based on the task block and task descriptor, performing non-uniform sampling to uniform grid mapping through Gaussian kernel convolution gridding, and obtaining the phase-compensated matrix using precise phase decoupling includes: Based on the task block and task descriptor, confirm the target baseband echo signal received by the radar and the corresponding initial ranging result; Construct a phase correction operator based on the initial distance measurement results; Based on the target baseband echo signal and the phase correction operator, Gaussian kernel lattice processing is performed based on the Gaussian kernel function to confirm the uniform sequence after mapping. Spectral recovery and deconvolution correction are performed on the uniform sequence to obtain the phase-compensated matrix.

[0008] In one embodiment of the present invention, the expression for the target baseband echo signal is as follows: ; in, Indicates the first The target baseband echo signal corresponding to each pulse. Represents the imaginary unit. Indicates the reference center frequency. Indicates the first Random frequency offset of each pulse Represents the standard vacuum speed of light. This indicates the distance between the detected target point and the radar receiver. This indicates the target's velocity when the radar waves detect it. This indicates the pulse repetition period.

[0009] In one embodiment of the present invention, the expression for the mapped uniform sequence is as follows: ; in, Represents a uniform frequency domain grid. Indicates the first The target baseband echo signal corresponding to each pulse. Indicates the first Phase correction operator corresponding to each pulse Indicates the total number of pulses. This indicates a non-uniform frequency hopping sampling position. This represents a Gaussian kernel function with narrow support characteristics. This represents the sampling density compensation factor.

[0010] In one embodiment of the present invention, the step of performing spectral recovery and deconvolution correction on the uniform sequence to obtain a phase-compensated matrix includes: Perform a 2x oversampling FFT on the uniform sequence to obtain the FFT result; The FFT processing results are deconvolutionally corrected in the frequency domain to eliminate the low-pass envelope effect introduced by the Gaussian kernel lattice processing, and the phase-compensated matrix is ​​output.

[0011] In one embodiment of the present invention, the step of using a cascaded detection algorithm to perform adaptive one-dimensional coarse screening and adaptive variable window topological convergence on the phase-compensated matrix to obtain a search window includes: The phase-compensated matrix is ​​subjected to sliding detection along the Doppler axis. By calculating the mean of the reference cells in the target neighborhood and coarsely screening with a preset false alarm rate threshold, the corresponding threshold points are obtained. For points exceeding the threshold, based on the local signal-to-noise ratio SNR Dynamically calculate the search window to achieve adaptive variable window topology convergence.

[0012] In one embodiment of the present invention, the step of determining the target center coordinates and magnitude based on the search window according to local extremum constraints and gradient convergence constraints includes: Based on local mechanism constraints, the center point is determined from the threshold points according to the search window; Calculate the Laplacian operator for the center point and the four surrounding pixels. If the value of the Laplacian operator exceeds a preset threshold, confirm the target center coordinates and magnitude.

[0013] In one embodiment of the present invention, in the DSP multi-core parallel architecture, task descriptors are distributed by the master core, and the slave cores use the L2 SRAM ping-pong buffer to prefetch the data corresponding to the next task block while calculating the data corresponding to the current task block.

[0014] The beneficial effects of this invention are: The solution provided in this invention significantly improves the scale of data processing by introducing the NUFFT algorithm combined with a multi-core parallel architecture of a DSP. The total processing time is reduced by 85% compared to traditional serial solutions, with a measured processing time of approximately 70ms~80ms, demonstrating a substantial improvement in processing efficiency. Through Gaussian kernel latticeization and precise phase decoupling, the echo consistency loss caused by frequency hopping is recovered to the maximum extent. Compared to traditional methods, the coherent gain is improved by 3dB~6dB. The adaptive variable window topology convergence intelligently suppresses false targets caused by residual sidelobes from frequency hopping, reducing the false alarm rate of output points by more than 30% compared to fixed window algorithms. This greatly reduces the enormous computational load and mistracking risk brought by the backend track association and tracking algorithms. Attached Figure Description

[0015] Figure 1 A schematic diagram illustrating the steps of a frequency-hopping echo parallel coherent accumulation method based on Chebyshev NUFFT provided in an embodiment of the present invention; Figure 2 The diagram shows the DSP multi-core parallel architecture in a frequency-hopping echo parallel coherent accumulation method based on Chebyshev NUFFT provided in an embodiment of the present invention. Figure 3 This diagram illustrates the overlap between the EDMA ping-pong buffer and the CPU core computation time in a frequency-hopping echo parallel coherent accumulation method based on Chebyshev NUFFT provided in an embodiment of the present invention. Figure 4 This is a comparison chart of the effects of conventional interpolation FFT and NUFFT provided in the embodiments of the present invention; Figure 5 The flowchart shows a cascaded detection algorithm in a frequency-hopping echo parallel coherent accumulation method based on Chebyshev NUFFT provided in an embodiment of the present invention. Detailed Implementation

[0016] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0017] Frequency-agile radar effectively avoids active interference and reduces the probability of interception by randomly changing the carrier frequency between pulses. However, its non-uniform carrier frequency sampling prevents the echo from obtaining coherent accumulation gain through conventional FFT. The embodiments of this invention aim to solve the following deep-seated technical problems: Carrier-range phase decoupling and energy divergence problems: due to carrier frequency Random jump, the target's phase term There exists a high-dimensional multiplicative coupling between carrier frequency and range and velocity. Under broadband conditions, this coupling causes severe cross-range migration and Doppler defocusing of the target energy in the range-Doppler (RD) plane, resulting in significant main lobe widening accompanied by artifact grating lobes in the high plane, which severely reduces the detection probability.

[0018] The high-complexity latency bottleneck of embedded real-time processing: The accumulation of frequency-agile echoes involves Non-Uniform Fast Fourier Transform (NUFFT). This algorithm, in its physical implementation, involves complex convolutional gridding mappings, a large number of unaligned memory accesses, and exponential operations. When processing large-scale RD matrices (such as 2048×1024) on a single-core embedded processor, the computational latency is typically on the order of hundreds of milliseconds, which cannot meet the 50ms real-time refresh requirements of radar systems.

[0019] Data throughput limitations in multi-core heterogeneous memory: In multi-core DSP architectures, the CPU core's access speed to external DDR3 memory is much lower than the speed of internal register operations, resulting in a "memory wall" effect. Without efficient data prefetching and transfer mechanisms, frequent non-contiguous memory accesses will cause the CPU core to remain in a static state for extended periods, with computational resource utilization typically below 30%, failing to leverage the peak performance of multi-core processors.

[0020] Track splitting and false alarm issues in complex sidelobe backgrounds: The spectral sidelobe structure after frequency hopping signal accumulation is complex and has a high level. Traditional fixed sliding window detection algorithms are prone to classifying a physical target as multiple discrete detection points (i.e., track splitting) under strong target interference or complex backgrounds. This will bring huge computational load and false tracking risk to subsequent track association and tracking algorithms.

[0021] To address the aforementioned problems, this invention provides a frequency-hopping echo parallel coherent accumulation method based on Chebyshev NUFFT, applied to a task space slicing-double-buffered DSP multi-core parallel architecture. The method is as follows: Figure 1 As shown, it may include: S1, the master core in the DSP multi-core parallel architecture dynamically obtains the dimension information of the input RD matrix, and outputs the task block and task descriptor corresponding to each slave core according to the number of slave cores in the DSP multi-core parallel architecture; In the S2 DSP multi-core parallel architecture, each slave core processes data through a double-buffered parallel pipeline. Based on the task block and task descriptor, the NUFFT algorithm is used to perform non-uniform sampling to uniform grid mapping through Gaussian kernel convolution. Precise phase decoupling is then used to obtain the phase-compensated matrix. A cascaded detection algorithm is then used to adaptively perform one-dimensional coarse screening and adaptive variable window topology convergence on the phase-compensated matrix to obtain the search window. Based on local extremum constraints and gradient convergence constraints, the target center coordinates and magnitude are confirmed according to the search window. S3: The main core outputs the final target list based on the target center coordinates and magnitudes obtained from the parallel computation of all slave cores.

[0022] Understandably, the DSP multi-core parallel architecture based on task space slicing-double buffering provided in this embodiment of the invention is illustrated in the following diagram: Figure 2 As shown, this DSP multi-core parallel architecture incorporates a sophisticated task scheduling and memory management mechanism. Distance gate dimension slicing logic: The master core (Core 0) acts as the management node, dynamically acquiring the dimension information (M distance gates × N pulses) of the input RD matrix. The master core evenly divides the N pulses into task blocks of size N / 7 according to the number of slave cores. Each task block is associated with a task descriptor. The task descriptor may include: the physical base address of the task block in DDR3, the step size, and the frequency hopping sequence index required for coherent accumulation. Optionally, the task slicing can be adjusted from the distance gate dimension to the Doppler dimension to minimize the EDMA row stepping.

[0023] Fine-grained configuration of EDMA3 PaRAM parameters: Within the slave core, EDMA3 transfer parameters are configured. Data is divided into three levels: ACNT is the number of complex data bytes per row (e.g., 2×4 bytes), BCNT is the number of rows transferred in a single ping-pong pass (defined as Block Size, e.g., 16 rows), and CCNT is the total number of task iterations. Lock-free synchronization is achieved by configuring the EDMA completion interrupt to be associated with the CPU's semaphore. Calculations are performed in the cache, and the results are stored in DDR. The data source is several target baseband echo signals allocated by the master core.

[0024] Double-buffered computation pipeline (Ping-Pong Flow): Initial loading: Initiate EDMA asynchronous transfer of Block_0 to L2_SRAM_Ping.

[0025] Steady-state execution: The CPU calls C66x instruction set optimized code (such as using cmpysp instructions for 4-way parallel complex multiplication) from the data in L2_SRAM_Ping to perform phase compensation and gridding.

[0026] Meanwhile, the background EDMA controller independently moves Block_1 to L2_SRAM_Pong.

[0027] Cyclic flip: After the calculation is complete, the CPU immediately flips the pointer to process Block_1 and triggers EDMA to move the next block of data. This mechanism ensures that the core processing unit (CPU Core) does not need to wait for DDR response throughout the entire RD processing cycle, achieving a near 100% duty cycle.

[0028] A diagram illustrating the overlap between EDMA ping-pong buffer and CPU core processing time, as shown below. Figure 3 As shown, in the DSP multi-core parallel architecture, task descriptors are distributed by the master core, and the slave cores use the L2 SRAM ping-pong buffer to prefetch the data corresponding to the next task block while calculating the data corresponding to the current task block.

[0029] Based on the task block and task descriptor, the NUFFT algorithm is used to perform non-uniform sampling to uniform grid mapping through Gaussian kernel convolution. Precise phase decoupling is then employed to obtain the phase-compensated matrix, which may include: Based on the task block and task descriptor, confirm the target baseband echo signal received by the radar and the corresponding initial ranging result; Construct a phase correction operator based on the initial distance measurement results; Based on the target baseband echo signal and the phase correction operator, Gaussian kernel lattice processing is performed based on the Gaussian kernel function to confirm the uniform sequence after mapping. Spectral recovery and deconvolution correction are performed on the uniform sequence to obtain the phase-compensated matrix.

[0030] For frequency-agile pulse trains, let the first... The carrier frequency of each pulse is ,in, Indicates the reference center frequency. Indicates the first The expression for the target baseband echo signal with a random frequency offset of pulses is as follows: ; in, Indicates the first The target baseband echo signal corresponding to each pulse. Represents the imaginary unit. Indicates the reference center frequency. Indicates the first Random frequency offset of each pulse Represents the standard vacuum speed of light. This indicates the distance between the detected target point and the radar receiver. This indicates the target's velocity when the radar waves detect it. This indicates the pulse repetition period.

[0031] Specifically, the NUFFT accumulation process and detailed compensation scheme designed in this embodiment of the invention are as follows: Multi-stage phase decoupling and alignment: First, using the initial ranging result R_r corresponding to the echo signal, a phase correction operator is constructed. By performing point-by-point complex multiplication on the original sampling sequence, the shift in the first-order distance term caused by frequency jumps is eliminated. This step is a crucial prerequisite for ensuring that subsequent Doppler coherent accumulation does not diverge.

[0032] Improved Gaussian kernel Gridding: To optimize the non-uniform frequency hopping sampling positions High-precision mapping to uniform frequency domain grid A Gaussian kernel function with narrow support characteristics is used. The expression for the mapped uniform sequence is as follows: ; in, Represents a uniform frequency domain grid. Indicates the first The target baseband echo signal corresponding to each pulse. Indicates the first Phase correction operator corresponding to each pulse Indicates the total number of pulses. This indicates a non-uniform frequency hopping sampling position. This represents a Gaussian kernel function with narrow support characteristics. This represents the sampling density compensation factor, used to balance the amplitude modulation caused by uneven sampling point density on the frequency axis. (Selection...) With a grid spacing of 2 to 3, the interpolation kernel support set is limited to a 4×4 or 6×6 neighborhood to maximize accuracy within the limited computing power of the DSP chip. Optionally, the gridded interpolation kernel can be replaced by a Gaussian kernel function with a Kaiser-Bessel kernel function.

[0033] Spectral recovery and deconvolution correction: Performing spectral recovery and deconvolution correction on a uniform sequence yields a phase-compensated matrix, which may include: Perform a 2x oversampling FFT on the uniform sequence to obtain the FFT result; The FFT processing results are deconvolutionally corrected in the frequency domain to eliminate the low-pass envelope effect introduced by the Gaussian kernel lattice processing, and the phase-compensated matrix is ​​output.

[0034] For the mapped uniform sequence A 2x oversampling FFT is performed, followed by deconvolution correction in the frequency domain, i.e., dividing by the Fourier transform response of the Gaussian kernel, to eliminate the low-pass envelope effect introduced by the gridding process and ensure the flatness of the output spectrum.

[0035] Comparison of the effects of conventional interpolation FFT and NUFFT, as shown in the image. Figure 4 As shown, the following points can be observed: Processing method: Conventional interpolation FFT requires first using an interpolation algorithm to force non-uniform sampling points to be mapped onto an equally spaced grid, while NUFFT directly processes non-uniformly distributed sampling points through mathematical transformations.

[0036] Computational accuracy: Conventional interpolation FFT is prone to introducing approximation errors and spectral leakage during the interpolation process, resulting in distortion of high-frequency components; NUFFT, on the other hand, can provide reconstruction accuracy close to the theoretical limit through precise weight compensation.

[0037] Computational complexity: Although conventional interpolation FFT combined with the Fast Fourier Transform algorithm is extremely fast, it requires extremely high grid density to ensure accuracy; NUFFT, although slightly more computationally intensive, still maintains the same order of magnitude in terms of algorithmic complexity as conventional interpolation FFT.

[0038] Spectral fidelity: When processing signals with drastic fluctuations or sparse sampling, conventional interpolation FFT often exhibits obvious spurious spectrum phenomena, while NUFFT can more realistically restore the original spectral characteristics of the signal.

[0039] Application scenarios: Conventional interpolation FFT is suitable for simple real-time processing where high accuracy is not required, while NUFFT is the preferred solution for high-speed targets in radar detection.

[0040] To address the artifacts that are prone to occur in frequency hopping systems, this invention proposes a cascaded detection algorithm. The flowchart of the cascaded detection algorithm is as follows: Figure 5 As shown, by using the cascaded detection algorithm to perform adaptive one-dimensional coarse screening and adaptive variable window topological convergence on the phase-compensated matrix, a search window can be obtained, which may include: The phase-compensated matrix is ​​subjected to sliding detection along the Doppler axis. By calculating the mean of the reference cells in the target neighborhood and coarsely screening with a preset false alarm rate threshold, the corresponding threshold points are obtained. For points exceeding the threshold, based on the local signal-to-noise ratio SNR Dynamically calculate the search window to achieve adaptive variable window topology convergence.

[0041] For adaptive one-dimensional CA-CFAR coarse screening, since NUFFT has already completed phase decoupling, the Doppler energy is highly concentrated. The system performs sliding detection along the Doppler axis, calculating the mean of the reference cells in the target neighborhood. In conjunction with the preset false alarm rate threshold Filtering out more than 90% of the background noise area significantly reduces the amount of subsequent computation.

[0042] For adaptive variable window topology convergence, for threshold points, based on the local signal-to-noise ratio... SNR Dynamic calculation of search window Understandably, SNR The higher the value, the sharper the main lobe of the target. Automatically shrinks; conversely, expands. The side length of the search window. The mapping relationship is refined as follows: ; in, This represents the morphological adjustment coefficient.

[0043] Based on local extremum constraints and gradient convergence constraints, the target center coordinates and magnitude are determined according to the search window, which may include: Based on local mechanism constraints, the center point is determined from the threshold points according to the search window; Calculate the Laplacian operator for the center point and the four surrounding pixels. If the value of the Laplacian operator (second derivative) exceeds a preset threshold, confirm the target center coordinates and magnitude.

[0044] By designing the value of the Laplace operator (second derivative) to exceed a preset threshold, flat sidelobe artifacts caused by frequency modulation residual phase shift can be effectively eliminated. This algorithm completes the "purification" of the target during the detection phase, outputting only the physically true centroid coordinates, greatly reducing the track management burden on the back-end data processing unit.

[0045] Optionally, the frequency-hopping echo parallel coherent accumulation method based on Chebyshev NUFFT provided in this embodiment of the invention can be ported to an FPGA + DSP platform, where the FPGA performs high-throughput NUFFT calculations at the front end.

[0046] The core data structure and parallel control of the frequency-hopping echo parallel coherent accumulation method based on Chebyshev NUFFT, as well as the corresponding pseudocode for EDMA ping-pong processing and cascade detection provided in this embodiment of the invention, are as follows: / / Core 1-7 (from core) parallel computing pipeline: void Slave_Process_Loop(SliceTask *myTask) { / / 1. Initialize the local L2 SRAM ping-pong buffer: float* Buffer_Ping = L2_SRAM_BASE; float* Buffer_Pong = L2_SRAM_BASE + BLOCK_SIZE; / / 2. Preload the first block of data (Initial EDMA Load): EDMA_Copy_Async(myTask->ddr_src_ptr, Buffer_Ping, BLOCK_SIZE); for(int block_idx = 0; block_idx <Total_Blocks; block_idx++) { / / Waiting for the current DMA transfer to complete semaphore: SEM_wait(edma_ping_done); / / Asynchronously trigger the transfer of the next block of data (Hide Memory Latency): if(block_idx <Total_Blocks - 1) { EDMA_Copy_Async(next_addr, Buffer_Pong, BLOCK_SIZE); } / / --- Core computing phase (CPU running at full load): float* active_data = Buffer_Ping; for(int i = 0; i <BLOCK_SIZE; i++) { / / A. NUFFT phase compensation and gridding; / / Call the C6678 built-in complex multiplication instruction _cmpysp to accelerate decoupling: NUFFT_Phase_Compensation(active_data[i], myTask->hop_seq_offset); Gridding_To_Uniform_Grid(active_data[i], &gridded_buffer); / / B. Phase 1: One-dimensional CA-CFAR coarse inspection; float threshold = Calculate_1D_CFAR_Threshold(gridded_buffer); if(gridded_buffer[i]>threshold) { / / C. Phase 2: Adaptive variable window convergence fine inspection: int adaptive_W = Compute_Adaptive_Window_Size(gridded_buffer, i); if(Is_Local_Maximum_With_Gradient(gridded_buffer, i, adaptive_W)) { / / Record the centroid of the high-precision target: Save_Target_Centroid(i,&result_buffer); } } } / / Switch the ping-pong pointers to prepare for the next round of the assembly line: Swap(Buffer_Ping, Buffer_Pong); SEM_post(edma_ready_for_next); } } The frequency-hopping echo parallel coherent accumulation method based on Chebyshev NUFFT provided in this invention has the following advantages: A generational leap in processing efficiency: By introducing the NUFFT fast algorithm and combining it with the hard parallelism of multi-core DSPs, the total processing time for 2048×1024 data is reduced by 85% compared to the serial solution, with a measured processing time of approximately 70ms~80ms.

[0047] Fully coherent accumulation gain: By using Gaussian kernel lattice and precise phase decoupling, the echo consistency loss caused by frequency hopping is recovered to the maximum extent. Compared with traditional methods, the coherent gain is improved by 3dB~6dB.

[0048] Extremely high purity of target points: The adaptive centroid algorithm can intelligently suppress false targets caused by residual sidelobes of frequency hopping, and the false alarm rate of the output points is reduced by more than 30% compared with the fixed window algorithm.

[0049] Excellent platform adaptability: The algorithm is implemented in C language with built-in instructions, which has good portability and can be quickly deployed on various embedded real-time processing platforms.

[0050] It should be noted that, in the description of this invention, the term "multiple" means two or more, unless otherwise explicitly specified.

[0051] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A Chebyshev NUFFT based frequency hopping echo parallel coherent accumulation method applied to a task space slice-double buffer based DSP multi-core parallel architecture, characterized in that, The method comprises the following steps: The master core in the DSP multi-core parallel architecture dynamically obtains the dimension information of the input RD matrix, and outputs the task block and the task descriptor corresponding to each slave core according to the number of slave cores in the DSP multi-core parallel architecture; Each slave core in the DSP multi-core parallel architecture processes through a double-buffer parallel pipeline, and adopts the NUFFT algorithm to perform mapping of non-uniform sampling to uniform grid points through Gaussian kernel convolution grid points, utilizes accurate phase decoupling to obtain a phase-compensated matrix, and utilizes a cascade detection algorithm to perform adaptive one-dimensional coarse screening and adaptive variable window topology centring on the phase-compensated matrix to obtain a search square window; and the target center coordinates and amplitude are confirmed based on local extreme value constraint and gradient convergence constraint according to the search square window. The master core outputs a final target list according to the target center coordinates and amplitude obtained by parallel operation of all slave cores.

2. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 1, characterized in that, The master core in the DSP multi-core parallel architecture serves as a management node, divides the pulses in the dimension information of the input RD matrix into task blocks of the same size according to the number of slave cores, and associates each task block with a task descriptor.

3. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 1, characterized in that, The task descriptor comprises: The physical base address, step length and frequency hopping sequence index required for coherent accumulation of the task block in the DDR3.

4. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 1, characterized in that, The NUFFT algorithm is adopted to perform mapping of non-uniform sampling to uniform grid points through Gaussian kernel convolution grid points, utilize accurate phase decoupling to obtain a phase-compensated matrix according to the task block and the task descriptor, which comprises: Confirming the target baseband echo signal received by the radar and the corresponding initial ranging result according to the task block and the task descriptor; Constructing a phase correction operator according to the initial ranging result; Performing Gaussian kernel grid point processing based on the Gaussian kernel function according to the target baseband echo signal and the phase correction operator to confirm the mapped uniform sequence; Performing frequency spectrum recovery and deconvolution correction on the uniform sequence to obtain the phase-compensated matrix.

5. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 4, characterized in that, The expression of the target baseband echo signal is as follows: ; wherein, represents the target baseband echo signal corresponding to the represents the imaginary unit, represents the reference center frequency, represents the random frequency offset of the represents the standard vacuum light speed, represents the distance between the detection target point and the radar receiver, represents the speed of the target when the radar wave is detected to the target, represents the pulse repetition period.​​ 6. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 4, characterized in that, The expression of the mapped uniform sequence is as follows: ; wherein, denotes a uniform frequency grid, denotes the target baseband echo signal corresponding to the denotes the target baseband echo signal corresponding to the denotes the phase correction operator corresponding to the denotes the phase correction operator corresponding to the denotes the total number of pulses, denotes the non-uniform frequency hopping sampling positions, denotes a Gaussian kernel function with narrow support, denotes a sampling density compensation factor.

7. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 4, characterized in that, The frequency spectrum recovery and deconvolution correction are performed on the uniform sequence to obtain the phase-compensated matrix, which comprises: Performing 2-fold oversampling FFT processing on the uniform sequence to obtain an FFT processing result; Performing deconvolution correction on the FFT processing result in the frequency domain to eliminate the low-pass envelope effect introduced in the Gaussian kernel grid point processing process, and outputting the phase-compensated matrix.

8. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 1, characterized in that, The cascade detection algorithm is utilized to perform adaptive one-dimensional coarse screening and adaptive variable window topology centring on the phase-compensated matrix to obtain a search square window, which comprises: Performing sliding detection on the phase-compensated matrix along the Doppler axis, performing coarse screening through calculation of the mean value of the reference unit of the target neighborhood, and cooperating with a preset false alarm rate threshold to obtain corresponding over-threshold points; For the threshold crossing point, the local signal-to-noise ratio SNR The search square window is dynamically calculated to realize adaptive variable window topology centring.

9. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 1, characterized in that, The target center coordinates and amplitude are confirmed based on local extreme value constraint and gradient convergence constraint according to the search square window, which comprises: Confirming the center point from the over-threshold points based on the search square window according to the local mechanism constraint; Calculating the Laplacian of the center point and the surrounding four pixels, and confirming the target center coordinates and amplitude if the value of the Laplacian exceeds a preset threshold.

10. The Chebyshev NUFFT based parallel coherent integration method of frequency hopping echo according to claim 1, characterized in that, In the DSP multi-core parallel architecture, task descriptors are distributed by the master core, and the slave cores use L2 SRAM ping-pong buffers to prefetch the data corresponding to the next task block while calculating the data corresponding to the current task block.