Real-time BP imaging processing device and method based on ZYNQMPSoC

By co-designing the hardware and software of the ZYNQMPSoC platform and optimizing Sinc interpolation reconstruction, combined with parallel pipeline design, the computational complexity and resource consumption issues of high-performance real-time BP imaging processing devices were solved, achieving efficient and flexible real-time SAR imaging processing.

CN120949236AActive Publication Date: 2025-11-14AEROSPACE INFORMATION RES INST CAS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511473905.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-11-14
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

In the existing technology, high-performance real-time BP imaging processing devices have high computational complexity, high resource consumption, and insufficient flexibility, making it difficult to meet the requirements of harsh environments such as airborne and spaceborne systems.

Method used

It adopts a hardware and software co-architecture based on ZYNQMPSoC, combined with Sinc interpolation reconstruction and parallel pipeline design, and optimizes computational efficiency and resource consumption by scheduling tasks through the Cortex-A53 processor and using the BP_IP core for imaging calculation.

Benefits of technology

It significantly improves processing speed and flexibility, is suitable for real-time processing of high-resolution SAR imaging, overcomes the shortcomings of high power consumption of GPU, long development cycle of ASIC and insufficient flexibility of FPGA, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949236A_ABST
    Figure CN120949236A_ABST
Patent Text Reader

Abstract

The invention discloses a ZYNQMPSoC-based real-time BP imaging processing device and a ZYNQMPSoC-based real-time BP imaging processing method, and belongs to the technical field of synthetic aperture radars. A software and hardware collaborative architecture of a ZYNQMPSoC platform is adopted, task scheduling and parameter configuration are carried out through a PS end, an optimized BP calculation core is integrated at a PL end, and efficient real-time imaging is achieved. In combination with a Sinc interpolation reconstruction method, the calculation is simplified by using a triangular identical equation, and the logic resource consumption is reduced; and meanwhile, a highly parallel pipeline structure is designed, and a ping-pong cache mechanism is combined, so that the data throughput and the processing efficiency are improved. In addition, parallel expansion of multiple assembly lines is supported, and the imaging time is further shortened. According to the invention, the defects of the existing GPU, ASIC and FPGA platforms in the aspects of power consumption, volume and flexibility are overcome, the method is particularly suitable for airborne, satellite-borne and other scenes with high requirements on real-time performance and integration level, and the performance and energy efficiency ratio of SAR imaging are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of synthetic aperture radar technology, specifically relating to a real-time BP imaging processing device and method based on ZYNQMPSoC. Background Technology

[0002] Synthetic Aperture Radar (SAR) is an active microwave remote sensing technology with all-weather, all-time imaging capabilities, widely used in topographic mapping, disaster monitoring, resource surveys, and battlefield awareness. SAR imaging algorithms are mainly divided into frequency domain algorithms and time domain algorithms. Frequency domain algorithms (such as the range-Doppler algorithm) have efficient decoupling capabilities under regular tracks and standard geometric configurations, but their adaptability to slant-out, wide swath, and complex motion conditions is poor, limiting their application in dynamic environments.

[0003] In contrast, the back projection (BP) method in the time domain algorithm does not rely on model assumptions, has higher imaging accuracy, and can adapt to complex tracks and motion conditions, making it an important means of achieving high-resolution SAR imaging. However, the BP algorithm requires point-by-point accumulation of all pulses for each imaging pixel, resulting in extremely high computational complexity and posing a severe challenge to real-time processing capabilities.

[0004] Currently, high-performance real-time backpropagation (BP) processing devices mainly rely on hardware platforms such as GPUs, ASICs, or FPGAs. While GPUs possess powerful parallel computing capabilities, their high power consumption and large size make them difficult to meet the stringent requirements of airborne and spaceborne scenarios. ASICs excel in performance and energy efficiency, but their long development cycles and poor flexibility make them difficult to adapt to rapidly changing task requirements. FPGAs offer advantages in low power consumption and reconfigurability, but they still face issues of insufficient flexibility and excessive resource consumption when implementing the BP algorithm, which restricts the system's scalability and real-time processing capabilities. Summary of the Invention

[0005] To address the aforementioned technical issues, this invention provides a real-time BP imaging processing device and method based on ZYNQMPSoC. By optimizing computational efficiency through a hardware-software co-architecture and combining Sinc interpolation reconstruction and parallel pipeline design, it significantly reduces resource consumption and improves processing speed and flexibility, making it suitable for the real-time processing needs of high-resolution SAR imaging.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A real-time BP imaging processing device based on ZYNQMPSoC, comprising:

[0008] On the processing system side, it is equipped with a Cortex-A53 processor and a DDR controller, and has an external DDR memory.

[0009] The programmable logic side integrates a pulse compression IP core, a BP_IP core, and a DMA controller;

[0010] The Cortex-A53 processor drives the DMA controller and DDR controller to transmit the raw echo data stream in the DDR memory to the pulse compression IP core for pulse compression preprocessing. The preprocessed data is sent to the BP_IP core for imaging calculation. The final imaging result is written back to the DDR memory through the DMA controller and DDR controller.

[0011] The BP_IP core includes a backward projection module and a parallel pipeline structure for performing backward projection calculations.

[0012] On the other hand, the present invention provides a real-time BP imaging processing method based on ZYNQMPSoC, comprising:

[0013] By processing system-side configuration parameters and scheduling tasks, the raw echo data is transmitted to the programmable logic terminal via the DMA controller.

[0014] Pulse compression preprocessing is performed at the programmable logic stage, and imaging calculations are performed using the BP_IP kernel;

[0015] The imaging results are written back to the DDR memory via the DMA controller;

[0016] The BP_IP core includes a backward projection module and a parallel pipeline structure for performing backward projection calculations.

[0017] Thirdly, the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned real-time BP imaging processing method based on ZYNQMPSoC.

[0018] Fourthly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned real-time BP imaging processing method based on ZYNQMPSoC.

[0019] The beneficial effects of this invention are as follows:

[0020] Hardware-software co-engineering architecture enhances system flexibility and energy efficiency: This invention, based on the ZYNQ MPSoC platform, employs a hardware-software co-engineering design. Task scheduling and parameter configuration are handled by the embedded processor (Cortex-A53) on the PS side, while the core BP computation task is efficiently executed by the hardware accelerator on the PL side. This architecture combines flexibility and high performance, overcoming the drawbacks of high power consumption of GPUs, long development cycles of ASICs, and insufficient flexibility of traditional FPGAs. It is particularly suitable for airborne or spaceborne SAR systems with stringent requirements for real-time performance, integration, and power consumption.

[0021] Optimize computational efficiency and reduce resource consumption through Sinc interpolation reconstruction: To address the high computational complexity of the BP algorithm, this invention combines an optimized Sinc interpolation reconstruction method. By using trigonometric identities, the original eight sine calculations required for each interpolation are compressed into one sine operation and several addition, subtraction, and multiplication operations. This significantly reduces the occupation of logic resources (such as LUTs and DSPs) while maintaining high-precision interpolation results, thus significantly improving hardware computational efficiency.

[0022] Parallel pipeline design and ping-pong caching mechanism improve throughput: This invention employs a highly parallel pipeline architecture, supporting multiple pipeline expansions. Each pipeline independently processes pixel data. Combined with a ping-pong caching mechanism, parallel execution of data loading and computation is achieved, effectively reducing access latency and improving system throughput. Experiments show that this design can significantly shorten imaging time and meet the requirements of real-time SAR processing.

[0023] Suitable for high-resolution SAR imaging and highly scalable: The architecture of this invention can flexibly adjust the number of parallel pipelines according to mission requirements, optimize resource allocation, adapt to SAR imaging tasks of different resolutions and complexities, and has good scalability, providing a technical foundation for future higher-performance real-time imaging systems. Attached Figure Description

[0024] Figure 1 This is a block diagram of the real-time BP imaging processing device based on ZYNQMPSoC of the present invention;

[0025] Figure 2 This is a circuit design diagram for a single production line.

[0026] Figure 3 This is a schematic diagram of two parallel pipeline circuits. Detailed Implementation

[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0028] like Figure 1The diagram shows a block diagram of the real-time back-wave imaging processing device based on the ZYNQMPSoC of this invention. It consists of a processing system (PS) and programmable logic (PL), achieving efficient collaboration via the AXI4 bus using AXI smartconnect. The PS is equipped with a Cortex-A53 processor and a DDR controller, while the PL integrates a pulse compression IP, a BP_IP, and a DMA controller. The Cortex-A53 processor is responsible for algorithm scheduling, parameter configuration, and system control; the pulse compression IP and BP_IP enable high-performance parallel computing. Within the acceleration core, functional sub-modules are connected via an AXI-Stream interface. The system interacts with external DDR memory through the DMA and DDR controllers to ensure high-speed data transmission. The Cortex-A53 processor drives the DMA controller to transmit the raw echo data stream from the DDR to the PL for pulse compression preprocessing. The preprocessed data is then sent to the BP_IP core for imaging calculations, and the final imaging result is written back to the DDR via the DMA controller. This hardware-software co-engineering architecture ensures both system flexibility and maximizes computational performance.

[0029] Under ideal pipelined conditions, the theoretical execution time of the BP algorithm can be estimated based on the image size (ix, iy), the number of pulses p, and the number of parallel pipelines DP. Ignoring the overhead of initial pipeline filling and final flushing, the required number of clock cycles BPcycles can be approximated as: BPcycles = ix × iy × p / DP. Therefore, with a fixed image size and number of pulses, the execution speed depends only on the parallelism DP and the system clock frequency.

[0030] The following section details the principles of the aforementioned BP_IP acceleration core and BP algorithm pipeline.

[0031] like Figure 2 The diagram shows the complete circuit structure of the BP_IP with the aforementioned single pipeline configuration. The following notation conventions are used to represent the various input parameters:

[0032] P: Index of each pulse;

[0033] R: Distance from the platform position to each pixel;

[0034] T min : Delay of the first point of the radar echo, a constant;

[0035] Fs: Range sampling rate;

[0036] x p , y p , z p Cartesian coordinates of the radar platform corresponding to each pulse;

[0037] x, y, z: The position of the pixel in Cartesian coordinates;

[0038] C: Speed ​​of light;

[0039] x k The index position of each pixel in the echo;

[0040] x f :x k The floor value;

[0041] g x,y : Echo value of each pixel;

[0042] m: Phase offset corresponding to each pixel echo value;

[0043] f(x,y): The final value of each pixel.

[0044] The entire hardware architecture adopts a fully pipelined design, with modules built in a streaming manner, allowing for continuous and unblocked data transmission between modules to achieve high throughput.

[0045] This invention employs a truncated 8-point sinc interpolation method to improve interpolation accuracy in BP imaging. In the interpolation module, the raw echo data is sequentially assigned an incrementing integer index x upon input. n Simultaneously, the system calculates the index x of the interpolation position based on the track and pixel coordinates. k The target interpolation index x is calculated. k Then, the sinc weights and multiply-accumulate operations need to be performed on the four data points on each side, for a total of eight data points, to obtain the interpolation result. The sine value is calculated in real time using the CORDIC module. This method has high accuracy, but because each interpolation operation involves eight sine calculations, and each channel in the pipeline needs to perform parallel and independent calculations, the CORDIC module cannot be reused, resulting in a large consumption of logic resources and severely restricting system scalability.

[0046] To improve the hardware efficiency of interpolation operations in BP imaging, this invention optimizes the interpolation module for the BP pipeline structure by combining the Sinc interpolation reconstruction formula. This scheme introduces trigonometric identities to algebraically reconstruct the Sinc expression:

[0047]

[0048] The original eight sine function calculations required for each interpolation are compressed into a combination of one sine calculation and several addition, subtraction and multiplication operations, thereby significantly reducing the consumption of logic resources (such as LUTs and DSPs) by sine calculations.

[0049] Unlike general interpolation methods, this invention is designed for on-chip buffered SAR echo processing systems, fully considering the problems caused by non-streaming data access and complex interpolation paths. Because the slant range relationship between pixels and the radar platform in the BP algorithm is not linear, the interpolation index x... k The changes are unpredictable. In order to find the four sampling points on the left and right sides required for interpolation in real time, the system needs to cache the entire frame of echo data in on-chip memory for repeated access.

[0050] Because the system needs to cache the entire frame of raw echo data on-chip, and in this architecture, the number of pixels to be processed is comparable to the number of sampled echo points, the data loading time is almost equal to the actual computation time. The loading and access latency of the raw echo directly accounts for about half of the entire processing time, severely restricting the overall processing efficiency of the system. To solve this problem, this invention introduces a ping-pong buffer structure, deploying two independent on-chip storage areas at the data receiving end to alternately execute data loading and computation tasks, thereby decoupling read and write operations and improving the system's continuous processing capability without introducing additional waiting cycles.

[0051] The following are Figure 2 The specific circuit design and pipeline principle are described below:

[0052] To meet the requirements of a fully pipelined process, the truncated eight-point Sinc interpolation process employs a highly parallel architecture in the circuit. The system combines eight small-capacity BRAMs into a large-capacity BRAM1 to cache the raw echo data, thus supporting parallel reading of the required sampling points from eight ports. First, based on the track information and image coordinates, the slant range value and its corresponding echo data index for each pixel are calculated in real time. The corresponding phase offset value is then calculated using the slant range value. Eight adjacent echo sampling points are extracted based on the integer part of the index, while eight corresponding Sinc interpolation coefficients are calculated in parallel based on the fractional part of the index. Subsequently, using a parallel multiply-accumulate array, the eight Sinc interpolation coefficients are multiplied one by one with the corresponding sampling points and accumulated to obtain the interpolation result. Finally, the interpolation result is multiplied by the phase offset value using a complex multiplication operation, and then accumulated and updated with the partially cached target pixel values. After each set of pixel values ​​is calculated, the system writes it back to external memory and loads new pixels to continue imaging processing, repeating this iterative process until the entire image is generated.

[0053] The system mainly includes the following on-chip BRAM modules:

[0054] 1) BRAM0: Used to cache platform track information corresponding to the raw echoes being processed on-chip. This memory has a small capacity and is refreshed in real time as the echo frames are updated.

[0055] 2) BRAM1: Used to buffer all sample points of the currently processed single pulse. The system combines eight small-capacity BRAMs into one large-capacity BRAM, and echo data is sequentially written into these eight BRAMs, allowing simultaneous access to eight consecutive sample points during processing. During the processing of the current pulse, the data for the next pulse is preloaded into another set of BRAM1 to implement a typical ping-pong mechanism.

[0056] 3) BRAM2: Used to store pixel values ​​currently being accumulated. Similar to BRAM1, BRAM2 also uses a ping-pong buffer structure: while a group of pixel outputs is being generated, the previous group of pixels is written back to external memory in parallel. It is important to note that the size of BRAM2 has a critical impact on system performance. To achieve a balance between data loading and computation latency, the space of BRAM2 should be comparable to that of BRAM1. When the capacity of BRAM2 is smaller than that of BRAM1, the data loading process will slow down the overall pipeline and cause congestion; while when BRAM2 is larger than BRAM1, although it can reduce the number of repeated reads of the original echo, since the original echo reading process itself is performed in a single pipeline manner (i.e., one sample point is read per clock cycle), the pressure on external memory bandwidth is very small, so it will not improve overall performance, but rather exacerbate the pressure on on-chip resources.

[0057] Considering that the output pixels are independent of each other, Figure 2 The circuit architecture shown supports parallel expansion through pipeline replication. Figure 3 A schematic diagram of the BP_IP core implementation with two parallel pipelines is shown. The designed architecture supports continuous data stream input and can perform high-throughput imaging computation by outputting the number of pixels per pipeline in parallel per cycle.

[0058] As the number of parallel pipelines increases, hardware resource and on-chip memory requirements will grow proportionally. Since echo sampling data cannot be shared, each pipeline needs its own complete copy of the echo data to meet the parallel access needs of multiple pipelines. Similarly, output pixels also require independent accumulation cache space due to the different processing areas of each pipeline.

[0059] The complete system was tested using real UAV echo data at a scale of 8 parallel pipelines and a clock frequency of 200MHz. Specific parameters of the real UAV echo are shown in Table 1.

[0060] Table 1

[0061]

[0062] The specific test results are shown in Table 2:

[0063] Table 2

[0064]

[0065] According to the aforementioned BP theory execution time calculation formula:

[0066] BPtime=4096×3840×264 / 8×5×10^(-9)=2.5952s;

[0067] The actual execution time was 2.5968s, which is highly consistent with the theoretical value, indicating that the system has reached a running speed close to the theoretical limit under the current configuration.

[0068] Calculated based on parameters:

[0069] Data acquisition time = Na / PRF = 409662.499925 ≈ 65.54s;

[0070] In this scenario, the ratio of data acquisition time to processing time is 25.24, indicating that data processing is much faster than acquisition, and the system has significant real-time processing capabilities.

[0071] On the other hand, the present invention provides a real-time BP imaging processing method based on ZYNQMPSoC, including:

[0072] By processing system-side configuration parameters and scheduling tasks, the raw echo data is transmitted to the programmable logic terminal via the DMA controller.

[0073] Pulse compression preprocessing is performed at the programmable logic stage, and imaging calculations are performed using the BP_IP kernel;

[0074] The imaging results are written back to the DDR memory via the DMA controller;

[0075] The BP_IP core includes a backward projection module and a parallel pipeline structure for performing backward projection calculations.

[0076] Thirdly, the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned real-time BP imaging processing method based on ZYNQMPSoC.

[0077] Fourthly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned real-time BP imaging processing method based on ZYNQMPSoC.

[0078] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A real-time BP imaging processing device based on ZYNQMPSoC, characterized in that, include: On the processing system side, it is equipped with a Cortex-A53 processor and a DDR controller, and has an external DDR memory. The programmable logic side integrates a pulse compression IP core, a BP_IP core, and a DMA controller; The Cortex-A53 processor drives the DMA controller and DDR controller to transmit the raw echo data stream in the DDR memory to the pulse compression IP core for pulse compression preprocessing. The preprocessed data is sent to the BP_IP core for imaging calculation. The final imaging result is written back to the DDR memory through the DMA controller and DDR controller. The BP_IP core includes a backward projection module and a parallel pipeline structure for performing backward projection calculations.

2. The real-time BP imaging processing device based on ZYNQMPSoC according to claim 1, characterized in that, It also includes the AXI4 bus, which connects the processing system and the programmable logic to enable data interaction.

3. The real-time BP imaging processing device based on ZYNQMPSoC according to claim 1, characterized in that, The backward projection module uses the truncated 8-point Sinc interpolation method to read 8 adjacent echo sampling points in parallel and complete the interpolation calculation through a multiply-accumulate array.

4. The real-time BP imaging processing device based on ZYNQMPSoC according to claim 3, characterized in that, Based on the track information and image coordinates, the slant range value and its corresponding echo data index for each pixel are calculated in real time. The corresponding phase offset value is then calculated using the slant range value. Eight adjacent echo sampling points are extracted based on the integer part of the index, while eight corresponding Sinc interpolation coefficients are calculated in parallel based on the fractional part of the index. Subsequently, using a parallel multiply-accumulate array, the eight Sinc interpolation coefficients are multiplied one by one with the corresponding sampling points and accumulated to obtain the interpolation result. Finally, the interpolation result is multiplied by the phase offset value and accumulated with the partial target pixel values ​​cached on the chip to update the result. After each set of pixel values ​​is calculated, the system writes it back to external memory and loads new pixels to continue imaging processing. This process is repeated iteratively until the entire image is generated.

5. The real-time BP imaging processing device based on ZYNQMPSoC according to claim 1, characterized in that, The number of parallel pipelines is configurable, and each pipeline independently caches echo data and accumulates pixel values.

6. The real-time BP imaging processing device based on ZYNQMPSoC according to claim 5, characterized in that, The production line includes: BRAM0 module: used to cache platform track information corresponding to the raw echo being processed on-chip; BRAM1 module: used to buffer all sampling points of the single pulse currently being processed; BRAM2 module: Used to store pixel values ​​that are being accumulated.

7. A real-time BP imaging processing device based on ZYNQMPSoC according to claim 6, characterized in that, The BRAM1 module includes eight BRAMs, and the raw echo data is sequentially written into the eight BRAMs. During the current pulse processing, the data of the next pulse will be preloaded into another set of BRAM1 to realize a typical ping-pong mechanism.

8. A real-time BP imaging processing method based on ZYNQMPSoC, characterized in that, include: By processing system-side configuration parameters and scheduling tasks, the raw echo data is transmitted to the programmable logic terminal via the DMA controller. Pulse compression preprocessing is performed at the programmable logic stage, and imaging calculations are performed using the BP_IP kernel; The imaging results are written back to the DDR memory via the DMA controller; The BP_IP core includes a backward projection module and a parallel pipeline structure for performing backward projection calculations.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the real-time BP imaging processing method based on ZYNQMPSoC as described in claim 8.

10. A computer-readable storage medium, characterized in that, It stores executable instructions, which, when executed by a processor, enable the processor to implement the real-time BP imaging processing method based on ZYNQMPSoC as described in claim 8.

Citation Information

Patent Citations

  • Polar coordinate format algorithm SAR imaging method based on FPGA

    CN116413724A

  • Efficient BP imaging method based on scheduling acceleration

    CN119959946A

  • Two-dimensional real-time back projection sensing imaging method, device, equipment and medium

    CN120085300A

  • Terahertz radar real-time imaging device and method based on dual-processor heterogeneous platform

    CN120652470A

  • Image processor and its method

    JP1999167627A