Video SAR real-time imaging processing system, method and equipment based on mercuric chloride NPU cluster and FPGA isomerism

Through Ascend NPU cluster and FPGA heterogeneous video SAR real-time imaging system, the problem of low imaging frame rate in traditional SAR systems is solved, high frame rate imaging is achieved, and system power consumption and development costs are reduced. It is suitable for real-time imaging of drone video SAR.

CN120334912APending Publication Date: 2025-07-18XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510489633.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional SAR systems have low imaging frame rates and are difficult to meet the real-time requirements of motion target detection. The existing processing platforms have problems such as high development difficulty, high cost, large volume and high power consumption.

Method used

Using a video SAR real-time imaging system based on Ascend NPU cluster and FPGA heterogeneous, FPGA is used for high-speed signal acquisition and transmission, and Ascend NPU performs parallel computing, and through pipeline architecture and task-level parallel optimization, combined with Ascend C programming, high frame rate imaging is achieved.

Benefits of technology

Real-time video SAR imaging with high frame rate is achieved, which shortens the development cycle, reduces system power consumption and volume, and provides a cost-effective solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120334912A_ABST
    Figure CN120334912A_ABST
Patent Text Reader

Abstract

The invention discloses a video SAR (Synthetic Aperture Radar) real-time imaging processing system, method and equipment based on mercuric chloride NPU (Network Processing Unit) cluster and FPGA (Field Programmable Gate Array) isomerism, and belongs to the field of radar imaging and signal processing, the system comprises a DDS (Direct Digital Synthesizer) signal generation module, a digital-to-analog conversion module, a signal acquisition module, a data processing module and a data storage module; according to the method, generation and acquisition of a linear frequency modulation signal, digital down-conversion processing and data communication between mercuric chloride NPU clusters are completed through an FPGA based on the system, realization of an RD algorithm and an MD algorithm which can be subjected to parallel calculation is implemented on the mercuric chloride NPU clusters, and through pipeline scheduling and task parallelization, the mercuric chloride NPU clusters are subjected to parallel calculation. The full-process processing from signal generation and acquisition, preprocessing to an imaging algorithm is realized by adopting a pipeline processing mode, and high-frame-rate video SAR real-time imaging processing is realized; the equipment is used for implementing the method. The invention has the advantages of short development period, small volume, low power consumption and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radar imaging and signal processing, and particularly relates to a video SAR real-time imaging processing system, method and device based on the heterogeneous combination of Ascend NPU clusters and FPGAs. Background Art

[0002] Synthetic Aperture Radar (SAR) can penetrate clouds, dust, vegetation, rain and snow, etc., and perform high-resolution imaging on the imaging area at any time of the day, without being restricted by external factors such as weather and light. According to different application scenario requirements, synthetic aperture radar can be deployed on various platforms such as missiles, satellites and unmanned aerial vehicles to achieve functions such as missile precision guidance, battlefield situation awareness, geographical resource observation and disaster monitoring and assessment, and play a huge role in modern military, scientific research and civilian fields. However, in order to achieve an ideal azimuth resolution, traditional SAR systems usually require a large accumulation angle. Due to the limited speed of the carrier platform, traditional SAR takes a long time in the process of synthetic aperture formation, resulting in a low imaging frame rate, so there are certain limitations in moving target detection.

[0003] Video SAR can generate high-frame-rate continuous frame images by continuously irradiating the target imaging area. Compared with traditional SAR, video SAR not only has the characteristics of high resolution, but also can perform real-time imaging at a high frame rate, intuitively presenting the dynamic change trend of the target area and providing rich dynamic information for subsequent data processing. This ability makes video SAR have important application value in the fields of geological structure interpretation, surface deformation monitoring and ecological environment assessment.

[0004] The core goal of a video SAR real-time imaging system is to achieve low-latency and high-frame-rate dynamic imaging while ensuring high resolution, which poses strict requirements on the system architecture design and the selection of the processing platform. Currently, the mainstream real-time signal processing platforms mainly include CPUs, DSPs, FPGAs and GPUs. Although the CPU has a relatively high main frequency, due to its serial computing characteristics, it is difficult to provide sufficient processing performance, especially limited in parallel computing tasks; the DSP system not only has the problem of limited storage bandwidth, but also its clock frequency is difficult to increase, resulting in low parallel operation efficiency and difficult to meet the real-time imaging requirements of video SAR; FPGAs have high data throughput capabilities, but the development cycle of FPGAs from algorithm implementation to hardware verification is long, and the debugging process is also relatively complex. These factors increase the development difficulty and cost, restricting its application in the field of video SAR; GPUs are suitable for tasks with high parallelism and high computational complexity, and at the same time have a short development cycle. However, traditional GPUs have the disadvantages of large volume and high power consumption, and are difficult to be directly applied to the field of real-time imaging processing. Summary of the Invention

[0005] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a video SAR real-time imaging processing system, method and device based on the heterogeneous combination of Ascend NPU clusters and FPGAs. This system integrates FPGAs and Ascend NPU clusters to achieve video SAR real-time imaging. The FPGA is responsible for high-speed signal acquisition and transmission, and the NPU gives full play to its parallel computing advantages to complete real-time processing. A pipeline architecture is adopted to perform task-level parallel optimization on data acquisition, preprocessing, and imaging algorithms, and the amount of computation is reduced through dynamic allocation of computing resources; at the hardware level, the rich interfaces of the FPGA are combined with the energy efficiency ratio advantages of the NPU, significantly reducing the volume of the device and lowering power consumption. At the software level, Ascend C language programming is used to accelerate algorithms, improving the imaging frame rate while shortening the development cycle; the present invention has the advantages of strong processing performance, high development efficiency, compact hardware, and low power consumption, providing a cost-effective solution for airborne real-time imaging.

[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0007] A video SAR real-time imaging processing system based on the heterogeneous combination of Ascend NPU clusters and FPGAs. The Ascend NPU cluster consists of five Ascend NPUs to form a data processing module, which are respectively denoted as Node #1, Node #2, Node #3, Node #4, and Node #5. Each Ascend NPU is connected to the FPGA through a 2.5G Ethernet, and Node #1 to Node #5 respectively complete the data processing tasks of the data processing module; the FPGA includes a DDS signal generation module, a digital-to-analog conversion module, a signal acquisition module, and a data storage module;

[0008] The DDS signal generation module is used to generate a chirp signal to be transmitted;

[0009] The digital-to-analog conversion module is used to convert the chirp signal generated by the DDS signal generation module from a digital signal to an analog signal through a DAC chip and radiate it through a transmitting antenna;

[0010] The signal acquisition module is used to convert the analog signal received by the receiving antenna into a digital signal and perform digital down-conversion processing to obtain a zero-IF baseband signal;

[0011] The data processing module is responsible for implementing the core algorithm of the system, performing RD+MD algorithm processing on the zero-IF baseband signal processed by the signal acquisition module, and realizing video SAR real-time imaging;

[0012] The data storage module is used to cache the data after digital down-conversion processing in the signal acquisition module, as well as the input and output data of the data processing module.

[0013] The signal acquisition module includes an analog-to-digital conversion sub-module and a preprocessing sub-module;

[0014] The analog-to-digital conversion sub-module is used to convert the analog signal received by the receiving antenna into a digital signal through an ADC chip;

[0015] The preprocessing sub-module is used to perform digital down-conversion processing on the digital signal converted by the analog-to-digital conversion sub-module to obtain a zero-intermediate-frequency baseband signal.

[0016] A method for a real-time imaging processing system based on video SAR includes the following steps:

[0017] Step 1, the DDS signal generation module generates a chirp signal, and then the signal acquisition module collects and preprocesses the received signal to obtain a zero-intermediate-frequency baseband signal;

[0018] Step 2, the zero-intermediate-frequency baseband signal data obtained after the preprocessing in Step 1 is sent to Node #1 through a 2.5G Ethernet, and Node #1 performs the RD algorithm processing;

[0019] Step 3, the data processed by the RD algorithm in Step 2 is transmitted to the FPGA through a 2.5G Ethernet;

[0020] Step 4, after the FPGA receives the data transmitted in Step 3, it starts to perform MD algorithm processing, performs sub-block segmentation, divides it into three data blocks in an overlapping segmentation manner and transmits them to Node #2, Node #3, and Node #4 respectively;

[0021] Step 5, after Node #2, Node #3, and Node #4 receive the data transmitted in Step 4, they perform parallel calculation of the error estimation of the three data blocks. At the same time, the FPGA sends the data processed by the RD algorithm transmitted in Step 3 to Node #5;

[0022] Step 6, the three groups of error estimation results calculated in Step 5 are transmitted to the FPGA for phase error splicing, the spliced phase error data is sent to Node #5, and then the data processed by the RD algorithm sent to Node #5 in Step 5 is subjected to phase error compensation and imaging Dechirp. Thus, the MD algorithm processing is completed, and finally the imaging result of a frame of image is obtained.

[0023] The above Steps 1 to 6 implement the processing flow of a frame of image. When the present invention implements multi-frame image processing, a processing method of pipeline scheduling and task parallelism is adopted. Specifically, when the nth frame of image processing executes Step 2, the (n + 1)th frame of image processing executes Step 1, and through the pipeline processing method, high-frame-rate video SAR real-time imaging is realized.

[0024] Specifically, Step 1 includes:

[0025] Step 1.1: The DDS signal generator generates a broadband linear frequency modulation signal, and the DAC chip of the digital-to-analog conversion module converts the digital signal into an analog signal.

[0026] Step 1.2: The ADC chip of the analog-to-digital conversion sub-module converts the received echo intermediate-frequency analog signal into a digital signal.

[0027] Step 1.3: The preprocessing sub-module preprocesses the digital signal converted in Step 1.2 and processes it into a zero-intermediate-frequency baseband signal through digital down-conversion.

[0028] Step 1.4: The data storage module stores the data preprocessed in Step 1.3.

[0029] The specific steps of Step 2 are as follows:

[0030] Step 2.1: Transmit the zero-intermediate-frequency baseband signal to Node #1 through a 2.5G Ethernet.

[0031] Step 2.2: Perform a range Fourier transform on the data transmitted to Node #1 in Step 2.1, then multiply it by the phase compensation function H1, and finally perform a matrix transpose operation on the result.

[0032] Step 2.3: Perform an azimuth Fourier transform on the matrix transpose result of Step 2.2, then perform a matrix transpose operation, and then multiply it by the second range pulse compression function and the inertial navigation compensation parameter H2.

[0033] Step 2.4: Perform an inverse range Fourier transform on the calculation result of Step 2.3, multiply it by the square root to quadratic function H3, and finally perform a matrix transpose operation.

[0034] Step 2.5: Perform an inverse azimuth Fourier transform on the matrix transpose result of Step 2.4, then perform a matrix transpose operation to complete the preliminary imaging of the RD algorithm.

[0035] The specific steps of Step 4 are as follows:

[0036] First, the FPGA caches the data into the data storage module, then reads the data sequentially, and in the way of overlapping and segmented blocks, reads three data blocks sequentially and transmits the data to Node #2, Node #3, and Node #4 in parallel.

[0037] The specific error estimation calculation in Step 5 is as follows:

[0038] Step 5.1: Perform azimuth Dechirp processing on the data transmitted to Node #2, Node #3, and Node #4 in Step 4 in parallel. At the same time, the FPGA transmits the data processed by the RD algorithm in Step 3 to Node #5.

[0039] Step 5.2: Parallelly perform chunking operations on the output data of Node #2, Node #3, and Node #4 in Step 5.1. The output result of each node is evenly divided into two sub-chunk data of the same size. Then, parallelly perform zero-padding operations to zero-pad the two sub-chunks of each node into data chunks of the same scale as the data processed by the RD algorithm transmitted in Step 3. Next, parallelly perform azimuth Fourier transform operations on the two sub-chunk data of each node;

[0040] Step 5.3: Parallelly perform operations on the azimuth frequency domain amplitudes of the results after the azimuth Fourier transform operations in Node #2, Node #3, and Node #4 in Step 5.2, and then parallelly perform azimuth Fourier transform operations on the three nodes;

[0041] Step 5.4: Parallelly perform conjugate operations on the results after the azimuth Fourier transform operations in Node #2, Node #3, and Node #4 in Step 5.3 to obtain a frequency domain cross-correlation matrix data chunk, and parallelly perform inverse azimuth Fourier transform operations on the cross-correlation matrix data chunks of the three nodes;

[0042] Step 5.5: Parallelly perform azimuth amplitude calculations on the results after the inverse azimuth Fourier transform operations in Node #2, Node #3, and Node #4 in Step 5.4, and accumulate along the range direction to generate a real cross-correlation sequence. Then, the three nodes respectively search for the maximum value in the cross-correlation sequence in parallel according to the azimuth direction, and perform error fitting estimation calculations to obtain an error estimation data chunk. Thus, the error estimation calculations of the three data chunks by Node #2, Node #3, and Node #4 are completed.

[0043] The specific steps of Step 6 include:

[0044] Step 6.1: Parallelly transmit the three groups of error estimation results calculated by Node #2, Node #3, and Node #4 in Step 5 to the FPGA, then perform phase error splicing in the FPGA, and finally send the spliced phase error data to Node #5;

[0045] Step 6.2: In Node #5, through differential derivative calculations and least squares fitting calculations, reconstruct the motion error data spliced in Step 6.1 into a data chunk of the same scale as the data processed by the RD algorithm transmitted in Step 3, and then compensate the reconstructed motion error data to the data processed by the RD algorithm received by Node #5;

[0046] Step 6.3: In Node #5, parallelly perform the product operation of the data after motion error compensation obtained in Step 6.2, the range window function, and the azimuth dechirp function to complete data compensation and obtain a frame of imaging result.

[0047] A video SAR real-time imaging processing device based on the heterogeneous combination of Ascend NPU clusters and FPGAs, comprising:

[0048] A memory: used to store computer programs for implementing the video SAR real-time imaging processing method;

[0049] A processor: used to implement the video SAR real-time imaging processing method when executing the computer program.

[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0051] The present invention designs a video SAR real-time imaging processing system based on the heterogeneous combination of Ascend NPU clusters and FPGAs, divides the processing process into multiple stages, optimizes the processing time of each stage, and realizes high-frame-rate video SAR real-time imaging through pipelining.

[0052] First, taking advantage of the rich high-speed interfaces and high data throughput of FPGAs, the generation and acquisition of chirp signals, digital down-conversion processing, and data communication between Ascend NPU clusters are implemented on the FPGA, reducing system latency and increasing data throughput.

[0053] Second, Ascend NPU is a real-time processing platform dedicated to high-performance computing, with significant advantages in parallel computing, high performance, low power consumption, and a shorter development cycle. Therefore, the implementation of the parallelizable RD algorithm and MD algorithm is carried out on the Ascend NPU cluster to utilize the parallel processing capability of its tensor computing core and reduce the pipeline latency of algorithm processing.

[0054] Finally, the system integrates the advantages of FPGAs and Ascend NPUs, and through pipeline scheduling and task parallelism, realizes the full-process processing from signal generation and acquisition, preprocessing to imaging algorithms, achieving high-frame-rate video SAR real-time imaging. Compared with completing imaging processing on FPGAs, the video SAR real-time imaging processing system based on the heterogeneous combination of Ascend NPU clusters and FPGAs has a shorter development cycle; compared with completing imaging processing on GPUs, the video SAR real-time imaging processing system based on the heterogeneous combination of Ascend NPU clusters and FPGAs has the advantages of lower power consumption, smaller volume, and lower cost. At the same time, as a feasible domestic alternative solution, this method can be applied to the field of airborne video SAR real-time imaging processing.

[0055] In summary, the video SAR real-time imaging processing system based on the heterogeneous combination of Ascend NPU clusters and FPGAs proposed by the present invention has the advantages of a shorter development cycle, smaller volume, lower power consumption, and lower cost. Description of the Drawings

[0056] Figure 1 It is a schematic diagram of the system structure of the present invention.

[0057] Figure 2 This is the flowchart of the video SAR real-time imaging algorithm of the present invention.

[0058] Figure 3 This is the flowchart of the RD algorithm of the present invention.

[0059] Figure 4 This is the flowchart of the MD algorithm of the present invention.

[0060] Figure 5 This is the schematic diagram of the task pipeline of each stage of the imaging processing method of the present invention.

[0061] Figure 6(a) is the data processing result diagram of the RD algorithm based on Ascend NPU of the present invention, and Figure 6(b) is the data processing result diagram of the RD algorithm based on MATLAB.

[0062] Figure 7 This is the absolute error diagram of the processing result of the RD algorithm based on Ascend NPU and the processing result of the RD algorithm of MATLAB of the present invention.

[0063] Figure 8(a) is the imaging result diagram of the RD+MD algorithm based on Ascend NPU, and Figure 8(b) is the imaging result diagram of the RD+MD algorithm based on MATLAB.

[0064] Figure 9 This is the absolute error diagram of the processing result of the RD+MD algorithm based on Ascend NPU and the processing result of the RD+MD algorithm of MATLAB of the present invention. Detailed implementation manners

[0065] The following describes the present invention in detail with reference to the accompanying drawings.

[0066] As Figure 1 shown, a video SAR real-time imaging processing system based on the heterogeneous combination of Ascend NPU clusters and FPGAs. The Ascend NPU cluster consists of five Ascend NPUs to form a data processing module, which are respectively denoted as Node #1, Node #2, Node #3, Node #4, and Node #5. Each Ascend NPU is connected to the FPGA through a 2.5G Ethernet. Node #1 to Node #5 respectively complete the data processing tasks of the data processing module. The FPGA includes a DDS signal generation module, a digital-to-analog conversion module, a signal acquisition module, and a data storage module;

[0067] The DDS signal generation module is used to generate the chirp signal to be transmitted;

[0068] The digital-to-analog conversion module is used to convert the chirp signal generated by the DDS signal generation module from a digital signal to an analog signal through a DAC chip and radiate it through a transmitting antenna;

[0069] The signal acquisition module is used to convert the analog signal received by the receiving antenna into a digital signal, perform digital down-conversion processing, and obtain a zero-intermediate-frequency baseband signal.

[0070] The data processing module is responsible for implementing the core algorithm of the system, performing RD+MD algorithm processing on the zero-intermediate-frequency baseband signal processed by the signal acquisition module, and realizing real-time video SAR imaging.

[0071] The data storage module is used to cache the data after digital down-conversion processing in the signal acquisition module, as well as the input and output data of the data processing module.

[0072] The signal acquisition module includes an analog-to-digital conversion sub-module and a preprocessing sub-module.

[0073] The analog-to-digital conversion sub-module is used to convert the analog signal received by the receiving antenna into a digital signal through an ADC chip.

[0074] The preprocessing sub-module is used to perform digital down-conversion processing on the digital signal converted by the analog-to-digital conversion sub-module to obtain a zero-intermediate-frequency baseband signal.

[0075] As Figure 2 shown, a method for a real-time video SAR imaging processing system includes the following steps:

[0076] Step 1, the DDS signal generation module generates a chirp signal, and then the signal acquisition module collects and preprocesses the received signal to obtain a zero-intermediate-frequency baseband signal, which is 2048*2048 single-precision floating-point complex data in this embodiment.

[0077] Step 2, the 2048*2048 single-precision floating-point complex data obtained after preprocessing in Step 1 is sent to Node #1 through a 2.5G Ethernet, and Node #1 performs RD algorithm processing. The processing duration of Node #1 for the RD algorithm is within 200 ms.

[0078] Step 3, the data processed by the RD algorithm in Step 2 is transmitted to the FPGA through a 2.5G Ethernet. The data size is 32 MB, and the transmission time is within 200 ms.

[0079] Step 4, after the FPGA receives the data transmitted in Step 3, it starts to perform MD algorithm processing, performs sub-block segmentation, divides it into three 2048*1024 data blocks in an overlapping segmentation manner, and transmits them to Node #2, Node #3, and Node #4 respectively. The data storage module can achieve a measured data throughput of 4 GB / s, and the time for writing and reading 32 MB of data is about 5.2 ms, meeting the real-time read and write requirements of single-frame data.

[0080] Step 5: After receiving the data transmitted in step 4, nodes #2, #3, and #4 calculate the error estimates of the three data blocks in parallel. This process takes less than 200ms. At the same time, FPGA sends the data processed by the RD algorithm transmitted in step 3 to node #5 to achieve the purpose of hiding the data transmission delay.

[0081] Step 6, the three sets of error estimation results calculated in step 5 are transmitted to FPGA for phase error splicing, and the spliced phase error data are sent to node #5. Then, the data processed by the RD algorithm sent to node #5 in step 5 is subjected to phase error compensation and imaging Dechirp. This process takes less than 200ms. At this point, the MD algorithm processing is completed, and finally the imaging result of one frame of image is obtained.

[0082] The steps 1 to 6 implement the processing flow of one frame of image. When the present invention implements multi-frame image processing, a processing method of pipeline scheduling and task parallelism is adopted, specifically: when the nth frame image processing executes step 2, the n+1th frame image processing executes step 1, and high frame rate video SAR real-time imaging is realized by pipeline processing.

[0083] The step 1 specifically includes:

[0084] Step 1.1, the DDS signal generator generates a broadband linear frequency modulation signal, and converts the digital signal into an analog signal through the DAC chip of the digital-to-analog conversion module;

[0085] Step 1.2, the ADC chip of the analog-to-digital conversion submodule converts the received echo intermediate frequency analog signal into a digital signal;

[0086] Step 1.3, the preprocessing submodule preprocesses the digital signal converted in step 1.2 into a zero intermediate frequency baseband signal through digital down-conversion;

[0087] Step 1.4, storing the data preprocessed in step 1.3 through the data storage module.

[0088] The step 2 specifically includes:

[0089] Step 2.1, send the zero-IF baseband signal to node #1 via 2.5G Ethernet;

[0090] Step 2.2, perform a distance Fourier transform on the data sent to node #1 in step 2.1, then multiply it with the phase compensation function H1, and finally perform a matrix transpose operation on the result;

[0091] Step 2.3, perform azimuth Fourier transform on the matrix transposition result of step 2.2, then perform matrix transposition operation, and then multiply it with the quadratic range pulse pressure function and the inertial navigation compensation parameter H2;

[0092] Step 2.4: Perform range inverse Fourier transform on the calculation result of Step 2.3, multiply it by the square root to quadratic function H3, and finally perform matrix transpose operation;

[0093] Step 2.5: Perform azimuth inverse Fourier transform on the matrix transpose result of Step 2.4, then perform matrix transpose operation to complete the preliminary imaging of the RD algorithm, as Figure 3 shown.

[0094] The specific steps of Step 4 include:

[0095] Firstly, the FPGA caches the data into the data storage module, then reads out the data sequentially, and in the block division method of overlapping segmentation, reads three data blocks sequentially and transmits the data to Node #2, Node #3, and Node #4 in parallel. The total time consumed by Step 4 is within 200 ms.

[0096] The error estimation calculation in Step 5 specifically includes:

[0097] Step 5.1: Parallelly perform azimuth Dechirp processing on the data transmitted to Node #2, Node #3, and Node #4 in Step 4. Meanwhile, the FPGA sends the data processed by the RD algorithm transmitted in Step 3 to Node #5;

[0098] Step 5.2: Parallelly perform block operation on the output data of Node #2, Node #3, and Node #4 in Step 5.1. The output result of each node is divided into two sub-block data of the same size. Then, perform zero-padding operation in parallel to zero-pad the two sub-blocks of each node into data blocks with the same scale as the data processed by the RD algorithm transmitted in Step 3. Then, parallelly perform azimuth Fourier transform operation on the two sub-block data of each node;

[0099] Step 5.3: Parallelly perform the operation of azimuth frequency domain amplitude on the results after the azimuth Fourier transform operation of Node #2, Node #3, and Node #4 in Step 5.2, and then parallelly perform azimuth Fourier transform operation on the three nodes;

[0100] Step 5.4: Parallelly perform conjugate operation on the results after the azimuth Fourier transform operation of Node #2, Node #3, and Node #4 in Step 5.3 to obtain the frequency domain cross-correlation matrix data block, and parallelly perform azimuth inverse Fourier transform operation on the cross-correlation matrix data blocks of the three nodes;

[0101] Step 5.5: Parallelly perform azimuth amplitude calculation on the results after azimuth inverse Fourier transform operation for Node #2, Node #3, and Node #4 in Step 5.4, accumulate along the range direction to generate a real cross-correlation sequence, and then the three nodes respectively search for the maximum value in the cross-correlation sequence in parallel according to the azimuth direction, and perform error fitting estimation calculation to obtain an error estimation data block. Thus, the error estimation calculation for the three data blocks by Node #2, Node #3, and Node #4 is completed.

[0102] The specific steps of Step 6 are as follows:

[0103] Step 6.1: Parallelly transmit the three groups of error estimation results calculated by Node #2, Node #3, and Node #4 in Step 5 to the FPGA, then perform phase error splicing in the FPGA, and finally send the spliced phase error data to Node #5.

[0104] Step 6.2: In Node #5, through differential derivative calculation and least squares fitting calculation, reconstruct the motion error data spliced in Step 6.1 into a data block with the same scale as the data processed by the RD algorithm transmitted in Step 3, and then compensate the reconstructed motion error data to the data processed by the RD algorithm received by Node #5.

[0105] Step 6.3: Parallelly perform the product operation of the data after motion error compensation obtained in Step 6.2, the range window function, and the azimuth dechirp function in Node #5 to complete the data compensation. Thus, the MD algorithm processing is completed, and an imaging result of one frame is obtained, as Figure 4 shown.

[0106] A video SAR real-time imaging processing device based on the heterogeneous combination of Ascend NPU clusters and FPGAs includes:

[0107] A memory: used to store a computer program for implementing the video SAR real-time imaging processing method;

[0108] A processor: used to implement the video SAR real-time imaging processing method when executing the computer program.

[0109] As Figure 5As shown, it presents a schematic diagram of the task pipeline operation at each stage of the imaging processing method of the present invention. Among them, orange represents the processing flow of the first-frame data. t1 is the calculation time of the RD algorithm by node #1. t2 is the time required to transmit the output to the FPGA after node #1 finishes the calculation. t3 is the time required for the FPGA to perform data chunking. t4 is the time required for the FPGA to transmit in parallel to three nodes (node #2, 3, 4) respectively. t5 is the time required for the three nodes (node #2, 3, 4) to perform parallel processing of the error estimation of the three segmented sub-blocks. At the same time, the FPGA also sends the processing result of the RD algorithm to node #5, which is represented as t6 in the figure. t7 is the time required for node #5 to calculate the error compensation and Dechirp. From the perspective of pipeline processing, the time difference between the two images is t8, and the value of t8 is equal to the longest time-consuming in each stage. The longest time-consuming of all the above steps is only 158.47 ms, which is far less than the 200 ms required for the video SAR imaging time interval and far exceeds the requirement of the video SAR imaging frame rate of 5 Hz.

[0110] Figure 6(a) and Figure 6(b) respectively present the comparison of the imaging results of the RD algorithm implemented based on the Ascend NPU and MATLAB platforms. Figure 7 shows the absolute error distribution of the two results. The experimental results show that the implementation result of the Ascend processor has good consistency with the MATLAB reference result, and its relative error is of the order of 10 -3 and remains within the acceptable range. Through the analysis of the phase error, it is found that the difference between the two mainly stems from the slight difference in the floating-point operation precision. This difference is within the engineering allowable range and will not have a substantial impact on the subsequent motion compensation processing of the MD algorithm. The comparison results verify the accuracy and reliability of the implementation scheme of the Ascend processor and provide an experimental basis for subsequent real-time processing.

[0111] After completing the RD algorithm processing, in order to further improve the imaging quality, it is necessary to perform motion error estimation and compensation on the data. This process effectively eliminates the imaging distortion caused by platform motion by accurately estimating the phase error introduced by motion and compensating it into the RD processing result. To verify the implementation effect of the RD+MD algorithm based on the Ascend processor, the data after RD processing is respectively input into the Ascend NPU and MATLAB platforms for motion compensation processing. Among them, the processing result of the Ascend NPU is transmitted to the PC side through a 2.5G high-speed Ethernet interface, and MATLAB is used for imaging display and result analysis. This verification method not only ensures the accuracy of the algorithm implementation but also provides a reliable basis for the system performance evaluation. Figure 8(a) and Figure 8(b) respectively show the imaging results of the RD algorithm + MD algorithm implemented based on the Ascend NPU and MATLAB platforms.

[0112] The absolute error graph between the imaging results of the Ascend NPU and those of MATLAB is as follows Figure 9 As shown, it can be seen that the maximum absolute error does not exceed 0.002, which is within the acceptable range. This indicates that the real-time imaging processing of video SAR based on the Ascend NPU can achieve high-precision imaging results and meet the actual application requirements.

[0113] To further quantify the accuracy of the imaging results of the Ascend processor, this section introduces the Peak Signal-to-Noise Ratio (PSNR) and the Structural Similarity Index (SSIM) to conduct a comparative analysis of the imaging results between the Ascend NPU and MATLAB. PSNR and SSIM are commonly used image quality assessment criteria. PSNR mainly judges the image quality by measuring the distortion degree between two images, while SSIM is used to evaluate the structural similarity of images. During the algorithm optimization process of video SAR images, in order to comprehensively understand the impact of the optimization method on the image quality, it is usually necessary to conduct a comparative analysis through these two metrics.

[0114] For the video SAR imaging results processed by the Ascend NPU and by MATLAB software, the PSNR and SSIM metrics are used to quantify and analyze the changes in image quality under the two processing methods. The specific results of PSNR and SSIM are shown in the following table.

[0115] Results of PSNR and SSIM

[0116] PSNR SSIM Image 47.1208dB 0.9659

[0117] From the analysis of the above table, it can be seen that the PSNR value of the result image of the RD+MD algorithm processed by the Ascend NPU is relatively high, indicating that the difference between the imaging results of the Ascend NPU and MATLAB is extremely small and the quality is good; at the same time, the SSIM value is close to 1, indicating that the two images are highly consistent in terms of structure, brightness, and contrast. The quantitative analysis of PSNR and SSIM further verifies the correctness of the processing results of the video SAR real-time imaging processing algorithm on the Ascend NPU, and its imaging accuracy meets the requirements of video SAR.

[0118] At present, the international situation is severe and complex, and the demand for domestic substitution in the field of technology is increasing day by day. High-performance computing chips such as FPGA and GPU are almost monopolized by foreign countries. The process of domestic substitution in the field of video SAR imaging processing is of crucial importance. The Ascend processor is a high-performance processor independently developed by Huawei. At present, its applications mostly focus on the field of artificial intelligence. However, relevant research work on radar signal processing based on the Ascend processor has been carried out in China. In the field of SAR imaging, the Ascend processor, with the characteristics of high performance and low power consumption, can meet the requirements of video SAR systems for processing performance, volume and power consumption. In addition, the development difficulty of the Ascend processor is not high, it has strong flexibility and low cost. Therefore, the research on the video SAR real-time imaging processing technology for the Ascend processor is of great significance.

Claims

1. A real-time video SAR imaging processing system based on the heterogeneous combination of Ascend NPU clusters and FPGAs, characterized in that, The Ascend NPU cluster is composed of five Ascend NPUs forming a data processing module, which are respectively denoted as node #1, node #2, node #3, node #4, and node #5. Each Ascend NPU is connected to the FPGA via 2.5G Ethernet. Node #1 to node #5 respectively complete the data processing tasks of the data processing module; the FPGA includes a DDS signal generation module, a digital-to-analog conversion module, a signal acquisition module, and a data storage module; The DDS signal generating module is used to generate a linear frequency modulation signal to be transmitted; The digital-to-analog conversion module is used to convert the linear frequency modulation signal generated by the DDS signal generation module into an analog signal through a DAC chip, and radiate it through a transmitting antenna; The signal acquisition module is used to convert the analog signal received by the receiving antenna into a digital signal and perform digital down-conversion processing to obtain a zero intermediate frequency baseband signal; The data processing module is responsible for the implementation of the core algorithm of the system, and performs RD+MD algorithm processing on the zero intermediate frequency baseband signal obtained by the signal acquisition module to realize real-time imaging of video SAR; The data storage module is used to cache the data after digital down-conversion processing in the signal acquisition module, and the input and output data of the data processing module.

2. The system according to claim 1, characterized in that The signal acquisition module includes an analog-to-digital conversion submodule and a preprocessing submodule; The analog-to-digital conversion submodule is used to convert the analog signal received by the receiving antenna into a digital signal through the ADC chip; The preprocessing submodule is used to perform digital down-conversion processing on the digital signal converted by the analog-to-digital conversion submodule to obtain a zero intermediate frequency baseband signal.

3. A real-time video SAR imaging processing method based on the system described in claim 1 or 2, characterized in that, The following steps are involved: Step 1: The DDS signal generation module generates a linear frequency modulation signal, and then the signal acquisition module collects and preprocesses the received signal to obtain a zero intermediate frequency baseband signal; Step 2: Send the zero intermediate frequency baseband signal data obtained after preprocessing in step 1 to node #1 via 2.5G Ethernet, and node #1 performs RD algorithm processing; Step 3, transmitting the data processed by the RD algorithm in step 2 to the FPGA via 2.5G Ethernet; Step 4, after receiving the data transmitted in step 3, FPGA starts to process the data using the MD algorithm, performs sub-block segmentation, divides the data into three data blocks in an overlapping segmentation manner, and transmits the data to node #2, node #3, and node #4 respectively; Step 5: After receiving the data transmitted in step 4, nodes #2, #3, and #4 calculate the error estimates of the three data blocks in parallel. At the same time, FPGA sends the data processed by the RD algorithm transmitted in step 3 to node #5. Step 6, transmit the three sets of error estimation results calculated in step 5 to FPGA, perform phase error splicing, send the spliced phase error data to node #5, and then perform phase error compensation and imaging Dechirp on the data processed by the RD algorithm sent to node #5 in step 5. At this point, the MD algorithm processing is completed, and finally the imaging result of a frame of image is obtained.

4. The method according to claim 3, wherein The above steps 1 to 6 implement the processing flow of one frame of image. When the present invention implements the processing of multiple frames of images, a processing method of pipeline scheduling and task parallelism is adopted. Specifically: when the processing of the nth frame of image executes step 2, the processing of the (n + 1)th frame of image executes step 1. Through the way of pipeline processing, real-time imaging of high-frame-rate video SAR is realized.

5. The method according to claim 3, characterized in that The above step 1 specifically includes: Step 1.1, a DDS signal generator generates a broadband chirp signal, and a DAC chip of the digital-to-analog conversion module converts the digital signal into an analog signal; Step 1.2, an ADC chip of the analog-to-digital conversion sub-module converts the received echo intermediate-frequency analog signal into a digital signal; Step 1.3, the preprocessing sub-module preprocesses the digital signal converted in step 1.2, and processes it into a zero-intermediate-frequency baseband signal through digital down-conversion; Step 1.4, the data after preprocessing in step 1.3 is stored through the data storage module.

6. The method according to claim 3, characterized in that, The above step 2 specifically includes: Step 2.1, the zero-intermediate-frequency baseband signal is sent to node #1 through a 2.5G Ethernet; Step 2.2, perform a range Fourier transform on the data sent to node #1 in step 2.1, then multiply it by the phase compensation function H1, and finally perform a matrix transpose operation on the result; Step 2.3, perform an azimuth Fourier transform on the matrix transpose result in step 2.2, then perform a matrix transpose operation, and then multiply it by the second range pulse compression function and the inertial navigation compensation parameter H2; Step 2.4, perform an inverse range Fourier transform on the calculation result in step 2.3, multiply it by the square root to quadratic function H3, and finally perform a matrix transpose operation; Step 2.5, perform an inverse azimuth Fourier transform on the matrix transpose result in step 2.4, then perform a matrix transpose operation to complete the preliminary imaging of the RD algorithm.

7. The method according to claim 3, characterized in that The above step 4 specifically includes: First, the FPGA caches the data into the data storage module, then reads out the data sequentially, and in the way of overlapping segmentation and block division, reads three data blocks sequentially, and transmits the data to node #2, node #3, and node #4 in parallel respectively.

8. The method according to claim 3, characterized in that The error estimation calculation in the above step 5 specifically includes: Step 5.1, perform azimuth Dechirp processing on the data transmitted to node #2, node #3, and node #4 in step 4 in parallel. At the same time, the FPGA sends the data processed by the RD algorithm transmitted in step 3 to node #5; Step 5.2, perform block division on the output data of node #2, node #3, and node #4 in step 5.1 in parallel. The output result of each node is divided into two sub-block data of the same size. Then perform zero-padding operation in parallel, pad the two sub-blocks of each node with zeros to data blocks of the same scale as the data processed by the RD algorithm transmitted in step 3. Then perform azimuth Fourier transform operation on the two sub-block data of each node in parallel; Step 5.3, perform the operation of azimuth frequency domain amplitude on the results after azimuth Fourier transform operation of node #2, node #3, and node #4 in step 5.2 in parallel, and then perform azimuth Fourier transform operation on the three nodes in parallel; Step 5.4: Perform conjugate operations on the results after azimuth Fourier transform operations for Node #2, Node #3, and Node #4 in Step 5.3 in parallel to obtain the frequency-domain cross-correlation matrix data blocks, and perform azimuth inverse Fourier transform operations on the cross-correlation matrix data blocks of the three nodes in parallel; Step 5.5: Perform azimuth amplitude calculations on the results after azimuth inverse Fourier transform operations for Node #2, Node #3, and Node #4 in Step 5.4 in parallel, accumulate along the range direction to generate a real cross-correlation sequence, and then the three nodes respectively search for the maximum value in the cross-correlation sequence in parallel according to the azimuth direction, and perform error fitting estimation calculations to obtain error estimation data blocks. Thus, the error estimation calculations for the three data blocks by Node #2, Node #3, and Node #4 are completed.

9. The method according to claim 3, characterized in that, The specific steps of Step 6 include: Step 6.1: Transmit the three groups of error estimation results calculated by Node #2, Node #3, and Node #4 in Step 5 to the FPGA in parallel, then perform phase error splicing in the FPGA, and finally send the spliced phase error data to Node #5; Step 6.2: In Node #5, through differential derivative calculations and least squares fitting calculations, reconstruct the motion error data spliced in Step 6.1 into a data block of the same scale as the data processed by the RD algorithm transmitted in Step 3, and then compensate the reconstructed motion error data to the data processed by the RD algorithm received by Node #5; Step 6.3: Perform the product operation of the data after motion error compensation obtained in Step 6.2, the range window function, and the azimuth dechirp function in parallel in Node #5 to complete the data compensation and obtain a frame of imaging result.

10. A video SAR real-time imaging processing device based on the heterogeneous combination of Ascend NPU clusters and FPGAs, characterized in that, It includes: A memory: used to store a computer program for implementing the video SAR real-time imaging processing method according to any one of claims 3 to 9; A processor: used to implement the video SAR real-time imaging processing method according to any one of claims 3 to 9 when executing the computer program.