A method and device for fast real-time sub-aperture imaging of ground-based synthetic aperture radar

By using GPU in the ground-based synthetic aperture radar to process echo signal data in parallel, using CUFFT function library and multiple parallel methods, the problem of low efficiency of traditional imaging algorithms is solved, and fast real-time imaging is achieved at high data rates, suitable for fast orbit SAR and array SAR systems.

CN115542320BActive Publication Date: 2025-07-08SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211232990.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-10
Publication Date
2025-07-08
Estimated Expiration
2042-10-10

AI Technical Summary

Technical Problem

The existing foundation synthetic aperture radar imaging algorithm is inefficient and cannot achieve rapid imaging at high data rates. It is especially difficult to solve the phase disintegration problem under complex meteorological conditions, which affects the real-time and accuracy of deformation monitoring.

Method used

The echo signal data is processed in parallel by using a graphics processor (GPU), and the fast Fourier transform and keystone transform are realized through the CUFFT function library. Combining two-dimensional grid index and thread index, distance compression, distance migration correction, orientation windowing and block imaging are carried out to improve imaging efficiency.

Benefits of technology

Real-time and efficient radar imaging are achieved, and are suitable for systems such as fast track SAR and array SAR, reducing data transmission time between CPU and GPU, improving imaging efficiency and reducing errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115542320B_ABST
    Figure CN115542320B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for fast real-time sub-aperture imaging of a ground synthetic aperture radar. The method includes: transmitting the echo signal data to be processed from a central processing unit to a graphics processing unit; windowing the echo signal data row by row according to a window function, and performing a fast Fourier transform on the rows of the matrix of the windowed echo signal data row by row through a CUFFT function library to obtain first signal data; performing keystone transform parallelization on the first signal data according to azimuth time domain and range frequency to obtain second signal data; windowing the second signal data in the azimuth direction in parallel using a two-dimensional grid indexing method to obtain target signal data; performing block imaging parallelization on the target signal data by combining thread indexing and block indexing to obtain a target radar image. This method can improve the radar imaging efficiency and has good imaging effects and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of radar imaging, and particularly to a ground-based synthetic aperture radar fast real-time sub-aperture imaging method and device. Background Art

[0002] Ground-Based Synthetic Aperture Radar (GB-SAR) is a ground-based active microwave remote sensing technology developed in the past decade or so. It can perform non-contact, high-precision, large-scale, and long-distance deformation monitoring, and is of great significance in the field of deformation monitoring. It is widely used in monitoring the deformation of large artificial buildings and natural disasters. The existing data acquisition rate of GB-SAR deformation is from dozens of seconds to several minutes. A high data rate plays an important role in monitoring the critical sliding stage or collapse of rapid deformation. At the same time, under complex meteorological conditions, a high data rate can also enable the radar to avoid the problem of phase unwrapping.

[0003] The general working process of a deformation monitoring radar includes multiple steps such as echo data acquisition, imaging processing, and deformation calculation. The radar systems of fast-track SAR and array SAR both have the characteristics of fast echo data acquisition, so they have the potential for a high deformation data rate. However, the amount of computation in the imaging link is huge, and the traditional imaging algorithm has low efficiency and long time consumption, which hinders the realization of a high deformation data rate. Therefore, realizing fast imaging of a large scene is one of the key problems of a high data rate slope deformation monitoring radar. Summary of the Invention

[0004] The present application provides a ground-based synthetic aperture radar fast real-time sub-aperture imaging method and device, which can improve the radar imaging efficiency and have better real-time performance.

[0005] In a first aspect, an embodiment of the present application provides a ground-based synthetic aperture radar fast real-time sub-aperture imaging method, which includes:

[0006] Transfer the echo signal data to be processed from the central processing unit to the graphics processing unit. The graphics processing unit is used to generate the azimuth time domain, range frequency, and window function, and a CUFFT function library for parallel implementation of fast Fourier transform and inverse fast Fourier transform is provided in the graphics processing unit;

[0007] Perform row-by-row windowing on the echo signal data according to the window function, and perform matrix row fast Fourier transform on the echo signal data after row-by-row windowing through the CUFFT function library to achieve range compression, and obtain first signal data;

[0008] Perform keystone transform parallelization on the first signal data according to the azimuth time domain and range frequency to achieve range migration correction, and obtain second signal data;

[0009] The azimuth windowing is performed on the second signal data in parallel using a two-dimensional grid indexing method to obtain target signal data;

[0010] The target signal data is block imaged in parallel by combining thread indexing and block indexing to achieve azimuth focusing, obtaining a target radar image;

[0011] The target radar image is sent to the central processing unit for display.

[0012] In one embodiment, before the echo signal data to be processed is transferred from the central processing unit to the graphics processing unit, the method further includes:

[0013] The row fast Fourier transform, column fast Fourier transform, and corresponding inverse row fast Fourier transform and inverse column fast Fourier transform are implemented through the CUFFT library, and the CUFFT library is encapsulated for the graphics processing unit to call.

[0014] In one embodiment, when windowing the echo signal data row by row according to the window function, it further includes:

[0015] The graphics processing unit performs global reads on both the echo signal data and the window function; or,

[0016] The graphics processing unit performs a global read on the echo signal data and reads the window function in a way that one thread reads multiple thread block data; or,

[0017] The graphics processing unit reads the echo signal data in a way that one thread block reads one column of echo signals, reads the window function in a way that one thread reads one data element, and uses block indexing to traverse the echo signal data.

[0018] In one embodiment, when performing keystone transform parallelization on the first signal data according to azimuth time domain and range frequency to achieve range migration correction, obtaining the second signal data, it includes:

[0019] The keystone transform is performed on the first signal data according to azimuth time domain and range frequency through two kernel functions to obtain the second signal data;

[0020] Among them, the two kernel functions include a first kernel function and a second kernel function;

[0021] The first kernel function is used to perform a scaling operation on each column of data in the first signal data to obtain the virtual time domain x of each column of data k and transfer the virtual time domain x of each column of data k to the virtual time domain matrix Axx, and the total number of points of the virtual time domain matrix Axx is Nr*Na;

[0022] The second kernel function is used to perform traversal interpolation of Nr*Na points in the direction of the virtual time-domain matrix Axx to obtain the second signal data.

[0023] In one embodiment, the azimuth windowing of the second signal data is performed in parallel using the indexing method of a two-dimensional grid to obtain the target signal data, including:

[0024] Read the window function and the second signal data using the indexing method of a two-dimensional grid;

[0025] Perform azimuth windowing operations on the read window function and the second signal data in parallel to obtain the target signal data.

[0026] In one embodiment, the target signal data contains multiple range gate data; the target signal data is block-imaged in parallel by combining thread indexing and block indexing to achieve azimuth focusing, and the target radar image is obtained, including:

[0027] Perform block imaging parallelization on the target signal data according to the azimuth time domain and range frequency through six kernel functions to obtain the target radar image;

[0028] Among them, the six kernel functions include the third kernel function, the fourth kernel function, the fifth kernel function, the sixth kernel function, the seventh kernel function, and the eighth kernel function;

[0029] The third kernel function is used to divide each range gate data into multiple sub-blocks, calculate the number of points of each sub-block of each range gate data, and establish a pointer vector according to the number of points of each sub-block of each range gate data;

[0030] The fourth kernel function is used to calculate the rectangular window center data of each sub-block of each range gate data and the central array index of the reference signal on the polar coordinate horizontal axis according to the pointer vector;

[0031] The fifth kernel function is used to generate a rectangular window function corresponding to each sub-block of each range gate data according to the central array index and the rectangular window center data of each sub-block of each range gate data, and transmit it to the window function matrix; multiply each sub-block by the corresponding rectangular window function, extract each sub-block signal and transmit it to the sub-block signal matrix; generate a demodulation reference signal corresponding to each sub-block signal and transmit it to the reference signal matrix;

[0032] The sixth kernel function is used to multiply the extracted sub-block signals by the corresponding demodulation reference signals to obtain the signals after dechirping of each sub-block signal;

[0033] The seventh kernel function is used to accumulate the signals after dechirping of each sub-block in each range gate data through a reduction algorithm to obtain the accumulation result corresponding to each range gate data;

[0034] The eighth kernel function is used to merge the accumulated results corresponding to the data of each range gate into a target radar image.

[0035] In one embodiment, the reduction algorithm is configured as a staggered pairing algorithm that includes the points lacking matching pairs in the pairing calculation of the previous step when the number of reduction points is odd.

[0036] In a second aspect, an embodiment of the present application provides a ground-based synthetic aperture radar fast real-time sub-aperture imaging device, which includes:

[0037] A data input module for inputting the echo signal data to be processed from the central processing unit into the graphics processing unit. The graphics processing unit is used to generate the azimuth time domain, range frequency, and window function, and a CUFFT function library for parallel implementation of the fast Fourier transform and the inverse fast Fourier transform is provided in the graphics processing unit;

[0038] A range compression parallelization module for windowing the echo signal data row by row according to the window function, and performing a matrix row fast Fourier transform on the windowed echo signal data row by row through the CUFFT function library to achieve range compression, obtaining first signal data;

[0039] A keystone transform parallelization module for performing keystone transform parallelization on the first signal data according to the azimuth time domain and range frequency to achieve range migration correction, obtaining second signal data;

[0040] An azimuth windowing processing parallelization module for performing azimuth windowing on the second signal data in parallel using the indexing method of a two-dimensional grid to obtain target signal data;

[0041] A block imaging parallelization module for performing block imaging parallelization on the target signal data by combining thread indexing and block indexing to achieve azimuth focusing, obtaining a target radar image;

[0042] An image output module for sending the target radar image to the central processing unit for display.

[0043] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it performs the steps of the sub-aperture imaging method of the ground-based synthetic aperture radar in any of the above embodiments.

[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the sub-aperture imaging method of the ground-based synthetic aperture radar in any of the above embodiments.

[0045] In summary, compared with the prior art, the beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0046] A method and device for fast real-time sub-aperture imaging of a ground-based synthetic aperture radar provided by an embodiment of the present application. This method can input the echo signal data detected by the ground-based synthetic aperture radar from the CPU memory to the GPU memory. The GPU calls a kernel function to process the radar data stored in the GPU memory, and finally transmits the processed target radar image from the GPU memory to the CPU memory. The CPU can image and display the target radar image; among them, the data transmission between the CPU and the GPU only occurs at the input of the data to be processed and the output of the final processing result, realizing the generation of GPU parameters in the processing process and the temporary storage of intermediate results, avoiding unnecessary transmission time; at the same time, by adopting a variety of parallel methods to parallelize the processing process, and its error is far less than the true value. Therefore, this method can improve the imaging efficiency, has better real-time performance, and is applicable to radar systems with fast processing capabilities, such as fast-track SAR and array SAR, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is an application flow chart of the sub-aperture imaging method of the ground-based synthetic aperture radar provided by an exemplary embodiment of the present application.

[0048] Figure 2 It is a flow chart of the sub-aperture imaging method of the ground-based synthetic aperture radar provided by an exemplary embodiment of the present application.

[0049] Figure 3 It is a schematic diagram of an indexing method for reading echo signal data provided by an exemplary embodiment of the present application.

[0050] Figure 4 It is a schematic diagram of the indexing method of the virtual time domain matrix Axx, the azimuth time domain x, and the range frequency f provided by an exemplary embodiment of the present application.

[0051] Figure 5 It is a schematic diagram of the number of sub-blocks corresponding to different range gates provided by an exemplary embodiment of the present application.

[0052] Figure 6 It is a schematic diagram of the indexing method of the central array index s_idx provided by an exemplary embodiment of the present application.

[0053] Figure 7 It is a schematic diagram of the indexing method of the sub-block signal matrix sig_rect and the reference signal matrix cankao provided by an exemplary embodiment of the present application.

[0054] Figure 8Schematic diagram of the indexing method for the window function matrix tem1 and the rectangular window timing na provided for an exemplary embodiment of the present application.

[0055] Figure 9 Schematic diagram of matrix indexing provided for an exemplary embodiment of the present application.

[0056] Figure 10 Schematic diagram of the interleaved pairing reduction algorithm provided for an exemplary embodiment of the present application.

[0057] Figure 11 Histogram of the amplitude relative error probability distribution of the processing results of the CPU and GPU provided for an exemplary embodiment of the present application.

[0058] Figure 12 Histogram of the phase relative error probability distribution of the processing results of the CPU and GPU provided for an exemplary embodiment of the present application.

[0059] Figure 13 Probability density map of the amplitude relative error of the processing results of the CPU and GPU provided for an exemplary embodiment of the present application.

[0060] Figure 14 Probability density map of the phase relative error of the processing results of the CPU and GPU provided for an exemplary embodiment of the present application.

[0061] Figure 15 Imaging result diagram of the parallelization algorithm and the traditional CPU algorithm provided for an exemplary embodiment of the present application.

[0062] Figure 16 Structure diagram of the sub-aperture imaging device of the ground-based synthetic aperture radar provided for an exemplary embodiment of the present application. Detailed implementation manners

[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0064] Please refer to Figure 1, the radar detects target deformation, and the collected data contains deformation information. To achieve a high deformation data rate, that is, a relatively high acquisition rate and transmission rate of deformation data, the embodiments of this application provide a method for fast real-time sub-aperture imaging of a ground-based synthetic aperture radar. This method can be implemented based on the central processing unit (CPU) and graphics processing unit (GPU) of the ground-based synthetic aperture radar. Among them, the CPU loads the radar data and passes the radar data into the GPU. At the same time, before data processing, the GPU generates parameters such as azimuth time domain x, range frequency f, and window function ham, and successively performs windowing on the radar data to achieve range compression, keystone transformation to achieve range cell migration correction (RCMC), azimuth windowing aggregation, and sub-block focusing processing according to the above parameters generated by itself, and transmits the finally obtained target radar image to the CPU for image display.

[0065] Taking the execution entity as the GPU as an example for illustration, please refer to Figure 2 , the method specifically includes the following steps:

[0066] Step S1, transfer the echo signal data to be processed from the central processor to the graphics processor, and the graphics processor is used to generate azimuth time domain, range frequency, and window function.

[0067] Among them, the graphics processor is equipped with a CUDA (Compute Unified Device Architecture) function library for parallel implementation of fast Fourier transform and inverse fast Fourier transform; CUDA is a general parallel computing architecture launched by NVIDIA. This architecture enables the GPU to solve complex computing problems. It includes the CUDA instruction set architecture (ISA) and the parallel computing engine inside the GPU. The echo signal data to be processed is the radar data detected by the ground-based synthetic aperture radar. As shown in Table 1, the size of the echo signal is Na×Nr (for example, 260×15360); the window function generated by the GPU can be a Hamming window function, etc.

[0068] Table 1 Data Parameters

[0069]

[0070] Specifically, the CUDA library can provide the encapsulated CUFFT library. The CUFFT library can implement the parallel operations of row and column FFT (Fast Fourier Transform) and IFFT (Inverse Fast Fourier Transform), that is, parallel FFT. Different from the ordinary FFT algorithm, parallel FFT is more convenient for organizing vector operations, with high efficiency and good performance. In specific implementation, the row FFT, column FFT, corresponding row IFFT, and column IFFT can be implemented through the CUFFT library and encapsulated into functions for the GPU to call. In practice, the CUDA library takes about 4.5 times less time to complete the operation of 260×15360-point row FFT than the CPU.

[0071] In the above steps, the GPU realizes the input of echo signal data and the generation of parameters. The generated parameters at least include azimuth time domain, range frequency, and window function.

[0072] In some embodiments of this embodiment, before step S1, the method further includes:

[0073] Implement the row fast Fourier transform, column fast Fourier transform, corresponding row inverse fast Fourier transform, and column inverse fast Fourier transform through the CUFFT library, and encapsulate the CUFFT library for the graphics processor to call.

[0074] Step S2: Window the echo signal data row by row according to the window function, and perform the matrix row fast Fourier transform on the windowed echo signal data through the CUFFT library to achieve range compression, obtaining the first signal data.

[0075] Specifically, the GPU can read a column of echo signals through a thread block, read a window function data by one thread, and complete the parallel implementation of range compression through row-by-row windowing and then matrix row FFT. That is, window the echo signal data in parallel according to the window function, perform matrix row FFT on the windowed echo signal data, and finally obtain the first signal data.

[0076] Among them, step S2 realizes the range compression processing of the echo signal data through windowing processing and fast Fourier transform. The obtained first signal data is the echo signal data after range compression. For the convenience of parallel calculation, the parallel implementation of range compression can be divided into row-by-row windowing and matrix row FFT. The matrix row FFT can be implemented by using the functions encapsulated in the CUFFT library.

[0077] In specific implementation, the GPU first reads the echo signal and the window function, and then performs dot multiplication on the read echo signal and window function in parallel to implement windowing processing, obtaining a windowed echo signal data matrix; then, the matrix row fast Fourier transform is used to perform parallel computing on the above windowed echo signal data matrix to obtain the Fourier transform result corresponding to each row in the data matrix, thereby obtaining the first signal data corresponding to the echo signal data.

[0078] For example, assume that ρ0 is the distance from the point target to the center of the radar aperture, the corresponding azimuth angle is θ0, and at the same time assume that the distance equation is in the form of a hyperbola. Then, the echo signal after range compression can be expressed as Equation (1):

[0079]

[0080] where t is the fast time in the azimuth direction, x is the azimuth time domain, λ c is the carrier wavelength, c represents the speed of light, L is the length of the radar aperture, Pr is the range compression envelope, and j is a complex number. It can be seen from the above equation that the range is coupled with the slow time, which causes the range cell where the target is located in the time-domain echo to change between different echoes, that is, the range cell migration of the target occurs. Moreover, the larger the bandwidth, the higher the range resolution, and the more serious the range migration phenomenon at the same speed.

[0081] Step S3: Perform keystone transform parallelization on the first signal data according to the azimuth time domain and range frequency to achieve range migration correction, obtaining the second signal data.

[0082] Among them, step S3 realizes the range migration correction of the first signal data through the keystone transform, and the obtained second signal data is the first signal data after range migration correction.

[0083] For example, the Fourier transform can be performed on Equation (1) along the range direction, and then the slow time is scaled by the scaling factor x = αx k transformation to perform RCMC, as shown in Equation (2) below:

[0084]

[0085] where x k represents the slow time virtual time domain.

[0086] After RCMC, the time-domain form of the obtained second signal data is as shown in Equation (3) below:

[0087]

[0088] where f represents the range frequency, f cDenote the carrier frequency. Since the synthetic aperture time of the sub-aperture is short, the linearly space-variant component is the main component of the residual space-variant range migration, and the keystone transform in the azimuth time can effectively remove this space-variant linear component. It can be seen from Equation (3) that the range frequency and the slow time have been decoupled.

[0089] In specific implementation, there are various implementation methods for the keystone transform, such as the interpolation method, the DFT-IFFT method, and the CZT method. In this step, the linear interpolation method that is easier to implement parallel processing is preferably used for the keystone transform.

[0090] First, calculate the virtual time-domain matrix Axx. By performing a scaling operation on each column of data in the first signal data, the virtual time-domain x of each column can be calculated, k and the x of each column k is transmitted to each row of the virtual time-domain matrix Axx; then linear interpolation is performed. The linear interpolation of the keystone transform is a triple loop of (Nr*Na)*Na, where (Nr*Na) is the total number of points in the virtual time-domain matrix Axx, and Na is the number of points in the real time-domain x. Therefore, the graphics processor has two directions to choose for parallelizing the interpolation: one is to perform traversal interpolation on (Nr*Na) points in the direction of the virtual time-domain matrix Axx; the other is to perform traversal interpolation on Na points in the real time-domain. In this embodiment, the traversal interpolation method in the direction of the virtual time-domain matrix Axx is adopted. From the perspective of the virtual time-domain, because (Nr*Na) points are interpolated simultaneously, as long as the number of points in the virtual time-domain does not exceed the number of threads limited by the GPU, the running time will theoretically remain unchanged, that is, the running time within the number of points not exceeding the number of threads limited by the GPU remains the same, thereby obtaining the second signal data.

[0091] Table 2 shows the running times of the two interpolation directions on the CPU and the GPU. It can be seen that the running time of the GPU is much less than that of the CPU. As the number of interpolation points increases, the running time of the CPU gradually increases, while the running times of the two interpolation methods on the GPU basically remain unchanged. From the perspective of the real time-domain, the running time only slowly increases when the number of points increases to the millions level; while from the perspective of the virtual time-domain, because (Nr*Na) points are interpolated simultaneously, as long as the number of points in the virtual time-domain does not exceed the number of threads limited by the GPU, the running time will theoretically remain unchanged.

[0092] Table 2 Comparison of CPU and Two GPU Interpolation Methods

[0093]

[0094] Step S4, perform azimuth windowing on the second signal data in parallel using the indexing method of a two-dimensional grid to obtain the target signal data.

[0095] Among them, the GPU can read the second signal data and the window function through the indexing method of a two-dimensional grid, and then perform windowing processing in the azimuth direction to obtain the target signal data; the target signal data contains multiple range gate data.

[0096] Step S3 completes the focusing in the range direction, and step S4 preliminarily completes the focusing in the azimuth direction. At this time, there is a linear relationship between the slow-time virtual time domain x k and the azimuth angle. According to the stationary phase principle, the second signal data after range migration correction is converted to the range-Doppler domain, as shown in the following formula (4):

[0097]

[0098] where K a is the azimuth frequency modulation parameter; the radar image is two-dimensional, the vertical axis in the rectangular coordinate system is the range direction ρ, and the horizontal axis is the azimuth time domain x; while in the polar coordinate system, the vertical axis is ρ and the horizontal axis is the sinθ axis. L sinθ (ρ0,θ0) represents the width of the signal on the sinθ axis. It can be seen from the above formula that the signal values of the target signal data are scattered on the sinθ axis, providing a basis for sub-block division.

[0099] Step S5 performs block imaging parallelization on the target signal data through a combination of thread indexing and block indexing to achieve azimuth focusing, and obtains the target radar image.

[0100] Specifically, the GPU divides each range gate data into multiple sub-blocks, demodulates each sub-block of each range gate data in parallel to obtain the demodulated signals of each sub-block; and accumulates the demodulated signals of each sub-block of each range gate data to obtain the accumulation result corresponding to each range gate data; combines the accumulation results corresponding to multiple range gate data to obtain the target radar image.

[0101] Among them, the GPU can divide each range gate data of the target signal into m(ρ) sub-blocks with a width of Δ′ sinθ (ρ), and the number of sub-block points is n(ρ). During the sub-block division process, the extraction of the i-th sub-block signal can be regarded as multiplying the target signal by a rectangular window with a specific center, as shown in the following formula (5):

[0102]

[0103] where W(sinθ; ρ,i) is the rectangular window function corresponding to the i-th sub-block signal, (ρ0,sinθ0) is the polar coordinate of the point target, and sinθ i represents the center position of the rectangular window of the i-th sub-block, L i sinθdenotes the width of the rectangular window, i.e., the signal width of the i-th sub-block.

[0104] Each sub-block is transformed to the ρ-x k time domain by using the stationary phase principle, and then multiplied by the reference function f shown in Equation (6) b to perform de-chirping, and the demodulated signals of each sub-block after de-chirping can be obtained, as shown in the following Equation (7):

[0105]

[0106]

[0107] Step S6: Send the target radar image to the CPU for display.

[0108] Specifically, the GPU outputs the target radar image to the CPU, and the CPU displays the target radar image.

[0109] A ground-based synthetic aperture radar fast real-time sub-aperture imaging method provided in the above embodiment can input the echo signal data detected by the ground-based synthetic aperture radar from the CPU memory to the GPU memory. The GPU calls the kernel function to process the radar data stored in the GPU memory, and finally transfers the processed target radar image from the GPU memory to the CPU memory. The CPU can image and display the target radar image. Among them, the data transmission between the CPU and the GPU only occurs in the input of the data to be processed and the output of the final processing result, realizing the generation of GPU parameters in the processing process and the temporary storage of intermediate results, avoiding unnecessary transmission time. At the same time, by using a variety of parallel methods to parallelize the processing process, and its error is far less than the true value. Therefore, this method can improve the imaging efficiency, has better real-time performance, and is applicable to radar systems with fast processing capabilities, such as fast-track SAR and array SAR, etc.

[0110] In some embodiments, step S2 further includes: The GPU reads the echo signal data and the window function, and can be any one of the following three reading methods:

[0111] (1) The GPU performs global reading on both the echo signal data and the window function.

[0112] (2) The GPU performs global reading on the echo signal data and reads the window function by using a method of one thread reading multiple thread block data.

[0113] (3) The GPU reads the echo signal data by using a method of one thread block reading a column of echo signals, reads the window function by using a method of one thread reading one data element, and uses the block index to traverse the echo signal data.

[0114] Among them, global reading means accessing and storing data in the global memory of the GPU. Usually, one thread reads one data element; the data element is the basic unit that makes up the data and is usually processed as a whole during the data processing process.

[0115] Specifically, the GPU can implement the above reading methods through three different kernel functions. Among them, the acceleration effects of the first and second reading methods are similar. Taking the number of points of the window function Nr = 15360 as an example, at this time, the amount of data of the echo signal reaches Na×Nr = 3993600, which is Na times that of the window function. Even if the efficiency of the second reading method for the thread to load and read the window function is increased by gridDim = 260 times compared to the first method, compared with the large amount of data reading and operation of the echo signal, the effect is minimal; the third method is to perform simultaneous access and operation on one column (15360 points) of the echo signal, and the acceleration effect is significant compared to the first two methods; as Figure 3 shown, the third method uses block indexing. Since the size of the echo signal is Na×Nr, and the maximum number of threads per thread is limited to no more than 1024, which is much smaller than Nr, but Na satisfies Na < 1024, so blockDim = Na can be set. Among them, gridDim represents the size of the thread, and blockDim represents the size of the thread block.

[0116] In the above embodiments, the method can parallelize and optimize the windowing process in the aperture imaging process by adopting three different loading and reading methods, thereby accelerating the windowing process and improving the radar imaging efficiency.

[0117] In some embodiments, step S3 may include the following steps:

[0118] Perform a keystone transform on the first signal data according to the azimuth time domain and range frequency through two kernel functions to obtain the second signal data; the two kernel functions include the first kernel function and the second kernel function.

[0119] Among them, in the range migration correction step, two pre-written kernel functions can be regarded as kernel function 1 and kernel function 2 used in the range migration correction stage. Kernel function 1 calculates the virtual azimuth time domain, that is, the virtual time domain matrix Axx, and kernel function 2 performs linear interpolation. Then in step S3:

[0120] Kernel function 1: Scale each column of data in the first signal data according to the azimuth time domain and range frequency to obtain the virtual time domain x of each column of data k and transmit the virtual time domain x of each column of data k to the virtual time domain matrix Axx. The total number of points in the virtual time domain matrix Axx is Nr*Na.

[0121] Among them, the indexing methods of the virtual time-domain matrix Axx, the azimuth time-domain x, and the range frequency f are as Figure 4 shown.

[0122] Kernel function 2: Traverse and interpolate Nr*Na points in the direction of the virtual time-domain matrix Axx to obtain the second signal data.

[0123] Specifically, the linear interpolation of the Keystone transform is a triple loop of (Nr*Na)*Na, where (Nr*Na) is the total number of points in the virtual time-domain matrix Axx, and Na is the number of points in the real time-domain x. Therefore, the graphics processing unit has two options for parallelizing the interpolation: one is to traverse and interpolate Nr*Na points in the direction of the virtual time-domain matrix Axx; the other is to traverse and interpolate Na points in the real time-domain. This embodiment adopts the method of traversing and interpolating in the direction of the virtual time-domain matrix Axx. From the perspective of the virtual time-domain, because the interpolation is performed on (Nr*Na) points simultaneously, as long as the number of points in the virtual time-domain does not exceed the number of threads limited by the GPU, the running time will theoretically remain unchanged, that is, the running time within the number of points not exceeding the number of threads limited by the GPU remains the same, thereby obtaining the second signal data.

[0124] In the above embodiment, the method can achieve parallel optimization of the Keystone transform part by taking steps, which can accelerate the range migration correction speed and thus improve the imaging efficiency.

[0125] After RCMC is completed through the Keystone transform, the linear coupling between range and azimuth has been well corrected, and the focusing in the range direction has been completed. However, further focusing is required in the azimuth direction to fully display the image of the radar detection target.

[0126] In some embodiments, step S4 may include the following steps:

[0127] Read the window function and the second signal data using the indexing method of a two-dimensional grid.

[0128] Perform azimuth windowing operations on the read window function and the second signal data in parallel to obtain the target signal data.

[0129] Specifically, the GPU reads the second signal data and the window function through the two-dimensional grid indexing method, and performs azimuth windowing operations on the window function and the second signal data read through the two-dimensional grid in parallel to obtain the target signal data.

[0130] In the above embodiment, the method can pre-generate the window function by the GPU before processing the data, which can effectively reduce the transmission time between the GPU and the CPU.

[0131] Although windowing in the azimuth direction can also achieve the focusing effect, the focusing effect is not ideal. Therefore, further partitioning is required to achieve better focusing in the azimuth direction.

[0132] Take Figure 5 the number of sub - blocks when the detection range of the radar data shown is from 31m to 300m as an example for illustration. When the number of divided sub - blocks is less than 2, demodulation is not required because the effect already meets the requirements of imaging. The larger the range gate, the fewer the number of sub - blocks mp(ρ) that can be divided. When the range gate is 31m, the number of sub - blocks is the largest, reaching 260 blocks. When the range gate is the maximum detection range of 300m, the number of sub - blocks is the smallest, only 4 blocks.

[0133] Since block - demodulation needs to be performed for each range gate, and the number of sub - blocks and points of each range gate are different, the parallelized block - imaging of the target signal data can be carried out by six kernel functions respectively.

[0134] The target signal data contains multiple range - gate data. Step S5 may include the following steps: parallelize the block - imaging of the target signal data according to the azimuth - time domain and range - frequency through six kernel functions to obtain the target radar image; among them, the six kernel functions include the third kernel function, the fourth kernel function, the fifth kernel function, the sixth kernel function, the seventh kernel function, and the eighth kernel function, which can also be regarded as kernel function 3, kernel function 4, kernel function 5, kernel function 6, kernel function 7, and kernel function 8 for implementing the parallelized block - imaging step.

[0135] Kernel function 3: Divide the data of each range gate into multiple sub - blocks, calculate the number of points of each sub - block of the data of each range gate, and establish a pointer vector according to the number of points of each sub - block of the data of each range gate.

[0136] Among them, kernel function 3 calculates the number of points of each sub - block of the data of each range gate, and establishes a pointer vector according to the number of points of each sub - block of the data of each range gate to mark the number of points of each sub - block.

[0137] Kernel function 4: Calculate the center data of the rectangular window of each sub - block of the data of each range gate and the central array index of the reference signal on the polar - coordinate horizontal axis according to the pointer vector.

[0138] Among them, kernel function 4 calculates the center data center of the rectangular window of each sub - block of the data of each range gate and the central array index s_idx of the reference signal on the polar - coordinate horizontal axis. Among them, the central array index s_idx is the indexing method of the center data of the rectangular window of the sub - block on the polar - coordinate horizontal axis (sinθ axis), which can be used to read the center data center of the rectangular window of the sub - block during parallel calculation on the GPU.

[0139] In specific implementation, set the number of sub - blocks of any distance gate as mρ(ρ), that is, the distance gate includes mρ(ρ) sub - blocks. For example, Figure 6 as shown, the indexing method of the central array index s_idx can be as follows:

[0140] Divide the central data of the rectangular window of mρ(ρ) sub - blocks into mp(ρ) / 8 thread blocks. mp(ρ) / 8 thread blocks share 8 threads, and each thread reads the central data of the rectangular window of mp(ρ) / 8 sub - blocks simultaneously.

[0141] Kernel function 5: Generate the rectangular window function corresponding to each sub - block of each distance - gate data according to the central array index and the central data of the rectangular window of each sub - block of each distance - gate data, and transmit it to the window function matrix; Multiply each sub - block by the corresponding rectangular window function, extract each sub - block signal and transmit it to the sub - block signal matrix; Generate the demodulation reference signal corresponding to each sub - block signal and transmit it to the reference signal matrix.

[0142] Specifically, kernel function 5 first reads the central data center of the rectangular window of each sub - block of each distance - gate data and the rectangular window timing na in parallel according to the central array index, calculates the rectangular window function corresponding to each sub - block of each distance - gate data, and transmits the rectangular window function corresponding to each sub - block into the window function matrix tem1, that is, obtains the window function matrix tem1 corresponding to each distance - gate data; Then, multiply each sub - block of each distance - gate data by its corresponding rectangular window function, obtain the above - mentioned sub - block signals and transmit them into the sub - block signal matrix sig_rect, and simultaneously generate the demodulation reference signal corresponding to the above - mentioned sub - block signals, and transmit the above - mentioned demodulation reference signal into the reference signal matrix cankao.

[0143] Among them, as Figure 7 shown, the sizes of the sub - block signal matrix sig_rect and the reference signal matrix cankao are both Na×mp(ρ). Use the thread - block index to read the matrix columns of the sub - block signal matrix sig_rect and the reference signal matrix cankao, and use the thread index to read the matrix rows of the sub - block signal matrix sig_rect and the reference signal matrix cankao; As Figure 8 shown, the sizes of the window function matrix tem1 and the rectangular window timing na are both mp(ρ)×Na. Use the thread index to read the matrix columns of the window function matrix tem1 and the rectangular window timing na, and use the thread - block index to read the matrix rows of the window function matrix tem1 and the rectangular window timing na.

[0144] Kernel function 6: Multiply each extracted sub - block signal by the corresponding demodulation reference signal to obtain the signal after de - chirping of each sub - block signal;

[0145] Specifically, the kernel function 6 extracts each sub-block signal and its corresponding demodulation reference signal, multiplies each sub-block signal by its corresponding demodulation reference signal, performs the de-chirp processing on each sub-block signal, and obtains the signal of each sub-block signal after the de-chirp processing.

[0146] The sub-block signal IFFT matrix sig_rect_ifft is the time-domain form of the sub-block signal matrix sig_rect, and it is itself a two-dimensional data element matrix; the sub-block signal IFFT matrix sig_rect_ifft and the reference signal matrix establish matrix indexes in the kernel function 4 by using a two-dimensional network and a two-dimensional thread block, and the indexing method is as Figure 9 shown. Three indexing methods need to be processed: thread index (threadIdx), thread block index (blockIdx), the element point coordinates (ix, iy) in the sub-block signal IFFT matrix, and the offset idx in the GPU global linear memory. Therefore, it is necessary to map the thread index (threadIdx) and the thread block index (blockIdx) to the element point coordinates (ix, iy) of the sub-block signal IFFT matrix, so as to realize the reading and writing of the data of each element point coordinate (ix, iy) in the sub-block signal IFFT matrix. The specific mapping process can be realized by the following formula:

[0147] First, map the thread index (threadIdx) and the thread block index (blockIdx) to the sub-block signal IFFT matrix:

[0148] ix = threadIdx.x + blockIdx.x · blockDim.x

[0149] iy = threadIdx.y + blockIdx.y · blockDim.y

[0150] Then, map from the sub-block signal IFFT matrix to the offset idx in the GPU global linear memory:

[0151] idx = iy · mp(ρ) + ix

[0152] Among them, the offset idx corresponding to each sub-block is the demodulation signal of the sub-block after the de-chirp processing.

[0153] Kernel function 7: Accumulate the signals after de-chirping of each sub-block in each range gate data through a reduction algorithm to obtain the accumulation result corresponding to each range gate data.

[0154] In some embodiments of this embodiment, as Figure 10 shown, the reduction algorithm is configured as a staggered pairing algorithm that includes the points lacking matching pairs in the pairing calculation in the previous step when the number of reduction points is odd.

[0155] For example, when the reduced number of points is 14. The first reduction is as follows: The pairing form is 0 - 7, 1 - 8, 2 - 9, 3 - 10, 4 - 11, 5 - 12, 6 - 13, and the corresponding thread IDs are 0 - 6 in sequence. The second reduction: The number of points is halved, and the number of points becomes 7, which is an odd number. The pairing form is 0 - 3, 1 - 4, 2 - 5. Point 6 lacks a paired point, so before pairing 2 - 5, 5 + 6 is done first, that is, 2 - (5 + 6), and the corresponding thread IDs are 0 - 2. The third reduction: The number of points is halved again, and the number of points becomes 3, which is an odd number of points and also lacks a paired point, so the pairing form is 0 - (1 + 2), and the corresponding ID threads are 0. At this point, the sum of the 14 points has been obtained, that is, the result of the third reduction is the result of adding the 14 points.

[0156] Specifically, the GPU accumulates the demodulated signals through kernel function 7. The indexing method of kernel function 7 is the same as the indexing method of the above reference signal matrix. The thread block dimension of the demodulated signal matrix is mp(ρ), and the thread dimension is Na, that is, the size of the demodulated signal matrix is mp(ρ) × Na.

[0157] Kernel function 8: Combine the accumulated results corresponding to each range gate data into the target radar image.

[0158] Specifically, kernel function 6 combines the accumulated results of each range gate to obtain the final target radar image.

[0159] In the above embodiments, the method can adopt a combination of thread indexing and block indexing to realize parallel processing of processing steps such as calculating the sub - block center, generating a rectangular window, and extracting sub - block signals. Then, two - dimensional grids and two - dimensional thread blocks are used for parallel processing of de - skewing, accumulation, and merging of each sub - block, which speeds up the data processing speed and further improves the imaging efficiency.

[0160] Figures 11 to 14 They are respectively the amplitude relative error probability distribution histogram, phase relative error probability distribution histogram, amplitude relative error probability density map, and phase relative error probability density map of the CPU and GPU processing results. From Figure 13 and Figure 14 The shown probability density maps, it can be found that the amplitude and phase errors are mainly distributed near 0, and the probability of deviating from 0 is lower; from Figure 11 and Figure 12 The shown distribution histograms, it can be known that the amplitude relative error and phase relative error are respectively in the range of 0 - 2.5*10 -13 interval and 0 - 2.5*10 -10Within the interval. Therefore, the amplitude and phase deviations between the CPU and the GPU are on the order of one trillionth and one billionth of the CPU amplitude and phase, respectively. This deviation mainly stems from the differences in the rounding and data truncation methods between the CPU and the GPU. Figure 15 The figure shows the imaging results of the parallel algorithm (left) and the traditional CPU algorithm (right). It is not difficult to find that the sub-aperture imaging algorithm based on the GPU can achieve the same effect as the traditional CPU imaging algorithm in less time, meeting the real-time requirements of the radar system. Table 3 shows the running times of the algorithms for different detection distance ranges. It can be seen that as the detection range expands, the running times of both the CPU and the GPU increase accordingly, and the acceleration effect of the GPU is further enhanced. When the detection distance exceeds 1500m, the running times of the two algorithms stabilize, and the maximum acceleration ratio reaches 27.6 at this time. Because when the detection distance exceeds 1500m, the minimum divisible sub-block number has been reached. Even if the detection distance further increases, the number of range gates that can be divided has stabilized, and the corresponding number of sub-blocks will not change, so the total workload of the block division will be fixed accordingly, and the acceleration ratio will also stabilize.

[0161] Table 3 Time consumption of the parallel algorithm and the serial algorithm at different detection distances

[0162]

[0163] In the above-described embodiments, a sub-aperture imaging method based on the GPU is provided. Various methods are used to perform targeted parallel optimization on the core links such as windowing, Keystone transformation, and block imaging in the sub-aperture imaging algorithm. Among them, three different loading and reading methods are compared for windowing, and the third method with the highest acceleration ratio is adopted; the Keystone transformation and the block part are parallelized step by step, and a combination of thread index and block index is used to calculate the center of the sub-block, generate a rectangular window, and extract the sub-block signal; then two-dimensional grids and two-dimensional thread blocks are used for de-skewing, accumulation, and merging of each sub-block. Based on the measured data, the relative errors of amplitude and phase of the present invention approach 0, the overall acceleration ratio is 18 - 27, and as the detection distance increases, the acceleration effect of the GPU is further enhanced and finally reaches saturation. Therefore, the present invention not only has good real-time performance but also can ensure the imaging effect, meeting the real-time requirements of systems such as GB-SAR.

[0164] Please refer to Figure 16 , an embodiment of the present application provides a ground-based synthetic aperture radar fast real-time sub-aperture imaging device, which includes:

[0165] The data input module 101 is used to transfer the echo signal data to be processed from the central processing unit to the graphics processing unit. The graphics processing unit is used to generate the azimuth time domain, range frequency, and window function, and a CUFFT function library for parallel implementation of fast Fourier transform and inverse fast Fourier transform is provided in the graphics processing unit;

[0166] The range compression parallelization module 102 is used to perform windowing on the echo signal data row by row according to the window function, and perform matrix row fast Fourier transform on the windowed echo signal data row by row through the CUFFT function library to achieve range compression, obtaining the first signal data;

[0167] The keystone transform parallelization module 103 is used to perform keystone transform parallelization on the first signal data according to the azimuth time domain and range frequency to achieve range migration correction, obtaining the second signal data;

[0168] The azimuth windowing processing parallelization module 104 is used to perform azimuth windowing on the second signal data in parallel by using the indexing method of a two-dimensional grid, obtaining the target signal data;

[0169] The block imaging parallelization module 105 is used to perform block imaging parallelization on the target signal data by combining thread indexing and block indexing to achieve azimuth focusing, obtaining the target radar image;

[0170] The image output module 106 is used to send the target radar image to the central processing unit for display.

[0171] For the specific limitations of the sub-aperture imaging device of the ground-based synthetic aperture radar provided in this embodiment, reference can be made to the embodiment of the sub-aperture imaging method of the ground-based synthetic aperture radar in the foregoing text, which will not be elaborated here. Each module in the above-mentioned sub-aperture imaging device of the ground-based synthetic aperture radar can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0172] An embodiment of the present application provides a computer device, which may include a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the processor executes the steps of the sub-aperture imaging method of the ground-based synthetic aperture radar in any of the above embodiments.

[0173] For the working process, working details, and technical effects of the computer device provided in this embodiment, reference may be made to the embodiments of the sub-aperture imaging method of the ground-based synthetic aperture radar in the foregoing text, and details will not be elaborated here.

[0174] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the sub-aperture imaging method of the ground-based synthetic aperture radar in any of the above embodiments. Among them, the computer-readable storage medium refers to a carrier for storing data, which may include, but is not limited to, a floppy disk, an optical disc, a hard disk, a flash memory, a USB flash drive, and / or a Memory Stick, etc. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices.

[0175] For the working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of the sub-aperture imaging method of the ground-based synthetic aperture radar in the foregoing text, and details will not be elaborated here.

[0176] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0177] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification. At the same time, the content not described in detail in the above embodiments belongs to the prior art well known to those skilled in the art.

[0178] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

Claims

1. A fast real-time sub-aperture imaging method for ground synthetic aperture radar, characterized in that The method includes: transferring the echo signal data to be processed from the central processing unit to the graphics processing unit, where the graphics processing unit is used to generate the azimuth time domain, range frequency, and window function, and a CUFFT function library for parallel implementation of fast Fourier transform and inverse fast Fourier transform is provided in the graphics processing unit; windowing the echo signal data row by row according to the window function, and performing matrix row fast Fourier transform on the windowed echo signal data row by row through the CUFFT function library to achieve range compression, obtaining first signal data; performing keystone transform parallelization on the first signal data according to the azimuth time domain and the range frequency to achieve range migration correction, obtaining second signal data; performing azimuth windowing on the second signal data in parallel using the indexing method of a two-dimensional grid to obtain target signal data; performing block imaging parallelization on the target signal data by combining thread indexing and block indexing to achieve azimuth focusing, obtaining a target radar image; sending the target radar image to the central processing unit for display; wherein, the target signal data includes multiple range gate data; the performing block imaging parallelization on the target signal data by combining thread indexing and block indexing to achieve azimuth focusing and obtaining a target radar image includes: performing block imaging parallelization on the target signal data according to the azimuth time domain and the range frequency through six kernel functions to obtain a target radar image; wherein, the six kernel functions include a third kernel function, a fourth kernel function, a fifth kernel function, a sixth kernel function, a seventh kernel function, and an eighth kernel function; the third kernel function is used to divide each range gate data into multiple sub-blocks, calculate the number of points of each sub-block of each range gate data, and establish a pointer vector according to the number of points of each sub-block of each range gate data; the fourth kernel function is used to calculate the rectangular window center data of each sub-block of each range gate data and the central array index of the reference signal on the polar coordinate horizontal axis according to the pointer vector; the fifth kernel function is used to generate a rectangular window function corresponding to each sub-block of each range gate data according to the central array index and the rectangular window center data of each sub-block of each range gate data, and transmit it to the window function matrix; multiply each sub-block by the corresponding rectangular window function, extract each sub-block signal and transmit it to the sub-block signal matrix; generate a demodulation reference signal corresponding to each sub-block signal and transmit it to the reference signal matrix; the sixth kernel function is used to multiply the extracted sub-block signals by the corresponding demodulation reference signals to obtain the signals after dechirping of each sub-block signal; the seventh kernel function is used to accumulate the signals after dechirping of each sub-block in each range gate data through a reduction algorithm to obtain an accumulation result corresponding to each range gate data; the eighth kernel function is used to combine the accumulation results corresponding to each range gate data into the target radar image.

2. The method according to claim 1, wherein Before transferring the echo signal data to be processed from the central processing unit to the graphics processing unit, the method further includes: Implement the row fast Fourier transform (FFT), column FFT, corresponding row inverse FFT, and column inverse FFT through the CUFFT library, and encapsulate the CUFFT library for the graphics processing unit (GPU) to call.

3. The method according to claim 1, wherein The step of windowing the echo signal data row by row according to the window function further includes: The GPU globally reads both the echo signal data and the window function; or, The GPU globally reads the echo signal data and reads the window function in a way that one thread reads data of multiple thread blocks; or, The GPU reads the echo signal data in a way that one thread block reads one column of the echo signal and reads the window function in a way that one thread reads one data element, and uses the block index to traverse the echo signal data.

4. The method according to claim 1, characterized in that The step of parallelizing the Keystone transform on the first signal data according to the azimuth time domain and the range frequency to achieve range migration correction and obtain the second signal data includes: Performing the Keystone transform on the first signal data according to the azimuth time domain and the range frequency through two kernel functions to obtain the second signal data; Wherein, the two kernel functions include a first kernel function and a second kernel function; The first kernel function is used to perform a scaling operation on each column of data in the first signal data to obtain the virtual time domain x of each column of data k , and transmit the virtual time domain x of each column of data k to the virtual time domain matrix Axx, and the total number of points of the virtual time domain matrix Axx is Nr*Na; The second kernel function is used to traverse and interpolate Nr*Na points from the direction of the virtual time domain matrix Axx to obtain the second signal data.

5. The method according to claim 1, wherein The step of parallelly performing azimuth windowing on the second signal data by using the indexing method of a two-dimensional grid to obtain the target signal data includes: Reading the window function and the second signal data by using the indexing method of a two-dimensional grid; Parallelly performing azimuth windowing operations on the read window function and the second signal data to obtain the target signal data.

6. The method according to claim 1, characterized in that, The reduction algorithm is configured as a staggered pairing algorithm that includes the point lacking a matching pair in the pairing calculation of the previous step when the number of reduction points is odd.

7. A fast real-time sub-aperture imaging device for ground synthetic aperture radar, characterized in that The device includes: A data input module for inputting the echo signal data to be processed from the central processing unit (CPU) into the GPU. The GPU is used to generate the azimuth time domain, range frequency, and window function, and a CUFFT library for parallelly implementing the FFT and inverse FFT is provided in the GPU; A range compression parallelization module for windowing the echo signal data row by row according to the window function and performing a matrix row FFT on the echo signal data after row-by-row windowing through the CUFFT library to achieve range compression and obtain the first signal data; A Keystone transform parallelization module for parallelizing the Keystone transform on the first signal data according to the azimuth time domain and the range frequency to achieve range migration correction and obtain the second signal data; An azimuth windowing processing parallelization module for parallelly performing azimuth windowing on the second signal data by using the indexing method of a two-dimensional grid to obtain the target signal data; The block imaging parallelization module is used to perform block imaging parallelization on the target signal data by combining thread indexing and block indexing to achieve azimuth focusing, so as to obtain a target radar image. The target signal data includes multiple range gate data. Specifically, the method for performing block imaging parallelization on the target signal data by combining thread indexing and block indexing to achieve azimuth focusing and obtaining a target radar image includes: Performing block imaging parallelization on the target signal data according to the azimuth time domain and the range frequency through six kernel functions to obtain a target radar image; Among them, the six kernel functions include a third kernel function, a fourth kernel function, a fifth kernel function, a sixth kernel function, a seventh kernel function, and an eighth kernel function; The third kernel function is used to divide each range gate data into multiple sub-blocks, calculate the number of points of each sub-block of each range gate data, and establish a pointer vector according to the number of points of each sub-block of each range gate data; The fourth kernel function is used to calculate the rectangular window center data of each sub-block of each range gate data and the central array index of the reference signal on the polar coordinate horizontal axis according to the pointer vector; The fifth kernel function is used to generate a rectangular window function corresponding to each sub-block of each range gate data according to the central array index and the rectangular window center data of each sub-block of each range gate data, and transmit it to the window function matrix; multiply each sub-block by the corresponding rectangular window function, extract each sub-block signal and transmit it to the sub-block signal matrix; generate a demodulation reference signal corresponding to each sub-block signal and transmit it to the reference signal matrix; The sixth kernel function is used to multiply each extracted sub-block signal by the corresponding demodulation reference signal to obtain the dechirped signal of each sub-block signal; The seventh kernel function is used to accumulate the dechirped signals of each sub-block in each range gate data through a reduction algorithm to obtain an accumulation result corresponding to each range gate data; The eighth kernel function is used to combine the accumulation results corresponding to each range gate data into the target radar image; The image output module is used to send the target radar image to the central processing unit for display.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.