Method for Rapidly Extracting Target Information for Spaceborne Large-Field-of-View Staring Remote Sensing Imaging

By optimizing the target information extraction framework and deep network structure, and using registration methods and network sparseness and parameter quantization technologies, the problem of large-scale deep learning network calculations in satellite-borne large-field gaze remote sensing imaging is solved, and efficient target information extraction and real-time processing are achieved.

CN114494852BActive Publication Date: 2025-06-20BEIJING RES INST OF SPATIAL MECHANICAL & ELECTRICAL TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111592478.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-23
Publication Date
2025-06-20
Estimated Expiration
2041-12-23

AI Technical Summary

Technical Problem

In the prior art, the deep learning network has a huge structure and a lot of computation, which cannot meet the real-time processing requirements of large-field gaze-remote remote sensing imaging on satellite-borne large field of view.

Method used

Optimize the target information extraction framework, adopt the registration method with low computational volume, simplify the deep network, reduce the number of convolution kernels and parameter quantization through network sparseness and parameter quantization, and reduce the computing and storage requirements of the network.

Benefits of technology

It realizes efficient target information extraction under satellite-based conditions, reduces the amount of calculation, optimizes the network structure, and is suitable for real-time processing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494852B_ABST
    Figure CN114494852B_ABST
Patent Text Reader

Abstract

A method for quickly extracting target information for spaceborne large-field staring remote sensing imaging belongs to the field of target detection technology. Aiming at the problems that the existing deep network structure for target information extraction is large in size and high in computational complexity, and at the same time, the staring camera has a large amount of data and many pixels in a single image, and the deep learning method cannot be directly used for quickly extracting target information for spaceborne large-field staring remote sensing imaging, the present invention optimizes the target information extraction framework, adopts a registration method with less computational complexity to quickly discover moving targets, and simplifies the deep network for target classification and recognition, so that the target information extraction method proposed by the present invention is suitable for the conditions of spaceborne large-field staring remote sensing imaging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for rapidly extracting target information for spaceborne large-field-of-view staring remote sensing imaging, and belongs to the technical field of target detection. Background Art

[0002] A staring remote sensing satellite can continuously observe a fixed area through the "staring" imaging method, and can obtain a sequence of images representing the changes within the "staring" area. Extracting the features of interesting targets from the massive sequence data of the staring remote sensing camera has become a key research focus at present.

[0003] Using deep learning methods to achieve feature extraction of interesting targets is a current research hotspot. Deep learning methods do not require artificial design of target features. After designing a network for detection and training the detection network with a large number of samples, good detection results can be obtained even when the target is in a complex background. Existing deep learning networks, such as Faster RCNN, YOLO, etc., have achieved good results in the research of target recognition accuracy. For example, in reference [1], the YOLOv2 network structure was appropriately improved, and a deep convolutional neural network was applied to the target detection of staring video satellites, obtaining an efficient target detection method suitable for staring video satellites. However, this method does not consider the optimization of the detection framework and algorithm. This is also a common problem of current deep learning methods in real-time target detection, that is, the network structure and parameters are relatively large, and the requirements for computing resources and storage resources are relatively high. Staring remote sensing satellites will generate a large amount of sequence images. If only deep learning methods are used to perform deep learning processing on each frame of the sequence images, it is difficult to meet the requirements of on-board real-time processing. Summary of the Invention

[0004] The technical problem solved by the present invention is: overcoming the deficiencies of the prior art, providing a method for rapidly extracting target information for spaceborne large-field-of-view staring remote sensing imaging, aiming at the contradiction that the existing deep network structure for target information extraction is large and the amount of calculation is large, while the staring camera has a large amount of data and many pixels in a single image, and it is impossible to directly apply deep learning methods to rapidly extract target information for spaceborne large-field-of-view staring remote sensing imaging. The present invention optimizes the target information extraction framework, adopts a registration method with less calculation amount, simplifies the deep network, and makes the target information extraction method proposed by the present invention suitable for the conditions of spaceborne large-field-of-view staring remote sensing imaging.

[0005] The technical solution of the present invention is: A method for rapidly extracting target information for spaceborne large-field-of-view staring remote sensing imaging, including the following steps:

[0006] Construct a training sample set and a neural network model for spaceborne remote sensing images, and use the training sample set to train the neural network model to complete the initialization of the neural network model;

[0007] Optimize the structure of the initialized neural network model, and retrain it using the training sample set to obtain a neural network model for target information extraction;

[0008] Obtain multiple frames of spaceborne remote sensing images, and perform image registration on two consecutive frames;

[0009] Determine whether there are moving targets in the registered spaceborne remote sensing images; if the spaceborne remote sensing images of a preset number of consecutive frames contain moving targets, only input the frame images containing moving targets into the neural network model for target information extraction to detect whether the images contain moving targets; if the spaceborne remote sensing images of a preset number of consecutive frames do not contain moving targets, divide the last frame of the spaceborne remote sensing image before registration into sub-blocks and input them into the neural network model for target information extraction to detect whether the images contain stationary targets.

[0010] Further, the method for optimizing the structure of the initialized neural network model includes network sparsification and network weight quantization.

[0011] Further, the method for network sparsification includes the following steps:

[0012] First step, for any layer in the initialized neural network model, calculate the sum of the absolute values of the weights of all convolutional kernels in this layer; input_channel is the number of input feature maps, size is the filter size in the convolutional kernel, for the j-th kernel K i,j in the i-th layer of the network, the sum of the absolute values of the convolutional kernel weights is

[0013]

[0014] Second step, sort the calculated sum of the absolute values of the weights in this layer in descending order from largest to smallest to obtain the sorted vector S;

[0015] Third step, using the sequence mean mean(S) as the benchmark, remove the convolutional kernels whose sum of the absolute values of the weights is less than the sequence mean.

[0016] Further, the method for network weight quantization includes the following steps:

[0017] First step, calculate the quantization benchmark Q: when the weight distribution of the filters in the convolutional kernel satisfies max(k) < mean(k) + 3std(k), select the maximum absolute value of the weights as the quantization benchmark; when the maximum absolute value of the weights satisfies max(k) > mean(k) + 3std(k), select the second largest value of the convolutional kernel weights as the quantization benchmark; max(), mean(), and std() are the functions for finding the maximum value, finding the average value, and finding the standard deviation respectively;

[0018] In the second step, complete the quantization process according to the calculated quantization benchmark: round() is the rounding function, Q is the quantization benchmark, and k is the weight value in the convolution kernel m,n The calculation process of quantization is as follows:

[0019]

[0020] Furthermore, the method for registering two consecutive frames of images includes the following steps:

[0021] Divide the image into blocks;

[0022] Calculate the translation offset of each sub-block image of two consecutive frames of spaceborne remote sensing through Fourier-Mellin transform;

[0023] Calculate the rotation offset θ0 and scaling offset k of each sub-block image of spaceborne remote sensing through Fourier-Mellin transform;

[0024] Average the calculated offsets of each sub-block to obtain the offset of the entire image, and complete the registration of two consecutive frames of images according to each offset.

[0025] Furthermore, the calculation of the translation offset of two consecutive frames of spaceborne remote sensing images includes the following steps:

[0026] In the first step, divide the input image into sub-blocks: Through division, four sub-blocks with a length and width of 128 are extracted from the input image; (x, y) are the coordinates of the image in the spatial domain. Through the previous frame of image IMAGE i The four sub-blocks divided are:

[0027] region i 1 = IMAGE i (1:128, 1:128)

[0028] region i 2 = IMAGE i (1:128, 129:256)

[0029] region i 3 = IMAGE i (129:256, 1:128)

[0030] region i 4 = IMAGE i (129:256, 129:256)

[0031] Through the next frame of image IMAGE i+1 The four divided blocks are:

[0032] region i+11 = IMAGE i+1 (1:128,1:128)

[0033] region i+1 2 = IMAGE i+1 (1:128,129:256)

[0034] region i+1 3 = IMAGE i+1 (129:256,1:128)

[0035] region i+1 4 = IMAGE i+1 (129:256,129:256)

[0036] Second step, (x, y) are the coordinates of the image in the spatial domain, and (u, v) are the coordinates of the image in the frequency domain: For the sub - blocks region i 1 and region i+1 1, perform a two - dimensional Fourier transform to obtain the two - dimensional frequency spectra Region i 1(u, v) and Region i+1 1(u, v);

[0037] Third step, calculate the cross - power spectrum of region i 1 and region i+1 1, where * represents the conjugate operation and |●| is the modulus operation:

[0038] Region i 1 * (u, v)Region i+1 1(u, v) / |Region i 1*(u, v)Region i+1 1(u, v)|

[0039] Fourth step, perform an inverse two - dimensional Fourier transform on the calculated cross - power spectrum to obtain a two - dimensional impulse function δ(x - X, y - Y), where the relative position (X, Y) of the impulse is the relative translation amounts x0 and y0 between the two registered images.

[0040] Furthermore, the method for registering two consecutive frames of images according to each offset includes the following steps:

[0041] First step: For the sub - blocks region i 2 and region i+1 2, sub - block region i 3 and region i+1 3, sub - block regioni 4 and region i+1 4 respectively performs the calculations of the translation offset, rotation offset θ0, and scaling offset k. For each pair of sub-blocks, the relative translation amount, rotation offset, and scaling offset are calculated; the average values of the relative translation amounts, rotation offsets, and scaling offsets calculated for the four groups are obtained to get the relative translation offsets x0 and y0, rotation offset θ0, and scaling offset k of the two frames of images;

[0042] The second step: image i is Image i The registered image, image i The coordinates are recorded as (x1, y1), and for Image i The coordinates are recorded as (x, y), and the registration process is completed according to the following formula:

[0043]

[0044] The third step: Process the pixels at non-integer coordinate positions after registration using the nearest neighbor interpolation method.

[0045] Furthermore, the method for determining whether there are moving targets in the registered spaceborne remote sensing image is the frame difference method.

[0046] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method for quickly extracting target information for spaceborne large field-of-view staring remote sensing imaging.

[0047] A device for quickly extracting target information for spaceborne large field-of-view staring remote sensing imaging includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and is characterized in that: when the processor executes the computer program, it implements the steps of the method for quickly extracting target information for spaceborne large field-of-view staring remote sensing imaging.

[0048] The advantages of the present invention compared with the prior art are as follows:

[0049] (1) The present invention optimizes the target information extraction framework. Different processing methods are adopted for moving targets and static targets. Instead of uniformly processing all through a deep learning network, only the images containing moving targets and a small number of images not containing moving targets are input into the deep network for processing, greatly reducing unnecessary operations at the input end and reasonably allocating resources.

[0050] (2) The present invention uses a Fourier transform method that is easy to implement in hardware to achieve image registration. In the specific implementation, instead of calculating the registration offset using the entire image, only a sub-block that occupies a small part of the staring image is used to complete it; at the same time, a simple inter-frame difference method is used to detect possible moving targets in consecutive images. The registration method and the moving target detection method of the present invention have simple principles and are easy to implement, greatly saving the amount of calculation and meeting the usage requirements under spaceborne conditions.

[0051] (3) The present invention optimizes the Tiny-YOLO network in terms of structure. By reducing the number of convolutional kernels, the number of parameters in the network is reduced and the operation speed of the network is increased; by using the method of quantifying the parameters of the convolutional kernels, the storage space occupied by the network parameters is reduced, further improving the calculation speed. This makes the present invention meet the spaceborne usage environment. Brief Description of the Drawings

[0052] Figure 1 The overall framework for extracting moving target information designed for the present invention.

[0053] Figure 2 The method for optimizing the deep network for target information extraction.

[0054] Figure 3 The detection result of the network for ship targets after adopting the structure optimization method proposed by the present invention.

[0055] Figure 4 The image registration effect of the present invention. Detailed Embodiment

[0056] In order to better understand the above technical solution, the technical solution of the present application will be described in detail below through the drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present application and the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0057] The following further describes in detail the method for quickly extracting target information for spaceborne large-field-of-view staring remote sensing imaging provided by the embodiments of the present application in conjunction with the drawings of the specification. The specific implementation manner may include (such as Figures 1 to 4As shown in the figure: First, process the moving targets that may exist in the sequence images of the staring remote sensing camera. Perform a block operation on the ultra-high-resolution image. Select multiple image blocks at different positions and use the Fourier-Mellin transform method to calculate the image registration parameters. Use the calculated statistical parameters of multiple image blocks as the inter-frame geometric transformation parameters and perform registration. Based on the image registration, detect whether there are moving targets through the inter-frame difference method. If there are moving targets, input the block images with moving targets into the Tiny-YOLO network after convolutional layer sparsification and weight quantization to extract the information of the moving targets. For the sequence image group data without moving targets, only extract one frame of the image as the reference frame, perform a block operation on the image, and input the block images into the Tiny-YOLO network after convolutional layer sparsification and parameter quantization for target detection processing to extract the information of the possible static targets. This method processes moving targets and static targets of the large field-of-view staring remote sensing camera under spaceborne conditions using different methods, effectively reducing the computational load, optimizing the monitoring network, meeting the requirements of real-time processing under spaceborne conditions, and achieving efficient extraction of target information.

[0058] In the solution provided by the embodiment of the present application, the present invention includes the following steps:

[0059] (1) Use images containing targets such as ships and airplanes as the sample set to train the initial Tiny-YOLO network to obtain the weights of the convolution kernels in each convolutional layer of the network. This step is specifically as follows:

[0060] First, construct the training sample set. The training sample set consists of images containing ship and airplane targets and the corresponding txt documents. In the txt document, the category to which the target belongs, as well as the coordinates of the upper left corner and the lower right corner of the rectangular bounding box enclosing the target, are stored.

[0061] (2) According to the initial Tiny-YOLO network and the training sample set, determine the Tiny-YOLO network for ship and airplane target information extraction, specifically as follows:

[0062] Keep the structure of the initial network unchanged, that is, the number of convolutional layers and the number of convolutional kernels in each layer remain unchanged, and use the trained convolutional kernels to replace the original convolutional kernels to complete the initialization work training of the Tiny-YOLO network for target information extraction.

[0063] (3) Realize the sparsification of the Tiny-YOLO network for target information extraction by reducing the number of convolutional kernels. The specific steps are as follows:

[0064] Step 1: For any layer among conv3 - conv7 in the network parameter set, calculate the sum of the absolute values of the weights of all convolution kernels within this layer. input_channel is the number of input feature maps, and size is the filter size in the convolution kernel. For the j - th kernel K in the i - th layer of the network i,j , the sum of the absolute values of the weights of the convolution kernel is:

[0065]

[0066] Step 2: Arrange the calculated sum of the absolute values of the weights for this layer in descending order to obtain the sorted vector S.

[0067] Step 3: Using the sequence mean mean(S) as a benchmark, remove the convolution kernels whose sum of the absolute values of the weights is less than the sequence mean. This operation reduces the number of convolution kernels in this convolution layer, sparsifying the structure of the Tiny - YOLO network for target information extraction.

[0068] (4) For the sparsified Tiny - YOLO network in terms of structure, repeat the operations in steps (1) and (2), and ensure that the target information extraction ability does not significantly decline through retraining.

[0069] (5) Convert the weights of the convolution kernels in the Tiny - YOLO network after step (4) from floating - point numbers to integers to obtain the final Tiny - YOLO network. The specific steps are as follows:

[0070] Step 1: Calculate the quantization benchmark Q. When the weight distribution of the filters within the convolution kernel is relatively uniform, i.e., max(k) < mean(k)+3std(k), select the maximum absolute value among the weights as the quantization benchmark. When the maximum absolute value among the weights deviates from other weights, i.e., max(k)>mean(k)+3std(k), select the second - largest value among the weights of the convolution kernel as the quantization benchmark.

[0071] Step 2: Complete the quantization process based on the calculated quantization benchmark. round is the rounding function, Q is the quantization benchmark, and k is the weight within the convolution kernel m,n The calculation process for quantization is:

[0072]

[0073] (6) Calculate the translation offset through the Fourier - Mellin transform.

[0074] After the satellite imaging device enters the orbit and starts working, there are certain deviations between the sequential images generated along with the changes in environmental factors such as space thermology and mechanics, and image registration needs to be completed. The staring camera is designed to image a fixed area at the same angle and resolution, and the deviations between the sequences are very small, mainly including tiny translation, scaling, and rotation transformations. Therefore, good results can be achieved through simple registration.

[0075] The Fourier-Mellin transform method belongs to the frequency-domain-based registration method, which has the advantages of small computational complexity and strong robustness. Therefore, the present invention uses the Fourier-Mellin transform method to realize the tiny translation, scaling, and rotation transformations existing between the sequential images of the staring camera and complete the registration.

[0076] i is the serial number of consecutive frame images. First, calculate the translation amounts x0 and y0 between two consecutive frames of images IMAGE i and IMAGE i+1 . The specific process is as follows:

[0077] Step 1: Divide the input image into sub-blocks. Through division, four sub-blocks with a length and width of 128 are extracted from the input image. (x, y) are the coordinates of the image in the spatial domain. The four sub-blocks divided from the previous frame of image IMAGE i are respectively:

[0078] region i 1 = IMAGE i (1:128, 1:128)

[0079] region i 2 = IMAGE i (1:128, 129:256)

[0080] region i 3 = IMAGE i (129:256, 1:128)

[0081] region i 4 = IMAGE i (129:256, 129:256)

[0082] The four blocks divided from the subsequent frame of image IMAGE i+1 are respectively:

[0083] region i+1 1 = IMAGE i+1 (1:128, 1:128)

[0084] region i+1 2 = IMAGE i+1 (1:128, 129:256)

[0085] region i+1 3 = IMAGE i+1 (129:256,1:128)

[0086] region i+1 4 = IMAGE i+1 (129:256,129:256)

[0087] Step 2: (x, y) are the coordinates of the image in the spatial domain, and (u, v) are the coordinates of the image in the frequency domain. For the sub-blocks region i 1 and region i+1 1, perform a two-dimensional Fourier transform to obtain the two-dimensional frequency spectrum Region i 1(u, v) and Region i+1 1(u, v).

[0088] Step 3: Calculate the cross-power spectrum of region i 1 and region i+1 1, where * represents the conjugate operation and |●| is the modulus operation:

[0089] Region i 1 * (u, v)Region i+1 1(u, v) / |Region i 1*(u, v)Region i+1 1(u, v)|

[0090] Step 4: Perform an inverse two-dimensional Fourier transform on the calculated cross-power spectrum to obtain a two-dimensional impulse function δ(x - X, y - Y). The relative position (X, Y) of this impulse is the relative translation amounts x0 and y0 between the two registered images.

[0091] (7) Calculate the rotation offset θ0 and the scaling offset k through the Fourier-Mellin transform. The specific process is as follows:

[0092] Step 1: Perform a two-dimensional Fourier transform on the consecutive two-frame images region i 1 and region i+1 1, take the modulus of the obtained two-dimensional frequency spectrum to obtain the power spectra |Region i 1(u, v)| and |Region i+1 1(u, v)|.

[0093] Step 2: Perform a log-polar coordinate transform on the power spectra to obtain the power spectra M i+1 (logr, θ), Mi (logr - logk, θ - θ0);

[0094] Step 3: M i+1 (logr, θ) and M i (logr - logk, θ - θ0) are equal and have the same representation form. Therefore, the rotation offset θ0 and the scaling offset k can be calculated by using the phase - correlation operation in step (6).

[0095] (8) Complete the registration of two consecutive frames of images according to the offsets calculated by the Fourier - Mellin transform

[0096] Step 1: For the divided sub - blocks region i 2 and region i+1 2, the sub - block region i 3 and region i+1 3, the sub - block region i 4 and region i+1 4, perform the operations in (6) and (7) respectively. For each pair of sub - blocks, the relative translation amount, rotation offset, and scaling offset can be calculated. Average the relative translation amounts, rotation offsets, and scaling offsets calculated for the four groups to obtain the relative translation offsets x0 and y0, rotation offset θ0, and scaling offset k of the two frames of images.

[0097] Step 2: image i is the registered image Image i The coordinates of image i are recorded as (x1, y1), and the coordinates of Image i are recorded as (x, y). The registration process is completed according to the following formula::

[0098]

[0099] Step 3: Use the nearest - neighbor interpolation method to process the pixels at non - integer coordinate positions after registration.

[0100] (9) Determine whether there is a moving target by the inter - frame difference method

[0101] Step 1: Calculate the gray - level difference between two adjacent frames of registered images.

[0102] d i (x, y) = image i+1 (x, y)-Image i (x, y)

[0103] Step 2: The target area is in d iThe value is non-zero numerically, while the background area approaches zero. To distinguish the target from the noise, a judgment threshold needs to be set. T is the threshold, and it can be taken as 10. The judgment method is shown in Equation 5-2:

[0104]

[0105] There is a moving target in the area where T = 1, and the coordinates of this position are (x t , y t ). Extract the sub-block IMAGE containing the target coordinates i Input the network trained in step (5) to complete the information extraction of the moving target.

[0106] If all the values in T i (x, y) are zero, it means that there is no moving target in the i-th frame or the (i + 1)-th frame of the sequence, and the i-th frame image of the sequence is not processed by the deep network.

[0107] (10) Extraction of stationary target information

[0108] If the sequence of 1000 consecutive frames of images does not contain a moving target, only the 1000th original image is extracted, divided into sub-blocks, and input into the recognition network obtained in step (5) to detect whether the image contains important stationary targets.

[0109] Furthermore,

[0110] In a possible implementation scheme,

[0111] In a possible implementation manner,

[0112] Optionally,

[0113] This application provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions. When the computer instructions run on a computer, the computer is made to execute Figure 1 The method described above.

[0114] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0115] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0116] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0118] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

[0119] The content not described in detail in this specification of the present invention belongs to the well-known technology of those skilled in the art.

Claims

1. A method for rapid extraction of target information for spaceborne large field-of-view staring remote sensing imaging, characterized in that, It includes the following steps: Construct a training sample set and a neural network model for spaceborne remote sensing images, and use the training sample set to train the neural network model to complete the initialization of the neural network model; Optimize the structure of the initialized neural network model and retrain it using the training sample set to obtain a neural network model for target information extraction; Obtain multiple frames of spaceborne remote sensing images and perform image registration on two consecutive frames; Judge whether there are moving targets in the registered spaceborne remote sensing images; if the spaceborne remote sensing images of a continuously preset number of frames contain moving targets, only input the frame images containing moving targets into the neural network model for target information extraction to detect whether the images contain moving targets; if the spaceborne remote sensing images of a continuously preset number of frames do not contain moving targets, divide the last frame of the spaceborne remote sensing image before registration into sub-blocks and input them into the neural network model for target information extraction to detect whether the images contain stationary targets; The method for performing image registration on two consecutive frames includes the following steps: Perform block division on the images; Calculate the translation offset of each sub-block image of two consecutive frames of spaceborne remote sensing through Fourier-Mellin transform; Calculate the rotation offset θ0 and scaling offset k of each sub-block image of spaceborne remote sensing through Fourier-Mellin transform; Average the calculated offsets of each sub-block to obtain the offset of the entire image, and complete the registration of two consecutive frames of images according to each offset.

2. The method for rapid extraction of target information for spaceborne large field-of-view staring remote sensing imaging according to claim 1, characterized in that: The method for optimizing the structure of the initialized neural network model includes network sparsification and network weight quantization.

3. The method for rapid extraction of target information for spaceborne large field-of-view staring remote sensing imaging according to claim 2, characterized in that, The method for network sparsification includes the following steps: First step, for any layer in the initialized neural network model, calculate the sum of the absolute values of the weights of all convolutional kernels in this layer; input_channel is the number of input feature maps, size is the filter size in the convolutional kernel, for the j-th convolutional kernel K in the i-th layer of the network i,j , the weight k m,n in the convolutional kernel has an absolute value sum of Second step, sort the sum of the absolute values of the weights calculated for this layer in descending order to obtain the sorted vector S; Third step, using the sequence mean mean(S) as a benchmark, remove the convolution kernels whose sum of the absolute values of the weights is less than the sequence mean.

4. The method for rapid extraction of target information for spaceborne large field-of-view staring remote sensing imaging according to claim 2, characterized in that, The method for network weight quantization includes the following steps: First step, calculate the quantization benchmark Q: when the weight distribution of the filters in the convolution kernel satisfies max(k) < mean(k) + 3std(k), select the maximum absolute value of the weights as the quantization benchmark; when the maximum absolute value of the weights satisfies max(k) > mean(k) + 3std(k), select the second largest value of the weights in the convolution kernel as the quantization benchmark; max(), mean(), and std() are functions for finding the maximum value, finding the average value, and finding the standard deviation respectively; Step 2: Complete the quantization process according to the calculated quantization benchmark: round() is the rounding function, Q is the quantization benchmark, and the weight k mn in the convolutional kernel. The calculation process of quantization is as follows:

5. The method for rapid extraction of target information for spaceborne large field-of-view staring remote sensing imaging according to claim 1, characterized in that, The calculation of the translation offset of each sub-block image of two consecutive frames of spaceborne remote sensing includes the following steps: Step 1: Divide the input image into sub-blocks: Through division, four sub-blocks with a length and width of 128 are extracted from the input image; (x, y) are the coordinates of the image in the spatial domain, and through the previous frame image IMAGE i The four sub-blocks divided are respectively: region i 1 = IMAGE i (1:128, 1:128) region i 2 = IMAGE i (1:128, 129:256) region i 3 = IMAGE i (129:256,1:128) region i 4 = IMAGE i (129:256, 129:256) Through the subsequent frame image IMAGE i+1 The four divided parts are respectively: region i+1 1 = IMAGE i+1 (1:128,1:128) region i+1 2 = IMAGE i+1 (1:128, 129:256) region i+1 3 = IMAGE i+1 (129:256,1:128) region i+1 4 = IMAGE i+1 (129:256, 129:256) Step 2, (x, y) are the coordinates of the image in the spatial domain, and (u, v) are the coordinates of the image in the frequency domain: Divide the sub-blocks region i 1 and region i+1 1 perform a two-dimensional Fourier transform to obtain the two-dimensional frequency spectrum Region i 1(u, v) and Region i+1 1(u, v); Step 3: Calculate the region i 1 and the region i+1 1's cross-power spectrum, where * represents the conjugate operation and |●| is the modulus operation: Region i 1*(u,v) Region i+1 1(u,v) / | Region i 1*(u,v) Region i+1 1(u,v)| Fourth step, perform inverse two-dimensional Fourier transform on the calculated cross-power spectrum to obtain a two-dimensional impulse function δ(x - X, y - Y), where the relative position (X, Y) of the impulse is the relative translation amounts x0 and y0 between the two registered images.

6. The method for rapid extraction of target information for spaceborne large field-of-view staring remote sensing imaging according to claim 1, wherein, The method for completing the registration of two consecutive frames of images according to each offset includes the following steps: Step 1: Perform calculations of translation offset, rotation offset θ0, and scaling offset k on the divided sub-blocks region i 2 and region i+1 2, the sub-block region i 3 and region i+1 3, the sub-block region i 4 and region i+1 4 respectively, and calculate the relative translation amount, rotation offset, and scaling offset for each pair of sub-blocks; Average the calculated relative translation amounts, rotation offset, and scaling offset of the four groups to obtain the relative translation offset x0 and y0, rotation offset θ0, and scaling offset k of the two frames of images; Step 2: image i is the registered Image i The registered image, image i The coordinates are recorded as (x1, y1), and for Image i the coordinates are recorded as (x, y). The registration process is completed according to the following formula: Third step: Use the nearest neighbor interpolation method to process the pixels at non-integer coordinate positions after registration.

7. The method for rapid extraction of target information for spaceborne large field-of-view staring remote sensing imaging according to claim 1, characterized in that: The method for determining whether there are moving targets in the registered spaceborne remote sensing images is the inter-frame difference method.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

9. A device for rapid extraction of target information for spaceborne large field-of-view staring remote sensing imaging, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Optimization method of Tiny-YOLO network for detecting ship target on satellite

    CN110647977A

  • Motion detection method and device based on satellite videos, equipment and storage medium

    CN111104870A