Method for Detecting Moving Vehicles Based on Spatiotemporal Motion and Model Compression in Satellite Videos
By designing a lightweight network, using 3D convolution module and model compression technology, the problems of false alarms and low recall in satellite videos are solved, and real-time and efficient mobile vehicle detection on mobile devices is achieved.
Patent Information
- Application Number
- CN202310838303.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-07-10
AI Technical Summary
The prior art is prone to false alarms when changing non-stationary satellite platforms and atmospheric lighting conditions in satellite videos. Appearance-based methods fail to effectively utilize motion information, resulting in low recall, and high-precision methods have high computational complexity and cannot detect mobile vehicles in real time.
A lightweight network is designed, using a 3D convolutional double-layer motion information extraction module, combining the underlying and high-level features, and real-time detection of mobile vehicles in satellite video is achieved by reducing the number of channels and using a small convolution kernel compression model.
Improves the robustness and recall of mobile vehicle detection in satellite video, reduces false alarm rates, and achieves efficient real-time detection on mobile devices.
Smart Images

Figure CN116894975B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to a method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite video in the technical field of satellite video target detection. The present invention can be used for real-time detection and recognition of ground moving vehicles in satellite video. Background Art
[0002] Satellite remote sensing data, as the core of the earth observation system, is of great significance for comprehensively grasping the global economic, resource, and social development trends. Compared with natural images, satellite video contains richer temporal information and a larger observation area, so it is widely used in fields such as disaster response, resource monitoring, and traffic monitoring. Especially in traffic monitoring, satellite video shows obvious advantages because it can record the complete trajectories of moving vehicles. Target detection is one of the most critical and common technologies in the field of computer vision. Its task is to determine whether there are objects to be detected in an image or video. If there are target objects, the position of the target is located and recognized. When applied to satellite video analysis, it can effectively identify and detect moving vehicles on the ground. This application technology can not only provide strong technical support for satellite remote sensing monitoring technology, but also provide valuable decision-making basis for fields such as military reconnaissance, urban traffic planning, and vehicle remote control driving.
[0003] Central South University disclosed a method for extracting dynamic vehicle targets in satellite video in its patent document "Method for Extracting Dynamic Vehicle Targets in Satellite Video by Fusing Luminance-Temporal Features" (Patent Application No.: 202011369826.7, Publication No.: CN112489055A). This method first extracts and simplifies the luminance gradient features of vehicle targets in the video frames of satellite video data; then detects the luminance gradient change rate of the extracted and simplified features; then uses the multi-frame multi-scale visual background extraction algorithm (Vibe) to detect the multi-level motion regions of vehicle targets; finally, extracts dynamic vehicle targets according to the detection results of the luminance gradient change rate and the multi-level motion region detection results. The disadvantage of this method is that although it is applicable to satellite video, it seriously relies on the prior knowledge that "the motion speed of foreground targets is relatively faster than the background environment", and does not fundamentally solve the problem that the background segmentation method is prone to false alarms when affected by factors such as non-stationary satellite platforms and changes in atmospheric illumination conditions. Therefore, this method has obvious deficiencies in terms of robustness and generalization ability.
[0004] Zhou et al. proposed a high-precision object detection method based on appearance, CenterNet, in their published paper "Objects as Points" (IEEE Conference on Computer Vision and Pattern Recognition, 2019). This method first preprocesses a single image, scales the image to a fixed size as the input of the network structure; then the image extracts the apparent features of the image through the backbone network; then the apparent features are passed through a deconvolution module and three upsampling operations to generate output features; finally, the output features are fed into three prediction branches respectively to obtain the category, center position and bounding box size of the object. The shortcoming of this method is that although it has good ability to solve the misdetection problem caused by changes in lighting conditions, it only uses the apparent features of the object and does not effectively extract the motion information between satellite video frames. Especially in satellite videos, moving vehicles on the ground severely lack appearance information such as texture and geometry, which greatly limits the performance of this method in detection accuracy and recall rate.
[0005] Xiao et al. proposed a high-precision moving object detection method in satellite videos in their published paper "DSFNet: Dynamic and static fusion network for moving object detection in satellite videos" (IEEE Geoscience and Remote Sensing Letters, 2022). This method combines static background information and the dynamic motion information of the object, and realizes high-precision detection of small moving objects in satellite video scenes. The shortcoming of this method is that although it can achieve relatively high detection accuracy for the characteristics of satellite video objects, it has the problems of excessive number of parameters and high computational complexity. If it is forcibly deployed on mobile devices, it will lead to a decline in real-time performance and cannot meet the requirements of real-time object detection. Summary of the Invention
[0006] The purpose of the present invention is to propose a moving vehicle detection method based on spatio-temporal motion and model compression in satellite videos in view of the above deficiencies of the prior art, so as to solve the problem of false alarms easily generated by traditional technologies when facing factors such as non-stationary satellite platforms and changes in atmospheric lighting conditions, the problem of low recall rate caused by the failure of appearance-based methods in deep learning technology to utilize motion information, and the problem that high-precision methods cannot detect moving vehicles in real time on mobile devices due to the need for a large number of parameters and complex calculations.
[0007] The technical idea for achieving the purpose of the present invention is to design a lightweight network for detecting moving vehicles in satellite videos by using deep learning technology. The network introduces a two-layer motion information extraction module based on 3D convolution to fully mine the spatial and temporal information of moving vehicles and effectively fuse the low-level and high-level features, thereby solving the artifact problems caused by the movement of the satellite platform and the change of atmospheric illumination conditions in traditional technologies, as well as the problem of low recall rate due to the failure of appearance-based detection methods to utilize spatio-temporal motion information. In addition, considering the limitations of memory and computing resources faced by actual deployment devices, the present invention adopts the strategy of reducing the number of channels and using smaller convolution kernels during the network design process to compress the model, thereby minimizing the model parameters and computational cost and solving the problem that high-precision detection methods cannot detect moving vehicles in real time on mobile devices.
[0008] The technical solution of the present invention is as follows:
[0009] Step 1, build an input sub-network composed of a scaling layer and an enhancement layer connected in series;
[0010] Step 2, build a feature extraction sub-network composed of a 3D transposed convolution layer and a 3D convolution layer connected in series;
[0011] Step 3, construct a lightweight spatio-temporal motion information extraction sub-network composed of a shallow spatio-temporal motion information extraction module, a deep spatio-temporal motion information extraction module, and a spatio-temporal motion information fusion module:
[0012] Step 3.1, build a shallow spatio-temporal motion information extraction module composed of a 3D convolution layer, a first expansion layer, a compression layer, a second expansion layer, a residual connection layer, a 3D max pooling downsampling layer, and a deformable convolution layer connected in series in sequence; set the convolution kernel sizes of the 3D convolution layer and the compression layer to 1, the strides to 1, the input channel numbers to 16 and 32 respectively, and the output channel numbers to 16; the residual connection layer takes the original data and the current data as inputs and adds them along the channel direction; set the convolution kernel size of the 3D pooling downsampling layer to 3×1×1 and the stride to 3×1×1; set the convolution kernel size of the deformable convolution layer to 3, the stride to 1, the input channel number to 64, and the output channel number to 32;
[0013] Step 3.2, construct a deep spatio-temporal motion information extraction module composed of a first 3D max pooling downsampling layer, a 3D convolutional layer, a first expansion layer, a compression layer, a second expansion layer, a residual connection layer, a second 3D max pooling downsampling layer, a deformable convolutional layer, and a 2D transposed convolutional layer connected in series in sequence; set the convolutional kernel size of the first 3D max pooling downsampling layer to 1×2×2 and the stride to 1×2×2; set the convolutional kernel sizes of the 3D convolutional layer and the compression layer to 1, the strides to 1, the input channel numbers to 32 and 64 respectively, and the output channel numbers to 32; the residual connection layer takes the original data and the current data as inputs and adds them along the channel direction; set the convolutional kernel size of the second 3D pooling downsampling layer to 3×1×1 and the stride to 3×1×1; set the convolutional kernel size of the deformable convolutional layer to 3, the stride to 1, and the input and output channel numbers to 32; set the convolutional kernel size of the 2D transposed convolutional layer to 4, the stride to 2, and the input and output channel numbers to 32;
[0014] Step 3.3, construct a spatio-temporal motion information fusion module composed of a connection layer and a deformable convolutional layer connected in series in sequence; set the convolutional kernel size of the deformable convolutional layer to 3, the stride to 1, and the input and output channel numbers to 32; the connection layer adds the input data along the channel direction;
[0015] Step 3.4, cascade the residual connection layer in the shallow spatio-temporal motion information extraction module with the transposed convolutional layer in the deep spatio-temporal motion information extraction module, and cascade the deformable convolutional layer in the shallow spatio-temporal motion information extraction module and the first 3D max pooling downsampling layer in the deep spatio-temporal motion information extraction module with the connection layer in the spatio-temporal motion information fusion module to obtain a lightweight spatio-temporal motion information extraction sub-network;
[0016] Step 4, construct a detection head sub-network composed of a key point heat map prediction layer, a center point offset prediction layer, and an object box size prediction layer connected in parallel;
[0017] Step 5, construct a lightweight moving vehicle detection network:
[0018] Cascade the enhancement layer of the input sub-network with the 3D transposed convolutional layer in the feature extraction sub-network; cascade the 3D convolutional layer in the feature extraction sub-network with the 3D convolutional layer in the shallow spatio-temporal motion information extraction sub-network; cascade the deformable convolutional layer in the spatio-temporal motion information fusion sub-network with the key point heat map prediction layer, the center point offset prediction layer, and the object box size prediction layer in the detection head sub-network to form a lightweight moving vehicle detection network;
[0019] Step 6, generate a training set:
[0020] Step 6.1, form a sample set by combining at least 300 frames of large-size satellite videos. In each video frame of the sample set, there are no less than 50 targets, and the moving vehicles lacking appearance information with a size range of 4 to 20 pixels;
[0021] Step 6.2, crop each frame of the sample set into a video frame of 512×512 pixels, and stack the video frames at intervals of 5 frames to form a training set;
[0022] Step 7, train a lightweight moving vehicle detection network:
[0023] Set the training parameters, and input the training set into the lightweight moving vehicle detection network. Use the Adam optimizer and the stochastic gradient descent method to iteratively update the weights of the network parameters until the loss function converges, and obtain a trained network model;
[0024] Step 8, detect moving vehicles in satellite videos:
[0025] After performing the same cropping and stacking operations on the satellite video frames to be detected as described in Step 6.2, input them into the trained network, output the moving vehicle detection results, and save them as a label file in the coco format.
[0026] The present invention has the following advantages compared with the existing technologies:
[0027] First, since the present invention adopts a deep learning method, it deeply excavates the spatio-temporal motion information of ground moving vehicles in satellite videos, overcomes the problem that traditional technologies rely too much on prior knowledge and fail to fundamentally solve the false alarm problems caused by factors such as non-stationary satellite platforms and changes in atmospheric illumination conditions. This enables the present invention to also exhibit strong detection performance when facing new satellite video data sets, has good robustness and generalization ability, and effectively reduces the false alarm rate of detection.
[0028] Second, since the present invention uses 3D convolution operations in feature extraction and introduces the time dimension on the basis of the spatial dimension, it can fully extract the spatio-temporal motion characteristics of moving vehicles, overcomes the problem of low recall rate caused by the failure of appearance-based object detection methods in deep learning technologies to utilize motion information, greatly improves the ability of the present invention to extract motion information, can more accurately detect ground vehicles in satellite videos, and further improves the recall rate of detection.
[0029] Thirdly, since the present invention designs a model compression strategy in accordance with the principle of "using fewer channels and smaller-sized convolutional kernels", it overcomes the problem in the prior art that high-precision detection methods cannot detect moving vehicles in real time on mobile devices due to the need for a large number of parameters and complex calculations. The present invention minimizes the model parameters and computational cost, thereby ensuring that this method can be efficiently deployed on mobile devices and achieving a high detection speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flowchart of the present invention;
[0031] Figure 2 is a lightweight mobile vehicle detection network constructed by the present invention;
[0032] Figure 3 is a schematic structural diagram of an expansion layer in the lightweight mobile vehicle detection network constructed by the present invention;
[0033] Figure 4 is a schematic structural diagram of a three-dimensional convolutional layer and a compression layer in the feature extraction sub-network and the lightweight spatio-temporal motion information extraction sub-network constructed by the present invention;
[0034] Figure 5 , Figure 6 , Figure 7 is a graph of the simulation experiment results of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following further describes the present invention in detail with reference to the drawings and embodiments.
[0036] Refer to Figure 1 to further describe in detail the implementation steps of the embodiments of the present invention.
[0037] The lightweight mobile vehicle detection network constructed by the present invention includes four sub-networks, namely: an input sub-network, a feature extraction sub-network, a lightweight spatio-temporal motion information extraction sub-network, and a detection head sub-network. Among them, the lightweight spatio-temporal motion information extraction sub-network is composed of a shallow spatio-temporal motion information extraction module, a deep spatio-temporal motion information extraction module, and a spatio-temporal motion information fusion module, as shown by the six dashed boxes in Figure 2 . The input sub-network performs adaptive image scaling and data augmentation tasks on the input data to implement preprocessing operations before training. The feature extraction sub-network is responsible for performing feature extraction operations on the data passed in from the input sub-network to obtain rich feature information. The lightweight spatio-temporal motion information extraction sub-network is responsible for deeply mining the features extracted by the feature extraction sub-network to obtain shallow and deep spatio-temporal motion information of the moving vehicle and performing feature fusion. The detection head sub-network is responsible for predicting the moving vehicle for the features passed in from the lightweight spatio-temporal motion information to obtain the output of the network.
[0038] Step 1: Build an input sub-network composed of a scaling layer and an enhancement layer connected in series;
[0039] The scaling layer is implemented using an image adaptive scaling algorithm, and the data enhancement layer is implemented using a random mirror flipping algorithm, an affine transformation algorithm, and a photometric distortion algorithm.
[0040] Step 2: Build a feature extraction sub-network composed of a 3D transposed convolutional layer and a 3D convolutional layer connected in series;
[0041] The input data size of the 3D transposed convolutional layer is set to 512×512×5×3, the output data size is set to 1024×1024×5×3, the convolutional kernel size is set to 1×3×3, and the stride is set to 1×2×2. Through this upsampling operation, the spatio-temporal information of the target is amplified to assist in improving the detection accuracy; the input data size of the 3D convolutional layer is 1024×1024×5×3, the output data size is 512×512×5×16, the convolutional kernel size is 3, and the stride is 1×2×2;
[0042] Step 3: Construct a lightweight spatio-temporal motion information extraction sub-network composed of a shallow spatio-temporal motion information extraction module, a deep spatio-temporal motion information extraction module, and a spatio-temporal motion information fusion module:
[0043] Step 3.1: Build a shallow spatio-temporal motion information extraction module composed of a 3D convolutional layer, a first expansion layer, a compression layer, a second expansion layer, a residual connection layer, a 3D max pooling downsampling layer, and a deformable convolutional layer connected in series in sequence; set the convolutional kernel sizes of the 3D convolutional layer and the compression layer to 1, the strides to 1, the input channel numbers to 16 and 32 respectively, and the output channel numbers to 16. By using a smaller convolutional kernel size and reducing the number of channels of the feature map, the parameter quantity and computational complexity of the 3D convolutional operation are reduced; the residual connection layer takes the original data and the current data as inputs and adds them along the channel direction to alleviate the gradient disappearance problem while enhancing the spatio-temporal modeling ability of the network; set the convolutional kernel size of the 3D pooling downsampling layer to 3×1×1 and the stride to 3×1×1; set the convolutional kernel size of the deformable convolutional layer to 3, the stride to 1, the input channel number to 64, and the output channel number to 32.
[0044] Refer to the appendix Figure 3 for a further description of the expansion layer of the shallow spatio-temporal motion information extraction module in the lightweight mobile vehicle detection network constructed by the present invention.
[0045] Figure 3Figure 1 shows the structure of the expansion layer. The first and second expansion layers are each composed of two parallel 3D convolutional layers connected in series. The convolution kernel sizes of the two parallel 3D convolutional layers are set to 1 and 3, respectively, with strides of 1. The number of input and output channels is set to 16. The concatenation layer stacks and concatenates the input data along the channel direction. The expansion layers are used to mine detailed spatiotemporal motion information of moving vehicles.
[0046] Step 3.2, build a deep spatiotemporal motion information extraction module consisting of the first 3D maximum pooling downsampling layer, 3D convolution layer, first expansion layer, compression layer, second expansion layer, residual connection layer, second 3D maximum pooling downsampling layer, deformable convolution layer and 2D transposed convolution layer in series; set the convolution kernel size of the first 3D maximum pooling downsampling layer to 1×2×2, and the step size to 1×2×2; set the convolution kernel size of the 3D convolution layer and compression layer to 1, the step size to 1, the number of input channels to 32 and 64 respectively, and the number of output channels to 1. The number of channels is set to 32; the residual connection layer takes the original data and the current data as input and adds them along the channel direction; the convolution kernel size of the second 3D pooling downsampling layer is set to 3×1×1, and the stride is set to 3×1×1; the convolution kernel size of the deformable convolution layer is set to 3, the stride is set to 1, and the number of input channels and the number of output channels are both set to 32; the convolution kernel size of the 2D transposed convolution layer is set to 4, the stride is set to 2, and the number of input channels and the number of output channels are both set to 32, so as to alleviate the problem of misalignment between shallow features and deep features;
[0047] Refer to the attached Figure 3 , further describes the extension layer of the deep spatiotemporal motion information extraction module in the lightweight mobile vehicle detection network constructed by the present invention.
[0048] Figure 3 Figure 1 shows the structure of the expansion layer. The first and second expansion layers each consist of two parallel 3D convolutional layers connected in series. The convolution kernel sizes of the two parallel 3D convolutional layers are set to 1 and 3, respectively, with strides of 1. The number of input and output channels is set to 32. The concatenation layer stacks and concatenates the input data along the channel direction. The expansion layers are used to mine the abstract semantic information of the spatiotemporal motion of moving vehicles.
[0049] Step 3.3: Build a spatiotemporal motion information fusion module consisting of a connection layer and a variable convolution layer connected in series. Set the convolution kernel size of the variable convolution layer to 3, the stride to 1, and the number of input and output channels to 32. The connection layer adds the input data along the channel direction.
[0050] Step 3.4: Cascade the residual connection layer in the shallow spatio-temporal motion information extraction module with the transposed convolutional layer in the deep spatio-temporal motion information extraction module, and cascade the deformable convolutional layer in the shallow spatio-temporal motion information extraction module and the first three-dimensional max-pooling downsampling layer in the deep spatio-temporal motion information extraction module with the connection layer of the spatio-temporal motion information fusion module to obtain a lightweight spatio-temporal motion information extraction sub-network;
[0051] Refer to Figure 4 for a further description of the three-dimensional convolutional layer and compression layer structures in the feature extraction sub-network and the lightweight spatio-temporal motion information extraction sub-network constructed in the present invention.
[0052] The structures of the three-dimensional convolutional layer and compression layer in the feature extraction sub-network and the lightweight spatio-temporal motion information extraction sub-network of the present invention are the same. As Figure 4 shown, the three-dimensional convolutional layer and compression layer sequentially include a three-dimensional convolutional operation, three-dimensional batch normalization, and an R e LU activation function. The parameter settings of the three-dimensional convolutional operation are set according to the layer number where the three-dimensional convolutional layer is located.
[0053] Step 4: Construct a detection head sub-network composed of a key point heat map prediction layer, a center point offset prediction layer, and a target box size prediction layer connected in parallel;
[0054] The key point heat map prediction layer, the center point offset prediction layer, and the target box size prediction layer have the same structure, and are all composed of a first convolutional layer and a second convolutional layer connected in series; set the convolutional kernel sizes of the first and second convolutional layers to 3 and 1 respectively, the stride to 1, and the input channel numbers to 32 and 256 respectively; set the output channel number of the second convolutional layer of the key point heat map prediction layer to 1, and the output channel numbers of the second convolutional layers in the center point offset prediction layer and the target box size prediction layer to 2.
[0055] Step 5: Construct a lightweight moving vehicle detection network:
[0056] Cascade the enhancement layer of the input sub-network with the three-dimensional transposed convolutional layer in the feature extraction sub-network; cascade the three-dimensional convolutional layer in the feature extraction sub-network with the three-dimensional convolutional layer in the shallow spatio-temporal motion information extraction sub-network; cascade the deformable convolutional layer in the spatio-temporal motion information fusion sub-network with the key point heat map prediction layer, the center point offset prediction layer, and the target box size prediction layer in the detection head sub-network to form a lightweight moving vehicle detection network;
[0057] Step 6: Generate a training set:
[0058] Step 6.1: Compose at least 300 frames of large-size satellite videos into a sample set T. Each video frame in the sample set T has no less than 50 targets, and the moving vehicles lack appearance information with a size range of 4 to 20 pixels;
[0059] Step 6.2, crop each frame of the sample set T into a video frame of 512×512 pixels, and stack the video frames at intervals of 5 frames to form the training set T = {t1, t2,..., t n ,..., t N}. Where N≥50, represents the nth training sample in the training set T, and W×H represents the number of row and column pixels of the video frame. In this embodiment, N = 60, W = H = 512. The stacking operation refers to stacking 5 consecutive video frames along the time dimension.
[0060] Step 7, train the lightweight moving vehicle detection network:
[0061] Convert the size of the training set to 512×512×5×3 through the adaptive scaling algorithm, and enhance the input data through the enhancement method; set the training parameters, and input the training set into the lightweight moving vehicle detection network. Use the Adam optimizer and the stochastic gradient descent method to iteratively update the weights of the network parameters until the loss function L converges to obtain the trained network model; the setting of the training parameters means setting the batch size of the network to 16. The loss function L is as follows:
[0062] L = L h + λ f L f + λ s L s
[0063]
[0064]
[0065]
[0066] Among them, L h represents the key point heat map loss function, L f represents the center point offset loss function, L s represents the target box size loss function; λ f and λ s represent the weight parameters with values of 1 and 0.1 respectively; N represents the number of targets in each frame of the training set video; ∑ represents the summation operation, represents the prediction score that the key point heat map coordinates (x, y) output by the key point heat map prediction layer belong to the cth category, Y xy,c represents whether the key point heat map coordinates (x, y) in each video frame of the training set belong to the cth category. If (x, y) belongs to the cth category, then Y xy,cTake the value of 1. If (x, y) does not belong to the c-th category, then Y xy,c Take the value of 0; α and β represent the weight parameters with values of 2 and 4 respectively; represents the center point offset of the p-th target output by the center point offset prediction layer, represents the center point offset of the p-th target in each frame of the training set video. R represents the proportionality coefficient of the video frame size divided by the key point heat map size, represents the floor operation, and |·| represents the absolute value operation; represents the size of the k-th target box output by the target box size prediction layer, S k represents the size of the k-th target box in each frame of the training set video.
[0067] Step 8, detect moving vehicles in satellite videos:
[0068] Crop each frame of the satellite video to be detected into a video frame of 512×512 pixels, stack the video frames at intervals of 5 frames, input them into the trained network, load the trained lightweight moving vehicle detection network parameter file, perform vehicle detection on the satellite video, obtain the key point heat map, center point offset, and the size of the target box of the vehicle, and output the moving vehicle detection result and save it as a label file in coco format.
[0069] The following further illustrates the effect of the present invention in combination with simulation experiments:
[0070] 1. Simulation experiment conditions:
[0071] The hardware platform for the simulation experiment of the present invention is: the processor is an Intel(R) Core(TM) i9-10980XE CPU with a main frequency of 3.0 GHz and a memory of 128 GB.
[0072] The software platform for the simulation experiment of the present invention is: Windows 10 operating system and python 3.8.
[0073] The input data used in the simulation experiment of the present invention are the Dubai satellite video dataset, the San Diego satellite video dataset, and the Las Vegas satellite video dataset. Among them, the Dubai and San Diego satellite video datasets were respectively collected by the Jilin-1 satellite from Dubai, United Arab Emirates and San Diego, California, USA, and the videos are both 960 frames; the Las Vegas satellite video dataset was collected by the SkySat satellite from Las Vegas, USA, and the video has a total of 1400 frames. These three satellite video datasets only contain the vehicle category, and the video format is mp4.
[0074] 2. Simulation Content and Result Analysis:
[0075] The simulation experiment of the present invention uses the present invention and three existing technologies (ViBe detection method, CenterNet detection method, DSFNet detection method) to detect moving vehicles in Dubai satellite video, San Diego satellite video, and Las Vegas satellite video respectively. The obtained detection results are Figure 5 , Figure 6 and Figure 7 as shown.
[0076] The existing technology ViBe detection method refers to the moving vehicle detection method proposed by Barnich et al. in "Vibe: A universal background subtraction algorithm for video sequences, IEEE Transactions on Image Processing, vol. 20, no. 6, pp. 1709–1724, 2010.", abbreviated as the ViBe detection method. The existing technology CenterNet detection method refers to the object detection method proposed by Zhou et al. in "Objects as points, arXiv preprint arXiv:1904.07850, 2019", abbreviated as the CenterNet detection method. The existing technology DSFNet detection method refers to the moving object detection method proposed by Chao et al. in "DSFNet: Dynamic and static fusion network for moving object detection in satellite videos, in IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1-5, 2022.", abbreviated as the DSFNet detection method.
[0077] Next, in combination with Figure 5 , Figure 6 and Figure 7 's simulation diagrams, the effect of the present invention will be further described.
[0078] Figure 5 (a), Figure 6 (a), Figure 7 (a) is the result diagram of using the existing technology ViBe detection method to detect vehicles in Dubai satellite video, San Diego satellite video, and Las Vegas satellite video respectively. Figure 5 (b), Figure 6 (b),Figure 7 (b) shows the results of vehicle detection on Dubai satellite video, San Diego satellite video, and Las Vegas satellite video respectively using the existing CenterNet detection method. Figure 5 (c), Figure 6 (c), Figure 7 (c) shows the results of vehicle detection on Dubai satellite video, San Diego satellite video, and Las Vegas satellite video respectively using the existing DSFNet detection method. Figure 5 (d), Figure 6 (d), Figure 7 (d) shows the results of vehicle detection on Dubai satellite video, San Diego satellite video, and Las Vegas satellite video respectively using the method of the present invention. In the result diagrams, the blue boxes represent correct detections, the yellow boxes represent missed detections, and the red boxes represent false detections.
[0079] From Figure 5 (a), Figure 6 (a), Figure 7 It can be seen from (a) that there are many false detections and missed detections in the existing ViBe detection method. The main reason is that this method belongs to a traditional method and is extremely sensitive to the jittery background in the satellite video frames, resulting in incorrect target recognition and prediction of the target box size. From Figure 5 (b), Figure 6 (b), Figure 7 It can be seen from (b) that although the existing CenterNet detection method has a high recall rate, its accuracy is low and it is prone to false detections. The main reason is that this method only utilizes the appearance features of the target and does not effectively extract the motion information between satellite video frames, resulting in easy occurrence of false detections. From Figure 5 (c), Figure 6 (c), Figure 7 It can be seen from (c) that the existing DSFNet detection method has achieved good detection results. The main reason is that this method takes into account both the appearance information and the motion information of the target when extracting features, so there are fewer false detections and missed detections. However, through network structure analysis, this method has the problem of a too large number of network parameters, and deploying it on mobile devices will not meet the requirements of real-time target detection. From Figure 5 (d), Figure 6 (d), Figure 7(d) It can be seen that compared with the classification results of the three existing technologies, the detection results of the present invention have fewer false detections and missed detections, and have good detection results on multiple data sets. In addition, through network structure analysis, the model of the present invention has fewer parameters and faster detection speed, proving that the detection effect and detection speed of the present invention are superior to the first three existing technology detection methods, and the detection effect is relatively ideal.
[0080] Use three precision evaluation indicators (precision, recall rate, F1 score) to evaluate the detection results of the four methods of the simulation experiment of the present invention respectively. Using the following formulas, calculate the accuracy rate, recall rate, and F1 score, and plot all the calculation results in Table 1:
[0081]
[0082]
[0083]
[0084] Table 1. Quantitative analysis table of the precision results of the present invention and each existing technology in the simulation experiment
[0085]
[0086] Combined with Table 1, it can be seen that the F1 scores of the present invention on the Dubai satellite video, San Diego satellite video, and Las Vegas satellite video are 92.22%, 88.55%, and 89.45% respectively, and these indicators are all higher than the three existing technology methods. In addition, the precision and recall rate of the present invention on the above three data sets are also higher than the three existing technology methods, proving that the present invention can obtain higher satellite video mobile vehicle detection accuracy.
[0087] Use three performance evaluation indicators (number of parameters, floating-point operations per second, frames per second) to evaluate the performance of the CenterNet detection method, DSFNet detection method, and the method of the present invention respectively. The number of parameters refers to the total sum of the parameters that need to be learned in the model. The floating-point operations per second refer to the number of floating-point operations executed by the model per second. The frames per second refer to the number of frames processed by the model in computer vision tasks. These three performance evaluation indicators can all be calculated by statistical methods. Since the ViBe detection method does not belong to the deep learning method and these indicators cannot be calculated, no comparison is made.
[0088] Table 2. Quantitative analysis table of the performance results of the present invention and each existing technology in the simulation experiment
[0089]
[0090] As can be seen from Table 2, the number of parameters of the present invention is 0.35M, the number of floating-point operations per second is 26.84G, and the number of frames per second is 66.84. Both the number of parameters and the number of frames per second are higher than those of the two existing deep learning techniques. Compared with the CenterNet detection method, the number of parameters of the present invention only accounts for 1.84% of it, while the number of frames per second is 6.88 times that of it, which proves that the present invention can not only detect moving vehicles in satellite videos with high precision, but also has a lighter and faster detection speed for moving vehicles in satellite videos.
[0091] The above simulation experiments show that: the method of the present invention uses the constructed lightweight moving vehicle detection network, which can fully exploit the spatial and temporal information of moving vehicles and effectively fuse the low-level and high-level features, thereby solving the problem of artifacts caused by the movement of the satellite platform and the change of atmospheric illumination conditions in traditional technologies, as well as the problem of low recall rate caused by the failure of appearance-based detection methods to utilize spatio-temporal motion information. In addition, considering the limitations of memory and computing resources faced by actual deployment devices, the present invention adopts the strategy of reducing the number of channels and using smaller convolutional kernels during the network design process to compress the model, thereby minimizing the model parameters and computational cost, and solving the problem that high-precision detection methods cannot detect moving vehicles in real time on mobile devices. It is a very practical method for detecting moving vehicles in satellite videos.
Claims
1. A method for detecting moving vehicles in satellite videos based on spatio-temporal motion and model compression, characterized in that, In the construction of the spatio-temporal motion information extraction sub-network, a model compression strategy with fewer channels and small-size three-dimensional convolutional layers is adopted; the steps of this detection method are as follows: Step 1, build an input sub-network composed of a scaling layer and an enhancement layer in series; Step 2, build a feature extraction sub-network composed of a three-dimensional transposed convolutional layer and a three-dimensional convolutional layer in series; Step 3, construct a lightweight spatio-temporal motion information extraction sub-network composed of a shallow spatio-temporal motion information extraction module, a deep spatio-temporal motion information extraction module, and a spatio-temporal motion information fusion module: Step 3.1, build a shallow spatio-temporal motion information extraction module composed of a three-dimensional convolutional layer, a first expansion layer, a compression layer, a second expansion layer, a residual connection layer, a three-dimensional max pooling downsampling layer, and a deformable convolutional layer in series; set the convolutional kernel sizes of the three-dimensional convolutional layer and the compression layer to 1, the strides to 1, the input channel numbers to 16 and 32 respectively, and the output channel numbers to 16; the residual connection layer takes the original data and the current data as inputs and adds them along the channel direction; Set the convolutional kernel size of the three-dimensional pooling downsampling layer to 3×1×1 and the stride to 3×1×1; set the convolutional kernel size of the deformable convolutional layer to 3, the stride to 1, the input channel number to 64, and the output channel number to 32; Step 3.2, build a deep spatio-temporal motion information extraction module composed of a first three-dimensional max pooling downsampling layer, a three-dimensional convolutional layer, a first expansion layer, a compression layer, a second expansion layer, a residual connection layer, a second three-dimensional max pooling downsampling layer, a deformable convolutional layer, and a two-dimensional transposed convolutional layer in series; set the convolutional kernel size of the first three-dimensional max pooling downsampling layer to 1×2×2 and the stride to 1×2×2; set the convolutional kernel sizes of the three-dimensional convolutional layer and the compression layer to 1, the strides to 1, the input channel numbers to 32 and 64 respectively, and the output channel numbers to 32; the residual connection layer takes the original data and the current data as inputs and adds them along the channel direction; Set the convolutional kernel size of the second three-dimensional pooling downsampling layer to 3×1×1 and the stride to 3×1×1; set the convolutional kernel size of the deformable convolutional layer to 3, the stride to 1, the input channel number and the output channel number to 32; set the convolutional kernel size of the two-dimensional transposed convolutional layer to 4, the stride to 2, the input channel number and the output channel number to 32; Step 3.3, build a spatio-temporal motion information fusion module composed of a connection layer and a deformable convolutional layer in series; set the convolutional kernel size of the deformable convolutional layer to 3, the stride to 1, the input channel number and the output channel number to 32; the connection layer adds the input data along the channel direction; Step 3.4, cascade the residual connection layer in the shallow spatio-temporal motion information extraction module with the transposed convolutional layer in the deep spatio-temporal motion information extraction module, and cascade the deformable convolutional layer in the shallow spatio-temporal motion information extraction module and the first three-dimensional max pooling downsampling layer in the deep spatio-temporal motion information extraction module with the connection layer of the spatio-temporal motion information fusion module to obtain a lightweight spatio-temporal motion information extraction sub-network; Step 4, construct a detection head sub-network composed of a key point heat map prediction layer, a center point offset prediction layer, and a target box size prediction layer in parallel; Step 5, construct a lightweight mobile vehicle detection network: Cascade the enhancement layer input to the sub-network with the 3D transposed convolutional layer in the feature extraction sub-network; Cascade the 3D convolutional layer in the feature extraction sub-network with the 3D convolutional layer in the shallow spatio-temporal motion information extraction sub-network; cascade the deformable convolutional layer in the spatio-temporal motion information fusion sub-network with the key point heat map prediction layer, the center point offset prediction layer, and the target box size prediction layer in the detection head sub-network to form a lightweight mobile vehicle detection network; Step 6, generate a training set: Form a sample set from at least 300 frames of large-size satellite videos, with no less than 50 targets in each video frame in the sample set, and mobile vehicles lacking appearance information with a size range of 4 to 20 pixels; crop each frame of the sample set into a video frame of 512×512 pixels, and stack the video frames at intervals of 5 frames to form a training set; Step 7, train the lightweight mobile vehicle detection network: Set the training parameters, input the training set into the lightweight mobile vehicle detection network, use the Adam optimizer and the stochastic gradient descent method to iteratively update the weights of the network parameters until the loss function L converges, and obtain a trained network model; Step 8, detect mobile vehicles in satellite videos: After performing the same cropping and stacking operations on the satellite video frames to be detected as in Step 6, input them into the trained network, output the mobile vehicle detection results and save them as a label file in coco format.
2. The method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite video according to claim 1, wherein, The scaling layer described in Step 1 is implemented using an image adaptive scaling algorithm, and the data enhancement layer is implemented using a random mirror flip algorithm, an affine transformation algorithm, and a photometric distortion algorithm.
3. The method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite video according to claim 1, wherein, For the 3D transposed convolutional layer in Step 2, the input data size is set to 512×512×5×3, the output data size is set to 1024×1024×5×3, the convolutional kernel size is set to 1×3×3, and the stride is set to 1×2×2; for the 3D convolutional layer, the input data size is set to 1024×1024×5×3, the output data size is set to 512×512×5×16, the convolutional kernel size is set to 3, and the stride is set to 1×2×2.
4. The method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite videos according to claim 1, wherein In Step 3.1, both the first expansion layer and the second expansion layer are composed of two parallel 3D convolutional layers in series with a splicing layer. The convolutional kernel sizes of the two parallel 3D convolutional layers are set to 1 and 3 respectively, the stride is set to 1 for both, and the number of input channels and output channels are both set to 16; the splicing layer stacks and splices the input data along the channel direction.
5. The method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite video according to claim 1, wherein In Step 3.2, both the first expansion layer and the second expansion layer are composed of two parallel 3D convolutional layers in series with a splicing layer. The convolutional kernel sizes of the two parallel 3D convolutional layers are set to 1 and 3 respectively, the stride is set to 1 for both, and the number of input channels and output channels are both set to 32; the splicing layer stacks and splices the input data along the channel direction.
6. The method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite video according to claim 1, wherein The structures of the key point heat map prediction layer, the center point offset prediction layer, and the target box size prediction layer described in step 4 are the same, and they are all composed of a first convolutional layer and a second convolutional layer connected in series. The convolutional kernel sizes of the first and second convolutional layers are set to 3 and 1 respectively, the stride is set to 1, and the number of input channels is set to 32 and 256 respectively. The number of output channels of the second convolutional layer of the key point heat map prediction layer is set to 1, and the number of output channels of the second convolutional layer in the center point offset prediction layer and the target box size prediction layer are both set to 2.
7. The method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite videos according to claim 1, characterized in that, The stacking operation described in step 6 refers to stacking 5 consecutive video frames along the time dimension.
8. The method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite video according to claim 1, wherein The setting of the training parameters described in step 7 means setting the batch size of the network to 16.
9. The method for detecting moving vehicles based on spatio-temporal motion and model compression in satellite videos according to claim 1, wherein The loss function L described in step 7 is as follows: L = L h + λ f L f + λ s L s Among them, L h represents the key-point heatmap loss function, L f represents the center-point offset loss function, L s represents the target box size loss function; λ f and λ s represent weight parameters with values of 1 and 0.1 respectively; N represents the number of targets in each frame of the training set video; ∑ represents the summation operation, represents the prediction score that the key-point heatmap coordinates (x, y) output by the key-point heatmap prediction layer belong to the c-th category, Y xy,c represents whether the key-point heatmap coordinates (x, y) in each video frame of the training set belong to the c-th category. If (x, y) belongs to the c-th category, then Y xy,c takes the value of 1. If (x, y) does not belong to the c-th category, then Y xy,c takes the value of 0; α and β represent weight parameters with values of 2 and 4 respectively; represents the center-point offset of the p-th target output by the center-point offset prediction layer, represents the center-point offset of the p-th target in each frame of the training set video. R represents the proportionality coefficient of the video frame size divided by the key-point heatmap size, represents the floor operation, and |·| represents the absolute value operation; represents the size of the k-th target box output by the target box size prediction layer, S k represents the size of the k-th target box in each frame of the training set video.
Citation Information
Patent Citations
Brightness-time sequence feature fused satellite video dynamic vehicle target extraction method
CN112489055A
Vehicle driving environment abnormity monitoring method and device, electronic equipment and storage medium
CN113287120A
Prediction method and device, intelligent driving system and vehicle
CN115273015A