Multi-moving-target tracking method for remote sensing video and storage medium
By establishing detection subsequences of multi-frame images in remote sensing video, extracting static and dynamic features and performing spatiotemporal fusion, combining Kalman filtering and optical flow matrix processing, the problem of low tracking accuracy of multiple motion targets in remote sensing video is solved, and higher tracking accuracy and stability of target trajectory are achieved.
Patent Information
- Application Number
- CN202511000937.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-21
AI Technical Summary
In remote sensing video, the tracking accuracy of multiple moving targets is low, the target size is small, the movement is slow and easily blocked, resulting in poor feature separability and difficult to accurately identify and track.
By establishing detection subsequences of multi-frame images, static and dynamic features are extracted for spatiotemporal fusion, state correction is performed using Kalman filtering and optical flow matrix, interpolation reconstruction and static target removal are performed, and camera motion information is stripped.
The tracking accuracy of multiple moving targets in remote sensing video is improved, the feature expression of weak targets is enhanced, camera motion interference is reduced, and the continuity and accuracy of the target trajectory is optimized.
Smart Images

Figure CN120510189A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer vision and pattern recognition, and in particular to a method for tracking multiple moving targets in remote sensing video and a storage medium. Background Art
[0002] With the rapid development of remote sensing Earth observation technology, remote sensing video has become an important means of efficiently collecting ground object information. Remote sensing video refers to video that records the electromagnetic radiation levels of various ground objects. Multi-moving target tracking based on remote sensing video is a core technology in remote sensing video processing, with significant application prospects and value in scenarios such as military inspection, battlefield analysis, and digital cities. Multi-moving target tracking technology aims to identify all moving targets in a video, track each target's trajectory, and assign a unique identifier (ID) to each target. However, in remote sensing video, targets are often small, move slowly, and are prone to severe occlusion and interference from similar targets. These characteristics weaken the target's spatial signature, reducing the ability to distinguish foreground and background features, making target detection and recognition extremely difficult. Furthermore, the low separability of features between similar targets poses significant challenges for trajectory tracking and ID assignment. Summary of the Invention
[0003] The purpose of this application is to provide a method and storage medium for tracking multiple moving targets in remote sensing videos, so as to solve the problem of low accuracy in tracking multiple moving targets in remote sensing videos.
[0004] To achieve the above objectives, the present application provides, in a first aspect, a method for tracking multiple moving targets in a remote sensing video, comprising: Acquire a remote sensing video to be processed, and establish multiple detection subsequences based on multiple frames of images included in the remote sensing video; Extracting static features and dynamic features of each detection subsequence respectively to obtain spatiotemporal fusion features of the detection subsequence, and obtaining a detection frame of each frame of the image based on the spatiotemporal fusion features, where each detection frame corresponds to an object; Building a target library corresponding to a starting frame in the remote sensing video based on an initial detection frame set of the starting frame, sequentially obtaining predicted states of multiple frames of the image through Kalman filtering based on the target library of the previous frame of each frame of the image after the starting frame, and calculating an optical flow matrix of the two adjacent frames of the image based on background feature points of the two adjacent frames of the image; Based on the optical flow matrix, the predicted state of each frame of the image is corrected respectively, and a first target library sequence consisting of a target library of each frame of the image is obtained by matching with the detection frame of each frame of the image, where the target library of each frame of the image includes a plurality of targets; Based on the optical flow matrix, a post-processing operation is performed on the first object library sequence to obtain a second object library sequence, wherein the post-processing operation includes interpolation reconstruction and stationary object elimination.
[0005] A second aspect of the present application provides a computer-readable storage medium, in which a program is stored. The program can be loaded by a processor and execute the above-mentioned method for tracking multiple moving targets in remote sensing videos.
[0006] The beneficial effects of this application are: This application establishes multiple detection subsequences based on multiple frames of remote sensing video. It then extracts spatiotemporal fusion features, including static and dynamic features, from each detection subsequence. Based on these spatiotemporal fusion features, it generates a detection bounding box for each frame, enabling multi-moving target detection. Using these spatiotemporal fusion features as input for multi-moving target detection enhances the feature representation of small and weak targets in remote sensing video. Then, based on the detection bounding box for each frame, a Kalman filter is applied sequentially to obtain the predicted state of the multiple frames. The optical flow matrix is then calculated based on the background feature points of the two adjacent frames. The predicted state of the multiple frames is then modified based on the optical flow matrix to obtain a first target library sequence consisting of target libraries for each frame. This optical flow matrix reduces the interference of global target motion introduced by camera motion and maintains the stability of small and weak targets. Finally, the first target library sequence is post-processed with interpolation reconstruction and stationary target elimination based on the optical flow matrix to obtain a second target library sequence. This method separates target motion information from camera motion information, enabling more accurate detection of missed targets and determination of target motion status. Based on the above technical solution, this application can improve the tracking accuracy of multiple moving targets in remote sensing videos.
[0007] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A flowchart of a method for tracking multiple moving targets in a remote sensing video provided in an embodiment of the present application is shown; Figure 2 A schematic diagram of background modeling and motion region extraction provided in an embodiment of the present application; Figure 3 A schematic diagram of static feature enhancement based on a motion region provided in an embodiment of the present application; Figure 4 A schematic diagram of a dual-branch temporal feature enhancement and fusion module provided in an embodiment of the present application; Figure 5 A schematic diagram of a multi-scale feature aggregation module provided in an embodiment of the present application; Figure 6A schematic diagram of the overall framework of a multi-moving target detection model provided in an embodiment of the present application; Figure 7 A flowchart of the multi-moving target tracking reasoning stage provided in an embodiment of the present application. DETAILED DESCRIPTION
[0009] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0010] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and are not to be construed as indicating or implying relative importance or implicitly specifying the number of the technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of this application, "plurality" means two or more, unless otherwise specifically qualified. In this application, the word "exemplary" is used to mean "serving as an example, illustration, or illustration." Any embodiment described in this application as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is provided to enable anyone skilled in the art to implement and use the present application. In the following description, details are listed for illustrative purposes. It should be understood that one of ordinary skill in the art will recognize that the present application can be implemented without these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0011] Traditional approaches to multi-moving target detection and tracking in remote sensing videos are still subject to numerous limitations in practical application due to scene complexity, the small size of the targets, and the imperfections of current algorithms. Typically, moving target extraction methods utilize inter-frame difference information and input video sequences to extract features, implicitly extracting motion information through ground-truth supervision. The two types of motion information are only fused through the final decision, making it impossible to train traditional methods and deep learning-based approaches simultaneously. Only simple networks are used for temporal frame feature extraction, making temporal feature extraction less prominent. Furthermore, in trajectory association, when processing small targets, traditional methods' updates to target size can lead to unstable updates and fail to account for the effects of camera motion.
[0012] Based on this, the embodiment of the present application proposes a method for tracking weak multiple moving targets in remote sensing videos, which includes two parts: multiple moving target detection and multiple moving target tracking. First, in the step of multiple moving target detection, the embodiment of the present application designs a static and temporal dual-branch extraction network. On the one hand, through motion area extraction in the static branch, the features of weak moving targets in the remote sensing video are fully enhanced, the accuracy of weak moving target detection is improved, and the false detection rate of stationary targets is reduced. On the other hand, in the dynamic branch, by introducing the ResNet50-3D network, the continuity and spatial context information of weak moving targets in the time dimension are fully captured. Through the careful design of the dual-branch temporal feature enhancement module, the static frame features and the temporal frame features are efficiently interacted and deeply aligned, so that the fused features fully express the temporal characteristics of the target. In terms of parameter pre-training, by designing a convolution parameter reconstruction scheme from 2D to 3D, the model does not need to introduce additional data other than COCO and ImageNet, and the initialization parameters of ResNet50-3D are constructed, so that the model training starts from a stable state, improving the robustness and performance of the model training. Secondly, to address the inadequate consideration of camera motion in traditional multi-target tracking algorithms for remote sensing video when tracking multiple moving targets, a Kalman state correction module for weakly moving targets, based on optical flow matrix estimation, was designed. This enables the model to accurately account for global target motion information introduced by camera motion. Furthermore, in the post-processing stage, a target interpolation and reconstruction scheme and a stationary target elimination scheme, based on the optical flow matrix, were designed to strip away target and camera motion information, resulting in more accurate target interpolation and reconstruction, as well as stationary target discrimination. This is explained in detail below.
[0013] Figure 1 Schematic diagram of a flow chart of a method for tracking multiple moving targets in a remote sensing video provided in an embodiment of the present application. Figure 1 As shown, the multi-moving target tracking method may include steps 101-105, which are described in detail below.
[0014] Step 101: Acquire a remote sensing video to be processed, and establish multiple detection subsequences based on multiple frames of images included in the remote sensing video.
[0015] In this embodiment of the present application, the detection subsequence is a multi-frame image combination constructed to extract spatiotemporal fusion features. It can be a set of continuous or non-continuous frame images, including a static frame and multiple time-series frames. Compared to traditional single-frame detection, which only uses spatial features and has difficulty capturing inter-frame motion relationships, resulting in a high missed detection rate, this embodiment of the present application enhances the feature expression of small and weak targets by expanding single-frame detection to multi-frame joint analysis.
[0016] Step 102: extract the static features and dynamic features of each detection subsequence respectively to obtain the spatiotemporal fusion features of the detection subsequence, and obtain the detection frame of each frame image based on the spatiotemporal fusion features, where each detection frame corresponds to one target.
[0017] In an embodiment of the present application, static features are spatial dimension features extracted from a single-frame image, do not involve changes in the time dimension, and are semantic expressions of the image's spatial structure. For example, they can be extracted from static frames in a detection subsequence. Dynamic features are temporal dimension features extracted from multiple frames, which can describe the target's motion pattern and trajectory changes between consecutive frames, and capture the dynamic relationship between frames. For example, dynamic features can be extracted based on multiple frames of a detection subsequence. Targets in remote sensing videos usually have the problem of a low proportion of weak targets, and static features are relatively weak. Therefore, it is necessary to enhance the expression of weak targets through dynamic features. In an embodiment of the present application, static features and dynamic features are fused to obtain spatiotemporal fusion features of the detection subsequence. The spatiotemporal fusion features are used as input for multi-motion target detection to obtain a detection frame for each frame, thereby detecting multiple targets in each frame, which can enhance the feature expression of weak targets in remote sensing videos and reduce the missed detection rate caused by a low proportion of target pixels.
[0018] Step 103: construct a target library corresponding to the starting frame based on the initial detection frame set of the starting frame in the remote sensing video, and sequentially obtain the predicted state of multiple frames of images based on the target library of the previous frame of each frame after the starting frame through Kalman filtering, and calculate the optical flow matrix of two adjacent frames of images based on the background feature points of the two adjacent frames of images.
[0019] The Kalman filter is an algorithm that uses linear system state equations and observation data from the system's input and output to optimally estimate the system state. The Kalman filter can generate a predicted state for the current frame based on the image preceding the current frame. Therefore, a Kalman filter model can be established based on the target library information for each frame, and the predicted state for each frame can be obtained. The target library can first be constructed based on the initial set of detection boxes in the starting frame of the remote sensing video. After the starting frame, Kalman predictions are then performed based on the target library of the previous frame. The recursive prediction of the Kalman filter can reduce trajectory loss caused by short-term occlusions. The optical flow matrix calculates the global motion between two adjacent frames using background feature points in the two frames, separating camera motion from actual target motion. This provides a basis for subsequent corrections and reduces the misinterpretation of background displacement as target motion.
[0020] Step 104: Based on the optical flow matrix, the predicted state of each frame image is corrected respectively, and by matching with the detection frame of each frame image, a first target library sequence consisting of the target library of each frame image is obtained, and the target library of each frame image includes multiple targets.
[0021] The Kalman filter's predicted state for each frame is modified based on the optical flow matrix. By matching the predicted state with the detection bounding box for each frame, a target library for each frame is generated. The first target library sequence consists of the target libraries for multiple frames after the corrected predicted state. Using the optical flow matrix to modify the Kalman filter's predicted state can reduce spurious displacement caused by camera motion and eliminate global motion artifacts. This method is suitable for scenarios where global image offset due to Earth's rotation is a factor in satellite remote sensing.
[0022] Step 105: Based on the optical flow matrix, perform post-processing operations on the first object library sequence to obtain a second object library sequence. The post-processing operations include interpolation reconstruction and static object elimination.
[0023] The first target library sequence may contain some spurious trajectories that have become invalid due to prolonged periods of absence. For example, the detection boxes corresponding to targets in the first target library sequence may be missed due to noise. For example, small targets may be misclassified as background, which can lead to broken target trajectories. Therefore, post-processing with interpolation and reconstruction can be performed on the first target library sequence. Interpolation and reconstruction can resolve the problem of broken and discontinuous target trajectories.
[0024] The first object library sequence after interpolation and reconstruction may still contain stationary objects with minimal motion for extended periods, representing false trajectories. For example, parked vehicles can be repeatedly detected as moving objects, generating redundant trajectories. Therefore, stationary objects can be removed from the first object library after interpolation and reconstruction. This can address the problem of stationary objects being mistakenly detected as moving objects.
[0025] Through interpolation reconstruction and stationary target elimination, the target trajectory can be made to conform to the real target trajectory of physical laws, the trajectory continuity and accuracy can be optimized, and the final second target library sequence can be generated. The second target library sequence includes the final target library for each frame image, and each target library includes multiple moving targets in the frame image.
[0026] The embodiment of the present application uses spatiotemporal fusion features as input for multi-motion target detection, thereby enhancing the feature expression of small and weak targets in remote sensing videos. Based on the optical flow matrix, the predicted state of multiple frames of images after Kalman filtering can be corrected to obtain a first target library sequence composed of the target library of each frame of image. The optical flow matrix reduces the interference of global target motion introduced by camera motion and maintains the stability of small and weak targets. Finally, the first target library sequence is post-processed by interpolation reconstruction and elimination of stationary targets based on the optical flow matrix to obtain a second target library sequence. In this way, the motion information of the target and the motion information of the camera can be stripped off, making the supplementation of missed targets and the judgment of the target motion state more accurate. Based on the above technical solution, the present application can improve the multi-motion target tracking accuracy of remote sensing videos.
[0027] In an embodiment of the present application, each detection subsequence may include static frames and sequential frames. The static frame refers to the reference frame in the detection subsequence, which is used to extract the spatial features of the target, such as appearance, texture, contour, etc. In the embodiment of the present application, the last frame of the detection subsequence is taken as an example of a static frame. It should be noted that in actual applications, other frame images can also be selected as static frames. Sequential frames refer to other frames in the detection subsequence other than static frames, which are distributed before the static frames in chronological order and are used to extract the motion features of the target in the time dimension, such as displacement, speed, and trajectory. The combination of static frames and sequential frames can solve the problem of "weak single-frame features" of weak targets in remote sensing videos, and provide higher quality feature input for subsequent monitoring and tracking of multiple moving targets.
[0028] In step 101, a set length of a detection subsequence set in advance can be obtained. The set length is the first number of images included in the detection subsequence. The first number is the number of images that are set in advance to make up a detection subsequence. Then, each frame of the image is sequentially used as a static frame, and the second number of images located before the static frame are used as sequential frames, and a detection subsequence corresponding to each frame of the image is constructed based on the static frame and the sequential frame. The second number is the number of sequential frames in the detection subsequence. Since the length of the detection subsequence is the first number and the static frame is the last frame of the sequence, the second number corresponding to the sequential frames is the first number minus one.
[0029] For example, suppose remote sensing video include Frame image, remote sensing video It can be expressed as . And each frame of the image As a static frame, combined with its corresponding time frame, the set length is The detection subsequence , then each detection subsequence can be expressed as The complete set of detection subsequences can be expressed as: Each subsequence It will be used as the input of the subsequent target detection network to extract spatiotemporal fusion features and perform target reasoning.
[0030] During the construction of the detection subsequence, the number of frames of the image preceding the static frame may be less than the second number, making it impossible to construct the required set length of the detection subsequence. Therefore, in an embodiment of the present application, a padding operation is performed using static frames, i.e., copying the static frames and using the static frames as sequential frames to pad the detection subsequence. Specifically, if a third number of images preceding the static frame in the remote sensing video is detected to be less than the second number, a fourth number of the static frames is copied and used as sequential frames to fill the detection subsequence. The third number refers to the number of images preceding the static frame, and the fourth number refers to the number of static frames that need to be copied. Therefore, the sum of the third and fourth numbers is the second number, which satisfies the set length required to construct the detection subsequence. For example, assuming the set length is 5 frames, the first number is 5. Therefore, it is necessary to take the first 4 frames of the current frame as sequential frames, and the second number is 4. Assuming that the current frame is the third frame in the remote sensing video, the number of images preceding it is 2 frames, and the third number is 2. The third number is smaller than the second, so subtracting 2 from 4 gives a fourth number of 2, indicating that two static frames need to be copied as time-series frames. Static padding improves data availability without significantly increasing computational cost, providing more reliable input for subsequent calculations and enhancing algorithm robustness.
[0031] In step 102, for each detection subsequence, the static frames in the subsequence are fed into a deep layer aggregation neural network (DLA-34) to extract spatial features at different semantic levels, generating static features for the subsequence. DLA-34 fuses multi-scale features through dense concatenation, preserving spatial details of the object, such as texture and contours. In this embodiment, the deep layer aggregation neural network DLA-34 is initialized using weights pre-trained on COCO and ImageNet.
[0032] Then, the temporal frames and the static frames are aligned using the ORB (Oriented FAST and Rotated BRIEF) algorithm, and the Gaussian mixture model (GMM) is used to model the background and separate the foreground of the temporal frames and the static frames to obtain the motion foreground mask of the static frame.
[0033] The ORB algorithm combines the speed advantage of FAST feature point detection with the binary nature of the BRIEF descriptor, and introduces directional information (calculating the main direction based on the image pyramid). This allows for pixel-level alignment of sequential frames with static frames, providing a spatially consistent foundation for subsequent background modeling. In remote sensing video, the alignment error for camera translation can be kept within a small pixel, meeting the requirements for detecting small and dim targets. The Gaussian mixture model models the grayscale value of each pixel in an image as a weighted sum of K Gaussian distributions (typically K = 3-5). Background patterns are learned by iteratively updating the Gaussian distribution parameters (mean, variance, and weights). When the pixel value matches all Gaussian distributions below a threshold, it is identified as a foreground (moving target), ultimately generating a binary motion foreground mask.
[0034] Figure 2 Schematic diagram of background modeling and motion region extraction provided in the embodiment of this application. Figure 2 As shown, the ORB algorithm can be used to extract keypoint features. A proximity algorithm, or K-Nearest Neighbor (KNN) classification, is then used to match features between frames to obtain a homography matrix for each frame. All time-sequential frames are then mapped to the coordinate system of the static frame, resulting in a unified image sequence. Based on these aligned frames, a Gaussian mixture model is used for background modeling and motion region extraction, resulting in a motion foreground mask that describes the moving target region.
[0035] Next, the motion foreground mask of the static frame is fused with the static features of the static frame to obtain the static enhanced features of the detection subsequence. The binary motion foreground mask is fused with the deep features of the static frame channel by channel. The "gating" effect of the mask enhances the characteristic response of the moving target, suppresses static background noise, and generates static enhanced features, providing a purer target representation for subsequent detection.
[0036] Figure 3 Schematic diagram of a static feature enhancement based on a motion region provided in an embodiment of the present application. Figure 3 As shown, we can first pool the motion foreground mask to the same resolution as the static frame features, then multiply the static frame features and the motion foreground mask point by point to retain the features of the motion region. This residual structure is then added to the original static frame features to enhance the motion region portion of the static features. The formula can be as follows: ; in, represents element-wise multiplication, Represents element-level addition, and finally obtains the feature set after foreground enhancement.
[0037] In the implementation of this application, for dynamic features, the detection subsequence can be input into the three-dimensional convolutional network ResNet50-3D to obtain the dynamic features of the detection subsequence. In the embodiment of this application, the three-dimensional convolutional network ResNet50-3D is initialized by replicating the convolution parameters of the two-dimensional convolutional network ResNet50 pre-trained on COCO and ImageNet along the time dimension a set number of times, and dividing the result by the set number of times as the initial weights.
[0038] In order to improve the training performance of the model, traditional model training methods all choose to load pre-trained parameters into the backbone network. The ResNet50-3D backbone network has the ability to process sequential images to extract spatiotemporal features, but its model cannot be trained to obtain parameters on datasets composed of static images such as COCO and ImageNet. In order to improve the training performance of the model, the embodiment of the present application provides a method for generating Resnet50-3D pre-trained parameters without adding additional datasets. Specifically: first, all the convolution parameters in the Resnet50 network are extracted, and they are copied a set number of times along the time dimension. times to keep the total number of parameters consistent with Resnet50-3D. Then, to avoid parameter expansion, the convolution parameters are divided by This method constructs the pre-training parameters. Without introducing any datasets other than COCO and ImageNet, it effectively constructs the initial training parameters of the backbone network, effectively improving the training performance of the model.
[0039] Finally, the static enhancement features and dynamic features are fused based on the attention mechanism of the convolutional neural network to obtain the spatiotemporal fusion features of the detection subsequence.
[0040] Traditional convolution has a large amount of computation and parameters when extracting spatial context information and is coupled with spatial and channel weights. The embodiment of the present application adopts a combination of depthwise separable convolution and channel attention mechanism, which can significantly reduce the number of parameters and computation while improving the local spatial modeling capability and channel adaptability of the convolutional network and enhancing the dynamic attention capability to the target area. In scenes with weak moving targets, adaptively enhancing useful channels and suppressing invalid channels can significantly improve the network's responsiveness to key features and effectively integrate spatiotemporal features.
[0041] Figure 4 This is a schematic diagram of a dual-branch temporal feature enhancement and fusion module provided in an embodiment of the present application. In this embodiment of the present application, the convolutional neural network may include a multi-layer neural network. By designing a dual-branch temporal feature enhancement and fusion module, it is used to achieve efficient fusion and enhancement of static enhancement features and dynamic features.
[0042] Specifically, for each neural network layer, the static enhancement features and dynamic features are concatenated along the channel dimension to obtain a joint feature. Then, depthwise separable convolution is used to extract local contextual information from the joint feature. This local contextual information is residually added to the joint feature and then batch normalized to obtain a normalized feature. The normalized feature is globally pooled to generate a channel description vector. This is converted into channel weights through a fully connected layer and activation operation. The normalized feature is then multiplied channel-by-channel by the channel weights to obtain a channel-weighted feature. A multilayer perceptron (MLP), such as one consisting of two layers of 1×1 convolution and a GELU activation function, performs a nonlinear transformation on the channel-weighted features and then residually adds them to the channel-weighted features to ensure information integrity and cross-channel expressiveness. This results in a fused feature for each neural network layer. Finally, the fused features from each neural network layer are aggregated to obtain the spatiotemporal fused features of the detection subsequence.
[0043] Finally, the spatiotemporal fusion features are input into the multi-scale feature aggregation block (FAB) module and the target detection module to obtain the detection box of the static frame. Figure 5 This is a schematic diagram of a multi-scale feature aggregation module provided in an embodiment of the present application. After the fused features pass through the FAB module, output features are obtained. The FAB module consists of a convolutional layer and an upsampling layer, and its structure is shown in the figure. Subsequently, the fused features can be passed through three parallel detection heads to generate the position of the detected target. Specifically, the three detection heads output the target score, the target center point position, and the target size respectively.
[0044] The above process can be used to detect multiple moving targets using a trained multi-moving target detection model. The main goal of the multi-moving target model training phase is to provide supervised learning for the detector, enabling it to extract spatial and temporal fusion features from the training sequence, enhance the representation of moving region features, and accurately predict the position of targets in the image. This process includes data preprocessing, feature extraction from static frames, feature extraction from time series frames, background modeling and motion region separation, time series feature fusion, target prediction, and loss function calculation. Figure 6 This is a schematic diagram of the overall framework of a multi-moving target detection model provided in the embodiment of this application. Figure 6 A detailed explanation of the model training for multi-moving target detection is given.
[0045] (1) Video frame preprocessing and training subsequence construction.
[0046] The dataset in the training phase consists of multiple remote sensing video sequences, denoted as: ; in, Indicates the video sequences; Indicates the total number of videos.
[0047] In order to construct a sample sequence for training, each video is first discretized into frames to obtain a frame-level representation.
[0048] .
[0049] Set the fixed length of the training subsequence to , the maximum allowed sampling interval is , that is, each training subsequence should be composed of The last frame is defined as a static frame, and the remaining frames are sequential frames.
[0050] The specific training subsequence construction process is as follows.
[0051] For each video sequence , traverse each frame one by one (in ), which is taken as a candidate static frame.
[0052] For each static frame , first in the interval Randomly select a sampling interval According to this interval Forward sampling Frame, generate a sequence of frame indices: .
[0053] If the sample index contains a frame number less than 1 (i.e., beyond the start boundary of the current video), a static frame is used. Repeat the padding to ensure that the training subsequence always contains frame.
[0054] The sampled Timed frames and static frames Combined into a complete training subsequence, recorded as: .
[0055] Perform the above operations on all video sequences and get A set of training subsequences of videos: .
[0056] Finally, the training subsequence sets generated by each video are merged to form the total training sample set: .
[0057] (2) Static frame feature extraction.
[0058] The above process constructs a complete set of training subsequences Next, for each training subsequence sample Extract static frame features.
[0059] First, the static frame of the training subsequence The DLA-34 backbone network is input to extract spatial features at different semantic levels. The output includes shallow texture information and deep semantic expression, which can be expressed as: ; in, Indicates that the static frame is The feature map extracted by the layer, They correspond to the multi-scale feature layers from shallow to deep in the network.
[0060] (3) Temporal frame feature extraction.
[0061] To meet the needs of detecting dynamic targets in scenes, we use ResNet50-3D as the backbone network for temporal frame feature extraction. Compared with the technology of using simple 3D convolution stacking, this network uses a more layered network structure to effectively model the continuous changes of targets in the temporal dimension and improve the ability to capture motion information. At the same time, the number of semantic levels of its output can be consistent with the number of semantic levels of static frame feature extraction. Specifically, the complete training subsequence Input into ResNet50-3D, all time series frames are jointly modeled to extract multi-scale features that combine spatial structure and temporal dynamics, expressed as: ; in, Indicates the The spatiotemporal fusion features of the layers, , whose resolution corresponds to the static frame features Keep consistent to facilitate subsequent fusion and alignment operations.
[0062] (4) Background modeling and motion area extraction module.
[0063] First, use the ORB algorithm to extract key point features, and use the KNN algorithm to match the features between frames to obtain the homography matrix of each frame: .
[0064] Will Mapping to the static frame coordinate system, we get a sequence of images in a unified space: .
[0065] Based on the above aligned frames, the Gaussian mixture model is used to perform background modeling and motion region extraction to obtain the motion foreground mask. , used to describe the motion target area.
[0066] (5) Static frame feature enhancement based on motion regions.
[0067] First, the motion foreground mask is pooled to the same resolution as the static frame features. The static frame features are multiplied point by point with the motion foreground mask to retain the motion region features. The residual structure is then added to the original static frame features to enhance the motion region portion of the static features. The formula is as follows: ; in represents element-wise multiplication, Represents element-level addition, and finally obtains the feature set after foreground enhancement: .
[0068] (6) Dual-branch temporal feature enhancement and fusion module.
[0069] The embodiment of the present application designs a dual-branch temporal feature enhancement and fusion module to realize static enhancement feature With dynamic features The processing flow of this module is as follows.
[0070] First, for the Layer input, respectively take static enhancement features and dynamic features , and concatenate in the channel dimension to form joint features: .
[0071] Secondly, the local context information is extracted from the spliced features through depth-wise separable convolution, and the residual is added to the original spliced features before batch normalization: .
[0072] Then, the normalized features are input into the channel adaptation module (Squeeze-and-Excitation, SE), which specifically includes the following steps.
[0073] First, yes Perform global average pooling to generate channel description vectors: Then, two fully connected layers are used to generate channel weights and activate Sigmoid: .
[0074] Finally, the input is scaled channel-wise to get a channel-weighted result: .
[0075] in, represents the ReLU activation function, represents the Sigmoid function, and is the learnable weight matrix.
[0076] Next, use two layers of 1 1. The MLP composed of convolution and GELU activation function performs nonlinear transformation on it, and finally adds it to the input feature residual to ensure information integrity and cross-channel expression to obtain the fusion feature: .
[0077] Through this process, a multi-layer enhanced fusion feature set can be obtained: .
[0078] (7) Multi-scale feature aggregation and target position prediction.
[0079] After the fused features pass through the FAB module, the output features are obtained The FAB module consists of a convolutional layer and an upsampling layer. The formula is: .
[0080] The features are then passed through three parallel detection heads to generate the location of the detected object. Specifically, the three detection heads output the object score, the location of the object center point, and the size of the object respectively.
[0081] (8) Loss function calculation and gradient backpropagation.
[0082] Match the predicted target position with the annotation box of the corresponding frame, calculate the IoU, L1 and FocalLoss losses, and pass the gradient back to update the network parameters.
[0083] (9) Generate Resnet50-3D backbone network pre-training parameters.
[0084] In order to improve the training performance of the model, existing model training methods all choose to load pre-trained parameters into the backbone network. The ResNet50-3D backbone network we use has the ability to process sequence images to extract spatiotemporal features. However, its model cannot be trained on datasets composed of static images such as COCO and ImageNet to obtain parameters. In order to improve the training performance of the model, we invented a method to generate Resnet50-3D pre-trained parameters without adding additional datasets. Specifically: First, all the convolution parameters in the Resnet50 network are extracted and copied along the time dimension. times to keep the total number of parameters consistent with Resnet50-3D. Then, to avoid parameter expansion, the convolution parameters are divided by This method constructs the pre-training parameters. Without introducing any datasets other than COCO and ImageNet, it effectively constructs the initial training parameters of the backbone network, effectively improving the training performance of the model.
[0085] Figure 7 This is a flow chart of the multi-target tracking inference stage provided in the embodiment of the present application. This stage performs multi-target tracking inference on a single input video. This process does not rely on any annotation information and completes the following key operations in sequence: frame preprocessing, detection subsequence construction, target detection, target library initialization and update, Kalman state prediction, target association matching, post-processing, etc. Its overall framework is as follows Figure 7 shown.
[0086] (1) Remote sensing video preprocessing and detection subsequence construction.
[0087] Assume that the image sequence of the remote sensing video to be processed is: ; in, Indicates the total number of frames of remote sensing video, Indicates the Frame image.
[0088] The construction method of the detection subsequence in this stage is consistent with that in the training stage. It is a static frame, combined with its corresponding time sequence frame, to form a fixed length of The detection subsequence Finally, we get a complete set of detection subsequences: .
[0089] Each detection subsequence It will be used as the input of the subsequent target detection network to extract spatiotemporal fusion features and perform target reasoning.
[0090] (2) Multi-target detection reasoning.
[0091] For each detection subsequence , the same structure as the training stage is used for target detection, which mainly includes the following modules: static frame feature extraction, temporal frame feature extraction, static frame feature enhancement based on motion area, dual-branch spatiotemporal feature enhancement and fusion, and multi-scale feature aggregation and position prediction.
[0092] Finally, the detection frame set of the frame is output: ; Among them, each detection box , indicating the center position, width, height and score. Only the confidence is retained The target is used for subsequent associations.
[0093] (3) Target library structure definition.
[0094] Set up the first The target libraries for the frames are: ; in, Indicates the number of active targets in the current frame.
[0095] Each goal Defined as an ordered tuple consisting of five parts: .
[0096] The meaning of each part is as follows.
[0097] Target Number: ; Indicates the unique identifier assigned to the target throughout the tracking process, used for cross-frame target association and result output.
[0098] Space state vector: ; Indicates the coordinates of the target center point in the current frame ,area and aspect ratio and first-order velocity information of coordinates and area.
[0099] Kalman filter parameters: ; in Shared for all targets, is the private state covariance matrix of each target.
[0100] Valid status flags: ; Indicates whether the target is valid in the current frame (1 indicates valid, 0 indicates temporary mismatch or tracking failure).
[0101] Historical frame index collection: ; Record the frame number in the video where the target was successfully observed.
[0102] In step 103, the initial detection frame set of the starting frame in the remote sensing video can be obtained first. The initial detection frame set includes multiple initial detection frames, each of which corresponds to an initial target. For example, the initial detection frame set can be obtained from the multiple frames of the video. 1 image as the starting frame, according to the detection result set Initialize the target library ,in Indicates the number of targets detected in the first frame.
[0103] Then, the initial target corresponding to each initial detection frame is assigned a target identifier and initialized Kalman filter, and the initial detection frame is converted into an initial state vector to construct the target library corresponding to the starting frame.
[0104] Specifically, for each detection box , perform the following initialization steps to construct the target tuple .
[0105] Assign ID: Assign a unique number to each target .
[0106] Calculate the spatial state vector: call coordinate transformation to detect the box Convert to ,in: .
[0107] After obtaining the position parameters, the velocity term is initialized to zero, and the final spatial state vector is: .
[0108] Initialize the Kalman filter: Set the filter parameters to their default values. The velocity parameter has a high uncertainty, so assign values to each matrix accordingly: ; in, is a fixed shared matrix, They represent the state covariance matrix, observation noise covariance matrix, and process noise covariance matrix respectively.
[0109] The specific structure is as follows: Set the target state: valid flag is set to: ; Initialize the history track: add the first frame as the target's first appearance frame: ; Finally, build the target library: .
[0110] Provides initialization support for subsequent frame matching and status updates.
[0111] Next, Kalman prediction is performed based on the target library of the previous frame of each frame image after the starting frame to obtain a first predicted state vector and a first predicted state covariance of each frame image.
[0112] Specifically, from Frame starts, and the detection results of each frame are processed in turn and with the target library of the previous frame Matching and updating are performed. The process includes the following steps.
[0113] For each target in the previous frame , its state vector is: ; Its Kalman filter parameters are: .
[0114] The state prediction steps are as follows: The state vector prediction (a priori estimate) obtains the first predicted state vector: ; The state covariance prediction obtains the first predicted state covariance: .
[0115] In step 104, the sparse optical flow between the current frame image and the previous frame image is estimated to obtain the optical flow matrix of the two-dimensional affine transformation. The optical flow matrix is then applied to the position and velocity components of the first predicted state vector of the current frame image to form a corrected second predicted state vector. Correcting only the position and velocity, while retaining the size parameter, reduces the perturbation of detection noise on small targets. Furthermore, through state prediction and updating using a Kalman filter, the target ID switching rate is reduced, which can resolve the confusion problem of similar targets (e.g., multiple vehicles of the same model) in remote sensing video.
[0116] For example, by estimating the sparse optical flow between the current frame and the previous frame, we can obtain the two-dimensional affine transformation matrix , expressed as: ; in, is the rotation and scaling matrix, is the translation vector.
[0117] In order to enhance the adaptability to the overall motion of the image, the affine matrix is applied to the predicted state vector The position and velocity components of are updated as follows: .
[0118] The remaining components (such as area, aspect ratio and its velocity) remain unchanged, forming the revised state prediction vector: .
[0119] Then expand the optical flow matrix into a multi-dimensional expansion matrix. For example, construct the expansion matrix of the radiation matrix in the 7-dimensional state space , defined as follows: Then, based on the expansion matrix, the first prediction covariance matrix is updated to obtain the modified second prediction state covariance matrix: .
[0120] Next, in the first predicted state vector On the left, extract the position state component to form the predicted observation vector Match the predicted observation vector of the current frame image with the actual observation vector corresponding to the detection box of the current frame image: .
[0121] Loss matrix calculation: Each predicted target With each detection box in the current frame Calculate the matching loss and construct the cost matrix: .
[0122] Hungarian Algorithm Matching: Matching Cost Matrix Use the Hungarian algorithm to perform minimum cost matching and obtain a set of matching pairs: .
[0123] If the predicted observation vector matches the actual observation vector, the first predicted state vector and the first predicted covariance matrix corresponding to the predicted observation vector are updated using the detection box corresponding to the actual observation vector to obtain the second predicted state vector and the second predicted covariance matrix of the detection box. If the predicted observation vector does not match the actual observation vector, the detection box corresponding to the unmatched actual observation vector is initialized as a new target, and the detection box corresponding to the unmatched predicted observation vector is marked as invalid.
[0124] For example, for all matching targets , using the detection box Target status and the Kalman filter state Update and set the frame number Add to History Collection The specific steps are as follows.
[0125] The detection frame Convert to center point form: Target The predicted status is , the predicted covariance is , the update steps are as follows: Kalman gain calculation: ; Status Update: ; Covariance update: .
[0126] Update target Kalman filter The state covariance matrix in Covariance with measurement noise (Optional adjustment) and append the historical observation frame number: .
[0127] For unmatched detection boxes , initialized as a new target Add to target library , where the valid state is initialized to: .
[0128] To improve the robustness and timeliness of the target library, two timing thresholds can be set: : Validity determination window length; : Maximum number of unobserved frames tolerated.
[0129] In the Frame, update each target according to the following rules Valid status and survival status.
[0130] like , which is the initial stage, all targets are valid by default: .
[0131] like , then determine whether it is valid according to the following rules: .
[0132] If the target is continuous The frame is not observed (that is, the maximum observed frame number is greater than the threshold from the current frame): ; The target is removed from the target library Removed.
[0133] Finally, based on the second predicted state vector, the second predicted covariance matrix and the new target, a target library for each frame image is constructed to obtain a first target library sequence.
[0134] go through Frame loop reasoning, complete output of the first target library sequence: ; Among them, each Indicates in A set of multiple targets after detection, tracking, and status update at the frame time.
[0135] In the embodiment of the present application, the optical flow matrix set can be obtained based on the optical flow matrix between the two adjacent frames of the image. Then, post-processing operations are performed on the first object library sequence based on the optical flow matrix set. The following describes interpolation reconstruction and stationary object elimination respectively.
[0136] For the interpolation reconstruction operation, in step 105, first delete the targets with invalid marks in the first target library sequence. , if the target's validity flag , it is removed from the library and does not participate in subsequent trajectory output and interpolation operations.
[0137] Then, for each target identifier in the first target library sequence, its state across frames is aggregated to form a time-ordered trajectory: .
[0138] If the target mark has a discontinuous trajectory, that is, there is , then the missing frame sequence of the target logo is interpolated based on the optical flow matrix.
[0139] Specifically, the frame before the missing frame sequence is taken as the starting frame, and the frame after the missing frame sequence is taken as the ending frame. The coordinate transformation matrix of the starting frame is obtained based on the optical flow matrix, and the starting frame is aligned to the coordinate system of the ending frame. Then, by means of linear interpolation, the position information of the target corresponding to the target identifier in the aligned starting frame and ending frame is used to interpolate the interpolation target position of the missing frame sequence, wherein the interpolation target position is the position of the target in the coordinate system of the ending frame. Finally, the coordinate transformation matrix of all missing frames in the missing frame sequence is obtained based on the optical flow matrix, and the interpolation target position of the missing frame sequence is transformed to the respective coordinate system of each missing frame.
[0140] For example, let the interpolation interval be , it is known that the two end frames are and .
[0141] definition: : For a single optical flow affine matrix , frame Decomposed into two corner points: .
[0142] Apply to each point individually: .
[0143] Then combine into a new box: .
[0144] : Affine matrix for a single optical flow , split the box into two corner points: .
[0145] Combined into the reversed box: .
[0146] Multi-frame sequence conversion: To convert frames Sequence of optical flow affine matrices over multiple frames , then multiple frames Defined as nested calls to single frames in frame order : .
[0147] Multi-frame reverse transformation: If you want to From Frame reverse push back Frame, then multiple frames Defined as nested calls to single frames in reverse order : .
[0148] Interpolation steps: .
[0149] in, ,and .
[0150] Will Marked as interpolation results, they are added to the trajectory to form a complete and continuous trajectory.
[0151] After the above optical flow-based trajectory interpolation and invalid target elimination, the interpolated target library sequence is obtained, which is recorded as: .
[0152] in, Indicates in In the frame, the set of valid targets after invalid targets are eliminated and trajectory interpolation is completed.
[0153] To eliminate stationary objects, in step 105, the trajectories of the objects in the first object library sequence are first compensated for camera motion using an optical flow matrix set and uniformly mapped to a reference coordinate system. For each trajectory, a sliding window is used to calculate the mean position of the objects in the front and back windows within the trajectory and compare it with the position of the object in the current frame. If the distance between the position of the object in the current frame and the mean position of the objects in the front and back windows is less than a set distance, and the cumulative displacement between frames within the sliding window is less than a set motion threshold, the object is marked as stationary in the current frame image. Finally, the marked objects are deleted.
[0154] First, using the aforementioned optical flow affine matrix set , perform global motion compensation on the detection frames of all target trajectories in each frame, and uniformly map them to the reference coordinate system of the last frame of the video to accurately estimate the actual moving distance of the target.
[0155] Then, for each trajectory, a sliding window length , calculate the mean position of the front and back windows within the trajectory and compare it with the current frame position: .
[0156] in, 、 Represent the mean boxes of the front and back windows respectively, Indicates the center point distance.
[0157] If the distance between the current frame and the mean of the windows before and after it is less than the threshold , and the cumulative displacement between frames in the sliding window is less than the minimum motion threshold , the frame is determined to be in a static state and marked as to be eliminated.
[0158] Traverse all trajectories and delete the detection frames that are continuously marked as static, retaining the remaining moving trajectories. If a trajectory is completely eliminated, it is considered an invalid trajectory. Finally, the second target library sequence after the static targets are eliminated is obtained: .
[0159] in, Indicates the The valid target set after the frame is interpolated and completed and static targets are eliminated.
[0160] Through this stationary target elimination step, the false trajectories introduced by detection noise or camera shake are significantly reduced, ensuring that the output target trajectory is more physically reasonable and practically usable.
[0161] After the above steps, the target library sequence is finally obtained after removing invalid targets, trajectory interpolation and completion, and eliminating stationary targets: .
[0162] In order to facilitate downstream analysis and standard multi-target tracking evaluation, the target library needs to be exported as a trajectory result file in a unified format. The output format is: .
[0163] In summary, the multi-moving target tracking method for remote sensing video provided in the embodiment of the present application includes a training phase and an inference phase.
[0164] For the training phase, first, a training subsequence construction strategy is provided to ensure the consistency of the sequence length of the samples. By randomly sampling the frame intervals, the modeling ability of target motion patterns in different intervals and different spatiotemporal scenarios is effectively improved, thereby enhancing the diversity and integrity of the training data. Then, spatiotemporal feature extraction of the time-series frame sequence is performed, and the ResNet50-3D network is used to perform spatiotemporal joint modeling of the training subsequences, which can fully capture the continuity and spatial context information of small targets in the temporal dimension, thereby improving the detectability of small targets and the semantic richness of feature expression. For the generation of Resnet50-3D backbone network pre-training parameters, without introducing other datasets except COCO and ImageNet, the pre-training of ResNet50-3D is effectively constructed, which significantly improves the rationality of network weight initialization, making the entire training process more stable and converging more efficiently. In addition, static frame feature enhancement based on motion regions uses motion region masks to perform weighted enhancement on static frame features, highlighting the moving target area while suppressing the characteristic response of static background or pseudo-targets. This effectively improves the ability to characterize real moving targets and reduces the false alarm rate caused by static targets. Next, a dual-branch temporal feature enhancement and fusion is provided, combining deep separable convolution with a channel attention mechanism to enhance the network's local spatial modeling capabilities and channel adaptability, achieving efficient interaction and deep alignment of static frame features with sequential frame spatiotemporal features. The fused features combine the spatial features of static frames and the temporal features of sequential frames, allowing for a more targeted description of weak moving targets in remote sensing videos.
[0165] For the inference stage, the Kalman filter is first initialized for the state components that cannot be directly observed, such as the center point velocity and area change rate. A large initial variance is given in the P matrix to reflect its initial uncertainty. At the same time, a small process noise (0.1) is set for the center point velocity and area velocity in the Q matrix. 2 , 0.01 2), which reflects the smooth and slow characteristics of small target movement and deformation, thereby improving the physical rationality of the tracking process. Then, the predicted state is corrected based on the optical flow matrix estimation. Given that the area and aspect ratio of small targets are more sensitive to detection noise, when correcting the optical flow affine matrix, only the center point position and velocity are updated, keeping the area and aspect ratio parameters unchanged, thereby improving the robustness and stability of the target state update under camera motion. Next, trajectory post-processing based on the optical flow matrix. For interpolation reconstruction, different from that based only on linear interpolation, the embodiment of the present application introduces the inter-frame optical flow matrix in the trajectory interpolation process, incorporating the camera's motion information into the interpolation calculation, ensuring that the completion result is not only consistent with the observed bounding box, but also more consistent with the actual scene motion, and the interpolation trajectory is more accurate and coherent. For stationary target elimination, in order to effectively reduce false detections caused by stationary targets, a motion amplitude determination strategy based on a local time window is designed, and the camera motion is compensated in combination with the optical flow matrix to eliminate non-real moving targets, significantly reducing the false target rate in tracking, and further improving the accuracy and reliability of overall tracking.
[0166] An embodiment of the present application further provides a computer-readable storage medium, in which a program is stored. The program can be loaded by a processor and execute any one of the methods for tracking multiple moving targets in remote sensing videos according to the embodiments of the present application.
[0167] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer program. When all or part of the functions in the above embodiments are implemented by computer program, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above functions can be implemented. In addition, when all or part of the functions in the above embodiments are implemented by computer program, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and saved in the memory of the local device by downloading or copying, or the system of the local device is updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be implemented.
[0168] The above specific examples are used to illustrate the present application, which is only used to help understand the present application and is not intended to limit the present application. For those skilled in the art of the present application, based on the concept of the present application, they can also make some simple deductions, modifications or substitutions.
Claims
1. A method for tracking multiple moving targets in remote sensing video, characterized in that: include: Acquire a remote sensing video to be processed, and establish multiple detection subsequences based on multiple frames of images included in the remote sensing video; Extracting static features and dynamic features of each detection subsequence respectively to obtain spatiotemporal fusion features of the detection subsequence, and obtaining a detection frame of each frame of the image based on the spatiotemporal fusion features, where each detection frame corresponds to an object; Building a target library corresponding to a starting frame in the remote sensing video based on an initial detection frame set of the starting frame, sequentially obtaining predicted states of multiple frames of the image through Kalman filtering based on the target library of the previous frame of each frame of the image after the starting frame, and calculating an optical flow matrix of the two adjacent frames of the image based on background feature points of the two adjacent frames of the image; Based on the optical flow matrix, the predicted state of each frame of the image is corrected respectively, and a first target library sequence consisting of a target library of each frame of the image is obtained by matching with the detection frame of each frame of the image, where the target library of each frame of the image includes a plurality of targets; Based on the optical flow matrix, a post-processing operation is performed on the first object library sequence to obtain a second object library sequence, wherein the post-processing operation includes interpolation reconstruction and stationary object elimination.
2. The multi-moving target tracking method according to claim 1, characterized in that: Each of the detection subsequences includes a static frame and a time sequence frame, and establishing multiple detection subsequences based on the multiple frames of images included in the remote sensing video includes: Obtaining a preset length of the detection subsequence, where the preset length is a first number of the images included in the detection subsequence; sequentially taking each frame of the image as a static frame, taking a second number of images preceding the static frame as sequential frames, and constructing a detection subsequence corresponding to each frame of the image based on the static frames and the sequential frames, wherein the second number is the first number minus one; In the process of constructing the detection subsequence, if it is detected that a third number of images located before the static frame in the remote sensing video is less than the second number, a fourth number of the static frames are copied and used as the timing frames to fill the detection subsequence, wherein the sum of the third number and the fourth number is the second number.
3. The multi-moving target tracking method according to claim 2, characterized in that: The extracting the static features and the dynamic features of each detection subsequence respectively to obtain the spatiotemporal fusion features of the detection subsequence, and obtaining the detection frame of each frame of the image based on the spatiotemporal fusion features, includes: For each detection subsequence, input the static frames in the detection subsequence into the deep aggregation neural network DLA-34 to obtain the static features of the detection subsequence; Aligning the time series frame with the static frame using an ORB algorithm, and performing background modeling and foreground separation on the time series frame and the static frame using a Gaussian mixture model to obtain a motion foreground mask of the static frame; Performing feature fusion on the motion foreground mask of the static frame and the static features of the static frame to obtain the static enhanced features of the detection subsequence; Inputting the detection subsequence into a three-dimensional convolutional network ResNet50-3D to obtain dynamic features of the detection subsequence; The static enhancement feature and the dynamic feature are fused based on the attention mechanism of the convolutional neural network to obtain the spatiotemporal fusion feature of the detection subsequence; The spatiotemporal fusion features are input into a multi-scale feature aggregation module and an object detection module to obtain a detection frame of the static frame.
4. The method for tracking multiple moving targets according to claim 3, wherein: The deep aggregation neural network DLA-34 is initialized using weights pre-trained on COCO and ImageNet; the three-dimensional convolutional network ResNet50-3D is initialized by copying the convolution parameters of the two-dimensional convolutional network ResNet50 pre-trained on COCO and ImageNet a set number of times along the time dimension and dividing the result by the set number of times as the initial weights.
5. The method for tracking multiple moving targets according to claim 3, wherein: The convolutional neural network includes a multi-layer neural network. The attention mechanism based on the convolutional neural network fuses the static enhancement features and the dynamic features to obtain the spatiotemporal fusion features of the detection subsequence, including: For each layer of the neural network, the static enhancement features and the dynamic features are concatenated according to the channel dimension to obtain joint features; Extracting local context information of the joint feature through depthwise separable convolution, and performing batch normalization after adding residuals of the local context information and the joint feature to obtain normalized features; Performing global pooling on the normalized features to generate a channel description vector, converting the channel description vector into a channel weight through a fully connected layer and an activation operation, and multiplying the normalized features channel by channel with the channel weight to obtain a channel weighted feature; Performing nonlinear transformation on the channel weighted features through a multi-layer perceptron and then performing residual addition with the channel weighted features to obtain the fusion features of the neural network in each layer; The fusion features of each layer of the neural network are summarized to obtain the spatiotemporal fusion features of the detection subsequence.
6. The method for tracking multiple moving targets according to claim 1, wherein: Building a target library corresponding to a starting frame in the remote sensing video based on an initial detection frame set of the starting frame, and sequentially obtaining predicted states of multiple frames of the image through Kalman filtering based on the target library of the previous frame of each frame of the image after the starting frame, including: Obtaining an initial detection frame set of a starting frame in the remote sensing video, wherein the initial detection frame set includes a plurality of initial detection frames, each initial detection frame corresponding to an initial target; Assigning a target identifier and initializing a Kalman filter to the initial target corresponding to each initial detection frame, and converting the initial detection frame into an initial state vector to construct a target library corresponding to the starting frame; Kalman prediction is performed sequentially based on the target library of the previous frame of the image of each frame after the starting frame to obtain a first predicted state vector and a first predicted state covariance of each frame of the image.
7. The method for tracking multiple moving targets according to claim 6, wherein: The method of correcting the predicted state of each frame of the image based on the optical flow matrix and obtaining a first target library sequence consisting of a target library of each frame of the image by matching with the detection frame of each frame of the image comprises: Obtaining an optical flow matrix of a two-dimensional affine transformation by estimating a sparse optical flow between the image of the current frame and the image of the previous frame; Applying the optical flow matrix to the position and velocity components of the first predicted state vector of the image in the current frame to form a modified second predicted state vector; Expanding the optical flow matrix into a multi-dimensional expanded matrix; updating the first prediction covariance matrix based on the extended matrix to obtain a modified second prediction state covariance matrix; Extracting the position state component from the first predicted state vector to form a predicted observation vector; Matching the predicted observation vector of the image of the current frame with the actual observation vector corresponding to the detection frame of the image of the current frame; If the predicted observation vector matches the actual observation vector, updating the first predicted state vector and the first predicted covariance matrix corresponding to the predicted observation vector through the detection frame corresponding to the actual observation vector to obtain a second predicted state vector and a second predicted covariance matrix of the detection frame; If the predicted observation vector does not match the actual observation vector, the detection box corresponding to the unmatched actual observation vector is initialized as a new target, and an invalid mark is added to the detection box corresponding to the unmatched predicted observation vector; Based on the second prediction state vector, the second prediction covariance matrix and the new target, a target library for each frame of the image is constructed to obtain the first target library sequence.
8. The method for tracking multiple moving targets according to any one of claims 1 to 7, characterized in that: The post-processing operation on the first target library sequence based on the optical flow matrix to obtain a second target library sequence includes: Obtaining the optical flow matrix set according to the optical flow matrix between the two adjacent frames of the image; Deleting the targets with invalid markers in the first target library sequence; For each target identifier in the first target library sequence, if the target identifier has a discontinuous trajectory, interpolating the missing frame sequence of the target identifier based on the optical flow matrix; The frame before the missing frame sequence is set as the start frame, and the frame after the missing frame sequence is set as the end frame; Obtaining a coordinate transformation matrix of the start frame based on an optical flow matrix, and aligning the start frame to a coordinate system of the end frame; interpolating the position of the target in the start frame and the end frame after the alignment by linear interpolation, wherein the interpolated target position is the position of the target in the coordinate system of the end frame; A coordinate transformation matrix of all missing frames in the missing frame sequence is obtained based on the optical flow matrix, and an interpolation target position of the missing frame sequence is transformed into a coordinate system of each missing frame.
9. The method for tracking multiple moving targets according to any one of claims 1 to 7, characterized in that: The post-processing operation on the first target library sequence based on the optical flow matrix to obtain a second target library sequence includes: Obtaining the optical flow matrix set according to the optical flow matrix between the two adjacent frames of the image; Performing camera motion compensation on the trajectories of the targets in the first target library sequence using the optical flow matrix set, and uniformly mapping them to a reference coordinate system; For each of the trajectories, a sliding window is used to calculate the average position of the target in the front and rear windows within the trajectories, and the average position is compared with the position of the target in the current frame; If the distance between the position of the target in the current frame and the average position of the targets in the front and back windows is less than a set distance, and the cumulative displacement between frames in the sliding window is less than a set movement threshold, a stationary mark is added to the image in the current frame; The target to which the stationary mark is added is deleted.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which can be loaded by a processor and executed by the method for tracking multiple moving targets in remote sensing videos according to any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-target tracking method based on optical flow method and Kalman filtering
CN106803265A
Multi-target tracking method based on multi-model fusion and data association
CN107292911A
Multi-target tracking method, system and device based on optical flow and Kalman filtering
CN110415277A
Multi-target tracking method and device and storage medium
CN111445501A
Target tracking method and system under dynamic background based on correlation filtering framework
CN115170621A
Cited By
Bird detection method based on video stream
CN120726306A
A Bird Detection Method Based on Video Stream
CN120726306B
Unmanned aerial vehicle video multi-target tracking method based on motion decoupling and related device
CN121438170A
Detection tracking algorithm for weak texture multiple targets under monocular camera and application system
CN122090091A