Ground multi-target tracking algorithm for aerial video of unmanned aerial vehicle
By building a target detection network and a target tracking network, the context attention mechanism CAM module is introduced, which solves the stability and accuracy of multi-objective tracking in aerial video of aerial video of aerial video of aerial video of aerial video of aerial video of a manned aerial video of the drone, and achieves continuous and stable tracking of multi-objectives on the ground.
Patent Information
- Application Number
- CN202510585114.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-19
AI Technical Summary
The existing target tracking algorithms are difficult to track multiple targets stably and accurately in drone aerial video scenarios, especially when occlusion and aggregation, and cannot meet the requirements of high stability and high accuracy.
Build a target detection network and a target tracking network, introduce a context attention mechanism CAM module, predict the target state through multi-scale feature fusion and Kalman filters, and achieve continuous and stable tracking of multiple targets on the ground.
It improves the precise recognition capability of the target tracking algorithm in occlusion and aggregation situations, reduces target loss, and achieves high stability and high accuracy multi-target tracking in drone aerial videos.
Smart Images

Figure CN120510408A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a multi-target tracking algorithm and belongs to the field of computer vision. Background Art
[0002] In the field of security monitoring, drone-based multi-target tracking can be used to monitor the activity trajectories of multiple targets such as people and vehicles on the ground in real time; in traffic management, drone-based multi-target tracking can help monitor traffic flow and identify illegal vehicles; in security neighborhoods, drone-based multi-target tracking can detect the number of criminals and track their movement trajectories; however, existing target tracking algorithms face many challenges in drone aerial video scenes; targets in drone aerial videos are often prone to occlusion and aggregation; traditional algorithms find it difficult to stably and accurately track multiple targets under these complex conditions, resulting in low tracking accuracy and easy target loss, which cannot meet the requirements of high stability and high accuracy multi-target tracking in practical applications; therefore, a drone aerial video-based multi-target tracking algorithm is proposed to achieve continuous and stable tracking of multiple targets on the ground.
[0003] CN109615641A discloses a multi-target pedestrian tracking system and tracking method based on the KCF algorithm. The system includes an initialization module, a single target tracking KCF module, a tracking and detection matching module, a target removal module, a printing module and a target addition module; the initialization module is used to initialize all variables; the single target tracking KCF module is used to track a single target; the tracking and detection matching module matches the tracking result of each target with the detected target in the picture; the target removal module is used to determine whether the target has left the picture; the printing module is used to draw the pedestrian's border and its ID information on the picture based on the matching result of the tracking and detection matching module; the target addition module is used to determine whether the detected target is a newly appeared target; the present invention designs a multi-target tracking system framework based on the single target tracking algorithm KCF, which provides the motion trajectory and ID information of each target in real time.
[0004] CN111860532A proposes an adaptive target tracking method based on two complementary tracking algorithms. First, the image is sharpened using a Laplace filter module for image preprocessing, which enhances image edge information and obtains better directional gradient histogram features. Then, when kernel correlation filtering is used for target tracking in situations such as target occlusion, rapid motion deformation, and "gradually changing" targets, model drift is likely to occur, making it difficult to track the target subsequently. Therefore, it is proposed to use the intersection-over-union ratio of the predicted box areas of the two complementary tracking algorithms to adaptively control the update of the position filter. Comparative experiments with other algorithms on a standard test set demonstrate that the present invention has good tracking accuracy.
[0005] CN107562837A provides a road network-based maneuvering target tracking algorithm. With the help of a priori road information database, a tracking algorithm model adaptation strategy using road information is proposed. A variable structure multi-model method is used to achieve ground multi-maneuvering target tracking. This can improve the state estimation accuracy of maneuvering target tracking and reduce the target loss rate. At the same time, it avoids the computational burden brought by the use of too many models in a fixed multi-model algorithm, greatly reducing the running time. This solution has practical value in the problem of ground target tracking.
[0006] CN109284677A discloses a Bayesian filtering target tracking algorithm. In the first step, the method of the present invention obtains a one-step prediction estimate of the target state at the next moment through a motion model based on the optimal estimate of the target state at the k-1 moment. In the second step, after obtaining the observation value of the target at the k moment from the radar observation station, the distance information and angle information of the target relative to the radar are converted into the Cartesian coordinate position information of the target using a random variable fixed point sampling nonlinear transformation method. In the third step, the two parts of information, namely the one-step prediction prior information of the target state and the radar observation reverse estimation likelihood function, are multiplied and fused using the probability likelihood product rule of the present invention to finally obtain a posterior estimate of the target state at the k moment. After storing the target state, the time is updated and the next round of iteration is entered. The present invention has the characteristics of higher accuracy, better robustness and simpler algorithm structure, and has high practical value in radar, multi-sensor, maneuver and multi-target tracking.
[0007] Among the above invention patents, CN109615641A builds a multi-target pedestrian tracking system based on the KCF algorithm, providing target motion trajectory and ID information; CN111860532A uses two complementary tracking algorithms to adaptively control position filter updates to improve tracking accuracy; CN107562837A uses a variable structure multi-model method with the help of road information to improve the accuracy of maneuverable target tracking and reduce the computational burden; CN109284677A uses Bayesian filtering to achieve high-precision and high-robustness target tracking, and the present invention focuses on ground multi-target tracking in drone aerial videos, using airborne cameras to acquire data, and by constructing a target detection network, a contextual attention mechanism CAM module in the target tracking network, and a target tracking network, it solves the problem of accurate distinction and stable tracking of ground targets under occlusion and aggregation, and returns the video with target serial number and tracking frame to the ground computer. Summary of the Invention
[0008] The present invention provides a multi-target tracking algorithm for drone aerial video, including a target detection network, a context attention mechanism CAM module in the network, and target tracking network target tracking; First, the ground aerial video is obtained as sample data through the drone's onboard camera; data preprocessing and data enhancement are performed based on the collected video information, and the training set, test set and validation set are divided; a target detection network is constructed to detect ground targets; a contextual attention mechanism CAM module is constructed in the target tracking network, and the contextual relationship between the target pixel and the surrounding ring pixels is obtained through multi-scale feature fusion and weight calculation, thereby enhancing the expressive ability of the target feature and improving the adaptability to target occlusion and target aggregation during the target tracking process; a target tracking network is constructed, and the target tracking network uses the output results of the target detection network to predict the target state through Kalman filtering, and uses the Hungarian algorithm to combine appearance feature matching with geometric position for optimal matching, so as to continuously track the motion trajectory of the target in the video sequence, thereby achieving accurate distinction of multiple ground targets and continuous, stable and accurate tracking; by inputting the aerial video into the multi-target tracking algorithm, the target detection network and the target tracking network process it, and finally returns the video with the target serial number and tracking frame to the ground computer;
[0009] This application provides a multi-target tracking algorithm for drone aerial video, including: 1. Obtain sample data: Use a drone-mounted camera to shoot ground videos, collect a large amount of training sample data, convert the videos into frame images, and finally label each sample data to complete the training dataset; 2. Perform data preprocessing and enhancement on sample data: perform operations such as 2x and 4x zooming, 90-degree, 180-degree, and 270-degree rotation, translation, and cropping on the image; 3. Build the target detection network: Build the backbone feature extraction network of the target detection network, perform preliminary feature extraction on the received video images, enhance feature extraction, predict and output accurate target location and category information; The backbone feature extraction network first slices and concatenates the input image, downsampling it to P1 / 2. It then continuously extracts features, gradually reducing the feature map resolution and increasing the number of channels, extracting features at P3 / 8, P4 / 16, and P5 / 32 scales. Finally, the SPP module performs pooling operations on the features at different scales to enhance their expressiveness. After the convolution operation, an activation function is used to increase the nonlinearity of the network, and the feature map is downsampled through the maximum pooling layer to reduce the number of parameters while retaining important features and improving the robustness of the model. The weighted bidirectional feature pyramid network adjusts the number of channels of the feature map through convolution operations, restores the feature maps of different resolutions to the appropriate size through upsampling operations, and then splices the feature maps of different layers in the channel dimension through splicing to fuse multi-scale feature information; The prediction head receives the multi-scale feature map output by the enhanced feature extraction network, and predicts the location, size, and category of the target based on the predefined anchor boxes and number of categories, ultimately outputting accurate target location and category information. 4. Constructing a contextual attention mechanism (CAM) module in the target tracking network: Build a contextual attention mechanism (CAM) module in the target tracking network. The CAM module consists of a channel attention submodule, a spatial attention submodule, and a feature fusion layer to extract appearance features of targets at different scales. Among them, the channel attention submodule compresses the feature map in the spatial dimension through global average pooling to obtain the global features of the channel. Then, through the fully connected layer and activation function, corresponding weights are generated for each channel to highlight the feature information of important channels. The spatial attention submodule performs average pooling and maximum pooling on the input feature map in the channel dimension, concatenates the two feature maps in the channel dimension, and then passes them through a convolutional layer and an activation function to generate a spatial attention map, thereby enhancing the features of key spatial locations. The feature fusion layer fuses the features processed by the channel attention submodule and the spatial attention submodule. It combines the output features of the two submodules through element-by-element multiplication, so that the features have both channel and spatial attention information, further improving the feature expression ability and enabling the target tracking network to accurately identify the same target. 5. Build a target tracking network: Build a target tracking network to extract target appearance features, perform feature cascade matching, and predict the target detected by the target detection network, and output accurate target location and class ID information; Among them, the appearance feature extraction network extracts features from the targets detected by the target detection network, converting the detected target image into a feature vector of fixed dimension as the appearance feature of the target; this feature is used for subsequent matching operations; Among them, feature cascade matching first distinguishes confirmed and unconfirmed trajectories, uses appearance features to perform cascade matching on confirmed trajectories, calculates the cost matrix and performs gating operations; for unmatched trajectories and unconfirmed trajectories, IOU is then used for matching; through these two matching steps, the detection results are associated with the trajectories; Among them, the accurate target location and class ID information is predicted and output. The Kalman filter is used to update the status of the matching trajectory, and the unmatched trajectory is marked or initialized accordingly, and finally the information containing the target location and class ID is output; 6. The aerial video captured by the drone is fed into the target detection network to obtain the location and category of ground targets. Finally, the target tracking network outputs the target's location and sequence number information, thereby achieving continuous and stable tracking of targets of different scales in the aerial video. Beneficial effects
[0010] This application uses the drone's onboard camera to collect real-time data from the observation area, and performs data preprocessing and data enhancement on the data to improve the efficiency of the algorithm; in addition, the data is first input into the target detection network to extract the target location, and then the target image is input into the target tracking network to obtain the target's precise location and target serial number; in addition, by introducing the contextual attention mechanism CAM module into the appearance feature extraction network in the target tracking network, the target's perception ability of targets of different scales is improved, the clustered targets are accurately identified, and the target tracking network performance is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This application provides a schematic diagram of the process of multi-target tracking on the ground by drone aerial video; Figure 2 This is a schematic diagram of the target detection network provided by this application; Figure 3 This is a schematic diagram of the contextual attention mechanism CAM module in the target tracking network provided by this application; Figure 4 It is the target tracking network flow chart provided by this application; Figure 5 This is the result of multi-target tracking using the algorithm proposed in this application on aerial ground video. DETAILED DESCRIPTION
[0012] In order to facilitate a more straightforward further understanding and use of the present application, the present application is clearly and completely described below in conjunction with the accompanying drawings and embodiments to enhance the understanding of the present application by technical personnel and users; this embodiment is only an illustrative embodiment of part of the present application. Any other embodiments of the present application obtained by an individual or unit without creative work shall fall within the scope of the rights of the present application and shall be protected.
[0013] Figure 1 A schematic diagram of the multi-target tracking process of drone aerial video is provided for this application.
[0014] Step S201 first saves the real-time ground image captured by the drone's onboard camera to the server and pre-processes the image on the server. The specific steps include: (1) Histogram equalization: By redistributing the pixel values of the image, the histogram of the image becomes more uniform, thereby enhancing the contrast and details of the image; (2) Sharpening: Through the sharpening algorithm, the edges and details of the image can be enhanced; (3) Image enhancement filter: By applying a specific enhancement filter, such as an edge enhancement filter, the edges or details of the image can be highlighted, thereby improving the clarity and detail visibility of the image; (4) Zoom the image by 2 or 4 times; (5) Perform operations such as rotating 90 degrees, 180 degrees, and 270 degrees, translating, and cropping the image.
[0015] Step S202 is to build a target detection network; Figure 2 is a schematic diagram of a target detection network provided in this application; Among them, the network consists of a backbone feature extraction network, an enhanced bidirectional feature extraction network, and a prediction head; The backbone feature extraction network extracts basic features from the input aerial video. Through a series of convolution operations and downsampling processes, the feature map size is gradually reduced and the number of channels is increased. The input features are then sliced and spliced, and then downsampled using convolutional layers, such as from P1 / 2 to P5 / 32, to extract feature information at different scales. At the same time, the C3 module is used to enhance feature propagation and reuse, and the SPP module is used to extract multi-scale features through pooling kernels of different sizes to enrich feature diversity. Among them, the weighted bidirectional feature pyramid network further processes the feature maps output by the backbone feature extraction network, fusing feature information at different levels through upsampling and feature splicing operations. For example, high-level features are upsampled and spliced with low-level features, allowing the network to combine feature information at different scales, enhance the feature expression ability, and better capture the characteristics of objects of different sizes. Among them, the prediction head uses preset anchor points based on the fused feature map to perform category prediction, position positioning and bounding box regression of targets in aerial videos. It predicts information of targets of different sizes based on feature maps of different scales, performs detection at three scales: P3 / 8, P4 / 16, and P5 / 32, and finally completes target detection and sends the target location to the target tracking network.
[0016] Step S203 is to construct a contextual attention mechanism CAM module in the target tracking network, which consists of an adaptive average pooling layer, a convolution layer, a bottleneck convolution layer, and a weight generation network; Figure 3 This is a schematic diagram of the contextual attention mechanism CAM module in the target tracking network provided by this application; The adaptive average pooling layer is used to generate multi-scale features. Through adaptive average pooling operations of different sizes, the input feature map is downsampled to different degrees, thereby generating feature maps with different spatial resolutions. These feature maps of different scales can capture the information of objects of different sizes in the input image, allowing the appearance feature extraction network to accurately identify and process objects of different scales. Among them, the convolution layer is used to extract and transform the feature map after adaptive average pooling. By defining the convolution layer, the feature map after adaptive average pooling is convolved to extract more representative features. The 1x1 convolution operation can effectively reduce or increase the number of channels of the feature map, while also fusing information between different channels. The bottleneck convolution layer is used to fuse multi-scale features with original features. This part concatenates the multi-scale features and original features in the channel dimension and performs a convolution operation. The number of channels in the concatenated feature map is adjusted to the specified number of output channels. This allows the fusion of feature information at different scales with the original feature information, resulting in a more comprehensive and richer feature representation. Among them, the weight generation network is used to generate weights of features at different scales; the difference features are obtained by subtracting the original features from the feature maps of each scale, and then the weight of each scale feature is obtained through the activation function; the weight of each scale feature can reflect the importance of each scale feature to the final result, so that the model can adaptively adjust the contribution of features at different scales according to the scale of the target.
[0017] Step S204 is to build a target tracking network; Figure 4 It is a flow chart of the target tracking algorithm provided in this application; Among them, the network consists of target appearance feature extraction network, feature cascade matching part, and prediction tracking result part; The target appearance feature extraction network is responsible for extracting the target's appearance features from the input image. When extracting features, the network scales the input image pixel values to between 0 and 1, resizes them to (64, 128), converts them into tensors, and normalizes them. Then, the preprocessed image is converted into a fixed-dimensional feature vector that represents the target's appearance features. The feature cascade matching part is used to associate detected targets with existing trajectories. First, the feature cascade matching divides the trajectories into confirmed and unconfirmed states, and then performs cascade matching on the confirmed trajectories. During this process, the cost matrix is obtained by calculating the cosine distance between the detection box features obtained by the target detection network and the trajectory features, and the cost matrix is corrected using a gating mechanism to ensure the accuracy of the matching. For targets that are not matched in the cascade matching, the intersection-over-union matching is performed. During the matching process, the distance between features is calculated using the cosine distance. At the same time, the Hungarian algorithm is used to complete the feature matching. Among them, the prediction and tracking result part uses the Kalman filter to predict and update the target position; in the target tracking network process, each trajectory has a corresponding Kalman filter, and completes the prediction and update of the target state; in the prediction stage, based on the mean and covariance of the target at the previous moment, the state transfer matrix and the process noise matrix are used to predict the state of the target at the current moment; in the update stage, combined with the target information detected by the target detection network, the predicted state is corrected by the Kalman gain to obtain a more accurate target position; at the same time, according to the matching results, a unique ID is assigned to each trajectory, and finally the accurate target position and target serial number information are output.
[0018] Step S205 is the video image with the target category and detection frame finally obtained by the system.
[0019] The drone aerial video-to-ground multi-target tracking algorithm process provided in this application includes: using the drone's onboard camera to collect aerial video of the ground, inputting the collected video image into the target detection network to detect ground target information, inputting the aerial video into the context attention mechanism CAM module of the appearance feature extraction network to obtain fine target appearance features of the target tracking network, inputting the video into the target tracking network to obtain the target location information and serial number, and finally transmitting the video image with the target location information and serial number information back to the ground computer.
[0020] The following combination Figure 1 The present application is further described in detail with reference to the accompanying drawings and implementation examples; the specific implementation described herein is only used to explain the relevant application, rather than to limit the application; it should also be noted that, for ease of description, only the parts related to the present application are shown in the accompanying drawings.
[0021] The server is equipped with two NVIDIA V100 graphics cards. The drone aerial video ground multi-target tracking algorithm processes the video image in two stages. The target detection network of this application consists of a backbone network, an enhanced feature extraction network including an enhanced bidirectional feature pyramid, and a prediction head. Figure 2 As shown, this is a schematic diagram of the target tracking network of the present application; the network consists of an appearance feature extraction network (301), a feature cascade matching part (302), and a prediction tracking result part (303); in order to enable the target tracking network to accurately identify the same target, a contextual attention mechanism CAM module is introduced into the appearance feature extraction network in the target tracking network; and then the target position and serial number information are obtained through the target tracking network.
[0022] The appearance feature extraction network in the target tracking network is used to extract the appearance features of the target. The appearance feature extraction network loads a pre-trained model and pre-processes the input image, including normalization and resizing, to ultimately obtain a fixed-dimensional feature vector. The feature matching part of the target tracking network can associate the detected target with the existing tracking trajectory; cascade matching is performed by calculating the distance between target features; The prediction and tracking results part of the target tracking network uses a Kalman filter to predict and update the target state. The Kalman filter predicts and updates the target state through the state transfer equation and observation equation. Based on the mean and covariance of the target at the previous moment, the state transfer matrix and process noise matrix are used to predict the target state at the current moment. Secondly, combined with the target information detected by the target detection network, the predicted state is corrected through the Kalman gain to obtain a more accurate target position. At the same time, based on the matching results, a unique ID is assigned to each trajectory, and the precise target position and target sequence information are finally output. The contextual attention mechanism (CAM) module in the appearance feature extraction network in the target tracking network first receives the preliminary feature map output by the first convolution layer of the original appearance feature extraction network, and then divides the feature map into four processing paths. Each path uses an average pooling operation of different sizes to divide the preliminary feature map into four different grid blocks of 1×1, 2×2, 3×3, and 6×6. Since the original network uses a fixed-scale convolution kernel, the receptive field of the network is the same. To solve the above problem, the CAN network uses spatial pyramid pooling to extract multi-scale context information in the image. The calculation formula is shown in Equation 2. s j =U bi (F j (P ave (f v ,j),θ j )) (2) Where, F sa j Represents convolution operation and activation function operation, θ j sa Represents the weight, and then w j With s jThe four feature maps are multiplied and then concatenated to obtain the final feature map containing information of different scales and contexts. This module enables the appearance feature extraction network to extract features at multiple scales and learn important features in the image, thereby solving the problem of difficult recognition of targets with diverse scales in drone aerial images, thereby improving the accuracy of the target tracking network in identifying the same target and reducing the occurrence of target number switching problems during target tracking.
[0023] To demonstrate the effectiveness of our proposed algorithm for tracking multiple ground targets in aerial video and to verify the effectiveness of our specific improvement measures, we conducted a comparative experiment using our proposed algorithm with other mainstream target tracking algorithms in aerial video of multiple ground targets. The overall results are shown in Table 1. Table 1 Target tracking results of different target tracking algorithms The experimental results are shown in Table 1. The algorithm proposed in this application is tested and compared with other common multi-target tracking algorithms. The experimental results are shown in Table 1. As can be seen from Table 1, the algorithm in this chapter has better processing speed and better performance than other one-to-one algorithms. There is no identity switching problem in the target tracking results after cross-occlusion of the ground target. The MOTA is greatly improved, reaching the highest value among similar algorithms: 86.59%.
[0024] In order to verify the effectiveness of the contextual attention module (CAM) used in this application for identifying objects of different scales in aerial videos, a comparative experiment was conducted on various appearance feature extraction networks; some images from the Market-1501 dataset were used as experimental objects; the experimental results are shown in Table 2. Table 2 Experimental results of appearance feature extraction network The appearance feature extraction network with the CAM module has a 5.459% higher classification accuracy on the validation set than the original network. This is due to the CAM module's ability to extract features from targets of different scales in multiple receptive fields, and its ability to adaptively learn important features in the image based on the weight distribution results, thereby making re-identification of the same target more accurate.
[0025] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used; those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept; for example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A multi-target tracking algorithm for drone aerial video, characterized by: The specific steps include: S1: Obtain ground aerial video as sample data; S2: Perform data preprocessing and enhancement on the sample data, and divide it into training set, test set and validation set; S3: Build a target detection network to detect ground targets; S4: Construct the contextual attention mechanism CAM module in the target tracking network; S5: Build target tracking network; S6: Input the aerial video to the drone aerial video ground multi-target tracking algorithm, and after being processed by the target detection network and the target tracking network, the video with the target sequence number and tracking frame is finally returned to the ground computer.
2. The multi-target tracking algorithm for drone aerial video according to claim 1 is characterized in that: The step S2 comprises: 1) Preprocess the video image; 2) Perform data augmentation on dataset images; 3) Randomly zoom in on the dataset images; 4) Divide the preprocessed dataset into training set, validation set and test set according to the set ratio.
3. The multi-target tracking algorithm for drone aerial video according to claim 1 is characterized in that: Wherein said step S3 comprises: 1) Use the backbone network to perform preliminary feature extraction on the input aerial video; 2) Using the weighted bidirectional feature pyramid network to enhance feature extraction of the initially extracted features; 3) The features extracted from the enhanced features are input into the prediction head, and the target category and location information are finally output.
4. The method of tracking multiple targets on the ground using drone aerial video according to claim 1, characterized in that: Wherein said step S4 comprises: 1) Introducing the contextual attention mechanism CAM module into the appearance feature extraction network in the target tracking network; 2) Using the contextual attention mechanism (CAM) module, we extract features from multiple receptive fields of the target feature map to obtain feature maps of different scales. 3) Using feature maps of different scales to calculate the weights of feature maps of different scales; 4) Combine the feature map with the weights and use it in the appearance feature extraction network to distinguish different targets on the ground.
5. The multi-target tracking algorithm for drone aerial video according to claim 1 is characterized in that: Wherein said step S5 comprises: 1) For the targets detected by the target detection network, the appearance feature extraction network is used to extract the target's appearance features to distinguish and track different targets; 2) For the first-time target, the target tracking network creates a new trajectory for the new target. For the trajectory of the existing target, the target tracking network uses the Kalman filter to predict the position of the trajectory of the existing target in the next frame. 3) Use the Hungarian algorithm to optimally match the predicted position predicted by the Kalman filter with the actual target detection results and update the trajectory status; 4) Output the tracking results of each target based on the target and trajectory matching results.
6. The multi-target tracking algorithm for drone aerial video according to claim 1 is characterized in that: The data enhancement in claim 2 includes: 1) Rotation and flipping: By randomly rotating and flipping the image (horizontally or vertically), the model can learn objects at different angles and directions; 2) Scaling and cropping: These operations can change the scale and size of the image, helping the model recognize objects of different sizes; 3) Image smoothing: Reduce image noise and make the image smoother by applying a smoothing filter; 4) Affine transformation: By combining linear transformation with translation, the model can adapt to complex changes such as image rotation, scaling, and shearing; 5) Translation: By moving the image linearly along a specific direction, the model can learn feature stability under position changes.
7. The method of tracking multiple ground targets from drone aerial video according to claim 1, characterized in that: The context attention mechanism CAM module in claim 4 comprises: 1) Receive the number of feature channels and multiple size parameters output by the first convolutional layer of the appearance feature extraction network, initialize the weight network, and generate feature weights; 2) Adaptively average pool the input feature map to the specified size, then perform convolution operation and return the convolution result; 3) Use the original feature map to subtract the scale feature map to obtain the weight feature, and input it into the weight network. Then use the activation function to calculate the weight value of each pixel position in each feature map.
8. The multi-target tracking algorithm for drone aerial video according to claim 1, characterized in that: The target tracking network of claim 5 comprises: 1) For each target detected by the target detection network, the target appearance feature extraction network in the target tracking network is used to extract the appearance features; the target features are converted into feature vectors that can be recognized and processed by the computer; with the help of the appearance feature vectors, the target tracking network can distinguish different targets; 2) At each time step, the Kalman filter is used in combination with the state transfer equation to predict and update the state distribution of each trajectory in the tracker, update its mean and covariance matrix, increase the trajectory survival time and the number of unmeasured update frames, and provide a basis for subsequent operations; 3) Using the Hungarian algorithm, the predicted position predicted by the Kalman filter is optimally matched with the actual target detection result, and the trajectory status is updated. After obtaining the current detection result, the confirmed and unconfirmed trajectories are distinguished. Confirmed trajectories are matched using the nearest neighbor cosine distance and cascade matching, and unconfirmed and partially unmatched trajectories are matched using the IoU matching. Combining appearance and geometric features, the results are merged to accurately associate detections with trajectories. 4) Update the trajectory set according to the matching results; the matched trajectory uses the Kalman filter to update its trajectory, and the unmatched trajectory is deleted when it times out; the unmatched detection box initializes a new trajectory, and finally filters the updated trajectory list to ensure accuracy and effectiveness.
Citation Information
Patent Citations
Maneuvering target tracking algorithm based on road network
CN107562837A
A Bayesian filter target tracking algorithm
CN109284677A
The invention discloses a mMulti-target pedestrian tracking system and method based on a KCF algorithm
CN109615641A
Self-adaptive target tracking method based on two complementary tracking algorithms
CN111860532A