A Multi-Target Tracking Method in Complex Scenarios Based on Adaptive Association
By building a new object detection network model and adaptive association mechanism, the accuracy and anti-interference problems of multi-objective tracking in complex scenarios are solved, and high-precision multi-objective detection and trajectory prediction are achieved.
Patent Information
- Application Number
- CN202510370363.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The existing multi-objective tracking method is difficult to effectively deal with long-distance small targets and scene changes in complex scenarios, and the data correlation stage is disturbed by multiple factors, making it difficult to choose a suitable confidence threshold, and changes in camera perspective affect the tracking effect.
A new object detection network model is built, an adaptive association and camera offset compensation mechanism is introduced, scene information is extracted through adaptive threshold matching and hybrid attention modules, and trajectory prediction is performed in combination with Kalman filters.
It improves the multi-object detection accuracy and trajectory prediction capabilities in complex scenarios, enhances anti-interference ability, and achieves high-precision multi-object tracking.
Smart Images

Figure CN119887850B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video analysis, and particularly relates to a multi-object tracking method in complex scenarios based on adaptive association. Background Art
[0002] Multi-object tracking (MOT) is an important research focus in the field of computer vision, and its purpose is to detect and estimate the spatio-temporal motion trajectories of multiple targets in an image video stream. With the rapid development of computer vision technology, multi-object tracking has been widely applied in fields such as video surveillance, autonomous driving, and intelligent transportation.
[0003] Currently, detection-based multi-object tracking has become an effective paradigm for MOT tasks. Detection-based multi-object tracking can be regarded as an extension of the object detection problem, or can be described as a matching problem of object position boxes in consecutive frames. With the rapid development of object detection technology, more and more methods have begun to use more powerful detectors to obtain higher tracking performance. For example, SORT first uses FRCNN to detect objects, and uses the Kalman filter algorithm and the Hungarian algorithm to predict the current frame position and update the position, which helps to reduce the missed detection rate of occluded targets, and the ID switching effect is very good when the object movement degree is small. DeepSORT uses the same detector and adds cascade matching and new track confirmation on the basis of SORT. These methods have promoted the development of multi-object tracking.
[0004] Although the above methods have achieved good results, there are still the following problems: (1) Most detection methods select targets for tracking from a parallel perspective, ignoring some small targets at a distance or other complex scenarios, and are not applicable to other scene changes, such as the drone's aerial perspective and adverse weather interference. (2) In addition, during the tracking process, the data association stage is interfered by various factors, such as the size and aggregation degree of the targets, and the change of the camera perspective. Most methods classify high-score and low-score bounding boxes through a single predefined threshold, and then perform data association. This threshold can be determined by repeatedly experimenting to obtain the threshold with the highest tracking accuracy to achieve a good tracking effect. However, in practical applications, the scene is usually not understood, making it difficult to select a suitable confidence threshold. Therefore, setting a fixed confidence threshold will limit the tracking performance. (3) In terms of the camera perspective, most methods consider the motion model of target pedestrians or cars in an ideal situation, that is, the camera has a fixed perspective. The tracked target can be approximated as linearly moving, and the motion model is input into the Kalman filter to estimate the trajectory of the object movement, and a good tracking effect can be achieved. However, whether it is handheld or drone shooting, it is difficult to avoid camera movement, or even severe shaking, which will undoubtedly affect the tracking effect. Summary of the Invention
[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a multi-object tracking method in complex scenarios based on adaptive association in view of the deficiencies of the prior art. On the basis of focusing on key information, the method of the present invention aims to effectively extract scene context information and individual features, and input the detection boxes into an improved data association method for trajectory prediction to achieve high-precision multi-object tracking.
[0006] The method of the present invention comprises the following steps:
[0007] Step 1, obtaining a video frame sequence containing multi-object motion and preprocessing the video frame sequence;
[0008] Step 2, constructing a new object detection network model to further process the multi-object images of consecutive frames; the new object detection network model takes the YOLO11 model as a reference model and includes a backbone network, a neck network and a head network;
[0009] The backbone network is used to extract object and scene context relationship information, and by using spatial and channel information, obtain fused features with enhanced representation ability; replace 4 C3k2 modules in the backbone network of the original YOLO11 model with 4 context information extraction modules CM (Context Module);
[0010] The neck network introduces a mixed channel attention module ML to synthesize features of different scales and extract local and global object features based on mixed attention;
[0011] The head network adopts the LIOU loss function to control the intersection over union while ensuring the diagonal ratio, and finally outputs the confidence, the position of the bounding box;
[0012] Step 3, performing overall multi-scale feature training and learning on the new object detection network model, using the loss function to accelerate the convergence of the model;
[0013] Step 4, performing performance evaluation and assessment on the object detection network model obtained after training in Step 3 to obtain the final object detection network model, and using the object detection network model to predict the labels of multi-object detection, namely the detection boxes and confidence scores (x, y, w, h, score), where x and y respectively represent the abscissa and ordinate coordinates of the object, w represents the width of the detection box, h represents the height of the detection box, and score represents the confidence score of the detection box (the value ranges from 0 to 1);
[0014] Step 5, inputting the detection box objects into the tracking algorithm, predicting the target trajectory by constructing more than two adaptive matches, dynamically adjusting the matching threshold according to the confidence, and performing camera offset compensation at the same time.
[0015] Step 1 includes: Using the VisDrone2019-MOT and VisDrone2019-Det drone datasets open-sourced by teams such as Tianjin University in 2019, collect the motion of small targets in multiple scenarios from the perspective of drones, providing 10 types of targets. This dataset was collected using different models of drones under different scenarios, lighting, and weather conditions, and ensuring that the image quality is suitable for input into deep learning models. According to the motion intensity in the scenario, set the frame interval, dynamically sparsely sample the original video, dynamically set the sampling ratio according to the number of targets and the scenario, and obtain a video frame sequence; annotate the detected objects, sort the video frames in chronological order, obtain multi-target images containing two or more consecutive frames, perform histogram block local equalization preprocessing to obtain a small target image dataset, and divide the small target image dataset into a training set, a validation set, and a test set. The test set is used to finally evaluate the performance of the model. Considering the relatively fast motion speed of targets from the perspective of drones, the frame interval is set to 1 to avoid losing important information. When computing resources are limited, the frame interval can be appropriately increased to 5 - 10 frames; considering the large number of targets and high scene complexity, the sampling ratio is dynamically set to 30% - 40% according to the number of targets and the scenario.
[0016] In Step 1, the histogram block local equalization preprocessing specifically includes the following steps:
[0017] Step 1-1, divide the image into non-overlapping region blocks tiles (usually 16×16 non-overlapping region blocks tiles);
[0018] Step 1-2, calculate the histogram for each region block tiles, and calculate the cumulative distribution function CDF according to the histogram;
[0019] Step 1-3, set the limit value frequency of the gray level (usually 2). If the histogram gray level frequency of a region block tiles exceeds the limit value frequency, then reduce the frequency and reassign the reduced frequency to other gray levels, keeping the histogram sum unchanged;
[0020] Step 1-4, perform local equalization using the cumulative distribution function CDF, and map the value of each pixel to a new value;
[0021] Step 1-5, perform smoothing processing using bilinear interpolation between region blocks tiles.
[0022] In Step 2, construct the backbone network through the following steps:
[0023] Step 2-1: Construct a convolutional layer (64, 3, 2) for the input image. Here, 64 represents the number of output channels (filters) of the convolutional layer, 3 represents the size of the convolutional kernel, i.e., a 3x3 convolutional kernel for processing local regions of the input image, and 2 represents the stride, i.e., the convolutional kernel slides on the input image with a stride of 2;
[0024] Then construct four consecutive convolutional downsamplings, where the size of the convolutional kernel for downsampling is 3×3 and the stride is 2;
[0025] Next, construct a context information extraction module for reducing the resolution and extracting high-level features;
[0026] Step 2-2: Construct a multi-scale feature utilization structure. Starting from the second context information extraction module CM, connect it to the neck network and perform a concat operation with the upsample module of the neck network (in deep learning, concat refers to the operation of concatenating multiple tensors along a specific dimension);
[0027] Step 2-3: Connect the backbone network to a mixed spatial channel attention module ML, and the mixed spatial channel attention module ML is connected to the neck network;
[0028] In Step 2-2, the context information extraction module CM includes a depthwise convolution with a large kernel and a pointwise convolution;
[0029] The number of groups of the depthwise convolution with a large kernel is the same as the number of channels;
[0030] The depthwise convolution with a large kernel first connects to a ReLU activation function and a normalization operation to extract global information for each channel; then a residual connection is made. Subsequently, the pointwise convolution is repeated twice, and then the activation function ReLU and the normalization operation are connected;
[0031] The two pointwise convolutions adopt an inverted bottleneck design, that is, the hidden dimension between the two pointwise convolutions is four times the input dimension, so as to realize the mixing of spatial and channel information; this structure can improve the detection ability of the model for small targets, especially in complex backgrounds;
[0032] In step 2-3, the hybrid spatial channel attention module ML includes a local average pooling module LAP, a global average pooling module GAP, an un-average pooling module UNGAP, and a one-dimensional convolutional module Conv1d;
[0033] For the input feature vector, the hybrid spatial channel attention module ML first transforms the input feature vector into a 1*C*ks*ks vector, and then extracts local spatial information through local average pooling LAP, where C represents the channel dimension and ks (kernel size) represents the convolution kernel size; the extracted local spatial information is transformed into a one-dimensional vector using two branches. The first branch uses the global average pooling module GAP operation, which contains global information, and the second branch directly performs a reshape operation, which contains local spatial information; the first branch passes through the one-dimensional convolutional module Conv1d, and the second branch passes through the reshape and then restores the original resolution of the two branches through the un-average pooling module UNGAP, and then performs information fusion, and finally performs residual connection with the original input feature vector to output a feature vector that combines global, local spatial, and channel attention; the convolution kernel size k of the one-dimensional convolutional module Conv1d is proportional to the channel dimension C, that is, when capturing local cross-channel interaction information, only the relationship between each channel and k adjacent channels is considered;
[0034] The convolution kernel size k is determined by the following formula:
[0035] ,
[0036] where b and are both hyperparameters, with default values of 2, and if k is even, 1 is added; is a proportional function.
[0037] In step 3, the overall multi-scale feature training and learning of the new object detection network model includes:
[0038] Step 3-1, for the new object detection network model, pre-train it using the COCO image recognition dataset and perform overall training using the VisDrone2019-Det drone perspective image recognition dataset;
[0039] Step 3-2, set the number of training epochs, batch size, input image size, learning rate, and intersection over union IoU, and use stochastic gradient descent SGD as the optimizer. In the present invention, the number of training epochs epochs = 100, the batch size batch = 4, the input image size is (640, 640), the learning rate is 0.01, and the IoU is 0.7.
[0040] In step 3, most commonly used loss functions (SIoU, DIoU, CIoU) measure the overlap between the predicted bounding box and the ground truth bounding box through IoU, Euclidean distance, aspect ratio, and angle. Since the aspect ratio is defined as a relative value, it may lead to errors, especially for small targets at a long distance of the drone in this embodiment. When the ground truth bounding box and the predicted bounding box have the same aspect ratio but different widths and heights, the value of the loss function is the same, which will make the loss function ineffective. Therefore, the present invention designs a more robust loss function to solve this problem. The loss function is as follows:
[0041] ,
[0042] where represents the IoU loss of the intersection over union, represents the Euclidean distance loss, represents the aspect ratio loss, is the coordinate of the ground truth bounding box, is the coordinate of the predicted bounding box, , , , , where respectively represent the abscissa, ordinate of the upper left corner, abscissa, and ordinate of the lower right corner of the ground truth bounding box, respectively represent the abscissa, ordinate of the upper left corner, abscissa, and ordinate of the lower right corner of the predicted bounding box;
[0043] where , , represents an adjustable scale factor, which can be freely adjusted according to actual needs and final effects; represents the Euclidean distance between two points. By introducing the logarithmic difference between the lengths and widths of the predicted bounding box and the ground truth bounding box and the lengths and widths of the minimum bounding box, the present invention can avoid the loss function from being ineffective in specific cases, thereby achieving a faster convergence speed and better positioning results.
[0044] Step 4 includes: evaluating the average detection accuracy and detection speed of the object detection network model obtained after training in Step 3. By setting the number of training epochs, batch size, input image size, learning rate, and intersection over union (IoU), input the sample images in the training set into the object detection network model to train and obtain the optimal weight file of the object detection network model. Continuously adjust the training direction of the object detection network model through the loss function, and verify whether the training reaches the expected effect by calculating the mean average precision (mAP) value of the validation set. Screen the trained weight files to select the optimal weight file, and load the optimal weight file into the object detection network model to obtain a model suitable for small object detection in the UAV scenario.
[0045] In Step 5, the tracking algorithm process includes:
[0046] Step 5-1: Starting from the first frame of the input video frame, input the video frame into the object detection network model trained in Step 4, and obtain the confidence threshold conf for distinguishing based on scene changes through the adaptive threshold matching method.
[0047] The adaptive threshold matching method includes: Given a list of confidence scores { } corresponding to the detection objects, sort { } in ascending order. If then ; Use the gradient descent method to find the confidence function discrimination point, find the confidence threshold by calculating the gradient, and calculate the first discrete weighted difference where and are the set weights, represents the confidence function, represents the confidence score of the i-th object, represents the confidence score of the j-th object;
[0048] Find the maximum value of the confidence gradient to obtain the threshold ;
[0049] Step 5-2: Set the state vector of the Kalman filter KF as where r represents the ratio of height h to width w, r = h / w;
[0050] respectively represent the velocity component in the x-axis direction and the velocity component in the y-axis direction of the abscissa, represents the velocity change component of the width, Represents the angular velocity component of rotation; use the detection box width to replace the aspect ratio of the original Kalman filter state vector, without changing the aspect ratio r of the quaternion, which can compensate for the aspect ratio change caused by the movement of the UAV or the relative movement of the object, and will not cause errors due to the small size of the object target;
[0051] Step 5-3: Use global motion estimation (GMC) to represent the background motion, extract the key points of the image, and perform feature tracking for local outlier suppression of translation using sparse optical flow; calculate the affine transformation matrix to transform the predicted bounding box from the coordinates of the t-1 frame to the coordinates of the next frame t;
[0052] The calculation of the affine transformation matrix includes:
[0053] The input data is the image point pairs , initialize the parameters: the maximum number of iterations N (usually 1000), the inlier threshold (usually 2 pixels), the minimum sample number s (usually 3), the best affine transformation matrix and the optimal inlier set ; respectively represent the abscissa and ordinate of the initial point in the i-th image point pair, respectively represent the abscissa and ordinate of the updated point in the i-th image point pair;
[0054] For each iteration t = 1, 2,..., N, perform the following steps:
[0055] Step 5-3-1: Random sampling: Randomly select s pairs of points from the image point pairs, and obtain the affine transformation matrix by solving the linear equations:
[0056] ,
[0057] where and respectively represent the scaling and rotation amount of the abscissa and the scaling and rotation amount of the ordinate, and respectively represent the shear amount of the abscissa and the shear amount of the ordinate, represents the translation amount of the abscissa, represents the translation amount of the ordinate; , respectively represent the abscissa and ordinate of the initial point in the s-th image point pair, , respectively represent the abscissa and ordinate of the updated point in the s-th image point pair;
[0058] Step 5-3-2: For all image point pairs P, calculate the error between the transformed point and the target point:
[0059] ,
[0060] If the error between the points of the image point pair and the target point is less than the threshold , then mark the image point pair as an inlier;
[0061] Step 5-3-3, if the number of inliers of the current affine transformation matrix is greater than the number of inliers in the optimal inlier set , then update , , where is the current inlier set;
[0062] After the iteration is completed, use all inlier sets to re-estimate the affine transformation matrix by the least squares method, and return the optimal affine transformation matrix and the inlier set ;
[0063] Step 5-4, divide according to the confidence threshold conf. For the detection boxes with a confidence higher than the confidence threshold conf, perform camera offset compensation, and then use the Kalman filter for the first complete intersection over union (CIoU) and re-identification (ReID) data association; for the detection boxes with a confidence lower than the confidence threshold conf, prepare for the second data CIoU association;
[0064] The camera offset compensation includes: predicting according to the state vector in Step 5-2 and performing translation and rotation to obtain the final state quantity:
[0065] ,
[0066] = ,
[0067] where is the matrix containing the scale and rotation parts of the affine matrix, and the superscript T represents the transpose, , is the transformation matrix describing the scaling and rotation in the two-dimensional plane, is the matrix of the translation part, , R represents the real number space, and Tr represents the translation matrix, ; is the state estimate value at time t predicted based on the information at time t - 1; is the result after Kalman prediction, represents the predicted covariance matrix at time t, represents the covariance matrix at time t.
[0068] Advantages: Compared with the prior art, the present invention is based on a deep learning benchmark model and a detection-based tracking framework, achieving multi-object detection and trajectory prediction with high accuracy in special scenarios. By introducing a context information extraction module CM, the ability to extract context information and global scenarios can be effectively improved. And by introducing a hybrid attention module ML to extract key objects in the scenario, good conditions are provided for multi-object tracking. At the same time, the present invention introduces an adaptive threshold and a camera offset compensation mechanism, enhancing the anti-interference ability of tracking in different environments. The model proposed by the present invention is subjected to ablation experiments and compared with common deep learning models. Under the same test conditions, the present invention has excellent detection and tracking performance. The present invention has broad prospects in practical applications, especially in the fields of intelligent monitoring, autonomous driving, robot navigation, etc. Description of the Drawings
[0069] Figure 1 It is a flowchart of the method of the present invention.
[0070] Figure 2 It is an architecture diagram of the detector network of the present invention.
[0071] Figure 3 It is a structural diagram of the tracking algorithm module in the present invention.
[0072] Figure 4 It is a structural diagram of the context information extraction module in the present invention.
[0073] Figure 5 It is a structural diagram of the hybrid spatial channel attention module in the present invention.
[0074] Figure 6 It is a visualization effect diagram of tracking from the perspective of an unmanned aerial vehicle in the present invention. Detailed Embodiments
[0075] The following further specific descriptions of the present invention are made in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0076] As Figure 1 shown, the embodiment of the present invention provides a multi-object tracking method in a complex scenario based on adaptive association, including the following steps:
[0077] Step 1, obtain a video frame sequence containing multi-object motion, and preprocess the video frame sequence; sort the video frames in chronological order to obtain a multi-object image containing multiple consecutive frames.
[0078] Step 2, construct a new object detection network model to further process the multi-object image of consecutive frames; the new object detection network model uses the YOLO11 model as the benchmark model, including a backbone network, a neck network, and a head network;
[0079] The backbone network is used to extract the relationship information between the target object and the scene context, and utilize the spatial and channel information to obtain the fused features with enhanced representation ability. Four context information extraction modules CM (Context Module) are used to replace the four C3k2 modules in the backbone network of the original YOLO11 model.
[0080] The neck network introduces a hybrid channel attention module ML to integrate features of different scales and extract local and global target features based on the hybrid attention.
[0081] The head network adopts the LIOU loss function to ensure the diagonal ratio while controlling the intersection over union, and finally outputs the confidence, bounding box position.
[0082] As Figure 2 shown, the backbone structure first processes the input image through a convolutional layer (Conv). Then, four consecutive convolutional downsamplings (Conv) and context information extractions (CM) are constructed to extract features. After that, a Spatial Pyramid Pooling - Fast (SPPF) module is used to further enhance the expression ability of multi - scale features. In the neck structure part, the operations of a feature extraction component (C3k2) are first applied, and then upsampling (Upsample) is performed to restore the spatial resolution of the feature map. Next, the feature maps from different paths are merged through a concatenation operation (Concat) to generate a richer feature representation. Finally, in the head structure, a head module (LIoU - Detect) with the LIoU loss function is used for the final detection prediction to output the labels of the targets. Through the effective connection of the backbone, neck, and head, the entire network structure can achieve efficient object detection tasks.
[0083] Step 3: Conduct overall multi - scale feature training and learning on the new object detection network model, and use the loss function to accelerate the model convergence;
[0084] Step 4: Conduct performance evaluation on the object detection network model obtained after training in Step 3 to obtain the final object detection network model. Use the object detection network model to predict the labels of multi - object detection, namely the detection boxes and confidence scores (x, y, w, h, score), where x and y represent the abscissa and ordinate coordinates of the target respectively, w represents the width of the detection box, h represents the height of the detection box, and score represents the confidence score of the detection box (ranging from 0 to 1);
[0085] Step 5: Input the detection box objects into the tracking algorithm, construct more than two adaptive matches, dynamically adjust the matching threshold according to the confidence, and at the same time perform camera offset compensation to predict the target trajectory.
[0086] In Step 2, the backbone network is constructed through the following steps:
[0087] Step 2-1: Construct a convolutional layer (64, 3, 2) for the input image, where 64 represents the number of channels (filters) output by the convolutional layer, 3 represents the size of the convolutional kernel, i.e., a 3x3 convolutional kernel for processing local regions of the input image, and 2 represents the stride, i.e., the convolutional kernel slides on the input image with a stride of 2;
[0088] Then, construct four consecutive convolutional downsamplings, where the size of the convolutional kernel for downsampling is 3×3 and the stride is 2.
[0089] Next, construct a context information extraction module for reducing the resolution and extracting high-level features;
[0090] Step 2-2: Construct a multi-scale feature utilization structure, starting from the second context information extraction module CM and connecting it to the neck network, and performing a concat operation with the upsample module of the neck network (in deep learning, concat refers to the operation of concatenating multiple tensors along a specific dimension);
[0091] Step 2-3: Connect the backbone network to a mixed spatial channel attention module ML, and the mixed spatial channel attention module ML is connected to the neck network;
[0092] In Step 2-2, the context information extraction module CM is as Figure 4 shown. The context information extraction module CM includes depthwise convolution with a large kernel and pointwise convolution;
[0093] The number of groups of the depthwise convolution with a large kernel is the same as the number of channels;
[0094] The depthwise convolution with a large kernel is first connected to a ReLU activation function and a normalization operation for extracting global information of each channel; then, a residual connection is made. Subsequently, two pointwise convolutions are repeated, and then connected to the activation function ReLU and the normalization operation;
[0095] The two pointwise convolutions adopt an inverted bottleneck design, i.e., the hidden dimension between the two pointwise convolutions is four times the input dimension, so as to achieve the mixing of spatial and channel information; this structure can improve the detection ability of the model for small targets, especially in complex backgrounds;
[0096] In step 2-3, the hybrid spatial-channel attention module ML is as Figure 5 shown. The hybrid spatial-channel attention module ML includes a local average pooling module LAP, a global average pooling module GAP, an un-average pooling module UNGAP, and a one-dimensional convolutional module Conv1d;
[0097] For the input feature vector, the hybrid spatial-channel attention module ML first transforms the input feature vector into a 1*C*ks*ks vector, and then extracts local spatial information through local average pooling LAP, where C represents the channel dimension and ks (kernel size) represents the convolution kernel size; the extracted local spatial information is transformed into a one-dimensional vector using two branches. The first branch uses the global average pooling module GAP operation, which contains global information, and the second branch directly performs a reshape operation, which contains local spatial information; the first branch passes through the one-dimensional convolutional module Conv1d, and the second branch passes through the reshape. Then, the original resolution of the two branches is restored through the un-average pooling module UNGAP, and information fusion is performed. Finally, the feature vector of the original input is connected by residual, so as to output the feature vector of the hybrid global, local spatial, and channel attention; the convolution kernel size k of the one-dimensional convolutional module Conv1d is proportional to the channel dimension C, that is, when capturing local cross-channel interaction information, only the relationship between each channel and k adjacent channels is considered;
[0098] The convolution kernel size k is determined by the following formula:
[0099] ,
[0100] where b and are both hyperparameters, the default value is 2, and if k is even, 1 is added; is a proportional function.
[0101] In step 3, the overall multi-scale feature training and learning of the new object detection network model includes:
[0102] Step 3-1, for the new object detection network model, pre-train it using the COCO image recognition dataset and perform overall training using the VisDrone2019-Det drone perspective image recognition dataset;
[0103] Step 3-2, set the number of training epochs, batch size, input image size, learning rate, and intersection over union IoU, and use Stochastic Gradient Descent SGD as the optimizer. In the present invention, the number of training epochs epochs = 100, the batch size batch = 4, the input image size is (640, 640), the learning rate is 0.01, and the IoU is 0.7.
[0104] In step 3, the loss function is as follows:
[0105] ,
[0106] where represents the Intersection over Union (IoU) loss, represents the Euclidean distance loss, represents the aspect ratio loss, is the coordinate of the ground truth bounding box, is the coordinate of the predicted bounding box, , , , where respectively represent the abscissa, ordinate of the upper left corner and the abscissa, ordinate of the lower right corner of the ground truth bounding box, respectively represent the abscissa, ordinate of the upper left corner and the abscissa, ordinate of the lower right corner of the predicted bounding box;
[0107] where , , represents an adjustable scale factor that can be freely adjusted according to actual needs and final effects; represents the Euclidean distance between two points. By introducing the logarithmic difference between the length and width of the predicted box and the ground truth box and the length and width of the minimum bounding box, the present invention can avoid the loss function from failing in specific cases, thereby achieving a faster convergence speed and better localization results.
[0108] In step 5, the overall framework of the tracking algorithm process is as Figure 3 shown. The tracking algorithm process includes:
[0109] Step 5-1, starting from the first frame of the input video frame, input the video frame into the object detection network model trained in step 4, and obtain the confidence threshold conf for distinguishing based on scene changes through the adaptive threshold matching method;
[0110] The adaptive threshold matching method includes: given a list of confidence scores { } corresponding to the detection object, sort { } in ascending order. If , then ; use the gradient descent method to find the discrimination point of the confidence function, find the confidence threshold by calculating the gradient, and calculate the first discrete weighted difference , where and are set weights, represents the confidence function, Represents the confidence score of the i-th target, Represents the confidence score of the j-th target;
[0111] Find the maximum value of the confidence gradient to obtain the threshold ;
[0112] Step 5-2, set the state vector of the Kalman filter KF as , where r represents the ratio of height h to width w, r = h / w;
[0113] Represent the velocity component in the x direction of the abscissa and the velocity component in the y direction of the ordinate respectively, Represents the velocity component of the width change, Represents the angular velocity component of rotation; use the detection box width to replace the aspect ratio of the original Kalman filter state vector, without changing the four-element aspect ratio r, which can compensate for the aspect ratio change caused by the movement of the UAV or the relative movement of the object, and will not cause errors due to the small size of the object target;
[0114] Step 5-3, use global motion estimation GMC to represent the background motion, extract image key points, and perform feature tracking for local outlier suppression of translation using sparse optical flow; calculate the affine transformation matrix to transform the predicted bounding box from the coordinates of the t-1 frame to the coordinates of the next frame t;
[0115] The calculation of the affine transformation matrix includes:
[0116] The input data is the image point pair , initialize the parameters maximum number of iterations N (usually 1000), inlier threshold (usually 2 pixels), minimum sample number s (usually 3), the best affine transformation matrix and the optimal inlier set ; Represent the abscissa and ordinate of the initial point in the i-th image point pair respectively, Represent the abscissa and ordinate of the updated point in the i-th image point pair respectively;
[0117] For each iteration t = 1, 2,..., N, perform the following steps:
[0118] Step 5-3-1, random sampling: randomly select s pairs of points from the image point pair, and obtain the affine transformation matrix :
[0119] ,
[0120] where and respectively represent the scaling and rotation amount of the abscissa and the scaling and rotation amount of the ordinate, and respectively represent the shearing amount of the abscissa and the shearing amount of the ordinate, represents the translation amount of the abscissa, represents the translation amount of the ordinate; , respectively represent the abscissa and the ordinate of the initial point in the s-th image point pair, , respectively represent the abscissa and the ordinate of the updated point in the s-th image point pair;
[0121] Step 5-3-2, for all image point pairs P, calculate the error between the transformed point and the target point :
[0122] ,
[0123] If the error error between the point of the image point pair and the target point is less than the threshold , then mark the image point pair as an inlier;
[0124] Step 5-3-3, if the number of inliers of the current affine transformation matrix is greater than the number of inliers in the optimal inlier set , then update , , where is the current inlier set;
[0125] After the iteration is completed, re-estimate the affine transformation matrix using the least squares method for all inlier sets and return the optimal affine transformation matrix and the inlier set ;
[0126] Step 5-4, divide according to the confidence threshold conf. For the detection boxes with a confidence higher than the confidence threshold conf, perform camera offset compensation, and then use the Kalman filter for the first complete intersection over union CIoU and re-identification ReID data association; for the detection boxes with a confidence lower than the confidence threshold conf, prepare for the second data CIoU association;
[0127] The camera offset compensation includes: predicting according to the state vector in Step 5-2 and performing translation and rotation to obtain the final state quantity:
[0128] ,
[0129] = ,
[0130] where is a matrix containing the scale and rotation parts of the affine matrix, and the superscript T represents the transpose. , is a transformation matrix describing the scaling and rotation in the two-dimensional plane. is the matrix of the translation part. , R represents the real number space, and Tr represents the translation matrix. ; is the state estimate value at time t predicted based on the information at time t - 1. is the result after Kalman prediction. represents the predicted covariance matrix at time t. represents the covariance matrix at time t.
[0131] In this embodiment, the datasets used are VisDrone-DET2019, Visdrone2019-MOT, and COCO. The VisDrone-DET2019 dataset uses a drone platform to collect moving targets at different locations and different altitudes, including 8599 images. More than 540k target bounding boxes are labeled with 10 predefined categories: pedestrian, person1, car, van, bus, truck, motor, bicycle, awning-tricycle, tricycle. The dataset is divided into three subsets: training, validation, and test (6471 images in the training subset, 548 images in the validation subset, and 1580 images in the test subset). The VisDrone-MOT2019 dataset also uses a drone platform to collect moving targets at different locations and different altitudes, consisting of 79 video clips and a total of 33366 frames. According to the feature differences of moving targets in the scene, this dataset focuses on 5 selected label categories: pedestrian, car, van, bus, truck. The COCO (Common Objects in Context) dataset is a large-scale image recognition dataset used for object detection, segmentation, and caption tasks. It contains more than 330,000 images, each annotated with 80 object categories and 5 captions describing the scene. The COCO dataset is widely used in computer vision research and has been used to train and evaluate many state-of-the-art object detection and segmentation models.
[0132] On the VisDrone-DET2019 dataset, the input image size is (640, 640). A total of 100 rounds of training and testing are set. The initial learning rate is set to 0.01, the batch size is 4, and the IoU is set to 0.7. The mAP0.5 value on the test set of VisDrone-DET2019 reaches 34.7%, which is 1.5% higher than the baseline model. The mAP0.5:0.95 value reaches 20%, which is 1.6% higher than the baseline model. On VisDrone-MOT2019, the HOTA value reaches 43.6%, the MOTP reaches 76.6%, and the IDF1 reaches 53.3%. Compared with the baseline model, they are increased by 7.1%, 2.1%, and 12.6% respectively. The results of the ablation experiment are shown in Table 1. The visualization effect of tracking from the drone perspective is as Figure 6 shown.
[0133] Table 1
[0134]
[0135] The results show that the present invention has excellent detection and tracking performance, realizes multi-object detection and trajectory prediction with high accuracy in special scenarios, can effectively improve the context information and global scene extraction ability, and extracts key figures in the scene through the attention mechanism, providing good conditions for multi-object tracking.
[0136] The present invention provides a multi-object tracking method in complex scenarios based on adaptive association. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.
Claims
1. A multi-target tracking method in complex scenarios based on adaptive association, characterized in that, It includes the following steps: Step 1: Obtain a video frame sequence containing multi-object motion and preprocess the video frame sequence; Step 2: Construct a new object detection network model to further process the multi-object images of consecutive frames; the new object detection network model uses the YOLO11 model as the benchmark model and includes a backbone network, a neck network, and a head network; The backbone network is used to extract the object and scene context relationship information, utilize the spatial and channel information, and obtain the fused features with enhanced representation ability; replace the 4 C3k2 modules in the backbone network of the original YOLO11 model with 4 context information extraction modules CM; The neck network introduces a hybrid spatial channel attention module ML to integrate features of different scales and extract local and global object features based on the hybrid attention; The head network adopts the loss function L LIoU , while ensuring the diagonal ratio and controlling the intersection over union, and finally outputs the confidence and the bounding box position; Step 3, perform overall multi-scale feature training and learning on the new object detection network model, and use the loss function L LIoU Accelerate model convergence; Step 4: Perform performance evaluation and assessment on the object detection network model obtained after training in Step 3 to obtain the final object detection network model, and use the object detection network model to predict the labels of multi-object detection, that is, the detection boxes and confidence scores (x, y, w, h, score), where x and y represent the abscissa and ordinate of the object respectively, w represents the width of the detection box, h represents the height of the detection box, and score represents the confidence score of the detection box; Step 5: Input the detection box objects into the tracking algorithm, construct more than two adaptive matches, dynamically adjust the matching threshold according to the confidence, and perform camera offset compensation at the same time to predict the target trajectory.
2. The method according to claim 1, wherein Step 1 includes: setting the frame interval according to the motion intensity in the scene, dynamically sparsely sampling the original video, dynamically setting the sampling ratio according to the number of objects and the scene to obtain the video frame sequence; annotating the detected objects, sorting the video frames in chronological order to obtain multi-object images containing more than two consecutive frames, performing histogram block local equalization preprocessing to obtain a small object image dataset, and dividing the small object image dataset into a training set, a validation set, and a test set.
3. The method according to claim 2, wherein In Step 1, the histogram block local equalization preprocessing specifically includes the following steps: Step 1-1: Divide the image into non-overlapping region blocks tiles; Step 1-2: Calculate the histogram for each region block tiles, and calculate the cumulative distribution function CDF according to the histogram; Step 1-3: Set the limit value frequency of the gray levels. If the gray level frequency of a region block tiles exceeds the limit value frequency, reduce the frequency and reallocate the reduced frequency to other gray levels while keeping the histogram sum unchanged; Step 1-4: Use the cumulative distribution function CDF to perform local equalization and map the value of each pixel to a new value; Step 1-5: Perform smoothing processing using bilinear interpolation between the region blocks tiles.
4. The method according to claim 3, characterized in that, In Step 2, the backbone network is constructed through the following steps: Step 2-1, construct a convolutional layer (64, 3, 2) for the input image, where 64 represents the number of channels filters output by the convolutional layer, 3 represents the size of the convolutional kernel, i.e., a 3x3 convolutional kernel for processing local regions of the input image, and 2 represents the stride, i.e., the convolutional kernel slides on the input image with a stride of 2; Then construct four consecutive convolutional downsamplings, where the size of the convolutional kernel for downsampling is 3×3 and the stride is 2; Then construct a context information extraction module for reducing the resolution and extracting high-level features; Step 2-2, construct a multi-scale feature utilization structure, starting from the second context information extraction module CM and connecting it to the neck network, and performing a concat operation with the upsample module of the neck network; Step 2-3, connect the backbone network to a mixed spatial channel attention module ML, and the mixed spatial channel attention module ML is connected to the neck network.
5. The method according to claim 4, wherein In Step 2-2, the context information extraction module CM includes depthwise convolution with a large kernel and pointwise convolution; The number of groups of the depthwise convolution with a large kernel is the same as the number of channels; The depthwise convolution with a large kernel is first connected to a ReLU activation function and a normalization operation for extracting global information of each channel; Then perform a residual connection, and subsequently, repeat the pointwise convolution twice, and then connect to the activation function ReLU and the normalization operation; The two pointwise convolutions adopt an inverted bottleneck design, that is, the hidden dimension between the two pointwise convolutions is four times the input dimension, so as to realize the mixing of spatial and channel information.
6. The method according to claim 5, wherein In Step 2-3, the mixed spatial channel attention module ML includes a local average pooling module LAP, a global average pooling module GAP, an anti-average pooling module UNGAP, and a one-dimensional convolutional module Conv1d; For the input feature vector, the hybrid spatial channel attention module ML first transforms the input feature vector into a 1*C*ks*ks vector, and then extracts local spatial information through the local average pooling module LAP, where C represents the channel dimension and ks represents the convolution kernel size; the extracted local spatial information is transformed into a one-dimensional vector using two branches. The first branch uses the global average pooling module GAP operation, which contains global information, and the second branch directly performs a reshape operation, which contains local spatial information; the first branch passes through the one-dimensional convolution module Conv1d, and the second branch passes through the reshape. After that, the original resolution of the two branches is restored through the un-average pooling module UNGAP, and then information fusion is performed. Finally, the feature vector of the original input is connected by residual, so as to output the feature vector of hybrid global, local spatial and channel attention; Among them, the convolution kernel size k of the one-dimensional convolution module Conv1d is proportional to the channel dimension C; The convolution kernel size k is determined by the following formula: where both b and γ are hyperparameters; is a proportional function.
7. The method according to claim 6, wherein In step 3, the overall multi-scale feature training and learning of the new object detection network model includes: Step 3-1, for the new object detection network model, pre-train it using the COCO image recognition dataset, and perform overall training using the VisDrone2019-Det UAV perspective image recognition dataset; Step 3-2, set the number of training epochs, batch size, input image size, learning rate, and intersection over union IoU, and use stochastic gradient descent SGD as the optimizer.
8. The method according to claim 7, characterized in that, In step 3, the loss function L LIoU is as follows: Among them, L IoU represents the Intersection over Union (IoU) loss, and L dis represents the Euclidean distance loss, and L asp represents the aspect ratio loss, B gt is the coordinate of the ground truth bounding box, and B prd is the coordinate of the predicted bounding box, where respectively represent the abscissa and ordinate of the upper left corner and the abscissa and ordinate of the lower right corner of the ground truth bounding box, respectively represent the abscissa and ordinate of the upper left corner and the abscissa and ordinate of the lower right corner of the predicted bounding box; wherein α represents an adjustable scale factor, and ρ represents the Euclidean distance between two points.
9. The method according to claim 8, characterized in that, Step 4 includes: evaluating the average detection accuracy and detection speed of the object detection network model obtained after training in step 3. By setting the number of training epochs, batch size, input image size, learning rate, and intersection over union IoU, input the sample images in the training set into the object detection network model, and train to obtain the best weight file of the object detection network model. Continuously adjust the training direction of the object detection network model through the loss function, and verify whether the training reaches the expected effect by calculating the mean average precision mAP value of the validation set; select the best weight file by screening the trained weight files, and load the best weight file into the object detection network model to obtain a model suitable for small object detection in the UAV scenario.
10. The method according to claim 9, wherein In step 5, the tracking algorithm process includes: Step 5-1: Starting from the first frame of the input video frames, input the video frames into the object detection network model trained in Step 4, and obtain the confidence threshold c for distinguishing based on scene changes through the adaptive threshold matching method. ith ; The adaptive threshold matching method includes: Given a list of confidence scores {c i} corresponding to the detection object, sort {c i} in ascending order. If c i <c j , then f(c i ) < f(c j ); Use the gradient descent method to find the confidence function discrimination point, find the confidence threshold by calculating the gradient, and calculate the first discrete weighted difference Δ w f(c i ) = w1f(c i+1 ) - w2f(c i ), where w1 and w2 are set weights, f() represents the confidence function, c i represents the confidence score of the i-th target, and c j represents the confidence score of the j-th target; Find the maximum value of the confidence gradient to obtain the confidence threshold c ith ; Step 5-2, set the state vector of the Kalman filter KF as [x, y, w, r, v x , v y , v w , v r , where r represents the ratio of the height h to the width w, r = h / w; v x , v y represent the velocity component in the x-axis direction of the abscissa and the velocity component in the y-axis direction of the ordinate respectively, and v w represents the velocity component of the width change, and v r represents the angular velocity component of rotation; use the detection box width to replace the aspect ratio of the original Kalman filter state vector, without changing the aspect ratio r; Step 5-3, use global motion estimation GMC to represent the background motion, extract image key points, and perform feature tracking of local outlier suppression of translation using sparse optical flow; calculate the affine transformation matrix, and transform the predicted bounding box from the coordinates of the t-1 frame to the coordinates of the next frame t; The calculation of the affine transformation matrix includes: The input data is the image point pairs P = {(x i , y i ), (x' i , y' i )}, initialize the parameters: the maximum number of iterations N, the inlier threshold ∈, the minimum sample number s, the best affine transformation matrix M best and the optimal inlier set I best ; x i , y i represent the abscissa and ordinate of the initial point in the i-th image point pair respectively, and x' i , y' i represent the abscissa and ordinate of the updated point in the i-th image point pair respectively; For each iteration m = 1, 2,..., N, perform the following steps: Step 5-3-1, Random Sampling: Randomly select s pairs of points from the image point pairs, and obtain the affine transformation matrix M by solving the linear equations t :[[]]END]] where a 11 and a 22 represent the scaling and rotation amounts of the abscissa and the scaling and rotation amounts of the ordinate respectively, and a 12 and a 21 represent the shearing amounts of the abscissa and the shearing amounts of the ordinate respectively, tr x represents the translation amount of the abscissa, and tr y represents the translation amount of the ordinate; x s , y s represent the abscissa and the ordinate of the initial point in the s-th image point pair respectively, and x′ s , y′ s represent the abscissa and the ordinate of the updated point in the s-th image point pair respectively; Step 5-3-2, for all image point pairs P, calculate the error error between the transformed point and the target point: If the error error between the point of the image point pair and the target point is less than the threshold ∈, then mark the image point pair as an inlier; Step 5-3-3, if the number of inliers of the current affine transformation matrix M t is greater than the number of inliers in the optimal inlier set I best , then update M best = M t , I best = I t , where I t is the current inlier set; After the iteration is completed, use all the inlier sets I best Re - estimate the affine transformation matrix using the least - squares method and return the optimal affine transformation matrix M best and the inlier set I best ; Step 5-4, according to the confidence threshold c ith Partition, for the detection boxes with a confidence higher than the confidence threshold c ith perform camera offset compensation, and then use the Kalman filter for the first Complete Intersection over Union (CIoU) and Re-Identification (ReID) data association; for the detection boxes with a confidence lower than the confidence threshold c ith prepare for the second data CIoU association; The camera offset compensation includes: predicting according to the state vector in step 5-2 and performing translation and rotation to obtain the final state quantity: where M′ t|t-1 is the matrix containing the scale and rotation parts of the affine matrix, and the superscript T represents the transpose, M′ t|t-1 = diag{M,M,M,M} ∈ R 8×8 , M ∈ R 2×2 is the transformation matrix describing the scaling and rotation of the two-dimensional plane, T′ t|t-1 is the matrix of the translation part, T′ t|t-1 = {Tr,0,0,0,0,0,0} ∈ R 8 , R represents the real number space, Tr represents the translation matrix, Tr ∈ R; is the state estimate value at time t predicted based on the information at time t-1; is the result after Kalman prediction, P′ t|t-1 represents the predicted covariance matrix at time t, P t|t-1 represents the covariance matrix at time t.
Citation Information
Patent Citations
AC overvoltage detection circuit, air conditioner, indoor unit of air conditioner, and control board thereof
CN109239446A
Charging and discharging control circuit for charging lamp and LED emergency lamp
CN112087044A