A Multi-Object Tracking Method and System Based on Deep Learning

By embedding the spatial pyramid module and adaptive feature update mechanism in the multi-objective tracking method, the problem of insufficient accuracy and robustness of multi-objective tracking in complex scenarios is solved, and higher accuracy and robustness are achieved.

CN114419101BActive Publication Date: 2025-06-17NANCHANG HANGKONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210079363.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-06-17
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

The existing multi-objective tracking methods lack accuracy and robustness in complex scenarios, brightness changes, noise interference and target occlusion.

Method used

Using a multi-objective tracking method based on deep learning, the multi-objective tracking network model is optimized by embedding the spatial pyramid module into a multi-scale feature pyramid network and combining the adaptive feature update mechanism.

Benefits of technology

It improves the accuracy and robustness of multi-target tracking in complex scenarios, and reduces mismatch and identity transformation during target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114419101B_ABST
    Figure CN114419101B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-object tracking method and system based on deep learning, including: embedding a spatial pyramid module into a multi-scale feature pyramid network to obtain an improved feature extraction network; acquiring images at T moments in a target scene and inputting them into the improved feature extraction network to obtain target position detection results and target feature vectors at T moments; calculating the intersection over union based on the target position detection results and prediction results at the t-th moment; calculating the cosine similarity of target features according to the target feature vectors at two consecutive moments; screening out multiple tracked target objects according to the intersection over union and cosine similarity, and assigning the same label to the same tracked target object; adaptively weighting and updating the target feature vector at the (t + 1)-th moment until each tracked target object in all images is assigned a label, and performing multi-object tracking according to the labels. The spatial pyramid and adaptive feature update mechanism are introduced to improve the accuracy of multi-object tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and particularly to a multi-object tracking method and system based on deep learning. Background Art

[0002] Multi-object tracking technology is one of the important branches in the field of computer vision. Its main function is to traverse and find the positions of individual moving objects of interest with certain significant visual features in the acquired video images, and track them in the subsequent video frames. Therefore, multi-object tracking technology is widely used in various fields, such as important tasks like intelligent human-computer interaction, surveillance and security, driverless, robot intelligent navigation and positioning, and cruise missile guidance.

[0003] Currently, traditional multi-object tracking methods and deep learning methods have high accuracy and few identity transformations in most scenarios. However, when the scene is complex, the crowd is crowded, there are too many interfering objects, and the target faces deformation and illumination changes, etc., the accuracy and robustness of existing multi-object tracking methods still need to be further improved. Therefore, aiming at the problem that the multi-object tracking effect is not ideal in complex scenarios, brightness changes, noise interference, occlusion between targets, etc. Therefore, the present invention provides a multi-object tracking method and system based on deep learning. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-object tracking method and system based on deep learning, which optimizes the multi-object tracking network model by using a spatial pyramid and an adaptive feature update mechanism to improve the accuracy and robustness of multi-object tracking in complex scenarios.

[0005] To achieve the above purpose, the present invention provides the following solutions:

[0006] A multi-object tracking method based on deep learning, comprising:

[0007] Embedding a spatial pyramid module into a multi-scale feature pyramid network to obtain an improved feature extraction network;

[0008] Obtaining images at T moments in a target scene; the images include multiple targets;

[0009] Inputting the images at T moments into the improved feature extraction network to obtain target position detection results at T moments and target feature vectors at T moments; T = 1, 2,..., t - 1, t, t + 1,...;

[0010] Based on the target position detection result at the t-th moment, using Kalman filtering to predict the target state at the (t + 1)-th moment to obtain the target position prediction result at the (t + 1)-th moment; t ∈ {1, 2, 3,..., T};

[0011] Calculate the intersection over union (IoU) of the predicted target position result at time t+1 and the detected target position result at time t+1 to obtain the target position IoU.

[0012] Calculate the cosine similarity of the target feature vectors based on the target feature vector at time t and the target feature vector at time t+1.

[0013] Select multiple tracked target objects according to the target position IoU and the cosine similarity of the target features, and assign the same label to the same tracked target object.

[0014] Adaptive weight the target feature vector at time t+1 according to the cosine similarity of the target features in combination with an adaptive weighting mechanism to obtain the weighted target feature vector at time t+1. Take the weighted target feature vector at time t+1 as the target feature vector at time t+1, and let t = t+1. Return to the step of "predicting the target state at time t+1 using the Kalman filter based on the detected target position result at time t" until t = T, and assign corresponding labels to multiple tracked target objects in each moment of the image.

[0015] Track multiple tracked target objects according to the labels.

[0016] A multi-object tracking system based on deep learning, comprising:

[0017] A feature extraction network acquisition module, configured to embed a spatial pyramid module into a multi-scale feature pyramid network to obtain an improved feature extraction network.

[0018] An image acquisition module, configured to acquire images at T moments in a target scene; multiple targets are included in the images.

[0019] A feature extraction module, configured to input the images at T moments into the improved feature extraction network to obtain the detected target position results at T moments and the target feature vectors at T moments; T = 1, 2,..., t-1, t, t+1,...

[0020] A target position prediction module, configured to predict the target state at time t+1 using the Kalman filter based on the detected target position result at time t to obtain the predicted target position result at time t+1; t ∈ {1, 2, 3,..., T}.

[0021] An IoU calculation module, configured to calculate the intersection over union (IoU) of the predicted target position result at time t+1 and the detected target position result at time t+1 to obtain the target position IoU.

[0022] A cosine similarity calculation module, which is used to calculate the cosine similarity of the target feature vectors at time t and at time t+1;

[0023] A target screening and labeling module, which is used to screen out multiple tracked target objects according to the target position intersection over union and the cosine similarity of the target features, and assign the same label to the same tracked target object;

[0024] An adaptive weighting and state update module, which is used to adaptively weight the target feature vector at time t+1 according to the cosine similarity of the target features in combination with an adaptive weighting mechanism to obtain the weighted target feature vector at time t+1, use the weighted target feature vector at time t+1 as the target feature vector at time t+1, and let t=t+1, then return to execute the target position prediction module; until t=T, corresponding labels are assigned to the multiple tracked target objects in the image at each moment;

[0025] A target tracking module, which is used to track the multiple tracked target objects according to the labels.

[0026] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0027] The present invention relates to a multi-object tracking method and system based on deep learning, including: embedding a spatial pyramid module into a multi-scale feature pyramid network to obtain an improved feature extraction network; acquiring images at T moments in a target scene; the images include multiple targets; inputting the images at T moments into the improved feature extraction network to obtain target position detection results at T moments and target feature vectors at T moments; T = 1, 2,..., t - 1, t, t + 1,...; predicting the target state at the (t + 1)-th moment using Kalman filtering based on the target position detection result at the t-th moment to obtain the target position prediction result at the (t + 1)-th moment; t ∈ 1, 2, 3,..., T; calculating the intersection over union of the target position prediction result at the (t + 1)-th moment and the target position detection result at the (t + 1)-th moment to obtain the target position intersection over union; calculating the target feature cosine similarity according to the target feature vector at the t-th moment and the target feature vector at the (t + 1)-th moment; screening out multiple tracking target objects according to the target position intersection over union and the target feature cosine similarity, and assigning the same label to the same tracking target object; adaptively weighting the target feature vector at the (t + 1)-th moment according to the target feature cosine similarity in combination with an adaptive weighting mechanism to obtain the weighted target feature vector at the (t + 1)-th moment, taking the weighted target feature vector at the (t + 1)-th moment as the target feature vector at the (t + 1)-th moment, and setting t = t + 1, returning to the step of "predicting the target state at the (t + 1)-th moment using Kalman filtering based on the target position detection result at the t-th moment", until t = T, and assigning corresponding labels to the multiple tracking target objects in the images at each moment; tracking the multiple tracking target objects according to the labels. Embedding the spatial pyramid module into the multi-scale feature pyramid network enhances the feature extraction ability of the network. An adaptive feature update mechanism is used to reduce mis-matching and identity transformation during target tracking. The spatial pyramid feature extraction network and the adaptive feature update mechanism are used to improve the accuracy and robustness of multi-object tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 It is a flowchart of a multi-object tracking method based on deep learning provided in Embodiment 1 of the present invention;

[0030] Figure 2 It is a structural diagram of the spatial pyramid module provided in Embodiment 1 of the present invention;

[0031] Figure 3 The image at a certain moment of the target scenario provided in Embodiment 1 of the present invention;

[0032] Figure 4 The improved feature extraction network structure diagram provided in Embodiment 1 of the present invention;

[0033] Figure 5 The image containing the target detection frame provided in Embodiment 1 of the present invention;

[0034] Figure 6 The block diagram of a multi-object tracking system based on deep learning provided in Embodiment 2 of the present invention. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0036] The purpose of the present invention is to provide a multi-object tracking method and system based on deep learning, which optimizes the multi-object tracking network model by using a spatial pyramid and an adaptive feature update mechanism to improve the accuracy and robustness of multi-object tracking in complex scenarios.

[0037] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0038] Embodiment 1

[0039] As Figure 1 shown, this embodiment provides a multi-object tracking method based on deep learning, including:

[0040] S1: Embed the spatial pyramid module into the multi-scale feature pyramid network to obtain an improved feature extraction network;

[0041] Embedding the spatial pyramid module into the multi-scale feature pyramid network enhances the feature extraction ability of the network. The spatial pyramid integrates feature maps with different receptive fields into one layer of feature map, enabling the feature extraction ability of the feature network to be enhanced while also having feature information of multiple-scale targets. As Figure 2 shows the pyramid module structure.

[0042] S2: Obtain the images at T moments in the target scenario; the images include multiple targets, as Figure 3 shown.

[0043] S3: Input the images at T moments into the improved feature extraction network to obtain the target position detection results and target feature vectors at T moments; T = 1, 2,..., t - 1, t, t + 1,...;

[0044] Specifically, S3 specifically includes:

[0045] S3-1: After performing a convolution operation on the image at each moment, obtain a downsampled feature map;

[0046] S3-2: Perform multi-layer convolution operations on the downsampled feature map to obtain a first-scale feature map, a second-scale feature map, and a third-scale feature map;

[0047] Use a convolution operation with a 3x3 convolution kernel for feature extraction, and the calculation formula is as follows:

[0048] X w = conv i (conv i-1 (…conv 1 (X))) n > 2, 0 < i ≤ n

[0049] In the formula: X represents the input image, conv i represents the i-th layer convolution operation for feature extraction of the image, and X w represents the feature map extracted through multi-layer convolution operations;

[0050] S3-3: Process the third-scale feature map through a spatial pyramid module to obtain a processed third-scale feature map; the expression of the spatial pyramid module is:

[0051] Y w = X w + X w (max5×5) + X w (max9×9) + X w (max13×13)

[0052] Among them, X w represents the input feature map (the third-scale feature map); max5x5, max9x9, and max13x13 respectively represent max-pooling operations with 5x5, 9x9, and 13x13 as pooling kernels; "+" represents the concat operation; Y w represents the output feature map obtained through the spatial pyramid module.

[0053] S3-4: Upsample the processed third-scale feature map, and at the same time perform a skip connection on the second-scale feature map to obtain a processed second-scale feature map;

[0054] The upsampling fusion calculation expression is as follows:

[0055] p i = conv i (x 1 , cat(…, conv 2 (cat(x i-2 , up(conv 1 (cat(x i-1 , up(x i )))))))) n > 2, 0 < i ≤ n

[0056] In the formula, x i represents the feature map of the i-th layer of the improved feature extraction network, up represents the operation of upsampling the feature map, cat represents the splicing of the feature maps on the feature channels, and conv i represents the i-th layer convolution operation for feature extraction after splicing the feature maps, and p i represents the upsampling fusion of the feature map of the i-th layer obtained by feature extraction, n represents the number of layers of the improved feature extraction network;

[0057] S3-5: Upsample the processed second-scale feature map, and at the same time perform a skip connection on the first-scale feature map to obtain a processed first-scale feature map;

[0058] S3-6: Splice and fuse the processed first-scale feature map, the processed second-scale feature map, and the processed third-scale feature map, and then extract the target position detection result and target feature vector of each target in each image.

[0059] Combined Figure 4 The process of image processing using the improved feature extraction network is introduced in detail: An image with a resolution of 608x608x3 is input. After convolution operations with a stride of 2, it is downsampled to a feature map of 304x304. After multiple convolution operations, feature maps of 76x76, 38x38, and 19x19 are finally obtained. Then the 19x19 feature map is processed by the spatial pyramid module SPP. The processed 19x19 feature map is sampled, and at the same time, a skip connection operation is performed on the 38x38 feature map to obtain a processed 38x38 feature map. The processed 38x38 feature map is upsampled, and at the same time, a skip connection operation is performed on the 76x76 feature map to obtain a processed 76x76 feature map. Finally, the feature map obtained by splicing and fusing the three feature maps of different scales (the processed 19x19 feature map, the processed 38x38 feature map, and the processed 76x76 feature map) is sent to the output end of the network, enabling the network to achieve good results for both large and small targets.Figure 4 In this, Track represents the target feature vector (the tracking branch output by the network); Det represents the target position detection result (the detection branch output by the network).

[0060] It should be noted that, as Figure 5 shown, anchor boxes are used to frame the target positions in the feature map, and non-maximum suppression (calculating the intersection over union) is used to filter the obtained target boxes. The target feature vector can be a one-dimensional vector with a length of 512.

[0061] S4: Based on the target position detection result at time t, use Kalman filtering to predict the target state at time t + 1, and obtain the target position prediction result at time t + 1; t ∈ 1, 2, 3..., T;

[0062] Among them, S4 specifically includes:

[0063] S4-1: Calculate the target position prediction value at time t + 1 according to the target position detection result at time t and the state transition matrix from time t to time t + 1; among them, the expression for calculating the target position prediction value at time t + 1 is:

[0064] x′ t+1 = Fx t + u

[0065] Among them, x′ t+1 represents the target position prediction value at time t + 1; F represents the state transition matrix, which contains time information; x t represents the target position detection result at time t, which contains target position information and speed information; u represents the noise in the Kalman filtering process;

[0066] In this step, also considering the uncertainty degree of the Kalman filtering system, it is necessary to predict the state covariance matrix representing the system uncertainty, and the expression for calculating the predicted state covariance matrix is:

[0067] P' = FPF T + Q

[0068] Among them, P is the state covariance matrix, representing the uncertainty degree of the system. This uncertainty degree is given a value during the Kalman filtering initialization. As more and more data is injected into the Kalman filter, the value of P will gradually decrease and finally stabilize within a range; P' is the predicted state covariance matrix; Q represents the noise that cannot be represented by u.

[0069] S4-2: Calculate the prediction error value according to the target position detection result at time t + 1 and the target position prediction value at time t + 1; among them, the expression for calculating the prediction error value is:

[0070] y = xt+1 -Hx′ t+1

[0071] where y represents the prediction error value; x t+1 represents the target position detection result at time t+1; H represents the measurement matrix, which is taken as the identity matrix in this embodiment;

[0072] S4-3: Calculate the predicted adjustment value of the target position at time t+1 according to the predicted value of the target position at time t+1, the prediction error value, and the Kalman gain; the predicted adjustment value of the target position at time t+1 is the predicted result of the target position at time t+1. Among them, the expression for calculating the predicted adjustment value of the target position at time t+1 is:

[0073] x″ t+1 =x′ t+1 +Ky

[0074] where x″ t+1 represents the predicted adjustment value of the target position at time t+1; K represents the Kalman gain.

[0075] In the formula, R is the observation noise matrix, representing the observation error, which is taken as the zero matrix in this embodiment;

[0076] Similarly, the adjustment value of the predicted state covariance matrix can be obtained, and its expression is:

[0077] P″=(I-KH)P', where I represents the identity matrix. Use this adjustment value for the prediction of the next cycle.

[0078] S5: Calculate the intersection over union of the predicted result of the target position at time t+1 and the detected result of the target position at time t+1 to obtain the target position intersection over union;

[0079] Specifically, the expression for calculating the intersection over union is:

[0080]

[0081] where IOU represents the intersection over union; A is the detection result box of the target output by the improved feature extraction network, B is the prediction result box of the target by Kalman filtering, ∩ is the symbol for taking the union, ∪ is the symbol for taking the union, and the intersection over union is the area of the intersection of AB divided by the area of the union of AB.

[0082] S6: Calculate the cosine similarity of the target features according to the target feature vector at time t and the target feature vector at time t+1;

[0083] Specifically, the expression for calculating the cosine similarity of the target features is:

[0084]

[0085]

[0086] Among them, α is the cosine similarity of the target feature; a represents the target feature vector at time t; |a| represents the modulus of the target feature vector at time t; b represents the target feature vector at time t + 1; |b| represents the modulus of the target feature vector at time t + 1.

[0087] S7: Screen out multiple tracked target objects according to the target location intersection over union and the target feature cosine similarity, and assign the same label to the same tracked target object;

[0088] Furthermore, screen out the targets with the target location intersection over union greater than the first preset value and the target feature cosine similarity greater than the second preset value as the tracked target objects. The first preset value and the second preset value are set according to requirements. The first preset value can be set to 0.6, and the second preset value can be set to 0.7.

[0089] The label can be an ID, that is, set the corresponding ID for the tracked target, and the ID values assigned to the same tracked target in the images at each moment are consistent. Use the Hungarian algorithm to match the targets in the front and rear frames (the targets in the image at time t and the targets in the image at time t + 1), assign a unique ID to each target, and keep the ID unchanged for each subsequent frame of the image.

[0090] S8: Adaptively weight the target feature vector at time t + 1 according to the target feature cosine similarity in combination with the adaptive weighting mechanism to obtain the weighted target feature vector at time t + 1, use the weighted target feature vector at time t + 1 as the target feature vector at time t + 1, and let t = t + 1, then return to step S4 until t = T, and assign the corresponding labels to the multiple tracked target objects in the images at each moment;

[0091] Among them, the expression for obtaining the weighted target feature vector at time t + 1 is:

[0092] F′ t+1 =(1 - α)F t +αF t+1

[0093] Among them, F′ t+1 represents the weighted target feature vector at time t + 1; F t represents the target feature vector at time t; F t+1 represents the target feature vector at time t + 1.

[0094] Adaptive weighting is performed on the target feature vector at time t and the target feature vector at time t + 1, and the weighted target feature vector at time t + 1 is used to replace the target feature vector at time t + 1 to participate in the calculation of the target feature cosine similarity between time t + 1 and time t + 2 in the subsequent process.

[0095] S9: Tracking multiple said tracked objects according to the said label.

[0096] In this embodiment, the spatial pyramid module is embedded into the multi-scale feature pyramid network to obtain an improved feature extraction network, and this network is used to extract image features, enhancing the network's ability to extract features, so that the large and small targets in the image can be accurately extracted. At the same time, during the target tracking process, the adaptive feature update mechanism is used to adaptively weight and update the target feature vector. The target feature vectors used during the tracking process are all adaptively weighted, which can reduce the mis-matching and identity transformation during target tracking, improve the target tracking effect, and enhance the accuracy and robustness of multi-target tracking.

[0097] Embodiment 2

[0098] As Figure 6 shown, this embodiment provides a multi-target tracking system based on deep learning, including:

[0099] A feature extraction network acquisition module M1, configured to embed a spatial pyramid module into a multi-scale feature pyramid network to obtain an improved feature extraction network;

[0100] An image acquisition module M2, configured to acquire images at T moments in a target scene; multiple targets are included in the images;

[0101] A feature extraction module M3, configured to input the images at T moments into the improved feature extraction network to obtain target position detection results at T moments and target feature vectors at T moments; T = 1, 2,..., t - 1, t, t + 1,...;

[0102] A target position prediction module M4, configured to predict the target state at time t + 1 using Kalman filtering based on the target position detection result at time t to obtain the target position prediction result at time t + 1; t ∈ 1, 2, 3..., T;

[0103] An intersection over union calculation module M5, configured to calculate the intersection over union of the target position prediction result at time t + 1 and the target position detection result at time t + 1 to obtain the target position intersection over union;

[0104] A cosine similarity calculation module M6, configured to calculate the target feature cosine similarity according to the target feature vector at time t and the target feature vector at time t + 1;

[0105] A target screening and labeling module M7 is configured to screen out multiple tracked target objects according to the target location intersection over union and the target feature cosine similarity, and assign the same label to the same tracked target object;

[0106] An adaptive weighting and state update module M8 is configured to adaptively weight the target feature vector at time t + 1 according to the target feature cosine similarity in combination with an adaptive weighting mechanism to obtain the weighted target feature vector at time t + 1, use the weighted target feature vector at time t + 1 as the target feature vector at time t + 1, and let t = t + 1, then return to execute the target position prediction module; until t = T, corresponding labels are assigned to multiple tracked target objects in each image at each moment;

[0107] A target tracking module M9 is configured to track multiple tracked target objects according to the labels.

[0108] For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For related parts, refer to the description in the method section.

[0109] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A multi-object tracking method based on deep learning, characterized in that, Including: Embedding a spatial pyramid module into a multi-scale feature pyramid network to obtain an improved feature extraction network; Obtaining images at T moments in the target scene; the images include multiple targets; Inputting the images at T moments into the improved feature extraction network to obtain target position detection results at T moments and target feature vectors at T moments; T = 1, 2,..., t - 1, t, t + 1,...; Based on the target position detection result at time t, using Kalman filtering to predict the target state at time t + 1 to obtain the target position prediction result at time t + 1; t ∈ 1, 2, 3..., T; Calculating the intersection over union of the target position prediction result at time t + 1 and the target position detection result at time t + 1 to obtain the target position intersection over union; Calculating the target feature cosine similarity according to the target feature vector at time t and the target feature vector at time t + 1; Screening out multiple tracking target objects according to the target position intersection over union and the target feature cosine similarity, and assigning the same label to the same tracking target object; Adaptive weighting the target feature vector at time t + 1 according to the target feature cosine similarity in combination with an adaptive weighting mechanism to obtain the weighted target feature vector at time t + 1, taking the weighted target feature vector at time t + 1 as the target feature vector at time t + 1, and setting t = t + 1, returning to the step "Based on the target position detection result at time t, using Kalman filtering to predict the target state at time t + 1", until t = T, and assigning corresponding labels to multiple tracking target objects in the images at each moment; Tracking multiple tracking target objects according to the labels; Wherein, the inputting the images at T moments into the improved feature extraction network to obtain target position detection results at T moments and target feature vectors at T moments specifically includes: After performing a convolution operation on the image at each moment, obtaining a downsampled feature map; Performing multi-layer convolution operations on the downsampled feature map to obtain a first-scale feature map, a second-scale feature map, and a third-scale feature map; Processing the third-scale feature map through a spatial pyramid module to obtain a processed third-scale feature map; Upsampling the processed third-scale feature map, and simultaneously performing a skip connection on the second-scale feature map to obtain a processed second-scale feature map; Upsampling the processed second-scale feature map, and simultaneously performing a skip connection on the first-scale feature map to obtain a processed first-scale feature map; Performing splicing and fusion on the processed first-scale feature map, the processed second-scale feature map, and the processed third-scale feature map, and then extracting the target position detection result and target feature vector of each target; Wherein, the expression of the spatial pyramid module is: Y w = X w + X w (max 5×5)+ X w (max 9×9)+ X w (max 13×13) Among them, X w represents the input feature map; max5×5, max9×9, and max13×13 respectively represent the maximum pooling operations with 5×5, 9×9, and 13×13 as the pooling kernels; "+" represents the concat operation; Y w represents the output feature map obtained through the spatial pyramid module.

2. The method according to claim 1, characterized in that, The using Kalman filtering to predict the target state at time t + 1 based on the target position detection result at time t to obtain the target position prediction result at time t + 1 specifically includes: Calculate the predicted value of the target position at time t+1 based on the target position detection result at time t and the state transition matrix from time t to time t+1; Calculate the prediction error value based on the target position detection result at time t+1 and the predicted value of the target position at time t+1; Calculate the predicted adjustment value of the target position at time t+1 based on the predicted value of the target position at time t+1, the prediction error value, and the Kalman gain; the predicted adjustment value of the target position at time t+1 is the predicted result of the target position at time t+1.

3. The method according to claim 2, characterized in that, The expression for calculating the predicted value of the target position at time t+1 based on the target position detection result at time t and the state transition matrix from time t to time t+1 is: x t ′ +1 = Fx t + u where x t ′ +1 represents the predicted value of the target position at time t + 1; F represents the state transition matrix; x t represents the detection result of the target position at time t; u represents the noise in the Kalman filtering process; The expression for calculating the prediction error value based on the target position detection result at time t+1 and the predicted value of the target position at time t+1 is: y = x t+1 -Hx t ′ +1 where y represents the prediction error value; x t+1 represents the target position detection result at time t + 1; H represents the measurement matrix; The expression for calculating the predicted adjustment value of the target position at time t+1 based on the predicted value of the target position at time t+1, the prediction error value, and the Kalman gain is: x t ″ +1 = x t ′ +1 + Ky Among them, x t ″ +1 represents the predicted adjustment value of the target position at time t + 1; K represents the Kalman gain.

4. The method according to claim 1, characterized in that, The expression for calculating the intersection over union of the predicted result of the target position at time t+1 and the detected result of the target position at time t+1 is: Where, IOU represents the intersection over union; A is the detection result box of the target output by the improved feature extraction network, B is the predicted result box of the target by Kalman filtering, ∩ is the symbol for taking the union, ∪ is the symbol for taking the union, and the intersection over union is the area of the intersection of AB divided by the area of the union of AB.

5. The method according to claim 1, wherein The expression for calculating the cosine similarity of the target features based on the target feature vector at time t and the target feature vector at time t+1 is: Among them, α is the cosine similarity of the target feature; represents the target feature vector at time t; represents the norm of the target feature vector at time t; b represents the target feature vector at time t + 1; |b| represents the norm of the target feature vector at time t + 1.

6. The method according to claim 1 or 4 or 5, wherein Filter out multiple tracking target objects based on the intersection over union of the target positions and the cosine similarity of the target features, specifically including: Filter out the target whose intersection over union of the target positions is greater than the first preset value and the cosine similarity of the target features is greater than the second preset value as the tracking target object.

7. The method according to claim 5, wherein The expression for adaptively weighting the target feature vector at time t+1 based on the cosine similarity of the target features and the adaptive weighting mechanism to obtain the weighted target feature vector at time t+1 is: F t ′ +1 = (1 - α)F t + αF t+1 Among them, F t ′ +1 represents the weighted target feature vector at time t + 1; F t represents the target feature vector at time t; F t+1 represents the target feature vector at time t + 1.

8. A multi - object tracking system based on deep learning, wherein Including: A feature extraction network acquisition module, used to embed a spatial pyramid module into a multi-scale feature pyramid network to obtain an improved feature extraction network; An image acquisition module, used to acquire images at T moments in the target scene; multiple targets are included in the images; A feature extraction module, used to input the images at T moments into the improved feature extraction network to obtain the target position detection results at T moments and the target feature vectors at T moments; T = 1, 2,..., t-1, t, t+1,...; Where, inputting the images at T moments into the improved feature extraction network to obtain the target position detection results at T moments and the target feature vectors at T moments specifically includes: After performing a convolution operation on the image at each moment, obtain a downsampled feature map; Perform a multi-layer convolution operation on the downsampled feature map to obtain a first-scale feature map, a second-scale feature map, and a third-scale feature map; Process the third-scale feature map through a spatial pyramid module to obtain a processed third-scale feature map; Upsample the processed third-scale feature map, and perform a skip connection on the feature map of the second scale at the same time to obtain a processed second-scale feature map; Upsample the processed second-scale feature map, and perform a skip connection on the feature map of the first scale at the same time to obtain a processed first-scale feature map; Perform splicing and fusion on the processed first-scale feature map, the processed second-scale feature map, and the processed third-scale feature map, and then extract the target position detection result and the target feature vector of each target; Among them, the expression of the spatial pyramid module is: Y w = X w + X w (max 5×5) + X w (max 9×9) + X w (max 13×13) Among them, X w represents the input feature map; max5×5, max9×9, and max13×13 respectively represent the maximum pooling operations with 5×5, 9×9, and 13×13 as the pooling kernels; "+" represents the concat operation; Y w represents the output feature map obtained through the spatial pyramid module; A target position prediction module, which is used to predict the target state at the t+1 moment by using Kalman filtering based on the target position detection result at the t moment to obtain the target position prediction result at the t+1 moment; t∈1,2,3...,T; An intersection over union calculation module, which is used to calculate the intersection over union of the target position prediction result at the t+1 moment and the target position detection result at the t+1 moment to obtain the target position intersection over union; A cosine similarity calculation module, which is used to calculate the target feature cosine similarity according to the target feature vector at the t moment and the target feature vector at the t+1 moment; A target screening and labeling module, which is used to screen out multiple tracked target objects according to the target position intersection over union and the target feature cosine similarity, and assign the same label to the same tracked target object; An adaptive weighting and state update module, which is used to adaptively weight the target feature vector at the t+1 moment by combining the adaptive weighting mechanism according to the target feature cosine similarity to obtain the weighted target feature vector at the t+1 moment, and use the weighted target feature vector at the t+1 moment as the target feature vector at the t+1 moment, and let t=t+1, and return to execute the target position prediction module; until t=T, corresponding labels are assigned to multiple tracked target objects in each moment of the image; A target tracking module, which is used to track multiple tracked target objects according to the label.

Citation Information

Patent Citations

  • Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field

    AU2020103901A4

  • Multi-target tracking method and system suitable for embedded terminal

    CN113034548A