Multi-target tracking method and system

Through the combination of the timing adaptive convolution module and the deep learning skeleton network, the timing information and depth features in the video data are extracted, and the two-part graph correlation model is constructed, which solves the problem of low accuracy of the existing multi-objective tracking method and achieves higher precision multi-objective tracking.

CN120107304AActive Publication Date: 2025-06-06AEROSPACE SCI & IND GRP INTELLIGENT TECH RES INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311639933.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06
Estimated Expiration
2043-12-04

AI Technical Summary

Technical Problem

The tracking accuracy of existing multi-objective tracking methods is low, and it is impossible to fully extract the differences between different targets and the commonality of the same target.

Method used

The timing adaptive convolution module is used to perform timing convolution operations on the video data, obtain a convolution kernel containing timing information, and input it into the deep learning skeleton network for multiple perception and feature extraction, and build a two-part graph correlation model to obtain multi-objective tracking correlation results.

Benefits of technology

It improves the accuracy of the multi-objective tracking algorithm in complex scenarios, and reduces the problems such as tracking failure and label switching caused by frequent interactions between targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107304A_ABST
    Figure CN120107304A_ABST
Patent Text Reader

Abstract

The invention provides a multi-target tracking method and system. The method comprises the following steps: taking a current frame and N-1 frames of images before the current frame in video data of a target to be tracked as a group of input images; n convolution kernels containing time sequence information are obtained; obtaining depth feature representation of a time sequence corresponding to the input image; obtaining a position thermodynamic diagram of the target center point; obtaining the position and size represented by the depth feature of the sequential sequence; constructing a target classification head and a target tracking feature head; training the deep learning skeleton network; obtaining the real position of the target; constructing a bipartite graph association model; and obtaining a multi-target tracking association result based on the bipartite graph association model. The technical problem that an existing tracking method is low in tracking precision can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision target tracking, and in particular to a multi-target tracking method and system. Background Art

[0002] As a research field with a long history and important application value, multi-target tracking plays a connecting role in the field of computer vision. Previous topics such as target segmentation and target detection distinguish the foreground and background of traffic video sequences, providing basic input information for target tracking; later more macro-level topics such as scene understanding and behavior analysis all need to be based on tracking information to carry out research.

[0003] Multi-target tracking methods usually consist of multiple sequentially executed steps, which mainly include input video sequence, detection and feature extraction, and feature association matching. According to the time series of input information required to generate the tracking sequence, it is divided into online tracking and offline tracking. Among them, online tracking can only use historical information and the current frame to obtain the association result, while offline tracking can use the entire video sequence for trajectory association. According to the way of generating trajectories for multi-target tracking, it can be divided into tracking methods based on probabilistic reasoning and tracking methods based on deterministic reasoning; according to the representation of multi-target tracking trajectories from coarse to fine, it can be divided into point-based tracking, contour-based tracking, posture-based tracking, and mask-based tracking. The framework of multi-target tracking is mainly divided into two categories: detection-based tracking and joint detection tracking. Among them, the detection-based tracking method is greatly restricted by the detection effect. The poor detection results not only bring great interference to the feature extraction of the tracked target, but also introduce a lot of noise to the association process. In addition, the detection and tracking processes under this framework both involve feature extraction, which is actually a repeated operation of the features. Therefore, in comparison, the multi-target tracking framework of joint detection can comprehensively utilize multiple feature information to realize the linkage of tracking and detection features, thereby realizing an end-to-end tracking method and improving tracking accuracy.

[0004] After the rise of deep learning, detection methods based on automatic learning have become a trend in target detection. Deep neural networks simulate the neural transmission mode of the human brain, can extract rich features at multiple levels, and have strong nonlinear fitting capabilities. According to the existing visualization results, deep learning has an absolute advantage in extracting abstract features. Subsequently, with the improvement of computer computing power, the application of deep learning in the field of computer vision has greatly improved the performance of related tasks, achieving performance that is superior to most traditional detection algorithms. However, for multi-target tracking tasks, temporal information and target spatial position play an important role in feature extraction and discrimination. The existing feature extraction method based on single-frame information and the general deep learning feature extraction mode cannot fully extract the differences between different targets and the commonalities of the same target, and the tracking accuracy needs to be improved. Summary of the invention

[0005] The present invention provides a multi-target tracking method and system, which can solve the technical problem of low tracking accuracy of existing tracking methods.

[0006] According to one aspect of the present invention, a multi-target tracking method is provided, the method comprising:

[0007] The current frame and N-1 frames of images before the current frame in the video data of the target to be tracked are taken as a group of input images;

[0008] Use the time-series adaptive convolution module to perform time-series convolution operation on the input image to obtain N convolution kernels containing time-series information;

[0009] Input N convolution kernels containing time series information into the deep learning skeleton network to obtain the deep feature representation of the time series corresponding to the input image;

[0010] Perform multiple perceptions on the deep feature representation of the time series to obtain the location heat map of the target center point;

[0011] The length and width of the detection bounding box of the target center point are obtained based on the position heat map of the target center point, and the position and size of the deep feature representation of the time series are obtained based on the length and width of the detection bounding box of the target center point;

[0012] Perform multiple perceptions on the deep feature representation of the time series and the size scaling relationship between the input image and the deep feature representation of the time series, and construct a target classification head and a target tracking feature head;

[0013] The deep learning skeleton network is trained. During the training process, the position offset of the target center point is obtained based on the scale relationship between the input image and the deep feature representation of the time series, the category of the target is obtained based on the target classification head, and the tracking characteristic representation of the target is obtained based on the target tracking feature head;

[0014] The real position of the target is obtained based on the position represented by the deep feature of the time series and the position offset of the target center point;

[0015] A bipartite graph association model is constructed based on the position and size of the deep feature representation of the time series, the position offset of the target center point, the target category, the tracking characteristic representation of the target and the historical trajectory;

[0016] Obtain multi-target tracking association results based on the bipartite graph association model.

[0017] Preferably, obtaining the multi-target tracking association result based on the bipartite graph association model includes:

[0018] The bipartite graph association model predicts the position of the historical trajectory in the current frame through Kalman;

[0019] Get the motion similarity between the target’s real position and the historical trajectory’s position in the current frame;

[0020] Obtain the appearance similarity between the tracking feature representation of the current frame target and the tracking feature representation of the historical trajectory;

[0021] Obtaining a feature similarity matrix based on motion similarity and appearance similarity;

[0022] The feature similarity matrix is ​​converted into a cost matrix, and the multi-target tracking association results are obtained by minimizing the cost matrix.

[0023] Preferably, training the deep learning skeleton network includes:

[0024] Obtain the heatmap loss of the deep learning skeleton network, the bias and size loss of the detection bounding box, and the target tracking feature acquisition loss;

[0025] The total loss of the deep learning skeleton network is obtained by weighted summing the heat map loss of the deep learning skeleton network, the loss of the detection bounding box bias and box size, and the target tracking feature acquisition loss;

[0026] The deep learning skeleton network is trained based on its total loss.

[0027] According to another aspect of the present invention, a multi-target tracking system is provided, the system comprising:

[0028] The detection module is used to take the current frame and N-1 frames of images before the current frame in the video data of the target to be tracked as a group of input images; to use the time-series adaptive convolution module to perform a time-series convolution operation on the input image to obtain N convolution kernels containing time-series information; to input the N convolution kernels containing time-series information into the deep learning skeleton network to obtain a deep feature representation of the time-series sequence corresponding to the input image; to perform multi-perception on the deep feature representation of the time-series sequence to obtain a position heat map of the target center point; to obtain the length and width of the detection bounding box of the target center point based on the position heat map of the target center point, and to obtain the length and width of the detection bounding box of the target center point based on the length and width of the detection bounding box of the target center point. The method is used to obtain the position and size of the deep feature representation of the time series; to perform multiple perceptions on the deep feature representation of the time series and the size scaling relationship between the input image and the deep feature representation of the time series, and to construct a target classification head and a target tracking feature head; to train the deep learning skeleton network, and during the training process, obtain the position offset of the target center point based on the size scaling relationship between the input image and the deep feature representation of the time series, obtain the category of the target based on the target classification head, and obtain the tracking characteristic representation of the target based on the target tracking feature head; and to obtain the true position of the target based on the position of the deep feature representation of the time series and the position offset of the target center point;

[0029] The tracking module is used to construct a bipartite graph association model based on the position and size of the deep feature representation of the time series, the position offset of the target center point, the target category, the tracking characteristic representation of the target and the historical trajectory; it is also used to obtain multi-target tracking association results based on the bipartite graph association model.

[0030] According to another aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a multi-target tracking program stored in the memory and executable on the processor, wherein the processor implements any of the above methods when executing the multi-target tracking program.

[0031] By applying the technical solution of the present invention, N convolution kernels containing temporal information are obtained through a temporal adaptive convolution module, so that the convolution kernels adaptively give different attention to images of different frames, so that spatial convolution can perform temporal reasoning; at the same time, the module pays more attention to the change state of the foreground position, obtains more attention weights for possible foreground positions, improves the model's sensitivity to the temporal position change of the target, and uses the image features of the first N-1 frames as attention inspiration to help the model obtain a more accurate feature representation of the last frame image. The present invention provides a multi-target tracking method based on a spatiotemporal attention mechanism in complex scenes, improves the accuracy of the multi-target tracking algorithm in complex scenes, and reduces problems such as tracking failure and label switching caused by frequent interactions between targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The included drawings are used to provide a further understanding of the embodiments of the present invention, which constitute a part of the specification, are used to illustrate the embodiments of the present invention, and together with the text description, explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0033] Figure 1 A block diagram of a multi-target tracking method provided according to an embodiment of the present invention is shown;

[0034] Figure 2 A flowchart of a multi-target tracking method provided according to an embodiment of the present invention is shown;

[0035] Figure 3 A flowchart of a timing adaptive convolution according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0036] It should be noted that, in the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present invention and its application or use. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0038] Unless otherwise specifically stated, the relative arrangement of the parts and steps described in these embodiments, numerical expressions and numerical values ​​do not limit the scope of the present invention. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the accompanying drawings are not drawn according to the actual proportional relationship. The technology, method and equipment known to ordinary technicians in the relevant field may not be discussed in detail, but in appropriate cases, the technology, method and equipment should be regarded as a part of the authorization specification. In all examples shown and discussed here, any specific value should be interpreted as being merely exemplary, rather than as a limitation. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once a certain item is defined in an accompanying drawing, it does not need to be further discussed in subsequent drawings.

[0039] like Figure 1-Figure 3 As shown, the present invention provides a multi-target tracking method, the method comprising:

[0040] S10, taking the current frame and N-1 frames of images before the current frame in the video data of the target to be tracked as a group of input images;

[0041] S20, using a time-series adaptive convolution module to perform a time-series convolution operation on the input image to obtain N convolution kernels containing time-series information;

[0042] S30, inputting N convolution kernels containing time series information into a deep learning skeleton network to obtain a deep feature representation of the time series sequence corresponding to the input image;

[0043] S40, performing multiple perceptions on the deep feature representation of the time series to obtain a position heat map of the target center point;

[0044] S50, obtaining the length and width of the detection bounding box of the target center point based on the position heat map of the target center point, and obtaining the position and size of the depth feature representation of the time series based on the length and width of the detection bounding box of the target center point; wherein the detection bounding box can be a rectangular bounding box;

[0045] S60, performing multi-perception on the deep feature representation of the time series and the size scaling relationship between the input image and the deep feature representation of the time series, and constructing a target classification head and a target tracking feature head;

[0046] S70, training the deep learning skeleton network. During the training process, obtaining the position offset of the target center point based on the size scaling relationship between the input image and the deep feature representation of the time series, obtaining the category of the target based on the target classification head, and obtaining the tracking characteristic representation of the target based on the target tracking feature head;

[0047] S80, obtaining the real position of the target based on the position represented by the depth feature of the time series and the position offset of the target center point;

[0048] S90, constructing a bipartite graph association model based on the position and size of the deep feature representation of the time series, the position offset of the target center point, the category of the target, the tracking characteristic representation of the target and the historical trajectory;

[0049] S100, obtaining multi-target tracking association results based on a bipartite graph association model.

[0050] The present invention obtains N convolution kernels containing temporal information through a temporal adaptive convolution module, so that the convolution kernel adaptively gives different attention to images of different frames, so that spatial convolution can perform temporal reasoning; at the same time, the module pays more attention to the change state of the foreground position, obtains more attention weights for possible foreground positions, improves the model's sensitivity to the temporal position change of the target, and uses the image features of the first N-1 frames as attention inspiration to help the model obtain a more accurate feature representation of the last frame image. The present invention provides a multi-target tracking method based on a spatiotemporal attention mechanism in complex scenes, improves the accuracy of the multi-target tracking algorithm in complex scenes, and reduces problems such as tracking failure and label switching caused by frequent interactions between targets.

[0051] According to an embodiment of the present invention, in S70 of the present invention, training the deep learning skeleton network includes:

[0052] S71, obtaining the heat map loss of the deep learning skeleton network, the bias and box size loss of the detection bounding box, and the target tracking feature acquisition loss;

[0053] S72, performing weighted summation on the heat map loss of the deep learning skeleton network, the loss of the bias and box size of the detection bounding box, and the target tracking feature acquisition loss to obtain the total loss of the deep learning skeleton network;

[0054] S73. Train the deep learning skeleton network based on the total loss of the deep learning skeleton network.

[0055] According to an embodiment of the present invention, in S100 of the present invention, obtaining multi-target tracking association results based on the bipartite graph association model includes:

[0056] S101, the bipartite graph association model predicts the position of the historical trajectory in the current frame through Kalman;

[0057] S102, obtaining the motion similarity between the real position of the target and the position of the historical trajectory in the current frame;

[0058] S103, obtaining the appearance similarity between the tracking characteristic representation of the current frame target and the tracking characteristic representation of the historical trajectory;

[0059] S104, obtaining a feature similarity matrix based on motion similarity and appearance similarity;

[0060] S105, the feature similarity matrix is ​​converted into a cost matrix, and the multi-target tracking association result is obtained by minimizing the cost matrix.

[0061] The present invention also provides a multi-target tracking system, the system comprising:

[0062] The detection module is used to take the current frame and N-1 frames of images before the current frame in the video data of the target to be tracked as a group of input images; to use the time-series adaptive convolution module to perform a time-series convolution operation on the input image to obtain N convolution kernels containing time-series information; to input the N convolution kernels containing time-series information into the deep learning skeleton network to obtain a deep feature representation of the time-series sequence corresponding to the input image; to perform multi-perception on the deep feature representation of the time-series sequence to obtain a position heat map of the target center point; to obtain the length and width of the detection bounding box of the target center point based on the position heat map of the target center point, and to obtain the length and width of the detection bounding box of the target center point based on the length and width of the detection bounding box of the target center point. The method is used to obtain the position and size of the deep feature representation of the time series; to perform multiple perceptions on the deep feature representation of the time series and the size scaling relationship between the input image and the deep feature representation of the time series, and to construct a target classification head and a target tracking feature head; to train the deep learning skeleton network, and during the training process, obtain the position offset of the target center point based on the size scaling relationship between the input image and the deep feature representation of the time series, obtain the category of the target based on the target classification head, and obtain the tracking characteristic representation of the target based on the target tracking feature head; and to obtain the true position of the target based on the position of the deep feature representation of the time series and the position offset of the target center point;

[0063] The tracking module is used to construct a bipartite graph association model based on the position and size of the deep feature representation of the time series, the position offset of the target center point, the target category, the tracking characteristic representation of the target and the historical trajectory; it is also used to obtain multi-target tracking association results based on the bipartite graph association model.

[0064] In order to further understand the present invention, the following Figure 1-Figure 3 The multi-target tracking method of the present invention is described in detail, and specifically comprises the following steps:

[0065] The first step is to set up a deep learning skeleton network structure, using a typical deep learning image feature extraction network DLA-34, whose main function is to output multi-level feature maps with multi-scale receptive fields. Figure 2 The receptive fields of the feature maps in are C1: 1 / 4, C2: 1 / 8, C3: 1 / 16, and C4: 1 / 32 respectively.

[0066] In the second step, based on the DLA-34 deep learning skeleton network structure, the temporal adaptive convolution module is added to each state of DLA-34, and the convolution kernel is adaptively adjusted along the time dimension, so that the model can adaptively assign different convolution kernels to different frames of the image sequence, so that the spatial convolution can perform temporal reasoning, and the model has both temporal and spatial attention with almost no additional computation.

[0067] In the third step, the current frame and the three frames of images before the current frame in the video data of the target to be tracked are input into the temporal adaptive convolution module as a group of input images.

[0068] The fourth step is to obtain a more interference-resistant deep feature representation based on DLA-34 and the time-series adaptive convolution feature extraction module. The features of the first three frames of images are used as attention as inspiration to help the model obtain a more accurate feature representation of the last frame of the image, that is, the deep feature representation of the time series.

[0069] The fifth step is to perform multi-perception on the deep feature representation of the time series to obtain the position heat map of the target center point.

[0070] Step 6: Based on the position heat map of the target center point, the length and width of the rectangular bounding box of the target center point are obtained, and based on the length and width of the rectangular bounding box of the target center point, the position and size of the deep feature representation of the time series are obtained;

[0071] Step 7: Perform multi-perception on the deep feature representation of the time series and the size scaling relationship between the input image and the deep feature representation of the time series to construct a target classification head and a target tracking feature head;

[0072] Step 8: Obtain the heat map loss of the deep learning skeleton network, the loss of the bias and size of the detection bounding box, and the target tracking feature acquisition loss;

[0073] Among them, the loss function of the position heat map of the target center point is as follows:

[0074]

[0075] Where, L heatmap Represents the loss of the location heat map of the target center point, is the predicted heat map, M xy is the true value of the heat map, N is the total number of targets in a single image, α and β are the first and second hyperparameters, x and y are the horizontal and vertical pixels of the target;

[0076] The loss function for detecting the bias and size of the bounding box is as follows:

[0077]

[0078] Where, L box represents the loss of the detection bounding box bias and box size, N is the total number of targets in a single image, i is the target number to be calculated, s represents the size of the predicted target detection box, represents the true value size of the target detection box, o represents the deviation of the predicted center point, represents the deviation of the true center point, ||·|| represents the sum of the absolute errors between the two vectors;

[0079] The target tracking feature acquisition loss function is as follows:

[0080]

[0081] Where, L identity represents the target tracking feature acquisition loss, N is the total number of targets in a single image, K is the number of categories, p(·) is the predicted target classification feature vector, and L i (·) is the true classification one-hot encoding vector of the i-th target.

[0082] In the ninth step, the heat map loss of the deep learning skeleton network, the bias and box size loss of the detection bounding box, and the target tracking feature acquisition loss are weighted and summed to obtain the total loss of the deep learning skeleton network. The deep learning skeleton network is trained based on the total loss of the deep learning skeleton network. During the training process, the position offset of the target center point is obtained based on the size scaling relationship between the input image and the deep feature representation of the time series, and the target category is obtained based on the target classification head. Here, the softmax loss training model is used to learn the target classification task; the tracking feature representation of the target is obtained based on the target tracking feature head.

[0083] Step 10: Get the real position of the target based on the position represented by the deep feature of the time series and the position offset of the target center point;

[0084] In the eleventh step, a bipartite graph association model is constructed based on the position and size of the deep feature representation of the time series, the position offset of the target center point, the target category, the tracking characteristic representation of the target and the historical trajectory; and the multi-target tracking association results are obtained based on the bipartite graph association model.

[0085] The present invention also provides a computer device, comprising a memory, a processor, and a multi-target tracking program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned methods when executing the multi-target tracking program.

[0086] In summary, the present invention provides a multi-target tracking method and system, which obtains N convolution kernels containing temporal information through a temporal adaptive convolution module, so that the convolution kernel adaptively gives different attention to images of different frames, so that spatial convolution can perform temporal reasoning; at the same time, the module pays more attention to the change state of the foreground position, obtains more attention weights for possible foreground positions, improves the model's sensitivity to the temporal position change of the target, and uses the image features of the first N-1 frames as attention inspiration to help the model obtain a more accurate feature representation of the last frame image. The present invention provides a multi-target tracking method based on a spatiotemporal attention mechanism in complex scenes, improves the accuracy of the multi-target tracking algorithm in complex scenes, and reduces problems such as tracking failure and label switching caused by frequent interactions between targets.

[0087] Parts of the present invention that are not described in detail are well known to those skilled in the art.

[0088] In the description of the present invention, it is necessary to understand that the directions or positional relationships indicated by directional words such as "front, back, up, down, left, right", "lateral, vertical, perpendicular, horizontal" and "top, bottom" are usually based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description. Unless otherwise specified, these directional words do not indicate or imply that the devices or elements referred to must have a specific direction or be constructed and operated in a specific direction. Therefore, they cannot be understood as limiting the scope of protection of the present invention. The directional words "inside and outside" refer to the inside and outside relative to the contours of each component itself.

[0089] For ease of description, spatially relative terms such as "above", "above", "on the upper surface of", "above", etc. may be used here to describe the spatial positional relationship between a device or feature and other devices or features as shown in the figure. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figure. For example, if the device in the accompanying drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below other devices or structures". Thus, the exemplary term "above" can include both "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatially relative descriptions used here are interpreted accordingly.

[0090] In addition, it should be noted that the use of terms such as "first" and "second" to limit components is only for the convenience of distinguishing the corresponding components. If not otherwise stated, the above terms have no special meaning and therefore cannot be understood as limiting the scope of protection of the present invention.

[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-target tracking method, It is characterized in that The method comprises: The current frame and N-1 frames of images before the current frame in the video data of the target to be tracked are taken as a group of input images; Use the time-series adaptive convolution module to perform time-series convolution operation on the input image to obtain N convolution kernels containing time-series information; Input N convolution kernels containing time series information into the deep learning skeleton network to obtain the deep feature representation of the time series corresponding to the input image; Perform multiple perceptions on the deep feature representation of the time series to obtain the location heat map of the target center point; The length and width of the detection bounding box of the target center point are obtained based on the position heat map of the target center point, and the position and size of the deep feature representation of the time series are obtained based on the length and width of the detection bounding box of the target center point; Perform multiple perceptions on the deep feature representation of the time series and the size scaling relationship between the input image and the deep feature representation of the time series, and construct a target classification head and a target tracking feature head; The deep learning skeleton network is trained. During the training process, the position offset of the target center point is obtained based on the scale relationship between the input image and the deep feature representation of the time series, the category of the target is obtained based on the target classification head, and the tracking characteristic representation of the target is obtained based on the target tracking feature head; The real position of the target is obtained based on the position represented by the deep feature of the time series and the position offset of the target center point; A bipartite graph association model is constructed based on the position and size of the deep feature representation of the time series, the position offset of the target center point, the target category, the tracking characteristic representation of the target and the historical trajectory; Obtain multi-target tracking association results based on the bipartite graph association model.

2. The method according to claim 1, It is characterized in that The multi-target tracking association results obtained based on the bipartite graph association model include: The bipartite graph association model predicts the position of the historical trajectory in the current frame through Kalman; Get the motion similarity between the target’s real position and the historical trajectory’s position in the current frame; Obtain the appearance similarity between the tracking feature representation of the current frame target and the tracking feature representation of the historical trajectory; Obtaining a feature similarity matrix based on motion similarity and appearance similarity; The feature similarity matrix is ​​converted into a cost matrix, and the multi-target tracking association results are obtained by minimizing the cost matrix.

3. The method according to claim 1 or 2, It is characterized in that Training a deep learning skeleton network involves: Obtain the heatmap loss of the deep learning skeleton network, the bias and size loss of the detection bounding box, and the target tracking feature acquisition loss; The total loss of the deep learning skeleton network is obtained by weighted summing the heat map loss of the deep learning skeleton network, the loss of the detection bounding box bias and box size, and the target tracking feature acquisition loss; The deep learning skeleton network is trained based on its total loss.

4. A multi-target tracking system, It is characterized in that The system comprises: The detection module is used to take the current frame and N-1 frames of images before the current frame in the video data of the target to be tracked as a group of input images; to use the time-series adaptive convolution module to perform a time-series convolution operation on the input image to obtain N convolution kernels containing time-series information; to input the N convolution kernels containing time-series information into the deep learning skeleton network to obtain a deep feature representation of the time-series sequence corresponding to the input image; to perform multi-perception on the deep feature representation of the time-series sequence to obtain a position heat map of the target center point; to obtain the length and width of the detection bounding box of the target center point based on the position heat map of the target center point, and to obtain the length and width of the detection bounding box of the target center point based on the length and width of the detection bounding box of the target center point. The method is used to obtain the position and size of the deep feature representation of the time series; to perform multiple perceptions on the deep feature representation of the time series and the size scaling relationship between the input image and the deep feature representation of the time series, and to construct a target classification head and a target tracking feature head; to train the deep learning skeleton network, and during the training process, obtain the position offset of the target center point based on the size scaling relationship between the input image and the deep feature representation of the time series, obtain the category of the target based on the target classification head, and obtain the tracking characteristic representation of the target based on the target tracking feature head; and to obtain the true position of the target based on the position of the deep feature representation of the time series and the position offset of the target center point; The tracking module is used to construct a bipartite graph association model based on the position and size of the deep feature representation of the time series, the position offset of the target center point, the target category, the tracking characteristic representation of the target and the historical trajectory; it is also used to obtain multi-target tracking association results based on the bipartite graph association model.

5. A computer device, It is characterized in that The method comprises a memory, a processor and a multi-target tracking program stored in the memory and executable on the processor, wherein the processor implements any one of the methods of claims 1 to 3 when executing the multi-target tracking program.

Citation Information

Patent Citations

  • Target tracking method, system, device and medium based on two-stream convolution neural network

    CN109410242A

  • Multi-target tracking method based on graph network

    CN111881840A

  • Target tracking method based on time sequence adaptive convolution and attention mechanism

    CN115147456A

  • Unmanned aerial vehicle target tracking prediction method and system

    CN116862948A

  • Methods and systems for object tracking

    EP4254267A1