Multi-target tracking method and device, equipment and storage medium
By obtaining the fusion feature map of adjacent image frames in drone videos, and combining embedded identity features and adaptive association modes, the problem of insufficient feature extraction in multi-objective tracker in multi-scale shooting videos is solved, and the accuracy of multi-objective tracking is improved.
Patent Information
- Application Number
- CN202510609187.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Prior Art In multi-scale drone videos shot at multi-scale, multi-target trackers cannot extract sufficient features, resulting in the impact of detection and tracking accuracy.
By obtaining the fusion feature map of adjacent image frames, the fusion feature map features fusion of the pyramid feature map by setting weights, updating the target detection feature with embedded identity features, and using the adaptive association mode for target association.
It improves the accuracy of multi-objective tracking, can effectively extract and match target features in drone videos, and enhances the detection and tracking capabilities of multi-scale targets.
Smart Images

Figure CN120147365A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a multi-object tracking method, device, equipment and storage medium. Background Art
[0002] The main task of multi-object tracking is to output the trajectories of all objects from a given video and maintain the identity information of each object. Among them, the tracking objects can be pedestrians, vehicles or other objects. With the development of computer vision technology, multi-object tracking has been widely applied in many fields, such as video intelligent monitoring and human-computer interaction. In addition, multi-object tracking is the basis for computer vision tasks such as pose estimation, behavior recognition, behavior analysis, and video analysis.
[0003] Traditional multi-object tracking methods mainly include Markov decision-making, joint probability association, particle filtering, etc. These methods have large prediction position errors and low robustness to occlusion and interference from similar objects. With the wide application of deep learning in the field of computer vision, multi-object tracking can be realized based on deep learning. Specifically, a multi-object tracker is used to identify the objects in each frame of the video sequence, extract the target features, and then perform association according to the feature data. However, due to factors such as irregular motion caused by the movement of the camera and view changes, it is difficult for the multi-object tracker to accurately extract and match the target features. Especially for the drone video taken at multiple scales, the multi-object tracker cannot extract enough features for reliable detection and tracking, which affects the accuracy of multi-object tracking. Summary of the Invention In view of this, this application provides a multi-object tracking method, device, equipment and storage medium, mainly aiming to solve the problem that in the prior art, for the drone video taken at multiple scales, the multi-object tracker cannot extract enough features for reliable detection and tracking, which affects the accuracy of multi-object tracking.
[0004] According to the first aspect of this application, a multi-object tracking method is provided, including: Obtain a fused feature map of adjacent image frames, where the fused feature map is obtained by fusing pyramid feature maps with set weights; Perform object detection on the fused feature map to obtain object detection features. The object detection identification at least includes object boundary features, object identity features, and object probability features. The object identity feature is a feature vector used to distinguish objects, and the object probability feature is a visual image feature used to locate objects. The visual image feature represents the object probability through the depth of color; Update the target detection feature using the embedded identity feature to obtain an embedded updated target detection feature. The embedded identity feature is extracted from the preset image frame, and the preset image frame is the image frame with a lower sequence number among adjacent image frames. Perform target association on the embedded updated target detection feature using an adaptive association mode to obtain the target correspondence in adjacent image frames.
[0005] Further, the obtaining of the fused feature map of adjacent image frames includes: Use a pre-trained feature extraction model to stack the feature maps obtained by sampling adjacent image frames using different paths in multiple layers to obtain a stacked image feature; During the multi-layer stacking process, add and fuse the image features with the same resolution in different layers through lateral connections to obtain a pyramid feature map; Perform weighted fusion on the pyramid feature map according to the set weights to obtain the fused feature map of adjacent image frames.
[0006] Further, before performing target detection on the fused feature map to obtain the target detection feature, the method further includes: Use a depthwise separable convolution composed of a depthwise convolution and a pointwise convolution for model training to obtain a feature detection model. One convolution kernel of the depthwise convolution is responsible for one channel, and the number of input channels and output channels corresponding to the feature detection model is the same. Through convolution operations, perform weighted summation on the fused feature map in terms of the number of channels, and correspondingly output target detection features consistent with the number of convolution kernels; Correspondingly, use the pre-trained feature detection model to perform target detection on the fused feature map to obtain the target detection feature.
[0007] Further, before using the embedded identity feature to update the target detection feature to obtain the embedded updated target detection feature, the method further includes: Use the image frame with a lower sequence number among adjacent image frames as the preset image frame, and extract the embedded identity feature in the preset image frame; Correspondingly, the using the embedded identity feature to update the target detection feature to obtain the embedded updated target detection feature includes: Select a preset number of key position points in the target probability feature. The key position points are the position points in the target probability feature whose probability ranks before a preset value; Extract a preset number of key point features from the embedded identity feature according to the key point positions in the target probability feature; After compressing the key point features, use the compressed key point features to update the target detection features to obtain the embedded and updated target detection features.
[0008] Further, the step of after compressing the key point features, using the compressed key point features to update the target detection features to obtain the embedded and updated target detection features includes: Perform an association operation on the adjacent image frames to obtain feature-enhanced attention weights, and the attention weights are used to guide the key point features that the image frame with a lower order in the adjacent image frames should focus on according to the image frame with a higher order in the adjacent image frames; After compressing the key point features, perform a weighted sum of the compressed key point features according to the feature-enhanced attention weights to obtain the attention feature of the image frame with a higher order in the adjacent image frames; Use the attention feature to update the target detection features to obtain the embedded and updated target detection features.
[0009] Further, the step of using the attention feature to update the target detection features to obtain the embedded and updated target detection features includes: Use the attention feature to perform feature association on the image frame with a higher order and the image frame with a lower order in the adjacent image frames to obtain the association feature of the adjacent image frames; Update the target detection features according to the association feature to obtain the embedded and updated target detection features.
[0010] Further, the step of using an adaptive association mode to perform target association on the embedded and updated target detection features to obtain the target correspondence in the adjacent image frames includes: Use Kalman filtering to calculate the target matching coefficient between the adjacent image frames; If the target matching coefficient is greater than a preset threshold, use the first association mode to perform target association on the embedded and updated target detection features to obtain the target correspondence in the adjacent image frames; If the target matching coefficient is less than or equal to the preset threshold, use the second association mode to perform target association on the embedded and updated target detection features to obtain the target correspondence in the adjacent image frames.
[0011] According to the second aspect of the present application, there is provided a multi-target tracking device, including: An acquisition unit, configured to acquire a fused feature map of adjacent image frames, where the fused feature map is obtained by performing feature fusion on a pyramid feature map with a set weight; The detection unit is used to perform object detection on the fused feature map to obtain object detection features. The object detection identification at least includes object boundary features, object identity features, and object probability features. The object identity feature is a feature vector used to distinguish objects, and the object probability feature is a visual image feature used to locate objects. The visual image feature characterizes the object probability through the depth of color. The update unit is used to update the object detection features by using the embedded identity features to obtain the embedded-updated object detection features. The embedded identity features are extracted from the preset image frame, and the preset image frame is the image frame with a previous sequence number among adjacent image frames. The association unit is used to perform object association on the embedded-updated object detection features by using an adaptive association mode to obtain the object correspondence in adjacent image frames.
[0012] Further, the obtaining unit is specifically used for: Using a pre-trained feature extraction model, stack the feature maps obtained by sampling adjacent image frames using different paths in multiple layers to obtain stacked image features; During the multi-layer stacking process, add and fuse the image features with the same resolution in different layers through lateral connections to obtain a pyramid feature map; Perform weighted fusion on the pyramid feature map according to the set weights to obtain the fused feature map of adjacent image frames.
[0013] Further, the device further includes: The training unit is used to perform model training using depthwise separable convolution composed of pointwise convolution and depthwise convolution before performing object detection on the fused feature map to obtain object detection features. One convolution kernel of the pointwise convolution is responsible for one channel. The number of input channels and output channels corresponding to the feature detection model is the same. Through convolution operations, perform weighted summation on the fused feature map in terms of the number of channels, and correspondingly output object detection features consistent with the number of convolution kernels; Correspondingly, the detection unit is specifically used to perform object detection on the fused feature map by using the pre-trained feature detection model to obtain object detection features.
[0014] Further, the device further includes: The extraction unit is used to extract the embedded identity features in the preset image frame with the image frame with a previous sequence number among adjacent image frames as the preset image frame before updating the object detection features by using the embedded identity features to obtain the embedded-updated object detection features. Correspondingly, the update unit is specifically used for: Select a preset number of key position points from the target probability features, where the key position points are the position points in the target probability features whose probability rankings are before a preset value; Extract a preset number of key point features from the embedded identity features according to the key point positions in the target probability features; After compressing the key point features, use the compressed key point features to update the target detection features to obtain the embedded and updated target detection features.
[0015] Further, the updating unit is specifically further configured to: Perform an association operation on the adjacent image frames to obtain an attention weight with enhanced features, where the attention weight is used to guide the key point features that the image frame with a later sequence number in the adjacent image frames should focus on according to the image frame with an earlier sequence number in the adjacent image frames; After compressing the key point features, perform a weighted sum of the compressed key point features according to the attention weight with enhanced features to obtain the attention feature of the image frame with an earlier sequence number in the adjacent image frames; Use the attention feature to update the target detection features to obtain the embedded and updated target detection features.
[0016] Further, the updating unit is specifically further configured to: Use the attention feature to perform feature association on the image frame with an earlier sequence number and the image frame with a later sequence number in the adjacent image frames to obtain the association feature of the adjacent image frames; Update the target detection features according to the association feature to obtain the embedded and updated target detection features.
[0017] Further, the association unit is specifically configured to: Calculate the target matching coefficient between the adjacent image frames by using Kalman filtering; If the target matching coefficient is greater than a preset threshold, perform target association on the embedded and updated target detection features by using a first association mode to obtain the target correspondence in the adjacent image frames; If the target matching coefficient is less than or equal to the preset threshold, perform target association on the embedded and updated target detection features by using a second association mode to obtain the target correspondence in the adjacent image frames.
[0018] According to the third aspect of the present application, a computer device is provided, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect are implemented.
[0019] According to the fourth aspect of the present application, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect above are implemented.
[0020] A multi-object tracking method, device, equipment and storage medium provided by the present application, compared with the current method of multi-object tracking based on deep learning in the prior art, the present application obtains a fused feature map of adjacent image frames, and the fused feature map is obtained by fusing the pyramid feature maps with set weights; performs object detection on the fused feature map to obtain object detection features. The object detection identification at least includes object boundary features, object identity features and object probability features. The object identity feature is a feature vector used to distinguish objects, and the object probability feature is a visual image feature used to locate objects. The visual image feature represents the object probability through the depth of color; updates the object detection features with the embedded identity features to obtain the embedded updated object detection features. The embedded identity features are extracted from a preset image frame, and the preset image frame is the image frame with a previous ordinal number among the adjacent image frames; uses an adaptive filtering mode to perform object association on the embedded updated object detection features to obtain the object correspondence relationship in the adjacent image frames, and obtains the object correspondence relationship in the adjacent image frames. The whole process shortens the information path based on the pyramid features, constructs a fused feature map containing semantic features and position information of different scales through weighted fusion, and this fused feature map can provide sufficient features for the multi-object tracker to improve the accuracy of the multi-object tracker. Since the object detection features will change with the change of the position of the acquisition device, by learning the embedded identity features of adjacent image frames, the object detection features can be adaptively updated, and the object association can be accurately completed using the adaptive filtering mode, improving the multi-object tracking ability.
[0021] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. Description of the Drawings
[0022] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings: Figure 1 is a schematic flowchart of a multi-object tracking method in an embodiment of the present application; Figure 2 is Figure 1 a schematic flowchart of a specific implementation manner of step 101 in Figure 3It is a schematic flowchart of a multi-object tracking method in another embodiment of the present application; Figure 4 It is a schematic flowchart of a multi-object tracking method in yet another embodiment of the present application; Figure 5 It is Figure 1 a schematic flowchart of a specific implementation manner of step 104 in Figure 6 It is a structural block diagram of a multi-object tracking method in an embodiment of the present application; Figure 7 It is a schematic flowchart of a multi-object tracking method in another embodiment of the present application; Figure 8 It is a pyramid feature map in an embodiment of the present application; Figure 9 It is a schematic diagram of feature fusion in an embodiment of the present application; Figure 10 It is a schematic diagram of embedding and updating target detection features in an embodiment of the present application; Figure 11 It is a schematic diagram of gradient-balanced focal loss in an embodiment of the present application; Figure 12 It is a schematic structural diagram of a multi-object tracking device in an embodiment of the present application; Figure 13 It is a schematic diagram of the device structure of a computer device provided by an embodiment of the present invention. Specific implementation manner
[0023] Now, the content of the present invention will be described with reference to several exemplary embodiments. It should be understood that these embodiments are described only to enable those of ordinary skill in the art to better understand and thus implement the content of the present invention, rather than to imply any limitation to the scope of the present invention.
[0024] As used herein, the term "comprising" and its variants are to be construed as open-ended terms meaning "including but not limited to". The term "based on" is to be construed as "at least partially based on". The term "one embodiment" and "an embodiment" are to be construed as "at least one embodiment". The term "another embodiment" is to be construed as "at least one other embodiment".
[0025] In the related art, due to factors such as irregular motion caused by the movement of the camera and view changes, it is difficult for a multi-object tracker to accurately extract and match target features. Especially for the video of an unmanned aerial vehicle taken at multiple scales, the multi-object tracker cannot extract sufficient features for reliable detection and tracking, affecting the accuracy of multi-object tracking.
[0026] To solve this problem, this embodiment provides a multi-object tracking method, as Figure 1As shown in the figure, it includes the following steps: 101. Obtain the fused feature map of adjacent image frames.
[0027] Among them, adjacent image frames refer to two frames of images that are temporally continuous and sequentially adjacent in a video sequence or image stream. For example, in a video captured by a drone, the image frames are arranged in chronological order, and the nth image frame and the (n + 1)th image frame here are adjacent image frames.
[0028] In this embodiment, the fused feature map is obtained by performing feature fusion on the pyramid feature map with set weights, where the set weights are the weights assigned to the feature maps of different layers in the pyramid feature map. Specifically, different weights can be assigned to the feature maps of different layers according to the characteristics and importance of the targets in the adjacent image frames. For example, for a scenario with many small targets, the weight of the bottom layer feature map can be appropriately increased; for a scenario that requires attention to the overall structure and context information, the weight of the upper layer feature map can be increased.
[0029] Specifically, an image pyramid can be constructed for adjacent image frames. An image pyramid is a multi-scale image representation method. By performing multiple downsampling and filtering operations on the original image, image layers with different resolutions are obtained. Then, a neural network is used to extract features from each layer of the pyramid to extract representative features from images at different scales. For the feature map of each layer, the feature maps corresponding to the two frames of images can be weighted and summed according to the weights to obtain the fused feature map. For example, let the and th layer pyramid feature maps of and be and respectively, and the corresponding weights be and , then the fused th layer feature map is
[0030] 102. Perform object detection on the fused feature map to obtain object detection features.
[0031] In this embodiment, the object detection identification at least includes object boundary features, object identity features, and object probability features. The object identity feature is a feature vector used to distinguish objects. Here, the object identity feature can be an ID feature used to represent the object identity. The object probability feature is a visual image feature used to locate the object. The visual image feature represents the object probability through the depth of color. Here, the visual image feature can be a heat map used to represent the position probability.
[0032] During the process of specifically performing object detection on the fused feature map, a pre-trained object detection model can be used, which is trained using a large-scale multi-object image dataset. The dataset needs to include bounding box annotations, class annotations, and unique ID annotations of the objects. During the training process, the model will learn to extract the features of the objects from the feature fusion map and perform supervised learning based on the annotation information, aiming to minimize the loss function to adjust the parameters of the model, so that the model can accurately predict the size, class, and ID features of the object's bounding box. Then, the object detection model is used to perform object detection on the feature map to obtain the object boundary features, object identity features, and object probability features.
[0033] For the object boundary features, the object detection model usually uses a regression algorithm to predict the position and size of the bounding box. For example, by processing the features of the candidate regions, a fully connected layer and a linear regression function are used to predict the offsets and scale changes of the bounding box relative to the candidate regions, so as to obtain the final bounding box coordinates. For the object probability features, the object detection model can generate visual image features by further processing the features within the bounding box. For the heat map, the average value or weighted average value of the features within the bounding box can be calculated and used as the heat value of this region, and the heat values corresponding to all bounding boxes are plotted on a heat map with the same size as the image. For the object identity features, when training the object detection model, appropriate network structures can be designed to learn the identity features of the objects. For example, a dedicated branch can be added to the model to extract the identity feature vector of the object. This branch can be composed of multiple convolutional layers and fully connected layers, taking the feature fusion map as the input and outputting an identity feature vector with a fixed length. During the training process, by comparing the identity feature vectors of different objects, the model can learn the differences between different objects, thus realizing the unique identification of the objects.
[0034] 103. Update the object detection features using the embedded identity features to obtain the embedded updated object detection features.
[0035] In this embodiment, the embedded identity features are extracted from a preset image frame. Here, the embedded identity features are feature vectors that can indicate the feature attributes or identities of the objects in the image, and can be obtained by mapping the object information in the image frame to a low-dimensional vector space. The preset image frame is the image frame with a lower ordinal number among adjacent image frames.
[0036] It can be understood that the embedded identity feature has the ability to uniquely identify the target in the image frame. Even if the target appears in different image frames, its embedded identity feature is relatively stable, facilitating the tracking and recognition of the target. For example, in a pedestrian tracking scenario, each person has unique appearance, posture and other characteristics. By extracting the embedded identity feature, these features can be transformed into a vector, so that the feature vectors of the same pedestrian in different frames have a high degree of similarity, while the feature vectors of different pedestrians have large differences.
[0037] Specifically, the feature similarity between the image frame with a lower order and the image frame with a higher order in adjacent image frames can be calculated using the embedded identity feature. According to the similarity result, the adjacent image frames are feature-associated, and then the feature of the associated image frame is used to update the target detection feature, obtaining the embedded-updated target detection feature. The process of updating the target feature here can use weighted average update or confidence update.
[0038] For the weighted average update method, if the feature association of the target in adjacent image frames is successful, the weighted average method can be used to update the target detection feature. Suppose the embedded identity feature of the image frame with a lower order in adjacent image frames is and the detection feature of the target in the current frame is and the updated target detection feature is , then the update formula is , where is the weight coefficient, and its value range is in . The value of can be adaptively adjusted according to factors such as the motion state of the target and the degree of appearance change. For example, when the target moves slowly and the appearance change is small, can take a larger value; when the target moves fast or the appearance change is large,
[0039] For the confidence update method, the confidence information of target detection can be combined to update the target detection feature. If the confidence of target detection in the current image frame is high, it indicates that the detection result is relatively reliable, and the weight of the embedded identity feature corresponding to the image frame with a lower order in adjacent image frames can be appropriately reduced; conversely, if the confidence is low, the weight of the embedded identity feature corresponding to the image frame with a lower order in adjacent image frames can be increased to improve the accuracy of target detection.
[0040] 104. Use the adaptive association mode to perform target association on the embedded-updated target detection feature to obtain the target correspondence in adjacent image frames.
[0041] In this embodiment, affected by the motion mode of the acquisition device, the target motion can be linear or non-linear. For example, the non-linear motion formed by the coupling of the acquisition device and the target itself. In different motion modes, the target in the video acquired by the acquisition device will present different motion states. It is difficult for traditional Kalman filters to handle such irregular motion modes. In this embodiment, an adaptive filtering mode is used to process the image frame features of the target in different motion states, and the target association can be accurately completed.
[0042] Specifically, the target matching value between two adjacent image frames can be calculated, and the filtering mode applicable to the image frame for target association can be determined according to the target matching value, and then the corresponding filtering mode is further used to perform target association on the image frame.
[0043] For example, the motion state of the target includes three types. The first motion state is applicable to the A filtering mode, the second motion state is applicable to the B filtering mode, and the third motion state is applicable to the C filtering mode. In the first motion state, the target matching value is less than m. In the second motion state, the target value is greater than m and less than n. In the third motion state, the target value is greater than n. Among them, m is less than n. Correspondingly, if the target matching value is less than m, the A filtering mode is used for target association. If the target matching value is greater than m and less than n, the B filtering mode is used for target association. If the target matching value is greater than n, the A filtering mode is used for target association.
[0044] The multi-object tracking method provided by the embodiments of the present application, compared with the existing method of multi-object tracking based on deep learning in the current art, the present application obtains a fused feature map of adjacent image frames, and the fused feature map is obtained by fusing the pyramid feature maps with set weights; performs object detection on the fused feature map to obtain object detection features, and the object detection identification at least includes object boundary features, object identity features, and object probability features. The object identity feature is a feature vector used to distinguish objects, and the object probability feature is a visual image feature used to locate objects. The visual image feature characterizes the object probability through the depth of color; uses the embedded identity feature to update the object detection feature to obtain an embedded updated object detection feature. The embedded identity feature is extracted from a preset image frame, and the preset image frame is the image frame with a previous order in the adjacent image frames; uses an adaptive filtering mode to perform object association on the embedded updated object detection feature to obtain the object correspondence relationship in the adjacent image frames, and obtains the object correspondence relationship in the adjacent image frames. The entire process shortens the information path based on the pyramid features, constructs a fused feature map containing semantic features and position information of different scales through weighted fusion, and the fused feature map can provide sufficient features for the multi-object tracker to improve the accuracy of the multi-object tracker. Since the object detection features will change with the change of the position of the acquisition device, by learning the embedded identity features of adjacent image frames, the object detection features can be adaptively updated, and the object association can be accurately completed using the adaptive filtering mode to improve the multi-object tracking ability.
[0045] In an actual application scenario, the pyramid feature map is a hierarchical feature representation formed by performing feature extraction and fusion on image frames at different scales. As the number of network layers increases, the receptive field gradually increases and can capture feature information at different scales. In the above embodiment, specifically, as Figure 2 shown, step 101 includes the following steps: 201. Use a pre-trained feature extraction model to stack the feature maps obtained by sampling adjacent image frames using different paths in multiple layers to obtain a stacked image feature.
[0046] 202. During the multi-layer stacking process, add and fuse the image features with the same resolution in different layers through lateral connections to obtain a pyramid feature map.
[0047] 203. Perform weighted fusion on the pyramid feature map according to the set weights to obtain a fused feature map of adjacent image frames.
[0048] In this embodiment, the pre-trained feature extraction model is used to perform feature sampling on adjacent image frames at different resolutions. The sampling path here can be top-down, bottom-up, and then top-down again. During the process of multi-layer stacking, lateral connection and convolutional fusion can be used to perform weighted fusion on the image features to obtain a pyramid feature map, where the pyramid feature map consists of image features at different scales.
[0049] Specifically, during the process of top-down sampling of the feature map, it usually starts from the high-level features of the network. High-level features have strong semantic information but low resolution. Through upsampling (such as deconvolution and other operations), the resolution of the high-level feature map is gradually increased to match the resolution of the low-level feature map. During this process, the semantic information of the high-level features can gradually spread to the layers with higher resolution, which helps to detect small targets. Because small targets may have richer location information but weaker semantic information in the low-level features, and the semantic information of the high-level features can make up for this.
[0050] Specifically, during the process of bottom-up sampling of the feature map, it starts from the bottom layer of the network. The bottom-layer feature map has a high resolution and can provide accurate target location information, but relatively less semantic information. Through operations such as pooling or convolution, the resolution of the feature map is gradually reduced, while increasing the semantic abstraction degree of the features, enabling the network to capture a larger range of semantic information, which is beneficial for detecting large targets.
[0051] Specifically, during the process of top-down sampling of the feature map again, performing the top-down operation again can further refine the features, spread more advanced semantic information to lower layers, strengthen the fusion and refinement of target features at different scales, and make the final feature map richer and more accurate in both semantic and location information.
[0052] It can be understood that during the top-down and bottom-up processes, feature maps with the same resolution can be fused through lateral connection. Lateral connection can directly combine the location information of the bottom-layer features and the semantic information of the high-level features, avoiding the loss of information during the transmission process, and effectively retaining the details and overall information of the target.
[0053] Correspondingly, after the feature map fusion, convolutional operations are usually used to process the fused feature map. Convolutional operations can extract and fuse features without adding too much computational complexity, enabling features from different sources to interact better and extract more representative features. Through appropriate convolutional kernel design and parameter adjustment, convolutional fusion can effectively integrate semantic information and location information, while reducing the influence of noise and redundant information.
[0054] Specifically, in the process of weighted fusion of pyramid feature maps, the weights are usually set according to actual needs, and the initial weights can be determined through experiments, experience, or based on some prior knowledge. For each level of feature maps, multiply it by the corresponding weight to obtain the weighted feature map, and then add the weighted feature maps of each level to obtain the fused feature map. When adding, it is necessary to ensure that the feature maps of different levels have the same size and number of channels. If they are not the same, some adjustment operations may be required first, such as cropping, padding, or convolutional transformation, etc., to enable them to be added correctly.
[0055] The above-mentioned fused image features adaptively learn the spatial weights for fusing feature maps at each scale based on the pyramid feature maps, which can make full use of features at different scales and improve the multi-scale object detection ability.
[0056] In the above embodiment, further, as Figure 3 shown, before step 102, the above Figure 1 multi-object tracking method further includes the following steps: 301. Use depthwise separable convolution composed of depthwise convolution and pointwise convolution for model training to obtain a feature detection model.
[0057] Correspondingly, in step 102, use the pre-trained feature detection model to perform object detection on the fused feature map to obtain object detection features.
[0058] In this embodiment, one convolution kernel of the depthwise convolution is responsible for one channel, and the number of input channels and output channels of the feature detection model is the same. Through convolution operation, weighted summation is performed on the fused feature map in terms of the number of channels, and the corresponding output is the object detection feature consistent with the number of convolution kernels. The advantage of this is that it greatly reduces the computational amount brought by the simultaneous operation of a large number of convolution kernels on multiple channels in the convolution operation.
[0059] For the obtained fused feature map, it can be processed using the convolution operation in the feature detection model. Here, the convolution operation first performs depthwise convolution on each channel of the fused feature map respectively to extract local feature information on each channel. Then, pointwise convolution is performed. Since the pointwise convolution uses a 1×1 convolution kernel, its role is to perform a linear combination on the feature map after depthwise convolution in the channel dimension to achieve the integration of information from different channels.
[0060] After performing per-channel convolution and pointwise convolution operations, the final object detection features are obtained by performing a weighted sum operation on the fused feature map in terms of the number of channels. Here, the weighted sum means assigning corresponding weights according to the degree of influence of the convolution kernel on features of different channels during the operation and performing a summation calculation. The number of the finally output object detection features is the same as the number of convolution kernels. These features contain the comprehensive information of each channel in the fused feature map and have been weighted, enabling them to be more effectively used for the object detection task.
[0061] Correspondingly, the fused feature map is input into a pre-trained feature detection model. Since the feature detection model has learned how to identify and extract information related to the target from image features during the previous training process, different branches or output layers of the feature detection model will respectively make predictions for the object detection features, outputting object boundary features, object identity features, and object probability features.
[0062] The specific feature detection model can predict the object boundary features through regression, that is, the position and size of the object in the image, usually a set of coordinate values to define the bounding box of the object. That is to say, the feature detection model will learn the mapping relationship between features such as the shape and position of the object and the bounding box coordinates.
[0063] The above object identity features can be heatmaps representing probabilities, used to indicate the possibility of the presence of an object at each position in the image. Correspondingly, the feature detection model will output a heatmap of one or more channels, and each channel may correspond to different object categories or attributes. The high-value regions in the heatmap indicate a higher probability of the presence of an object, and the low-value regions indicate a lower probability of the presence of an object. Through the heatmap, the distribution of the object in the image can be intuitively understood.
[0064] To distinguish different object instances, the identity features can help accurately track and identify each object in a multi-object scenario. The feature detection model will predict the identity features of each object, usually by learning the unique features of the object, such as the appearance and texture of the object.
[0065] In a multi-object tracking scenario, different objects may have similar appearances or features, making it difficult to accurately distinguish. Embedding identity features can capture the unique identity information of the object, provide a unique identifier for each object, and significantly enhance the ability to distinguish different objects. Further, as Figure 4 shown, before step 103, the above Figure 1 multi-object tracking method further includes the following steps: 401. Use the image frame with a lower sequence number among the adjacent image frames as the preset image frame, and extract the embedded identity features in the preset image frame.
[0066] Correspondingly, step 103 includes the following steps: 402. Select a preset number of key position points from the target probability features.
[0067] 403. Extract a preset number of key point features from the embedded identity features according to the key point positions in the target probability features.
[0068] 404. After compressing the key point features, use the compressed key point features to update the target detection features to obtain the embedded and updated target detection features.
[0069] In this embodiment, a convolutional neural network or a pre-trained deep learning model with the ability to extract identity features can be used. Input a preset image frame into the deep learning model. The last layer or a specific intermediate layer of the deep learning model will output a fixed-length feature vector, and this vector is the extracted embedded identity feature.
[0070] Among them, the key position points are the position points in the target probability features whose probability rankings are before a preset value, and correspondingly, a preset number of key point features are extracted from the embedded identity features. For example, k position points are selected for the probability feature ranking, and k key point features are extracted from the embedded identity features. Specifically, the embedded identity features can be input into a pre-trained model to extract the key point features of the embedded identity features through the model. It is also possible to construct a template library containing different key point features, match the embedded identity features with the templates in the template library, and determine the key point features of the embedded identity features according to the key point positions of the most matching template.
[0071] Specifically, during the process of updating the target detection features, perform an association operation on adjacent image frames to obtain attention weights with enhanced features. The attention weights are used to guide the key point features that the image frame with a later sequence number in the adjacent image frames should focus on according to the image frame with an earlier sequence number in the adjacent image frames; after compressing the key point features, perform a weighted sum on the compressed key point features according to the attention weights with enhanced features to obtain the attention features of the image frame with an earlier sequence number in the adjacent image frames; use the attention features to update the target detection features to obtain the embedded and updated target detection features.
[0072] Specifically, the attention features can be used to perform feature association on the image frame with an earlier sequence number and the image frame with a later sequence number in the adjacent image frames to obtain the association features of the adjacent image frames; update the target detection features according to the association features to obtain the embedded and updated target detection features.
[0073] In this embodiment, during the update process of the target detection features, the features of the image frames with earlier ordinal numbers and the features of the image frames with later ordinal numbers among adjacent image frames are actively extracted and associated to obtain the attention weights with enhanced features. The specific association process comprehensively considers factors such as the similarity of features, positional relationship, and motion trend. For example, if a certain target in the image frame with an earlier ordinal number has a position shift, pose change, or partial occlusion in the image frame with a later ordinal number, a complex matching and reasoning mechanism can be used to determine whether the target in the image frame with a later ordinal number is the same as that in the image frame with an earlier ordinal number.
[0074] Furthermore, for each element of the compressed key-point features, it is weighted according to the attention weights. Specifically, each weight value in the attention weight matrix is multiplied by the corresponding key-point feature element, and then the results of all multiplications are summed. Through the method of weighted summation, the key-point features with higher importance will occupy a larger proportion in the attention features, highlighting the information that is significant for target detection. Since the target detection features have different levels of features, in the shallow-level features, the attention features can help enhance the details and edge information of the target, and in the deep-level features, the attention features pay more attention to fusing semantic information and improving the ability to judge the target category and overall structure. Through the multi-level update of the target detection features, the target detection features can have the ability to represent targets at different scales.
[0075] In the actual application scenario, considering the dynamic change features of the targets in adjacent image frames, in order to track the dynamic change features of the targets in real time, an adaptive filtering mode can be used to effectively extract and associate the target features. Specifically, as Figure 5 shown, step 104 includes the following steps: 501. Calculate the target matching coefficient between the adjacent image frames using the Kalman filter.
[0076] 502. If the target matching coefficient is greater than the preset threshold, use the first association mode to perform target association on the embedded-updated target detection features to obtain the target correspondence in the adjacent image frames.
[0077] 503. If the target matching coefficient is less than or equal to the preset threshold, use the second association mode to perform target association on the embedded-updated target detection features to obtain the target correspondence in the adjacent image frames.
[0078] In this embodiment, the Kalman filter is an algorithm that uses a linear system state equation to optimally estimate the system state through system input and output observation data. Specifically, when calculating the target matching coefficient between adjacent image frames, the image frame with a higher sequence number among the adjacent image frames is used as the previous image frame, and the image frame with a higher sequence number is used as the current image frame. First, the target state and covariance of the current image frame can be predicted based on the target state and state transition matrix of the previous image frame. Then, after obtaining the target measurement value of the current image frame, the measurement residual and residual covariance are calculated to obtain the Kalman gain, which is used to update the target state and covariance. Finally, by calculating the Euclidean distance or Mahalanobis distance between the target states of adjacent image frames, the distance is converted into a target matching coefficient. The smaller the distance, the larger the target matching coefficient, and the target matching degree can be measured by the target matching coefficient.
[0079] It should be noted that in the UAV video sequence, the object motion is no longer linear motion, but non-linear motion formed by the coupling of UAV motion and the object itself. In the normal motion mode, the UAV flies smoothly and normally in the sky, and the objects in the video can be regarded as approximately linear motion. In the abnormal motion mode, the UAV will suddenly rotate or accelerate, and the objects in the video show a non-linear motion. For this irregular motion, different association modes can be adaptively switched according to the motion mode of the UAV for target association.
[0080] Specifically, if the target matching coefficient is greater than the preset threshold, it indicates that the UAV motion is in the normal motion mode. The first association mode can be used to perform target association on the embedded and updated target detection features to obtain the target correspondence relationship in adjacent image frames. Here, the first association mode can use the intersection over union (IoU) for target association. This association mode measures the overlap degree and similarity between two targets by calculating the ratio of the intersection area to the union area of two bounding boxes, and realizes the association matching of targets between different frames or different detection results. If the target matching coefficient is less than or equal to the preset threshold, it indicates that the UAV motion is in the abnormal motion mode, and the second association mode can be used. Here, the second association mode can use local relation filtering for association. This association mode can model and analyze local features and relationships, and can capture the essential features of the target and its relative position relationship in the scene. In this way, even if the target undergoes some pose changes or small displacements in different image frames, its local features and surrounding local relationships still have a certain degree of stability and coherence, so that the association of the target between different frames can be realized based on this information.
[0081] In the actual application scenario, Figure 6A structural block diagram of a multi-object tracking method is provided, which includes a feature extraction part, a detection part, and a tracking part. The main purpose is to track the object correspondence from adjacent image frames of a given video sequence and obtain the identity information of each object. Among them, the object can be a pedestrian, a vehicle, or other objects in the image frame. The feature extraction part adaptively learns the spatial weights of the fusion of feature maps at each scale based on the pyramid feature map, performs feature fusion on adjacent image frames, and obtains a fused feature map. This fused feature map can make full use of features at different scales and improve the detection ability of multi-scale objects. The detection part uses a network framework of depth-separable convolution to perform object detection on the fused feature map and obtain object detection features. The tracking part includes an embedded feature update module and an adaptive association module. For the embedded feature update module, the object detection features can be updated using the embedded identity features learned from adjacent image frames, so that the updated object detection features are fused with the object identity features to improve the object tracking ability. For the adaptive association module, different association modes can be adaptively selected according to the motion mode of the drone. For a drone in normal motion mode, the IoU association mode can be used, that is, the intersection over union is used for object association on the object detection features updated by embedding. For a drone in abnormal motion mode, the local relationship filtering association mode is used, that is, the local relationship filtering is used for object association on the object detection features updated by embedding, and the object association of adjacent image frames is completed.
[0082] Correspondingly, in combination with Figure 6 the structural block diagram of the multi-object tracking method in, a flow schematic diagram of a multi-object tracking method is provided. For details, see Figure 7 As shown, first, adjacent image frames are input into the convolutional neural network of the object detection algorithm. Different-scale pyramid feature maps are constructed by the way of addition and fusion in the convolutional neural network, and feature fusion is performed with a certain weight to obtain a fused feature map containing multi-scale features. Then, the feature fusion map is input into the detection part for object detection to obtain the object bounding box, heat map, and identity features. Further, the object bounding box, heat map, and identity features are updated with embedded identities to obtain the updated object bounding box, heat map, and identity features. Finally, the adaptive association mode is used to perform object association on the updated object bounding box, heat map, and identity features, and the object size, object identity, and class information are output.
[0083] Correspondingly, in the process of constructing the pyramid feature map, as Figure 8 shown, the convolutional neural network upsamples the small feature map at the top layer to the same size as the feature map of the previous layer, and once again transmits the first semantic information from bottom to top, from low dimension to high dimension. In this way, both the strong semantic features at the top layer and the high-resolution information at the bottom layer are utilized. Lateral connections are used to obtain three different-scale pyramid feature maps from the features of the previous layer that have the same resolution as the current layer after downsampling.
[0084] Correspondingly, during the process of feature fusion of the pyramid feature map, as Figure 9 shown, after sampling three pyramid feature maps of different scales, the three pyramid feature maps of different scales can be respectively input into three adaptive feature fusion blocks ASFF-1, ASFF-2, and ASFF-3, and the pyramid feature maps are subjected to feature fusion with certain weights to obtain a fused feature map containing multi-scale features.
[0085] Correspondingly, during the process of embedding and updating the object detection features, as Figure 10 shown, this process includes three stages. First, for a preset image frame with a lower ordinal number among adjacent image frames, the embedded identity features of the preset image frame are extracted , and here the corresponding TopK points in the heat map can be selected to extract the embedded identity features of K key point features. The K key point features of the embedded identity features are compressed from 128 dimensions to 16 dimensions to obtain the compressed key point features , which are used for subsequent feature updates. Secondly, through the correlation operation of adjacent image frames, an attention weight with enhanced features is obtained , and this attention weight can guide where the network should focus in the current image frame. Here, the current image frame is a preset image frame with a higher ordinal number among adjacent image frames. After multiplying with , the attention features of the preset image frame with a lower ordinal number among adjacent image frames are obtained through ReLU and linear transformation . Finally, the attention features are incorporated into the embedded identity features of the current image frame , and convolution is performed on the embedded identity features and the object detection features to complete the embedding update of the object detection features.
[0086] It should be noted that considering the characteristic of class imbalance in UAV videos during the object tracking process, in order to alleviate class imbalance during model training, specifically, gradient-balanced focal loss can be used to supervise the heat map learning ability of the feature detection model to pay more attention to the learning of small object classes. Specifically, referring to Figure 11 shown, here the gradient-balanced focal loss is an improvement based on the cross-entropy loss , and two adaptive weights are designed to reload the loss of heat map learning. Here is used to measure the difference between the predicted heat map and the true heat map. These two adaptive weights include the class balance weight and the small-scale object attention weight . Among them, For balancing classes, For focusing on small targets, these two adaptive loss weights are self-adjusted according to the unbalanced classes and sizes of the objects. Specifically, the gradient-balanced focal loss Can be defined as the following formula:
[0087] The above-mentioned small-target attention weight Focus on small targets and give them greater weights. The size of the target is measured by the bounding box area. Specifically, the small-target attention weight Can be defined as the following formula:
[0088] Where w and h represent the width and height of the target respectively, = 5.
[0089] The above-mentioned class-balancing weight Assign different weights to positive and negative samples according to the corresponding gradients. Specifically, the class-imbalance weight Can be defined as the following formula:
[0090] Where, And Represent the weights of positive and negative samples respectively, and they will be adaptively updated with the network training, Is a heat map used to represent the probability that a pixel is the center point of the target.
[0091] Furthermore, as Figure 1-5 A specific implementation of the method, an embodiment of the present application provides a multi-object tracking device, as Figure 12 Shown, the device includes: an acquisition unit 61, a detection unit 62, an update unit 63, and an association unit 64.
[0092] The acquisition unit 61 is used to acquire a fused feature map of adjacent image frames, and the fused feature map is obtained by fusing pyramid feature maps with set weights; The detection unit 62 is used to perform object detection on the fused feature map to obtain object detection features. The object detection identification at least includes object boundary features, object identity features, and object probability features. The object identity feature is a feature vector used to distinguish objects, and the object probability feature is a visual image feature used to locate objects. The visual image feature characterizes the object probability through the depth of color; An update unit 63 is configured to update the target detection features by using the embedded identity features to obtain the embedded updated target detection features. The embedded identity features are extracted from the preset image frame, and the preset image frame is the image frame with a lower sequence number among adjacent image frames. An association unit 64 is configured to perform target association on the embedded updated target detection features by using an adaptive association mode to obtain the target correspondence in adjacent image frames.
[0093] The multi-target tracking device provided by the embodiment of the present invention, compared with the existing multi-target tracking method based on deep learning in the prior art, the present application obtains the fused feature map of adjacent image frames, and the fused feature map is obtained by fusing the pyramid feature maps with set weights; performs target detection on the fused feature map to obtain target detection features. The target detection identification at least includes target boundary features, target identity features, and target probability features. The target identity feature is a feature vector used to distinguish targets, and the target probability feature is a visual image feature used to locate targets. The visual image feature characterizes the target probability through the depth of color; updates the target detection features by using the embedded identity features to obtain the embedded updated target detection features. The embedded identity features are extracted from the preset image frame, and the preset image frame is the image frame with a lower sequence number among adjacent image frames; performs target association on the embedded updated target detection features by using an adaptive filtering mode to obtain the target correspondence in adjacent image frames, and obtains the target correspondence in adjacent image frames. The entire process shortens the information path based on the pyramid features, constructs a fused feature map containing semantic features and position information of different scales through weighted fusion. This fused feature map can provide sufficient features for the multi-target tracker and improve the accuracy of the multi-target tracker. Since the target detection features will change with the change of the position of the acquisition device, by learning the embedded identity features of adjacent image frames, the target detection features can be adaptively updated, and the target association can be accurately completed by using the adaptive filtering mode, improving the multi-target tracking ability.
[0094] In a specific application scenario, the obtaining unit is specifically configured to: Use a pre-trained feature extraction model to stack the feature maps obtained by sampling adjacent image frames using different paths in multiple layers to obtain stacked image features; During the multi-layer stacking process, add and fuse the image features with the same resolution in different layers through lateral connections to obtain a pyramid feature map; Perform weighted fusion on the pyramid feature map according to the set weights to obtain the fused feature map of adjacent image frames.
[0095] In a specific application scenario, the device further includes: A training unit, which is used to perform model training using depthwise separable convolution composed of depthwise convolution and pointwise convolution before performing object detection on the fused feature map to obtain object detection features. One convolution kernel of the depthwise convolution is responsible for one channel. The number of input channels and output channels corresponding to the feature detection model is the same. Through convolution operations, the fused feature map is weighted and summed in terms of the number of channels, and object detection features consistent with the number of convolution kernels are correspondingly output. Correspondingly, the detection unit is specifically configured to use a pre-trained feature detection model to perform object detection on the fused feature map to obtain object detection features.
[0096] In a specific application scenario, the device further includes: An extraction unit, which is used to extract the embedded identity features in the preset image frame before updating the object detection features using the embedded identity features to obtain the embedded updated object detection features. The image frame with a smaller ordinal number among the adjacent image frames is used as the preset image frame. Correspondingly, the updating unit is specifically configured to: Select a preset number of key position points in the object probability features. The key position points are the position points in the object probability features where the probability ranks before a preset value. Extract a preset number of key point features from the embedded identity features according to the key point positions in the object probability features. After compressing the key point features, use the compressed key point features to update the object detection features to obtain the embedded updated object detection features.
[0097] In a specific application scenario, the updating unit is specifically further configured to: Perform an association operation on the adjacent image frames to obtain feature-enhanced attention weights. The attention weights are used to guide the key point features that the image frame with a larger ordinal number among the adjacent image frames should focus on according to the image frame with a smaller ordinal number among the adjacent image frames. After compressing the key point features, perform weighted summation on the compressed key point features according to the feature-enhanced attention weights to obtain the attention features of the image frame with a smaller ordinal number among the adjacent image frames. Use the attention features to update the object detection features to obtain the embedded updated object detection features.
[0098] In a specific application scenario, the updating unit is specifically further configured to: Use the attention features to perform feature association on the image frame with a smaller ordinal number and the image frame with a larger ordinal number among the adjacent image frames to obtain the association features of the adjacent image frames. Update the target detection feature according to the associated feature to obtain an updated target detection feature with embedding.
[0099] In a specific application scenario, the association unit is specifically configured to: Calculate the target matching coefficient between adjacent image frames using Kalman filtering; If the target matching coefficient is greater than a preset threshold, use the first association mode to perform target association on the updated target detection feature with embedding to obtain the target correspondence in adjacent image frames; If the target matching coefficient is less than or equal to the preset threshold, use the second association mode to perform target association on the updated target detection feature with embedding to obtain the target correspondence in adjacent image frames.
[0100] It should be noted that for other corresponding descriptions of each functional unit involved in the multi-target tracking device provided in this embodiment, reference can be made to Figure 1 - Figure 5 the corresponding description in, which will not be elaborated here.
[0101] Based on the method as described above in Figure 1 - Figure 5 Accordingly, an embodiment of the present application further provides a storage medium on which a computer program is stored, and when the program is executed by a processor, it implements the multi-target tracking method as described above in Figure 1 - Figure 5 shown.
[0102] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various implementation scenarios of the present application.
[0103] Based on the method as described above in Figure 1 - Figure 5 shown, and Figure 12 the virtual device embodiment shown, for the purpose of achieving the above object, an embodiment of the present application further provides an entity device for multi-target tracking, which can specifically be a computer, smart phone, tablet computer, smart watch, server, or network device, etc. The entity device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the multi-target tracking method as described above in Figure 1 - Figure 5 shown.
[0104] Optionally, the entity device may further include a user interface, a network interface, a camera, a Radio Frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display and an input unit such as a keyboard, etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0105] In an exemplary embodiment, refer to Figure 13 , the above-mentioned entity device includes a communication bus, a processor, a memory, and a communication interface, and may further include an input / output interface and a display device. Among them, each functional unit can complete mutual communication through the bus. The memory stores a computer program, and the processor is used to execute the program stored on the memory to execute the multi-object tracking method in the above-mentioned embodiment.
[0106] Those skilled in the art can understand that the structure of the entity device for multi-object tracking provided in this embodiment does not limit the entity device, and it may include more or fewer components, or combine some components, or have different component arrangements.
[0107] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned entity device for multi-object tracking, and supports the operation of the information processing program and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication between other hardware and software in the information processing entity device.
[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or by hardware. By applying the technical solution of this application, compared with the current existing methods, this application shortens the information path based on the pyramid feature, constructs a fusion feature map containing semantic features and position information of different scales through weighted fusion, and this fusion feature map can provide sufficient features for the multi-object tracker, improving the accuracy of the multi-object tracker. Since the target detection features will change with the position change of the acquisition device, by learning the embedded identity features of adjacent image frames, the target detection features can be adaptively updated, and the target association can be accurately completed using the adaptive filtering mode, improving the multi-object tracking ability.
[0109] Those skilled in the art can understand that the attached drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the attached drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more devices different from this implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.
[0110] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenarios. The above-disclosed are only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present application.
Claims
1. A multi-target tracking method, characterized in that: include: Obtaining a fused feature map of adjacent image frames, wherein the fused feature map is obtained by performing feature fusion on the pyramid feature map with a set weight; Performing target detection on the fused feature map to obtain target detection features, wherein the target detection features at least include target boundary features, target identity features, and target probability features, wherein the target identity features are feature vectors used to distinguish targets, and the target probability features are visual image features used to locate targets, wherein the visual image features represent target probability by color depth; The target detection feature is updated by using an embedded identity feature to obtain an embedded updated target detection feature, wherein the embedded identity feature is extracted from a preset image frame, and the preset image frame is an image frame that is in the previous position sequence among adjacent image frames; The embedded updated target detection features are subjected to target association using an adaptive association mode to obtain target correspondences in adjacent image frames.
2. The method according to claim 1, characterized in that: The step of obtaining a fusion feature map of adjacent image frames includes: Using the pre-trained feature extraction model, the feature maps obtained by sampling adjacent image frames using different paths are superimposed in multiple layers to obtain superimposed image features; In the process of multi-layer superposition, image features with consistent resolution in different layers are added and fused through lateral connections to obtain a pyramid feature map; The pyramid feature maps are weighted fused according to the set weights to obtain fused feature maps of adjacent image frames.
3. The method according to claim 1, characterized in that Before performing target detection on the fused feature map to obtain target detection features, the method further includes: A deep separation convolution consisting of channel-by-channel convolution and point-by-point convolution is used for model training to obtain a feature detection model, wherein one convolution kernel of the channel-by-channel convolution is responsible for one channel, and the number of input channels and output channels corresponding to the feature detection model are the same. The fusion feature map is weighted summed on the number of channels through a convolution operation, and the corresponding output is a target detection feature with the same number of convolution kernels; Correspondingly, a pre-trained feature detection model is used to perform target detection on the fused feature map to obtain target detection features.
4. The method according to claim 1, characterized in that Before updating the target detection feature by using the embedded identity feature to obtain the embedded updated target detection feature, the method further includes: Taking the image frame that is preceding in order among the adjacent image frames as a preset image frame, and extracting the embedded identity feature in the preset image frame; Correspondingly, the updating of the target detection feature by using the embedded identity feature to obtain the embedded updated target detection feature includes: Selecting a preset number of key position points in the target probability feature, wherein the key position points are position points in the target probability feature whose probability ranking is before a preset value; According to the key point positions in the target probability features, a preset number of key point features are extracted from the embedded identity features; After compressing the key point features, the object detection features are updated using the compressed key point features to obtain embedded updated object detection features.
5. The method according to claim 4, characterized in that After compressing the key point features, the target detection features are updated using the compressed key point features to obtain the embedded updated target detection features, including: Performing an association operation on the adjacent image frames to obtain an attention weight for feature enhancement, wherein the attention weight is used to guide the key point features that the image frame that is in the previous position in the adjacent image frames should pay attention to according to the image frame that is in the previous position in the adjacent image frames; After compressing the key point features, weighted summing is performed on the compressed key point features according to the feature-enhanced attention weights to obtain the attention features that are in the front position in the adjacent image frames; The target detection feature is updated using the attention feature to obtain an updated target detection feature embedded therein.
6. The method according to claim 5, characterized in that The updating of the target detection feature by using the attention feature to obtain the embedded updated target detection feature includes: Using the attention feature, feature association is performed on an image frame that is in the front position and an image frame that is in the back position in the adjacent image frames to obtain association features of the adjacent image frames; The target detection feature is updated according to the associated feature to obtain an embedded updated target detection feature.
7. The method according to any one of claims 1 to 6, characterized in that The method of using the adaptive association mode to perform target association on the embedded updated target detection features to obtain target correspondence in adjacent image frames includes: Using Kalman filtering to calculate the target matching coefficient between the adjacent image frames; If the target matching coefficient is greater than a preset threshold, the first association mode is used to perform target association on the embedded updated target detection features to obtain a target correspondence relationship in adjacent image frames; If the target matching coefficient is less than or equal to a preset threshold, a second association mode is used to perform target association on the embedded updated target detection features to obtain target correspondence in adjacent image frames.
8. A multi-target tracking device, characterized in that: include: An acquisition unit, used to acquire a fused feature map of adjacent image frames, wherein the fused feature map is obtained by performing feature fusion on the pyramid feature map with a set weight; A detection unit, configured to perform target detection on the fused feature map to obtain a target detection feature, wherein the target detection feature at least includes a target boundary feature, a target identity feature, and a target probability feature, wherein the target identity feature is a feature vector used to distinguish a target, and the target probability feature is a visual image feature used to locate a target, wherein the visual image feature represents the target probability by color depth; An updating unit, configured to update the target detection feature by using an embedded identity feature to obtain an embedded updated target detection feature, wherein the embedded identity feature is extracted from a preset image frame, and the preset image frame is an image frame that is in the previous position sequence among adjacent image frames; The association unit is used to use an adaptive association mode to perform target association on the embedded updated target detection features to obtain target correspondence in adjacent image frames.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multi-target tracking method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-target tracking method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Improved SSD target detection method based on self-attention and feature fusion
CN113743505A
Multi-target tracking method and system based on spatial-temporal trajectory association
CN114913200A
Visual target tracking method and terminal based on attention mechanism
CN117372721A
Multi-target association and tracking method based on motion and appearance feature adaptive fusion
CN117576530A
Multi-target tracking algorithm based on feature enhancement multi-scale fusion strategy and angular momentum mechanism
CN118522034A