Multi-target tracking method and system for gigabit-pixel-level resolution video

By introducing feature enhancement models into the feature extraction network, the problem of poor feature extraction effect in ultra-high resolution videos is solved, and the accuracy and robustness of multi-objective tracking are improved, which is suitable for multi-objective tracking of gigapixel-level resolution videos.

CN120355897APending Publication Date: 2025-07-22SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510434635.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing multi-objective tracking method has poor feature extraction effect in ultra-high resolution video scenarios, resulting in poor tracking accuracy and robustness, and it is difficult to deal with complex situations such as large differences in target scales, occlusion and appearance changes.

Method used

Feature enhancement models are introduced into the feature extraction network, including cross-scale context-aware feature module, element-level adaptive feature enhancement module and local correlation fine tuning module to enhance feature extraction capabilities and improve tracking effects in ultra-high resolution scenarios.

Benefits of technology

It improves the accuracy and robustness of feature recognition, improves the target regression, detection and occlusion problems in ultra-high resolution scenarios, and improves the accuracy and robustness of multi-objective tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355897A_ABST
    Figure CN120355897A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target tracking method and system for a gigabit-pixel-level resolution video, and the method comprises the steps: inputting each video frame into a feature extraction network for feature extraction, and obtaining a feature enhancement feature and a fine tuning feature; the feature extraction network comprises a convolutional neural network, a feature pyramid network and a feature enhancement model which are connected in sequence, and the feature enhancement model comprises a cross-scale context sensing feature module, an element-level adaptive feature enhancement module and a local correlation fine tuning module; and performing data association on the feature enhancement feature and the fine tuning feature with the track of the previous video frame, performing matching to obtain the track of the tracked target, and performing track updating. A feature enhancement model is introduced into a feature extraction network, extracted features are enhanced, useful features with richer information are obtained, it is guaranteed that the method is better suitable for an ultrahigh-resolution scene, similarity and difference between targets are calculated more accurately, and therefore the accuracy of feature recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-object tracking, and in particular to a multi-object tracking method and system for gigapixel-resolution video. Background Art

[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] Multi-object tracking (MOT) is a key research branch in the field of computer vision. With the continuous improvement of the resolution of surveillance videos, exploring the application of ultra-high-resolution videos in computer vision, especially in the multi-object tracking scenario, has become the focus of researchers. Current research mainly focuses on processing videos with ordinary resolution, and there is insufficient attention to the ultra-high-resolution video scenario. Existing MOT methods for ordinary resolution first perform object detection operations on each frame of the video sequence. In this step, by assigning a detection box to each recognized object, precise cropping and positioning of all objects in the image are achieved. Then, the object tracking problem is transformed into a task of inter-frame object association, and a similarity matrix is constructed through IoU matching and appearance feature matching, and solved by methods such as the Hungarian algorithm or greedy algorithm to find the optimal object matching scheme in the given similarity matrix and achieve trajectory association of objects.

[0004] The inventors found in their research that existing MOT methods pay insufficient attention to the ultra-high-resolution video scenario, and the tracking accuracy depends on the effect of feature extraction. When the feature extraction effect is poor, the tracking effect will be poor. In the ultra-high-resolution complex multi-object tracking scenario, the scale difference of each pedestrian is much larger than that in the ordinary resolution, and the regression of its width and height becomes more difficult, making it difficult for the model to accurately regress its position and scale. At the same time, in the ultra-high-resolution video, the scene is larger, the number of pedestrians is more and they are usually gathered together. This dense distribution increases the mutual occlusion and overlap of objects, making it difficult for traditional detection algorithms to distinguish each object. Secondly, the occlusion between objects will lead to a reduction in the visible part of the object, thus affecting the extraction of appearance features. The appearance of the object may also be affected by factors such as light and perspective, further resulting in poor feature extraction effect and affecting the accuracy of tracking. In summary, existing MOT methods must ensure the accuracy and robustness of video tracking to adapt to their applications in fields such as intelligent monitoring and autonomous driving. However, due to insufficient attention to the ultra-high-resolution scenario and poor appearance feature extraction effect, the accuracy and robustness of video tracking are poor. Summary of the Invention

[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a multi-object tracking method and system for gigapixel-level resolution videos. A feature enhancement model is introduced into the feature extraction network, enhancing the network's feature extraction ability, improving the accuracy of the similarity between the captured target features, and improving the regression problem, detection problem, and tracking effect in complex situations such as target occlusion and appearance changes in ultra-high resolution scenarios, thereby improving the accuracy and robustness of tracking.

[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0007] In a first aspect, the present invention provides a multi-object tracking method for gigapixel-level resolution videos, including:

[0008] Obtain the video frames of the video sequence to be tracked and the trajectories of the previous video frame.

[0009] Input each video frame into a feature extraction network for feature extraction to obtain feature-enhanced features and fine-tuned features; the feature extraction network includes a convolutional neural network, a feature pyramid network, and a feature enhancement model connected in sequence. The feature enhancement model includes a cross-scale context-aware feature module, an element-level adaptive feature enhancement module, and a local correlation fine-tuning module; the video frame is sequentially subjected to feature extraction and fusion through the convolutional neural network and the feature pyramid network to obtain a fused feature map, which is input into the feature enhancement module for feature enhancement; the cross-scale context-aware feature module receives the fused feature map and performs multi-level feature extraction and global dependence modeling on it to obtain a perceptual feature map, which is respectively input into the element-level adaptive feature enhancement module and the local correlation fine-tuning module; the element-level adaptive feature enhancement module dynamically enhances the perceptual feature map by combining spatial and channel features to obtain a feature-enhanced feature map, and the local correlation fine-tuning module performs multi-scale feature fusion and dynamic attention modeling on the perceptual feature map to obtain a fine-tuned feature map.

[0010] Associate the feature-enhanced features and fine-tuned features with the trajectories of the previous video frame, match to obtain the trajectories of the targets to be tracked, and update the trajectories.

[0011] In a further technical solution, the cross-scale context-aware feature module includes a multi-scale spatial feature extraction module and a global context modeling module; the multi-scale spatial feature extraction module performs spatial weighting processing on the fused feature map through hierarchical convolutional operations to obtain a spatial feature map.

[0012] In a further technical solution, the global context modeling module performs pooling fusion and normalization on the fused feature map to obtain a context feature map, and adds the context feature map to the spatial feature map to obtain a perceptual feature map.

[0013] For a further technical solution, the element-level adaptive feature enhancement module includes a spatial enhancement mechanism module and a channel optimization mechanism module; the spatial enhancement mechanism module performs grouped feature reshaping and normalization on the perception feature map to obtain a spatially enhanced feature map.

[0014] For a further technical solution, the channel optimization mechanism module performs channel weight allocation on the perception feature map to obtain a channel-optimized feature map, and adds the channel-optimized feature map to the spatially enhanced feature map to obtain a feature-enhanced feature map.

[0015] For a further technical solution, the local correlation fine-tuning module first captures the spatial information of the perception feature map, and sequentially passes through a convolutional layer and an activation function to obtain an attention feature map. After performing multi-scale convolutional operations on the attention feature map and adding them together, a fused feature is obtained, and dynamic attention modeling is performed on the fused feature to obtain a fine-tuned feature map.

[0016] For a further technical solution, the data association of the feature-enhanced feature and the fine-tuned feature with the trajectory of the previous video frame is specifically as follows: the fine-tuned feature is first associated with the trajectory of the previous video frame. Trajectories that do not match successfully are predicted using Kalman filtering, and a second data association is performed in combination with the feature-enhanced feature.

[0017] In a second aspect, the present invention provides a multi-object tracking system for gigapixel-resolution video, including:

[0018] A preprocessing module, which is configured to: obtain the video frames of the video sequence to be tracked and the trajectories of the previous video frame;

[0019] A feature extraction module, which is configured to: input each video frame into a feature extraction network for feature extraction to obtain a feature-enhanced feature and a fine-tuned feature; the feature extraction network includes a convolutional neural network, a feature pyramid network, and a feature enhancement model connected in sequence. The feature enhancement model includes a cross-scale context-aware feature module, an element-level adaptive feature enhancement module, and a local correlation fine-tuning module; the video frames are sequentially subjected to feature extraction and fusion through the convolutional neural network and the feature pyramid network to obtain a fused feature map, which is input into the feature enhancement module for feature enhancement; the cross-scale context-aware feature module receives the fused feature map and performs multi-level feature extraction and global dependence modeling on it to obtain a perception feature map, which is respectively input into the element-level adaptive feature enhancement module and the local correlation fine-tuning module; the element-level adaptive feature enhancement module dynamically enhances the perception feature map by combining spatial and channel features to obtain a feature-enhanced feature map, and the local correlation fine-tuning module performs multi-scale feature fusion and dynamic attention modeling on the perception feature map to obtain a fine-tuned feature map;

[0020] A matching and association module, which is configured to: perform data association between the feature enhancement feature and the fine-tuning feature and the trajectory of the previous video frame, match to obtain the trajectory of the target to be tracked, and perform trajectory update.

[0021] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the multi-object tracking method for gigapixel-level resolution video as described in the first aspect.

[0022] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the multi-object tracking method for gigapixel-level resolution video as described in the first aspect.

[0023] The above one or more technical solutions have the following beneficial effects:

[0024] The present invention innovatively introduces a feature enhancement model into the feature extraction network, enhances the extracted features, obtains more useful features with richer information, ensures better applicability to ultra-high resolution scenarios, and more accurately calculates the similarity and difference between targets, thereby improving the accuracy of feature recognition. The feature enhancement model innovatively designs a cross-scale context-aware feature module, an element-level adaptive feature enhancement module, and a local correlation fine-tuning module, which jointly improve the accuracy of the similarity between the captured target features, improve the regression problem, detection problem, and tracking effects in complex situations such as target occlusion and appearance changes in ultra-high resolution scenarios, thereby improving the accuracy and robustness of tracking.

[0025] The cross-scale context-aware feature module enables the model to capture rich context information at different feature levels through multi-level feature extraction and global dependency modeling, and obtain high-quality features suitable for both detection and ReID tasks. Compared with existing methods, the cross-scale context-aware feature module significantly alleviates the impact of task competition on the model performance and improves the discriminability and robustness of feature expression.

[0026] The element-level adaptive feature enhancement module combines a spatial enhancement mechanism and a channel weight optimization mechanism to dynamically adjust the feature expression, significantly enhance the detection robustness, and reduce false detections and missed detections. Compared with existing methods, the element-level adaptive feature enhancement module realizes the enhancement of feature expression ability while maintaining a lightweight design to meet the real-time detection requirements.

[0027] The local correlation fine-tuning module combines a multi-scale spatial attention mechanism and a dynamic convolution weight allocation strategy, selectively focuses on key features, fuses multi-scale information, and enhances the robustness of feature representation. Compared with existing methods, the local correlation fine-tuning module exhibits stronger object identification ability and interference information suppression effect in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which form a part of this specification, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0029] Figure 1 is a flowchart of the multi-object tracking method for gigapixel-resolution video according to an embodiment of the present invention;

[0030] Figure 2 is a schematic structural diagram of the cross-scale context-aware feature module according to an embodiment of the present invention;

[0031] Figure 3 is a schematic structural diagram of the element-level adaptive feature enhancement module according to an embodiment of the present invention;

[0032] Figure 4 is a schematic structural diagram of the local correlation fine-tuning module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0034] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0035] In the case of no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0036] Embodiment 1

[0037] As Figure 1 shown, this embodiment discloses a multi-object tracking method and system for gigapixel-resolution video, and the method includes the following steps:

[0038] S1: Obtain the video frames of the video sequence to be tracked and the trajectories of the previous video frame.

[0039] In this embodiment, the video sequence uses the publicly available PANDA, the first gigapixel-level human-centered video dataset for large-scale, long-term, and multi-object visual analysis. The videos in PANDA are captured by gigapixel cameras and cover real-world scenes with a wide field of view (about 1 square kilometer area) and high-resolution details (about gigapixel level per frame). The scenes may contain 4k head counts and scale changes over 100 times. PANDA provides rich and hierarchical real annotations, including 15,974.6k bounding boxes, 111.8k fine-grained attribute labels, 12.7k trajectories, 2.2k groups, and 2.9k interactions.

[0040] Furthermore, the processing of the video sequence includes independently extracting each detected target in the video to be tracked based on the annotation information and constructing a target-centered tracking sequence. The trajectory of the previous video frame is the trajectory of the previous video frame obtained through the previous tracking.

[0041] S2: Input each video frame into the feature extraction network for feature extraction to obtain feature-enhanced features and fine-tuned features.

[0042] In this embodiment, the feature extraction network includes a convolutional neural network, a feature pyramid network, and a feature enhancement model connected in sequence. The feature enhancement model includes a cross-scale context-aware feature module, an element-level adaptive feature enhancement module, and a local correlation fine-tuning module connected. The video frame is sequentially subjected to feature extraction and fusion through the convolutional neural network and the feature pyramid network to obtain a fused feature map, which is input into the feature enhancement module for feature enhancement; the cross-scale context-aware feature module receives the fused feature map and performs multi-level feature extraction and global dependency modeling on it to obtain a perceptual feature map and inputs it into the element-level adaptive feature enhancement module and the local correlation fine-tuning module respectively; the element-level adaptive feature enhancement module dynamically enhances the perceptual feature map by combining spatial and channel features to obtain a feature-enhanced feature map, and the local correlation fine-tuning module performs multi-scale feature fusion and dynamic attention modeling on the perceptual feature map to obtain a fine-tuned feature map.

[0043] Furthermore, the cross-scale context-aware feature module is used for multi-level feature extraction and global dependency modeling to obtain high-quality features suitable for both detection and ReID tasks and alleviate the impact of task competition on model performance; the dynamic enhancement of the element-level adaptive feature enhancement module by combining spatial and channel features is used to address the challenges under dynamic backgrounds and complex scenes; the local correlation fine-tuning module is used for multi-scale feature fusion and dynamic attention modeling to improve the target identification performance and scene adaptability of the model.

[0044] The convolutional neural network is an efficient convolutional neural network. Each video frame serves as an input image and is processed through the efficient convolutional neural network to extract general features in the input image, generating multi-scale feature maps. Each feature map contains feature information with different resolutions.

[0045] The Feature Pyramid Network is used to fuse feature maps with different resolutions, namely multi-scale feature maps, to obtain a richer feature representation. The multi-scale feature maps output by the convolutional neural network are input into the Feature Pyramid Network, where the feature maps with different resolutions are fused to obtain a richer feature representation, and the fused multi-scale feature maps, i.e., the fused feature maps, are output, providing feature information for subsequent object detection and re-identification tasks.

[0046] The feature enhancement model includes a cross-scale context-aware feature module, an element-level adaptive feature reinforcement module, and a local correlation fine-tuning module. The cross-scale context-aware feature module is used for multi-level feature extraction and global dependency modeling. The element-level adaptive feature reinforcement module combines the dynamic enhancement of spatial and channel features. The local correlation fine-tuning module is used for multi-scale feature fusion and dynamic attention modeling.

[0047] Specifically, the feature enhancement model is added after the last convolutional layer of the Feature Pyramid Network to obtain enhanced feature maps, namely feature-reinforced feature maps and fine-tuned feature maps.

[0048] (1) Cross-scale context-aware feature module

[0049] The cross-scale context-aware feature combines multi-level feature weighting and global dependency modeling, simultaneously capturing local details and global context information, improving the robustness of feature expression and task adaptability to effectively cope with complex object distributions and numerous interfering backgrounds. As Figure 2 shown, the cross-scale context-aware feature module includes a multi-scale spatial feature extraction module (used to extract multi-scale spatial information of features) and a global context modeling module (used to enhance the global context modeling ability).

[0050] The multi-scale spatial feature extraction module first performs spatial weighting processing on the fused feature map through hierarchical convolutional operations. The input is the fused feature map X ∈ R B×C×H×W , where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map; the output is the processed feature map, i.e., the spatial feature map.

[0051] First, the fused feature map is divided into n blocks along the channel dimension. Each block represents a subset of features within a specific channel range, with the dimension (batch_size, chunk_dim, height_width), where chunk_dim = dim / n. The feature block is shown in formula (1):

[0052] xc = X.chunk(n, dim = 1) (1)

[0053] Among them, X.chunk means dividing the input tensor evenly into n chunks along the specified dimension, and dim represents the channel dimension.

[0054] Next, process the n chunks respectively. The value of i ranges from 0 to n - 1. For the chunks with i < 0, directly use depthwise separable convolution to process the chunk xc[i] to obtain the weighted feature map s, as shown in formula (2):

[0055] s = Conv2d(xc[i]) (2)

[0056] Among them, Conv2d represents depthwise separable convolution.

[0057] For the chunks with i > 0, first use adaptive max pooling to downsample the feature map and convert it to a size of (h / / 2 i , w / / 2 i ), where h and w represent the height and width of the input feature map, and 2 i represents the scale factor, which determines the degree of downsampling of the feature map. (h / / 2 i , w / / 2 i ) represents integer division of the height and width of the input feature map, that is, the floor operation, rounding down. Then process each chunk through depthwise separable convolution to obtain the feature map , and finally restore it to the size of the original feature map through interpolation upsampling to obtain the weighted feature map s, as shown in formulas (3) and (4):

[0058]

[0059] Among them, p_size represents a variable used to specify the target size after downsampling of the feature map, which is obtained by dividing the height and width of the input feature map by 2 i respectively.

[0060] Concatenate all the obtained weighted feature maps in the channel dimension and aggregate them through 1×1 convolution, as shown in formula (5):

[0061] out = Conv(concat(s0, s1,..., s n-1 )) (5)

[0062] Among them, concat represents concatenation.

[0063] Finally, use the GELU activation function to activate the aggregated feature map and multiply it with the original input feature map X to obtain the feature map of the first part, i.e., the spatial feature map X1, as shown in Equation (6):

[0064] X1 = GELU(out) ⊙ X (6)

[0065] Where, ⊙ represents the position dot product.

[0066] The global context modeling module introduces a pooling fusion and normalization mechanism. First, perform a dimensional transformation and rearrangement on the input feature map, i.e., the fused feature map, as shown in Equation (7):

[0067] X r = X.reshape(0, 2, 1, 3), X r ∈ R B×C×H×W (7)

[0068] Extract the significant features of the local region through two operations: max pooling and average pooling, and concatenate them to generate an intermediate feature containing rich context, as shown in Equation (8):

[0069] Y1 = Concat(MaxPool(X r ), AvgPool(X r )) (8)

[0070] Process the obtained feature map through depthwise separable convolution to obtain weighted features, then perform normalization and pass through a ReLU activation function to suppress the interference of background noise and obtain the context feature map Y2, as shown in Equation (9):

[0071] Y2 = ReLU(Normalization(Conv2d(Y1))). (9)

[0072] Multiply the obtained feature map Y2 with the feature map Xr obtained by performing dimensional rearrangement on the fused feature map element-wise to obtain the context feature map X2, as shown in Equation (10):

[0073] X2 = X r · Y2 (10)

[0074] Finally, add the spatial feature map X1 and the context feature map X2 to obtain the final feature map, i.e., the perceptual feature map X', as shown in Equation (11):

[0075] X' = X1 + X2. (11)

[0076] (II) Element-level Adaptive Feature Enhancement Module

[0077] The element - level adaptive feature enhancement module combines a spatial enhancement mechanism and a channel weight optimization mechanism to improve the feature expression ability, enhance the detection robustness, reduce false detections and missed detections, and is designed to be lightweight to meet the requirements of real - time detection. As Figure 3 shown, the element - level adaptive feature enhancement module includes a spatial enhancement mechanism module and a channel optimization mechanism module.

[0078] The input to the element - level adaptive feature enhancement module is the perceptual feature map \(X'\in R^{B\times C\times H\times W}\), B×C×H×W where \(B\) is the batch size, \(C\) is the number of channels, and \(H\) and \(W\) are the height and width of the feature map.

[0079] In the spatial enhancement mechanism module, to better capture multi - level information in the spatial dimension, a grouped feature reshaping and normalization mechanism is used to enhance the spatial feature expression. First, the feature map is subjected to grouped feature reshaping. The input feature map is divided into \(G\) groups along the channel dimension, and the dimension of each group of features is so as to achieve independent processing of local features at the grouped scale, as shown in formula (12):

[0080] \(x_{n}\) Group \(=X'.view(B\cdot G,-1,H,W)\ (12)\)

[0081] where \(X'.view\) represents reshaping the tensor.

[0082] Next, average pooling is performed on the feature map to extract the local response of each group of features. The shape of \(x_{n}\) remains as shown in formula (13).

[0083] \(x_{n}\) n \(=x_{n}\) Group \(\odot AvgPool(x_{n}\) Group )\ (13)\)

[0084] where \(\odot\) represents the dot product at the position, and \(x_{n}\) n represents the spatially enhanced local feature, which is used to capture local information while reducing the channel dimension.

[0085] The reshaped feature and the average - pooled feature are summed along the channel dimension, as shown in formula (14):

[0086]

[0087] where \(x_{n}\) n becomes \((B\cdot G,1,H,W)\).

[0088] Calculate the mean of the feature map. Reshape the feature map \(x_{n}\) n into a two - dimensional form \(t\) with a shape of \((B\cdot G,H\cdot W)\), as shown in formula (15):

[0089] t = x n .view(B·G, -1) (15)

[0090] Among them, -1 indicates automatically calculating this dimension to make its size H·W. x n .view means reshaping the tensor and reshaping the feature map into its original shape.

[0091] The mean μ is calculated as shown in formula (16):

[0092] μ = mean(t, dim = 1) (16)

[0093] Formula (16) indicates calculating the mean for each group, and the shape of the result is (B·G, 1).

[0094] The standard deviation is calculated as shown in formula (17):

[0095] σ = std(t, dim = 1) + ε (17)

[0096] Among them, ε is a small constant used to prevent division by zero. The calculation of the standard deviation is to measure the dispersion degree of the features.

[0097] Subsequently, normalization is performed, and then the normalized feature map is reshaped into its original shape to suppress the influence of outliers on feature expression, as shown in formulas (18) and (19):

[0098]

[0099] x1 = t norm .view(B, G, H, W) (19)

[0100] Among them, x1 represents the spatially enhanced feature map, and t norm .view means reshaping the normalized feature map into its original shape, and t norm represents the normalized feature map.

[0101] The channel optimization mechanism module, in order to effectively extract global context information, enhances the channel feature expression through the channel weight allocation mechanism. First, the perceptual feature map X′ is subjected to global average pooling to obtain a new feature map y ∈ R B×C , as shown in formula (20):

[0102]

[0103] Among them, h represents the height of the input feature map, w represents the width of the input feature map, i represents the row index (in the height direction) of the feature map, j represents the column index (in the width direction) of the feature map, and X′ :,:,i,jRepresents the value of a certain pixel point in the channel dimension, that is, all channels in the channel dimension share the value of this pixel position (i, j).

[0104] Next, calculate the channel weights through the fully connected layer, as shown in formula (21):

[0105] y′ = Sigmoid(W2·RELU(W1·y)) (21)

[0106] Among them, W1 and W2 represent the weights of the fully connected layer, Sigmoid represents the Sigmoid activation function, and RELU represents the RELU activation function.

[0107] Next, obtain the feature output of channel processing through weighting, as shown in formula (22):

[0108] x2 = X′·y′ (22)

[0109] Among them, x2 represents the channel-optimized feature map.

[0110] Finally, combine the spatial enhancement feature map x1 and the channel-optimized feature map x2 to obtain the final output feature map, that is, the feature enhancement feature map, as shown in formula (23):

[0111] x′ = x1 + x2. (23)

[0112] (III) Local Correlation Fine-Tuning Module

[0113] As Figure 4 shown, the local correlation fine-tuning module combines the multi-scale spatial attention mechanism and the dynamic convolution weight allocation strategy. By selectively focusing on key features and fusing multi-scale information, it enhances the robustness of feature expression, improves the target recognition ability in complex scenarios, and the suppression effect on interference information.

[0114] The input of the local correlation fine-tuning module is the perceptual feature map X′ ∈ R B×C×H×W , where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map; the output is the processed feature map, that is, the fine-tuned feature map.

[0115] The feature relationship in the spatial dimension is captured by the combination of max pooling and average pooling. First, perform max pooling and average pooling on the input feature map respectively to help extract the spatial information of the features and reduce the computational amount. Subsequently, splice the two obtained feature maps. Max pooling, average pooling, and splicing are shown in formulas (24), (25), and (26) respectively:

[0116] MaxPool(X′) = max i (X′[i]) (24)

[0117]

[0118] Q = concat(MaxPool(X′), AvgPool(X′)) (26)

[0119] Wherein, X′[i] represents the feature map of the i-th channel in the input feature map, C represents the number of channels, i represents the i-th channel, and Q represents the new feature map obtained by concatenating the max pooling and average pooling. After max pooling and average pooling, the number of channels becomes 1, and the number of channels of the new feature map Q after concatenation is 2. This process selectively focuses on key local regions and enhances the expression ability of local features.

[0120] The feature map Q will pass through a convolutional layer for calculating spatial attention. Next, the feature map will pass through the sigmoid activation function to generate attention weights, as shown in formula (27):

[0121] Q′ = σ(s_attention(Q)) (27)

[0122] Wherein, Q′ represents the attention feature map, the s_attention function represents calculating spatial attention through a convolutional layer, and σ represents the sigmoid activation function.

[0123] Next, in order to capture spatial features at different scales, a multi-scale convolutional mechanism is adopted. Apply convolutional operations with different convolutional kernels to the feature map Q′, as shown in formula (28):

[0124] conv_out[j] = Conv(Q′, Kernel_size = k j ) (28)

[0125] Wherein, k j represents the convolutional kernel size of the j-th convolution, and the output conv_out is a list containing all the convolutional results.

[0126] Add all the convolutional outputs to obtain the fused feature U. Then, take the average of the fused feature U in the spatial dimension. Finally, obtain the feature map Z through the fully connected layer fc. The formulas are shown in (29) and (30) respectively:

[0127]

[0128] Wherein, l and m represent spatial positions. K represents the number of convolutional branches, that is, before fusion, multiple feature maps are obtained through convolutional operations at different scales; U[:,:,l,k] represents the channel feature of the fused feature U at the spatial position (l,m), and its shape is B×C.

[0129] Apply multiple fully connected layers to the feature map Z to generate the weights corresponding to each convolutional kernel, and finally normalize them. As shown in Formulas (31) and (32):

[0130] weight[n]=fc n (Z) (31)

[0131] attention_weights=σ(weights) (32)

[0132] By dynamically generating the convolutional kernel weights, the contribution of local features can be adaptively adjusted, so as to better fuse multi-scale information and improve the robustness of local features.

[0133] Apply the attention weights to the feature maps feats of different convolutions to obtain the feature map V, as shown in Formula (33):

[0134]

[0135] Finally, add the attention feature map Q′ and the feature map V to obtain the final feature map, that is, the fine-tuned feature map F, as shown in Formula (34):

[0136] F=Q′+V (34)

[0137] It can explicitly model the relationships between local features, thereby enhancing the correlation of local regions and improving the discriminability of feature expressions.

[0138] The feature extraction network also includes a loss function unit for calculating the loss function in the process of the feature extraction network. The loss function is a multi-task joint loss function, which specifically includes the following three core parts:

[0139] The classification loss is used to distinguish the foreground (pedestrians) and the background to solve the problem of class imbalance;

[0140] The regression loss is used to optimize the position and size of the bounding box;

[0141] The cross-entropy loss is used to model the pedestrian re-identification task as a classification task to learn pedestrian identity features.

[0142] S3: Perform data association between the feature-enhanced feature and the fine-tuned feature and the trajectory of the previous video frame, match to obtain the trajectory of the target to be tracked, and perform trajectory update.

[0143] In this embodiment, the fine-tuned feature map and the trajectories of the previous video frame are subjected to the first data association through an appearance feature matching algorithm. The matching scores are calculated by cosine similarity or other measurement methods, and the targets and trajectories with high similarity are preferentially associated. The trajectories that fail to match successfully are predicted using Kalman filtering, and the second data association is performed through the IoU matching algorithm in combination with the feature-enhanced feature map to obtain the trajectories of the tracked targets and update the trajectories.

[0144] Furthermore, the use of the two data association methods can compensate for short-term occlusion or blurred appearance while ensuring appearance consistency, and solve the limitations in traditional methods, that is, over-reliance on appearance features, which fails when the target is occluded or moving rapidly, and the inability to handle long-term occlusion with only motion information.

[0145] The specific steps are as follows:

[0146] S301: Obtain the trajectories of the previous video frame to be matched, perform the first data association with the fine-tuned features using an appearance feature matching algorithm, and obtain the trajectories that are successfully matched for the first time and the detections and trajectories that are not successfully matched for the first time.

[0147] S302: After performing Kalman filtering prediction on the detections and trajectories that are not successfully matched for the first time, perform the second data association through IOU matching in combination with the feature-enhanced feature to obtain the trajectories that are not successfully matched, the detections that are not successfully matched, and the trajectories that are successfully matched for the second time.

[0148] S303: Update the trajectories that are successfully matched for the first time and the trajectories that are successfully matched for the second time to obtain the updated trajectories.

[0149] S304: For the trajectories that are not successfully matched, remove the trajectories with the termination status, and mark the remaining trajectories as "inactive" trajectories.

[0150] S305: Initialize the detections that are not successfully matched as new trajectories, and together with the "inactive trajectories" and the updated trajectories, they are used as the trajectories to be matched and tracked in the next frame, which are also the trajectories obtained from this data association.

[0151] The updated trajectories are the trajectories of the tracked targets obtained through matching in the current video frame.

[0152] The present invention introduces a feature enhancement model into the feature extraction network, namely, introducing a cross-scale context-aware feature module, an element-level adaptive feature reinforcement module, and a local correlation fine-tuning module to optimize the generated features, thereby obtaining a more informative and discriminative feature representation. This method can better adapt to the ultra-high resolution scenario, calculate the similarities and differences between targets more accurately, and significantly improve the accuracy of feature recognition. Specifically, the cross-scale context-aware feature module generates high-quality features suitable for detection and re-identification tasks through multi-level feature extraction and global dependence modeling, alleviating the impact of task competition on the model performance; the element-level adaptive feature reinforcement module combines the dynamic adjustment of spatial and channel features to effectively cope with the challenges of dynamic backgrounds and complex scenes; the local correlation fine-tuning module further improves the target identification ability and scene adaptability of the model through multi-scale feature fusion and dynamic attention modeling. In addition, through two data association matching strategies, the accuracy of target matching is improved, thereby generating a more accurate target tracking trajectory.

[0153] Embodiment 2

[0154] This embodiment discloses a multi-target tracking system for gigapixel-resolution videos, including:

[0155] A preprocessing module, configured to: obtain video frames of a video sequence to be tracked and trajectories of the previous video frame;

[0156] A feature extraction module, configured to: input each video frame into a feature extraction network for feature extraction to obtain feature-enhanced features and fine-tuned features; the feature extraction network includes a convolutional neural network, a feature pyramid network, and a feature enhancement model connected in sequence, and the feature enhancement model includes a cross-scale context-aware feature module, an element-level adaptive feature reinforcement module, and a local correlation fine-tuning module; the video frames are sequentially subjected to feature extraction and fusion through the convolutional neural network and the feature pyramid network to obtain a fused feature map, which is input into the feature enhancement module for feature enhancement; the cross-scale context-aware feature module receives the fused feature map and performs multi-level feature extraction and global dependence modeling on it to obtain a perceptual feature map, which is respectively input into the element-level adaptive feature reinforcement module and the local correlation fine-tuning module; the element-level adaptive feature reinforcement module dynamically enhances the perceptual feature map by combining spatial and channel features to obtain a feature-enhanced feature map, and the local correlation fine-tuning module performs multi-scale feature fusion and dynamic attention modeling on the perceptual feature map to obtain a fine-tuned feature map;

[0157] A matching and association module, configured to: perform data association on the feature-enhanced features and the fine-tuned features with the trajectories of the previous video frame, match to obtain the trajectories of the targets to be tracked, and update the trajectories.

[0158] Embodiment 3

[0159] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method in Embodiment 1 are implemented.

[0160] Embodiment 4

[0161] The purpose of this embodiment is to provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method in Embodiment 1 are executed.

[0162] The steps involved in the devices in the above Embodiments 3 and 4 correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0163] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0164] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0165] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that on the basis of the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative labor are still within the protection scope of the present invention.

Claims

1. A multi-object tracking method for gigapixel-resolution video, characterized in that, Including: Obtain the video frames of the video sequence to be tracked and the trajectories of the previous video frames; Input each video frame into a feature extraction network for feature extraction to obtain feature enhancement features and fine-tuning features; the feature extraction network includes a convolutional neural network, a feature pyramid network, and a feature enhancement model connected in sequence according to the order. The feature enhancement model includes a cross-scale context-aware feature module, an element-level adaptive feature enhancement module, and a local correlation fine-tuning module; the video frames are sequentially passed through the convolutional neural network and the feature pyramid network for feature extraction and fusion to obtain a fused feature map, and are input into the feature enhancement module for feature enhancement; the cross-scale context-aware feature module receives the fused feature map and performs multi-level feature extraction and global dependence modeling on it to obtain a perceptual feature map and inputs it into the element-level adaptive feature enhancement module and the local correlation fine-tuning module respectively; the element-level adaptive feature enhancement module dynamically enhances the perceptual feature map by combining spatial and channel features to obtain a feature enhancement feature map, and the local correlation fine-tuning module performs multi-scale feature fusion and dynamic attention modeling on the perceptual feature map to obtain a fine-tuning feature map; Perform data association between the feature enhancement features and the fine-tuning features and the trajectories of the previous video frames, match to obtain the trajectories of the targets to be tracked, and update the trajectories.

2. The multi-object tracking method for gigapixel-resolution video according to claim 1, characterized in that The cross-scale context-aware feature module includes a multi-scale spatial feature extraction module and a global context modeling module; the multi-scale spatial feature extraction module performs spatial weighting processing on the fused feature map through hierarchical convolutional operations to obtain a spatial feature map.

3. The multi-object tracking method for gigapixel-level resolution video according to claim 2, wherein The global context modeling module performs pooling fusion and normalization on the fused feature map to obtain a context feature map, and adds the context feature map to the spatial feature map to obtain a perceptual feature map.

4. The multi-object tracking method for gigapixel-resolution video according to claim 1, characterized in that, The element-level adaptive feature enhancement module includes a spatial enhancement mechanism module and a channel optimization mechanism module; The spatial enhancement mechanism module performs grouped feature reshaping and normalization on the perceptual feature map to obtain a spatial enhancement feature map.

5. The multi-object tracking method for gigapixel-resolution video according to claim 4, characterized in that, The channel optimization mechanism module performs channel weight assignment on the perceptual feature map to obtain a channel optimization feature map, and adds the channel optimization feature map to the spatial enhancement feature map to obtain a feature enhancement feature map.

6. The multi-object tracking method for gigapixel-resolution video according to claim 1, characterized in that, The local correlation fine-tuning module first captures the spatial information of the perceptual feature map, and sequentially passes through a convolutional layer and an activation function to obtain an attention feature map, performs multi-scale convolutional operations on the attention feature map and then adds them to obtain a fused feature, and performs dynamic attention modeling on the fused feature to obtain a fine-tuning feature map.

7. The multi-object tracking method for gigapixel-resolution video according to claim 1, wherein Performing data association between the feature enhancement features and the fine-tuning features and the trajectories of the previous video frames specifically includes: performing the first data association between the fine-tuning features and the trajectories of the previous video frames, predicting the trajectories of the unmatched ones using Kalman filtering, and performing the second data association in combination with the feature enhancement features.

8. A multi-object tracking system for gigapixel-resolution video, characterized in that, Including: A preprocessing module, which is configured to: obtain the video frames of the video sequence to be tracked and the trajectories of the previous video frames; A feature extraction module, which is configured to: input each video frame into a feature extraction network for feature extraction to obtain a feature enhancement feature and a fine-tuning feature; the feature extraction network includes a convolutional neural network, a feature pyramid network, and a feature enhancement model connected in sequence, and the feature enhancement model includes a cross-scale context-aware feature module, an element-level adaptive feature enhancement module, and a local correlation fine-tuning module; the video frame passes through the convolutional neural network and the feature pyramid network in sequence for feature extraction and fusion to obtain a fused feature map, and is input into the feature enhancement module for feature enhancement; the cross-scale context-aware feature module receives the fused feature map and performs multi-level feature extraction and global dependency modeling on it to obtain a perceptual feature map and inputs it into the element-level adaptive feature enhancement module and the local correlation fine-tuning module respectively; the element-level adaptive feature enhancement module dynamically enhances the perceptual feature map by combining spatial and channel features to obtain a feature enhancement feature map, and the local correlation fine-tuning module performs multi-scale feature fusion and dynamic attention modeling on the perceptual feature map to obtain a fine-tuning feature map; A matching association module, which is configured to: perform data association between the feature enhancement feature and the fine-tuning feature and the trajectory of the previous video frame, match to obtain the trajectory of the target to be tracked, and perform trajectory update.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the multi-object tracking method for gigapixel-level resolution videos according to any one of claims 1-7.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the multi-object tracking method for gigapixel-level resolution videos according to any one of claims 1-7.