Method and device for tracking single target in multi-view video
By extracting face and body features in multi-view videos, using 3D convolution kernel generator and short-term correlation modeling technology, the tracking difficulties caused by target appearance changes and environmental differences are solved, and stable target tracking across videos is achieved, accuracy and robustness are improved, and suitable for real-time applications.
Patent Information
- Application Number
- CN202510429151.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When tracking single targets in multi-view videos, the prior art faces problems such as changes in target appearance, environmental differences and short-term disappearance of targets, which makes it difficult for tracking algorithms to accurately identify and locate targets, especially in complex scenarios.
Using a method based on short-term modeling technology, the face and body features are extracted, the position offset is calculated using the 3D convolution kernel generator, combined with local and global feature fusion, a time series of fusion is constructed, and the trajectory correlation is captured through the short-term correlation modeling module, and the multi-task loss function is optimized to achieve cross-video target correlation.
It improves tracking accuracy under target appearance changes and environmental differences, solves the problem of the target temporarily disappearing, and realizes stable target tracking across videos. It is suitable for multi-view video scenarios, has high computing efficiency, and is suitable for real-time applications.
Smart Images

Figure CN120339332A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for single-object tracking in multi-view videos based on short-term modeling technology, belonging to the field of computer vision technology. Background Art
[0002] From the surveillance system for safeguarding urban security to the free driving of autonomous vehicles on the road, object tracking technology plays a crucial role in the field of computer vision. In the prior art, most efforts focus on single-object tracking in a single video or multi-object tracking in a single video, with less research on single-object tracking in multi-videos; moreover, the prior art has obvious limitations in dealing with problems such as changes in object appearance, environmental differences, and temporary disappearance of objects.
[0003] In the prior art, single-video object tracking algorithms based on pure appearance features often fail to track due to unstable feature expression when dealing with problems such as object occlusion, deformation, and illumination changes. For example, traditional object tracking methods based on correlation filtering, although having advantages in computational efficiency, will significantly degrade in performance when facing complex scenes. In addition, single-video object tracking methods based on deep learning (such as Siamese networks), although improving the tracking accuracy by learning the deep features of the object, still have difficulty maintaining stable tracking effects in the case of drastic changes in object appearance.
[0004] In the prior art, cross-video tracking mainly relies on object appearance feature matching, but this method is prone to false matching when there are large differences in perspectives, illuminations, and resolutions between different videos. For example, some studies attempt to use pre-trained deep learning models to extract object features and then perform feature matching between different videos, but this method requires high generalization ability of the model, and the matching accuracy will drop significantly when the object appearance changes greatly.
[0005] In terms of data association, existing multi-object tracking algorithms usually adopt data association methods based on the Hungarian algorithm or greedy algorithm. These methods have high computational complexity and limited association accuracy when dealing with large-scale datasets and complex scenes. For example, in traffic surveillance scenarios, when the vehicle density is high and there are frequent occlusions and intersections, traditional data association algorithms are difficult to accurately associate object trajectories in different videos.
[0006] In summary, in the existing video object tracking technology, the switching between different videos will cause changes in object appearance, environment, and perspective, making it difficult for tracking algorithms to accurately identify and locate objects; in addition, in the same video, the object may temporarily disappear due to occlusion, loss, etc., resulting in tracking interruption. Summary of the Invention
[0007] Objective of the Invention: In order to overcome the deficiencies existing in the prior art, the present invention provides a method and device for single-object tracking in multi-view videos based on short-term modeling technology, which can effectively handle problems such as target appearance changes, environmental differences, and temporary disappearance of the target, achieve cross-video target association, perform continuous and stable tracking of a single object in multi-view videos, improve the accuracy and robustness of tracking, and meet the requirements of practical applications.
[0008] Technical Solution: To achieve the above objective, the technical solution adopted by the present invention is as follows:
[0009] A method for single-object tracking in multi-view videos, comprising the following steps:
[0010] Step1: Input videos of each view, and use an encoder to extract the face feature vector F and body feature vector A of the target, as well as the background feature B of the corresponding view;
[0011] Step2: Calculate the face position offset vector O according to the face feature vectors F of adjacent frames; (f) Calculate the body position offset vector O according to the body feature vectors A of adjacent frames; (a) ;
[0012] Step3: Use a trajectory kernel generator to respectively obtain the 3D convolution kernels K of the face position offset vector O (f) and the body position offset vector O (α) ; (f) and K (a) ;
[0013] Step4: Perform 3D convolution on the face feature vector F using the 3D convolution kernel K (f) to obtain a local feature vector; Perform 3D convolution on the body feature vector A using the 3D convolution kernel K (a) to obtain a global feature vector;
[0014] Step5: Use a feature aggregator to perform feature fusion on the local feature vector and the global feature vector to obtain a fused feature vector F (0) ;
[0015] Step6: Distinguish different-view videos based on the background feature B, and identify the specified single target in all view videos based on the fused feature vector F (0) Reshape and concatenate the fused feature vectors F of all video frames that identify the specified single target according to the time axis to form a fused feature time series F (0) ; (BT) ;
[0016] Step 7. Statistically analyze the fused feature time series F (T) for missing trajectory segments. For missing trajectory segments with a time length less than the set threshold, use the short-term association modeling module to capture the correlation between adjacent front and rear trajectory segments, and track the missing trajectory segments;
[0017] Step 8. Train the cross-video single-object tracking model composed of Step 1 to Step 7 by jointly optimizing the face recognition loss, body recognition loss, background recognition loss, object recognition loss, and object tracking loss;
[0018] Step 9. Input the multi-view video into the trained cross-video single-object tracking model to identify and predict the start and end times and the positions where the specified single object appears in each view video.
[0019] This method realizes cross-view robust tracking through multi-modal feature extraction and spatio-temporal association modeling. Face features and body features are extracted to capture the local details and global morphology of the target respectively, and the background features are combined to distinguish the perspective differences. The position offset vector is calculated through the feature difference between adjacent frames, and the 3D convolutional kernel is dynamically generated to adapt to the movement changes of the target in the spatio-temporal dimension. Local convolution is performed on the face features to retain the detailed information, and global convolution is performed on the body features to capture the movement trend. The local and global features are complemented and enhanced through the feature aggregator. The perspective information is distinguished when constructing the fused feature time series, and the missing frames are filled with zeros to ensure the temporal continuity. The short-term association modeling establishes the cross-view trajectory segment association through the self-attention mechanism to solve the trajectory breakage caused by occlusion or perspective switching. The model parameters are optimized by jointly using the multi-task loss function, enabling the feature extraction, trajectory prediction, and object recognition to form a collaborative optimization. Finally, through the spatio-temporal feature sequence analysis, the target trajectory prediction and missing frame compensation across video views are realized.
[0020] Specifically, in the Step 3, the trajectory kernel generator is composed of a 3D convolutional layer and a pooling layer:
[0021] K (f) =pool(conv3d(o (f) ,K))
[0022] K (a) =pool(conv3d(O (a) ,K))
[0023] where: pool(·) represents the pooling operation, conv3d represents the 3D convolutional operation, and K represents the weight parameter of the 3D convolutional layer.
[0024] This technical solution realizes spatio-temporal feature modeling of the face and body position offset vectors by constructing a trajectory kernel generator that includes 3D convolutional layers and pooling layers. The pooling operation compresses the features of the position offset vectors, effectively reducing the redundant information of high-dimensional data while retaining the key displacement features, providing an effective input after dimensionality reduction for the subsequent generation of 3D convolutional kernels; the 3D convolutional operation, based on the pooled features, dynamically captures the correlations in the spatio-temporal dimensions through learnable weight parameters, converting the displacement vectors into 3D convolutional kernels that can represent the motion laws of the target. Among them, the weight parameters of the 3D convolutional kernels are adaptively adjusted through end-to-end training, enabling the generated convolutional kernels to not only reflect the continuous motion patterns of the target on the time axis but also capture the correlation characteristics of spatial displacements between different perspectives, thus providing convolutional kernel support with physical significance for the subsequent separate extraction of local and global features.
[0025] Specifically, in the aforementioned Step4, the 3D convolutional kernel K (f) is used to perform 3D convolution on the face feature vector F to obtain a local feature vector The 3D convolutional kernel K (a) is used to perform 3D convolution on the body feature vector A to obtain a global feature vector
[0026] This technical solution establishes a local and global feature expression system for face features and body features respectively through a differential three-dimensional convolution processing mechanism. Specifically, applying the 3D convolutional kernel to the face feature vector to generate a local feature vector makes use of the fine spatio-temporal variation characteristics of face features. Through the sliding convolutional operation of the three-dimensional convolutional kernel in the time dimension, it can effectively capture micro-dynamic features such as facial expressions and head postures, enhancing the continuous expression of local features in the time series. Applying the 3D convolutional kernel to the body feature vector to generate a global feature vector is based on the overall motion characteristics of the human body posture. Through the spatio-temporal joint modeling ability of the three-dimensional convolutional kernel, it captures the motion trajectory and body change laws of the target in three-dimensional space, forming a spatio-temporally coherent global representation. This divide-and-conquer feature processing mechanism not only avoids information loss caused by single-feature modeling but also effectively overcomes the tracking interference caused by target appearance changes and perspective differences through the synergistic effect of local and global features.
[0027] Specifically, in the aforementioned Step5, a feature aggregator is used to perform feature fusion on the local feature vector and the global feature vector (0) to obtain a fused feature vector F
[0028] (51) Randomly initialize a local prompt vector L through a local feature aggregator, copy the local prompt vector L into T copies along the time dimension, and denote the t-th copy of the local prompt vector as Lt , where \(t = 1, 2, \ldots, T\);
[0029] (52) Perform cross - attention operation on the local prompt vector \(L\) t and local features to selectively aggregate valid information from the local features by the local prompt vector \(L\); Concatenate the valid information aggregated at all times and perform self - attention operation to fuse the local prompt vectors from different video frames: t from local features ; where CrossAttn and SelfAttn represent cross - attention mechanism and self - attention mechanism respectively, and MLP represents multi - layer perceptron;
[0030]
[0031] where \(F_{t}\) is the local feature at time \(t\); represents the enhanced local feature vector, and \(F_{t}^{enhanced}\) is the enhanced local feature at time \(t\);
[0032] (53) Calculate the attention map between the enhanced local feature vector and the local feature , convert the attention map into attention weights using the softmax function, and fuse the local feature vector and local feature
[0033]
[0034] F (l) = Concat(F l1 , F l2 , \ldots, F lt , \ldots, F lT )
[0035] where Attn ft represents the attention map between the enhanced local feature and the local feature , \(F_{t}^{enhanced}\) represents the feature after fusing the enhanced local feature lt and the local feature , \(W\) is a learnable matrix, \(C\) is the number of channels, Concat represents the concatenation operation in the time dimension, and \(F_{t}^{aggregated}\) represents the output of the local feature aggregator;
[0036] (l)
[0036] (54) For the global feature vector Perform spatial average pooling operation and temporal average pooling operation respectively to obtain spatial global feature P space and temporal global feature P time ,
[0037]
[0038] where: AvgPool hw represents the average pooling operation along the spatial dimension, and AvgPool t represents the average pooling operation along the temporal dimension;
[0039] (55) Randomly initialize a global cue vector G through the global feature aggregator; perform cross-attention operation on the global cue vector G and the spatial global feature P space along the temporal dimension, and perform channel adjustment through a multi-layer perceptron; continue to perform cross-attention operation with the temporal global feature P time along the spatial dimension, and perform channel adjustment through a multi-layer perceptron to obtain the enhanced global feature
[0040]
[0041] where: TemporalAttn and SpatialAttn represent the cross-attention mechanism in the temporal dimension and the cross-attention mechanism in the spatial dimension respectively, and the temporal dimension is T;
[0042] (56) Use the enhanced global feature to adjust the global feature to obtain the output of the global cue module:
[0043]
[0044] F (g) = Concat(F g1 , F g2 , …, F gt , …F gT )
[0045] where: Attn at represents the attention map between the enhanced global feature and the global feature , and F gt represents the feature after fusing the enhanced global feature and the local feature , is the global feature at time t, and F (g) represents the output of the local feature aggregator;
[0046] (57) Dynamically weight F (l) and F (g) to obtain the fused feature vector F (0) :
[0047] F (0) = α × F (l) + β × F (g)
[0048] where: α and β are learnable weights, is the fused feature at time t.
[0049] This technical solution solves the problem of complementary fusion of local face features and global body features in multi-view scenarios by constructing a hierarchical feature fusion mechanism. Specifically, through a randomly initialized local cue vector, the dynamic screening of local features is realized by combining the cross-attention mechanism, and the information fusion across time frames is realized by using the self-attention mechanism, so that the local features enhance the temporal correlation while retaining the detailed information. In the processing of global features, the spatial distribution characteristics and temporal motion laws of the target are respectively extracted through the dual-path pooling operation in the spatial and temporal dimensions, and then the channel adaptive adjustment of the global features is realized through the multi-level cross-attention mechanism. Finally, a dynamic weighting strategy with learnable weights is adopted to realize the collaborative optimization of local detailed features and global motion features. This fusion method not only avoids the sensitivity of a single feature to view changes, but also suppresses the interference of irrelevant background features through the attention mechanism, significantly improving the robustness of cross-view target recognition.
[0050] Specifically, in the Step6, after reshaping and concatenating along the time axis, the missing frames are filled with zeros, and the fused feature time series is represented as represents the fused feature at time t in the b-th view video, B represents the total number of views of the video, b = 1, 2, …, B, t = 1, 2, …, T, and the missing frames include the video frames of a single target that are theoretically invisible in a certain view video, and also include the video frames of a single target that are theoretically visible in a certain view video but not recognized; input into a fully connected layer to obtain the target recognition probability The weighted binary cross-entropy is used to calculate the target recognition loss.
[0051] This technical solution constructs a time series by reshaping and concatenating fusion features to solve the problem of target trajectory breakage in multi-view videos. The zero-padding method is used to fill in the missing frames, which not only maintains the complete structure of the time series but also avoids interference from invalid data in model learning. Differentiated processing is carried out specifically for two types of missing scenarios (theoretically invisible and mis-identification) to enhance the model's adaptability to complex scenarios. The fusion features are mapped to target recognition probabilities through a fully connected layer, and combined with a weighted binary cross-entropy loss function to effectively balance the weight differences between positive and negative samples and strengthen the model's attention to difficult-to-recognize samples. Among them, the missing frames that are theoretically visible but not recognized are processed separately to avoid misjudgment by the model due to perspective differences; probability prediction is performed on the time series of the fusion features after zero-padding to establish a cross-video target association mechanism, providing reliable feature support for subsequent trajectory repair. The weighted binary cross-entropy loss function adjusts the sample weights to improve the model's robustness in scenarios with data imbalance, especially for the problem of large differences in the appearance frequencies of targets in multi-view videos.
[0052] Specifically, in the above Step 7, the missing trajectory segments of the fused feature time series F (T) are counted, and a short-term association modeling module is used to capture the correlation between adjacent front and rear trajectory segments covering all view videos, and track the theoretically visible missing frames, including the following steps:
[0053] Given the fused feature at time t in the b-th view video Calculate the face position offset from time t - 2 to time t - 1 in the b-th view video Concatenate all the in all view videos and input them into the short-term association modeling module, and introduce the fused feature time series F (BT ) as supplementary information, and use the self-attention mechanism to obtain the self-attention matrix to achieve the correlation measurement between different view videos:
[0054]
[0055] G = pool(δ(conv2d(F (BT) , K (g) )))
[0056]
[0057] Where: represents the concatenation operation, φ(·, ·) represents the linear transformation; conv2d represents the 2D convolution operation, K (g) is the weight parameter of the convolutional layer, δ(·) represents the non-linear activation function, pool(·) represents the pooling operation; Qt represents the query of the self-attention mechanism, K t represents the key value of the self-attention mechanism, A (att)Represents the asymmetric directed association between videos from different perspectives, D represents the dimension value; W (E) 、W (Q) and W (K) represent the weight matrices of linear transformations;
[0058] Perform association modeling through R cascaded 2D convolutional layers:
[0059] A r =δ(conv2d(A r-1 , K 1×K ) + conv2d(A r-1 , K K×1 ))
[0060] where: A l represents the output of the r-th 2D convolutional layer, r = 1, 2,..., R, A0 is initialized as A (att) ; conv2d represents the 2D convolution operation, K 1×K and K K×1 are 2D convolutional kernels with sizes 1×K and K×1 respectively;
[0061] Retain significant attention through the asymmetric adjacency association matrix A (adj) :
[0062] A (mask) =sgn(Sigmoid(A R ) - ξ)
[0063] A (adj) =A (maxk) ⊙A (att)
[0064] where: ξ ∈ [0, 1] is a set threshold; sgn(·) represents the sign function; ⊙ represents element-wise multiplication;
[0065] Calculate the prediction of the face position offset from time t - 1 to time t
[0066]
[0067] where: W (A) 、W (O) are the weight matrices of linear transformations;
[0068] According to the fused feature at time t Convert the face position offset prediction into the fused feature prediction at time t + 1 Furthermore, trace the missing frame at time t + 1 to complete object tracking, and use the IoU loss to calculate the object tracking loss.
[0069] This technical solution constructs a cross-view spatio-temporal correlation model to solve the problem of trajectory breakage caused by the temporary disappearance of the target in multi-view videos. By calculating the face position offset between consecutive frames in each view video, the multi-view offset information is serially input into the short-term correlation modeling module. Combining the context information of the fused feature time series, the self-attention mechanism is used to capture the potential correlations between different views. A cascaded 2D convolutional layer is used to deeply model the correlation information, and the strong-correlation cross-view trajectory segments are screened out through an asymmetric adjacency matrix, effectively avoiding noise interference. Based on linear transformation, the predicted value of the face position offset is generated, and the spatio-temporal features of the missing frames are deduced backward in combination with the current fused features, realizing the trajectory completion of the theoretically visible missing frames. By introducing the IoU loss function, the geometric matching accuracy between the predicted position and the actual position is ensured, forming a closed-loop optimized tracking mechanism.
[0070] A single-object tracking device in multi-view videos, comprising a feature extraction unit, a trajectory kernel generator, a local-global feature fusion unit, a fused feature reshaping and concatenation unit, a short-term correlation modeling module, and a parameter optimization unit;
[0071] The feature extraction unit is used to extract the face feature vector F, the body feature vector A of the target in each view video, and the background feature B of the corresponding view;
[0072] The trajectory kernel generator calculates the face position offset vector O according to the face feature vector F of adjacent frames (f) , and then obtains the 3D convolutional kernel K (f) ; calculates the body position offset vector O according to the body feature vector A of adjacent frames (α) , and then obtains the 3D convolutional kernel K (a) ;
[0073] The local-global feature fusion unit uses the 3D convolutional kernel K (f) to perform 3D convolution on the face feature vector F to obtain the local feature vector uses the 3D convolutional kernel K (a) to perform 3D convolution on the body feature vector A to obtain the global feature vector uses a feature aggregator to perform feature fusion on the local feature vector and the global feature vector to obtain the fused feature vector F (0) ;
[0074] The fused feature reshaping and concatenation unit differentiates different view videos based on the background feature B, and identifies the specified single target in all view videos based on the fused feature vector F (0) , and for the fused feature vectors F of all video frames where the specified single target is identified (0)Reshape and concatenate according to the time axis, and use the zero-padding method to supplement the missing frames to form the fused feature time series F (BT) ;
[0075] The short-term correlation modeling module counts the missing trajectory segments of the fused feature time series F (T) For the missing trajectory segments with a time length less than the set threshold, capture the correlation between the adjacent front and rear trajectory segments covering all perspective videos, and track the theoretically visible missing frames;
[0076] The parameter optimization unit optimizes the cross-video single-object tracking model composed of the feature extraction unit, the trajectory kernel generator, the local-global feature fusion unit, the fused feature reshaping and concatenation unit, and the short-term correlation modeling module by jointly optimizing the face recognition loss, the body recognition loss, the background recognition loss, the object recognition loss, and the object tracking loss.
[0077] This technical solution realizes a systematic solution for multi-perspective video single-object tracking by constructing a device containing a multi-level feature processing module. The feature extraction unit provides multi-dimensional feature support for subsequent cross-perspective correlation by separately extracting two biometric features of face and body and combining the perspective background features. The trajectory kernel generator generates a 3D convolutional kernel with spatio-temporal perception ability by calculating the position offset of adjacent frame feature vectors, which can effectively capture the motion trajectory of the target in consecutive frames. The local-global feature fusion unit extracts local detail features and global motion features respectively through 3D convolutional operations, and then performs cross-modal fusion through a feature aggregator to solve the problem of insufficient representation of a single feature in complex scenarios. The fused feature reshaping and concatenation unit uses time axis reshaping and zero-padding processing, which not only maintains the time continuity of multi-perspective videos, but also can effectively handle the feature loss caused by the temporary disappearance of the target. The short-term correlation modeling module realizes spatio-temporal correlation modeling of cross-perspective videos by introducing a self-attention mechanism and a 2D convolutional cascade structure. Especially for the missing frames that are theoretically visible but not recognized, the correlation between the front and rear trajectory segments is used for prediction and compensation. The parameter optimization unit unifies the optimization objectives of each module to the overall tracking performance improvement through a multi-task joint training strategy, enhancing the robustness and generalization ability of the model. The collaborative work of each module breaks through the limitations of traditional single-video tracking and constructs a complete technical link for cross-perspective spatio-temporal feature correlation.
[0078] Specifically, input the multi-perspective video into the trained cross-video single-object tracking model to identify and predict the start and end times and the positions where the specified single object appears in each perspective video.
[0079] This technical solution realizes the spatio-temporal positioning of the target by constructing a complete cross-video single-target tracking model system and inputting multi-view videos into the trained model. Its core lies in using the local-global feature fusion mechanism, multi-view background feature discrimination mechanism, and short-term correlation modeling ability of missing trajectory segments established in the previous model training stage to solve three key problems in cross-view tracking: First, by directly inheriting the view background feature discrimination ability formed in the training stage through multi-view video input, it solves the problem of target positioning deviation caused by the view difference of different cameras; Second, based on the recognition mechanism of the fused feature time series, it can effectively overcome the feature mismatch problem caused by occlusion or deformation of the target in a single view; Third, through the trajectory missing compensation mechanism built into the trained model, it can perform spatio-temporal trajectory prediction for the scenario where the target disappears briefly. The model application link in this step is essentially to systematically integrate the core modules such as multi-modal feature extraction, spatio-temporal convolution kernel generation, dynamic feature fusion, and cross-view correlation modeling jointly optimized in the training stage, and finally achieve the accurate prediction of the spatio-temporal trajectory of the specified target in multi-view videos. Among them, "identifying the start and end times" corresponds to the model's ability to judge the time sequence of the target appearance moment, and "appearance position" reflects the optimization effect of the model in cross-view spatial positioning.
[0080] Advantages: The method and device for single-target tracking in multi-view videos based on short-term modeling technology provided by the present invention have the following advantages compared with the prior art: 1. Multi-feature fusion: By fusing face features and body features, it effectively improves the accuracy of target recognition. Especially in the case of large changes in the target appearance, it can still maintain a stable tracking effect; 2. Short-term correlation modeling: By capturing the correlation between adjacent front and rear trajectory segments through the short-term correlation modeling module, it effectively solves the problem of the target disappearing briefly and improves the robustness of tracking; 3. Cross-video target association: By distinguishing different view videos through background features, it realizes the accurate association of cross-video targets and is applicable to multi-view video scenarios; 4. Joint optimization: By jointly optimizing multiple loss functions, it comprehensively improves the performance of the model and ensures high precision in target recognition and tracking; 5. Efficient calculation: Using technologies such as 3D convolution and attention mechanism, it has high calculation efficiency and is applicable to real-time application scenarios. Description of the Drawings
[0081] Figure 1 It is the implementation flowchart of the method of the present invention. Detailed Embodiments
[0082] The following specifically introduces the present invention in combination with the drawings and specific embodiments.
[0083] In the prior art, target tracking technology faces challenges in single target tracking in multi-view video scenarios in the field of computer vision. Traditional methods are mainly designed for single video single target or single video multi-target scenarios. When the target switches between different view videos, due to problems such as view differences, environmental changes, and temporary disappearance of the target, it is easy to cause tracking interruption or cross-video association failure. Cross-video tracking methods based on appearance feature matching are limited by feature distortion caused by view changes, while single-video tracking algorithms are difficult to handle spatio-temporal association problems between multiple views. For example, in a traffic monitoring scenario, when a vehicle passes through areas covered by multiple cameras, traditional methods cannot effectively associate the discontinuous trajectories captured by different cameras.
[0084] To solve the above problems, considering the complementarity between the local features and global motion patterns of the target under different views, spatio-temporal association is constructed through joint modeling of multi-modal features. First, analyze the core contradictions in target tracking interruption in multi-view videos: the sensitivity of single-view features to target appearance changes, and the weak coupling of cross-view spatio-temporal association. Therefore, a mechanism is proposed to separately extract the local features of the face and the global features of the body, establish a spatio-temporal kernel generation mechanism based on dynamic convolution, and construct a robust tracking model by fusing multi-view feature sequences. For the problem of trajectory breakage, a short-term association module based on the attention mechanism is designed to use the adaptive association of cross-view trajectory segments to compensate for missing frames.
[0085] Specifically, this case provides a device for single target tracking in multi-view videos, which mainly includes a feature extraction unit, a trajectory kernel generator, a local-global feature fusion unit, a fused feature reshaping and concatenation unit, a short-term association modeling module, and a parameter optimization unit; among them, the feature extraction unit, the trajectory kernel generator, the local-global feature fusion unit, the fused feature reshaping and concatenation unit, and the short-term association modeling module constitute a cross-video single target tracking model, and the parameter optimization unit trains and optimizes the cross-video single target tracking model through a joint loss. The process of implementing single target tracking in multi-view videos based on this device is as Figure 1 shown, and the specific implementation steps are described below to illustrate the present invention in detail.
[0086] Step1. Feature Extraction
[0087] Input videos of each view to the feature extraction unit, and use an encoder to extract the face feature vector F and body feature vector A of the target, as well as the background feature B of the corresponding view; the face feature and body feature respectively capture the facial and body information of the target, and the background feature is used to distinguish different view videos.
[0088] The face feature vector F refers to the deep features of the facial region extracted by a pre-trained convolutional neural network, which can be specifically implemented using the ResNet-50 network and is used to capture the detailed changes in the face of the target. The body feature vector A refers to the features of the body contour and motion pattern extracted by a three-dimensional convolutional network, which can be specifically implemented using the SlowFast network and is used to characterize the overall shape and motion trend of the target.
[0089] Step 2, Position Offset Calculation
[0090] According to the face feature vector F of adjacent frames, calculate the face position offset vector O (f) ; According to the body feature vector A of adjacent frames, calculate the body position offset vector O (a) . The position offset vector can reflect the motion trajectory of the target in the spatio-temporal domain and provides a basis for the subsequent generation of 3D convolutional kernels.
[0091] Step 3, Trajectory Kernel Generation
[0092] Use the trajectory kernel generator to obtain the 3D convolutional kernels K (f) and K (α) of the face position offset vector O (f) and K (a) ; The 3D convolutional kernel performs local convolution on the face features to retain detailed information and global convolution on the body features to capture the motion trend; The 3D convolutional kernel can capture the motion pattern of the target in the spatial and temporal dimensions and provides dynamic information for the subsequent fused feature vector.
[0093] The trajectory kernel generator consists of a 3D convolutional layer and a pooling layer; The pooling layer reduces the dimensionality of the input face position offset vector and body position offset vector, aggregates the displacement information of adjacent frames along the time axis, eliminates noise interference and retains the macroscopic trend of the target motion; The 3D convolutional layer then performs deep feature extraction on the pooled features, and its weight parameters are adaptively adjusted through end-to-end training, so that the generated 3D convolutional kernel can not only characterize the continuous displacement pattern of the target in a single-view video, but also capture the spatial correlation of the displacement vectors between different views.
[0094] K (f) = pool(conv3d(O (f) , K))
[0095] K (a) = pool(conv3d(o (a) , K))
[0096] Where: pool(·) represents the pooling operation, conv3d represents the 3D convolutional operation, and K represents the weight parameters of the 3D convolutional layer.
[0097] Step4. Extract local features and global features
[0098] Use the 3D convolution kernel K (f) Perform 3D convolution on the face feature vector F Obtain the local feature vector Use the 3D convolution kernel K (a) Perform 3D convolution on the body feature vector A to obtain the global feature vector
[0099] Apply the 3D convolution kernel K to the face feature vector F (f) When performing the sliding window convolution operation along the time dimension, through the 3D convolution kernel K (f) In the weight sharing mechanism between consecutive frames, capture the subtle feature changes caused by facial muscle movements. For example, within the time window of three adjacent frames, the 3D convolution kernel K (f) Can learn the spatio-temporal correlation pattern of the position changes of facial features during head rotation.
[0100] Apply the 3D convolution kernel K to the body feature vector A (a) When performing cross-joint feature fusion in the spatial dimension and trajectory modeling along the time dimension. For example, within the time span of five consecutive frames, the 3D convolution kernel K (a) Can capture the spatio-temporal correlation relationship between the limb swing amplitude and the movement direction.
[0101] Through this divide-and-conquer processing strategy, the local stability of facial features is enhanced through continuity modeling in the time dimension, while the global motion pattern of body features is fully expressed through correlation modeling in the spatial dimension.
[0102] Step5. Feature fusion
[0103] Use the feature aggregator to perform feature fusion on the local feature vector And the global feature vector To obtain the fused feature vector F (0) ; The feature aggregator is a multi-modal fusion module that combines cross-attention and self-attention mechanisms. It selects effective local information through cross-attention, fuses temporal features by combining self-attention, and extracts global features through spatial and temporal pooling at the same time to achieve multi-modal feature complementarity; The feature aggregator can be specifically implemented using a multi-head attention structure for integrating local details and global motion information; It includes the following steps:
[0104] (51) Randomly initialize a local prompt vector L through the local feature aggregator, copy the local prompt vector L into T copies along the time dimension, and denote the t-th local prompt vector as L t , t = 1, 2,..., T;
[0105] (52) Perform cross-attention operation on the local prompt vector L t and the local features to selectively aggregate valid information from the local features by the local prompt vector L t ; Concatenate the valid information aggregated at all times and perform self-attention operation to fuse the local prompt vectors from different video frames:
[0106]
[0107] where: CrossAttn and SelfAttn represent cross-attention mechanism and self-attention mechanism respectively, and MLP represents multi-layer perceptron; is the local feature at time t; represents the enhanced local feature vector, is the enhanced local feature at time t;
[0108] (53) Calculate the attention map between the enhanced local feature vector and the local feature , convert the attention map into attention weights using the softmax function, and fuse the local feature vector and the local feature
[0109]
[0110] F (l) = Concat(F l1 , F l2 , …, F lt , …, F lT )
[0111] where: Attn ft represents the attention map between the enhanced local feature and the local feature , F lt represents the feature after fusing the enhanced local feature and the local feature , is a learnable matrix, C is the number of channels, Concat represents the concatenation operation in the time dimension, and F (l) represents the output of the local feature aggregator;
[0112] (54) Perform spatial average pooling operation and temporal average pooling operation on the global feature vector respectively to obtain the spatial global feature P space and the temporal global feature P time ,
[0113]
[0114] Among them: AvgPool hw represents an average pooling operation performed along the spatial dimension, and AvgPool t represents an average pooling operation performed along the temporal dimension;
[0115] (55) Randomly initialize a global prompt vector G through a global feature aggregator; perform cross-attention operation on the global prompt vector G and the spatial global feature P in the temporal dimension space and perform channel adjustment through a multi-layer perceptron; continue to perform cross-attention operation with the temporal global feature p in the spatial dimension time and perform channel adjustment through a multi-layer perceptron to obtain an enhanced global feature
[0116]
[0117] Among them: TemporalAttn and SpatialAttn respectively represent the cross-attention mechanism in the temporal dimension and the cross-attention mechanism in the spatial dimension, and the temporal dimension is T;
[0118] (56) Use the enhanced global feature to adjust the global feature to obtain the output of the global prompt module:
[0119]
[0120] F (g) = Concat(F g1 , F g2 , …, F gt , …F gT )
[0121] Among them: Attn at represents the attention map between the enhanced global feature and the global feature , and F gt represents the feature after fusing the enhanced global feature and the local feature ; is the global feature at time t, and F (g) represents the output of the local feature aggregator;
[0122] (57) Dynamically weight F (l) and F (g) to obtain a fused feature vector F (0) :
[0123] F (0) = α × F (l) + β × F (g)
[0124] where: α and β are learnable weights, and F (0) = {F (0)1 , F (0)2 , …, F (0)t , …, F (0)T}, and F (0)t is the fused feature at time t.
[0125] Step6. Feature reshaping and concatenation
[0126] Distinguish videos from different perspectives based on the background feature B, and identify the specified single target in all perspective videos based on the fused feature vector F (0) Reshape and concatenate the fused feature vectors F of all video frames that identify the specified single target along the time axis, and fill in the missing frames with the zero-padding method to form the fused feature time series F (0) . (BT) .
[0127] Among them, the zero-padding method refers to filling the feature vectors of the missing frames with zero values in the time series dimension, which can be specifically implemented by tensor concatenation operations to maintain the integrity of the time series and avoid interference of invalid data on the model; theoretically invisible missing frames refer to video frames in which the target cannot be observed due to physical occlusion or perspective range limitation under a specific perspective, which can be pre-annotated through a geometric constraint model; theoretically visible but unrecognized missing frames refer to video frames in which the target exists in the video frame but is not captured by the object detection algorithm, which can be distinguished by the object detection confidence threshold.
[0128] Represent the fused feature time series as represents the fused feature at time t in the b-th perspective video, B represents the total number of perspectives of the video, b = 1, 2, …, B, t = 1, 2, …, T, and the missing frames include video frames in which the specified single target is theoretically invisible in a certain perspective video, and also include video frames in which the specified single target is theoretically visible but not recognized in a certain perspective video; input into a fully connected layer to obtain the target recognition probability Calculate the target recognition loss using weighted binary cross-entropy; the weighted binary cross-entropy loss function is realized by adjusting the positive and negative sample weight coefficients, and the weights can be dynamically set based on the sample occurrence frequency to alleviate the problem of unbalanced data distribution.
[0129] Step7. Short-term correlation modeling
[0130] Statistical fusion feature time series F (T) For the missing trajectory segments of (T) , for the missing trajectory segments with a time length less than the set threshold, the short-term correlation modeling module is used to capture the correlation between the adjacent front and rear trajectory segments covering all perspective videos, and track the theoretically visible missing trajectory segments.
[0131] The short-term correlation modeling module, a cross-perspective trajectory correlation unit based on an asymmetric adjacency matrix, can be specifically implemented by cascading 2D convolution and self-attention mechanism, and is used to capture the temporal continuity of adjacent trajectory segments. When a trajectory missing is detected in a certain perspective video, first extract the fusion features of the effective trajectory segments before and after the missing time window in this perspective. Construct a motion trajectory model by calculating the face position offset between consecutive valid frames, and splice the multi-perspective offset sequences in the channel dimension to form a three-dimensional tensor and input it into the correlation modeling module. This module uses the self-attention mechanism to establish the potential correlation between trajectory segments of different perspectives, and discovers cross-perspective trajectories with similar motion patterns through key-value pair matching. The cascaded 2D convolutional layer performs multi-scale analysis on the spatio-temporal correlation features. For example, the first layer uses a 3×3 convolutional kernel to extract local correlation patterns, and the subsequent layers use 1×1 convolutional kernels for feature recombination. The asymmetric adjacency matrix retains the strong correlation relationship with physical space consistency between cross-perspective trajectory segments through element-wise multiplication and threshold screening. Finally, the correlation features are mapped to the face position offset prediction value through a linear transformation, and the target space coordinates of the missing frames are deduced reversely in combination with the fusion features at the current moment to complete the trajectory filling. The execution process of the short-term correlation modeling module includes the following steps:
[0132] Given the fusion feature at time t in the b-th perspective video Calculate the face position offset from time t - 2 to time t - 1 in the b-th perspective video For all perspective videos After concatenation, input it into the short-term correlation modeling module, and introduce the statistical fusion feature time series F (BT) As supplementary information, use the self-attention mechanism to obtain the self-attention matrix to realize the correlation measurement between different perspective videos:
[0133]
[0134] G = pool(δ(conv2d(F (BT) , K (g) )))
[0135]
[0136] Where: represents the concatenation operation, φ(·, ·) represents the linear transformation; conv2d represents the 2D convolution operation, K (g)is the weight parameter of the convolutional layer, δ(·) represents the non-linear activation function, and pool(·) represents the pooling operation; Q t represents the query of the self-attention mechanism, K t represents the key value of the self-attention mechanism, A (att) represents the asymmetric directed association between videos from different perspectives, and D represents the dimension value; W (E) 、W (Q) and W (K) represent the weight matrices of the linear transformation;
[0137] Perform correlation modeling through R cascaded 2D convolutional layers:
[0138] A r = δ(conv2d(A r-1 , K 1×K ) + conv2d(A r-1 , K K×1 ))
[0139] where: A l represents the output of the r-th 2D convolutional layer, r = 1, 2,..., R, and A0 is initialized as A (att) ; conv2d represents the 2D convolution operation, K 1×K and K K×1 are 2D convolutional kernels with sizes of 1×K and K×1 respectively;
[0140] Retain significant attention through the asymmetric adjacency correlation matrix A (adj) :
[0141] A (mask) = sgn(Sigmoid(A R ) - ξ)
[0142] A (adj) = A (mask) ⊙ A (att)
[0143] where: ξ ∈ [0, 1] is the set threshold; sgn(·) represents the sign function; ⊙ represents element-wise multiplication;
[0144] Calculate the prediction of the face position offset from time t-1 to time t
[0145]
[0146] where: W (A) 、w (O) are the weight matrices of the linear transformation;
[0147] According to the fused feature at time t Predict the face position offset Convert to the fusion feature prediction at time t+1 Furthermore, the missing frame at time t+1 is traced to complete object tracking, and the IoU loss is used to calculate the object tracking loss.
[0148] Step8, Joint optimization
[0149] By jointly optimizing the face recognition loss, body recognition loss, background recognition loss, object recognition loss, and object tracking loss, train the cross-video single-object tracking model composed of Step1~Step7.
[0150] Step9, Cross-video single-object tracking model test
[0151] Input the multi-view video into the trained cross-video single-object tracking model to identify and predict the start and end times and the positions where the specified single object appears in each view video.
[0152] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form. Any technical solutions obtained by using equivalent replacements or equivalent transformations fall within the protection scope of the present invention.
Claims
1. A method for single-object tracking in multi-view video, characterized in that: It includes the following steps: Step 1: Input videos from each perspective, and use an encoder to extract the face feature vector F and body feature vector A of the target, as well as the background feature B of the corresponding perspective. Step 2. Calculate the face position offset vector O based on the face feature vector F of adjacent frames (f) ; Calculate the body position offset vector O based on the body feature vector A of adjacent frames (a) ; Step 3. Use the trajectory kernel generator to obtain the face position offset vector O (f) and the body position offset vector O (a) of the 3D convolution kernel K (f) and K (a) ; Step4. Use the 3D convolution kernel K (f) Perform 3D convolution on the face feature vector F to obtain a local feature vector Use the 3D convolution kernel K (a) Perform 3D convolution on the body feature vector A to obtain a global feature vector Step 5. Use a feature aggregator to perform feature fusion on the local feature vector and the global feature vector to obtain a fused feature vector F (0) ; Step 6: Distinguish videos from different perspectives based on the background feature B and based on the fusion feature vector F (0) Identify a specified single target in all perspective videos, and for the fusion feature vectors F of all video frames where the specified single target is identified (0) Reshape and concatenate them along the time axis to form a fusion feature time series F (BT) ; Step 7. Statistically analyze the fused feature time series F (T) For the missing trajectory segments, for those with a time length less than the set threshold, use the short-term correlation modeling module to capture the correlation between adjacent front and back trajectory segments and track the missing trajectory segments. Step 8: Train the cross-video single-object tracking model composed of Step 1 to Step 7 by jointly optimizing the face recognition loss, body recognition loss, background recognition loss, object recognition loss, and object tracking loss. Step 9: Input the multi-perspective videos into the trained cross-video single-object tracking model to identify and predict the start and end times and the positions where the specified single object appears in each perspective video.
2. The method for single-object tracking in multi-view video according to claim 1, characterized in that: In the said Step 3, the trajectory kernel generator is composed of a 3D convolutional layer and a pooling layer: K (f) = pool(conv3d(O (f) , K)) K (a) = pool(conv3d(o (a) , K)) Where: pool(·) represents the pooling operation, conv3d represents the 3D convolutional operation, and K represents the weight parameter of the 3D convolutional layer.
3. The method for single-object tracking in multi-view video according to claim 2, wherein: In the said Step 4, the 3D convolution kernel K is used (f) to perform 3D convolution on the face feature vector F to obtain a local feature vector Using the 3D convolution kernel K (α) to perform 3D convolution on the body feature vector A to obtain a global feature vector 4. The method for single-object tracking in multi-view video according to claim 1, characterized in that: In the said Step 5, a feature aggregator is used for the local feature vector and the global feature vector to perform feature fusion and obtain a fused feature vector F (0) , which includes the following steps: (51) Randomly initialize a local prompt vector L through a local feature aggregator, and copy the local prompt vector L into T copies along the time dimension. Denote the t-th local prompt vector as L t , where t = 1, 2, …, T; (52)For the local prompt vector L t and the local features perform a cross-attention operation to allow the local prompt vector L t to selectively aggregate valid information from the local features ; concatenate the valid information aggregated at all times and perform a self-attention operation to fuse the local prompt vectors from different video frames: Wherein, CrossAttn and SelfAttn respectively represent the cross-attention mechanism and the self-attention mechanism, and MLP represents the multi-layer perceptron; is the local feature at time t; represents the enhanced local feature vector, is the enhanced local feature at time t; (53) Calculate enhanced local feature vectors and local features to obtain an attention map between them, use the softmax function to convert the attention map into attention weights, and use the attention weights to fuse the local feature vectors and local features Among them: Attn ft represents the enhanced local feature and the local feature the attention map between them, F lt represents the enhanced local feature and the local feature the fused feature, is a learnable matrix, C is the number of channels, Concat represents the concatenation operation in the time dimension, F (l) represents the output of the local feature aggregator; (54)Perform spatial average pooling operation and temporal average pooling operation on the global feature vector respectively to obtain the spatial global feature P and the temporal global feature P space time , Where: AvgPool hw represents an average pooling operation along the spatial dimension, and AvgPool t represents an average pooling operation along the temporal dimension; (55) Randomly initialize a global prompt vector G through a global feature aggregator; perform cross-attention operations on the global prompt vector G and the spatial global feature P in the time dimension space Perform cross-attention operations and adjust the channels through a multi-layer perceptron; continue to perform cross-attention operations with the temporal global feature P in the spatial dimension time Perform cross-attention operations and adjust the channels through a multi-layer perceptron to obtain an enhanced global feature Where: TemporalAttn and SpatialAttn respectively represent the cross-attention mechanism in the time dimension and the cross-attention mechanism in the spatial dimension, and the time dimension is T; (56)Utilize enhanced global features Adjust global features Obtain the output of the global hint module: Among them: Attn at represents the enhanced global feature and the global feature the attention map between them, F gt represents the enhanced global feature and the local feature the fused feature, is the global feature at time t, F (g) represents the output of the local feature aggregator; (57) Dynamically weight F (l) and F (g) to obtain the fused feature vector F (0) : F (0) = α × F (l) + β × F (g) Where: α and β are learnable weights, F (0) ={F (0)1 , F (0)2 , …, F (0)t , …, F (0)T}, F (0)t is the fused feature at time t.
5. The method for single-object tracking in multi-view video according to claim 1, characterized in that: In the said Step 6, after reshaping and concatenating according to the time axis, the missing frames are filled by the zero-padding method, and the fused feature time series is expressed as represents the fused feature at the t-th moment in the b-th perspective video, B represents the total number of perspectives of the video, b = 1, 2, …, B, t = 1, 2, …, T, the missing frames include the video frames of a specified single target that are theoretically invisible in a certain perspective video, and also include the video frames of a specified single target that are theoretically visible but not recognized in a certain perspective video; is input into a fully connected layer to obtain the target recognition probability The weighted binary cross-entropy is used to calculate the target recognition loss.
6. The method for single-object tracking in multi-view video according to claim 5, wherein: In the said Step 7, count the missing trajectory segments of the fused feature time series F (T) and use the short-term correlation modeling module to capture the correlation between the adjacent front and rear trajectory segments covering all perspective videos, and track the theoretically visible missing frames, including the following steps: Given the fused feature at time t in the b-th perspective video Calculate the face position offset in the b-th perspective video from time t-2 to time t-l b = 1, 2, …, B; for all perspective videos After concatenation, input into the short-term correlation modeling module, and introduce the fused feature time series F (BT) As supplementary information, use the self-attention mechanism to obtain the self-attention matrix and realize the correlation measurement between different perspective videos: Wherein: represents a series operation, φ(·, ·) represents a linear transformation; conv2d represents a 2D convolution operation, K (g) is the weight parameter of the convolutional layer, δ(·) represents a non-linear activation function, pool(·) represents a pooling operation; Q t represents the query of the self-attention mechanism, K t represents the key-value of the self-attention mechanism, A (att) represents the asymmetric directed correlation between videos of different perspectives, D represents the dimension value; W (E) 、W (Q) and W (K) represent the weight matrices of the linear transformation; Perform correlation modeling through a cascaded R 2D convolutional layers: A r = δ(conv2d(A r-1 , K 1×K )) + conv2d(A r-1 , K K×1 )) Where: A l represents the output of the r-th 2D convolutional layer, r = 1, 2, …, R, and A0 is initialized as A (att) ; conv2d represents the 2D convolution operation, K 1×K and K K×1 are 2D convolution kernels with sizes of 1×K and K×1, respectively; Through the asymmetric adjacency correlation matrix A (adj) Retain significant attention: A (mask) = sgn(Sigmoid(A R ) - ξ) A (adj) = A (mask) ⊙A (att) Where: ξ∈[0, 1] is the set threshold; sgn(·) represents the sign function; ⊙ represents element-wise multiplication; Calculate the predicted face position offset from time t-1 to time t Where: W (A) and W (O) are weight matrices of linear transformations; According to the fused feature at time t , the face position offset prediction is transformed into the fused feature prediction at time t+1 Furthermore, the missing frame at time t+1 is traced, and the target tracking is completed. The IoU loss is used to calculate the target tracking loss.
7. A single-object tracking device in multi-view video, characterized in that: It includes a feature extraction unit, a trajectory kernel generator, a local-global feature fusion unit, a fusion feature reshaping and concatenation unit, a short-term correlation modeling module, and a parameter optimization unit; The said feature extraction unit is used to extract the face feature vector F and body feature vector A of the target in each perspective video, as well as the background feature B of the corresponding perspective. The trajectory kernel generator calculates a face position offset vector O based on the face feature vectors F of adjacent frames (f) , and further obtains a 3D convolutional kernel K (f) ; calculates a body position offset vector O based on the body feature vectors A of adjacent frames (a) , and further obtains a 3D convolutional kernel K (a) ; The local-global feature fusion unit uses a 3D convolution kernel K (f) to perform 3D convolution on the face feature vector F to obtain a local feature vector uses the 3D convolution kernel K (a) to perform 3D convolution on the body feature vector A to obtain a global feature vector uses a feature aggregator to perform feature fusion on the local feature vector and the global feature vector to obtain a fused feature vector F (0) ; The fusion feature reshaping and concatenation unit distinguishes videos from different perspectives based on the background feature B and based on the fusion feature vector F (0) identifies a specified single target in all perspective videos, and for the fusion feature vectors F of all video frames in which the specified single target is identified (0) is reshaped and concatenated according to the time axis, and the missing frames are filled by the zero-padding method to form a fusion feature time series F (BT) ; The short-term correlation modeling module statistically fuses the missing trajectory segments of the feature time series F (T) For the missing trajectory segments with a time length less than the set threshold, capture the correlation between the adjacent front and rear trajectory segments covering all perspective videos, and track the theoretically visible missing frames; The said parameter optimization unit optimizes the cross-video single-object tracking model composed of the feature extraction unit, the trajectory kernel generator, the local-global feature fusion unit, the fusion feature reshaping and concatenation unit, and the short-term correlation modeling module by jointly optimizing the face recognition loss, body recognition loss, background recognition loss, object recognition loss, and object tracking loss.
8. The single-object tracking device in multi-view video according to claim 7, characterized in that: Input the multi-perspective videos into the trained cross-video single-object tracking model to identify and predict the start and end times and the positions where the specified single object appears in each perspective video.
Citation Information
Cited By
Multi-sensor target trajectory fusion modeling method, fusion method, equipment and storage medium
CN120524315A