Intelligent visual tracking method fusing memory distillation and space-time alignment
By integrating memory distillation with space-time alignment technology, the stability and real-time problems of visual tracking in complex scenarios are solved, and an intelligent visual tracking method with high precision, robustness and efficient computing is realized.
Patent Information
- Application Number
- CN202510153444.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing visual tracking methods are difficult to maintain stable tracking performance in complex scenarios, especially in the case of large apparent changes in the target, high motion uncertainty and long-term occlusion, and the computing resources consume too much, making it difficult to meet the real-time requirements.
The intelligent visual tracking method that integrates memory distillation and space-time alignment is adopted to enhance the adaptability and robustness of the tracking algorithm through the adaptive space-time consistency attention mechanism and the feature management mechanism based on dynamic memory bank, and optimize the performance and computational efficiency of the tracking model through space-time consistency loss function and knowledge distillation technology.
It significantly improves tracking accuracy, long-term tracking performance, timing consistency and computing efficiency, and can maintain stable and real-time tracking performance in complex scenarios. It is suitable for video surveillance, driverless driving and robot vision.
Smart Images

Figure CN120070503A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a visual object tracking method based on deep learning, especially an intelligent visual tracking method that integrates memory distillation and spatio-temporal alignment. By innovatively combining the memory distillation mechanism and spatio-temporal feature alignment technology, the present invention realizes accurate, stable and real-time tracking of moving objects. This method has strong robustness in complex scenarios and can effectively handle challenging problems such as target occlusion, deformation, and motion blur, and has important application value in fields such as video surveillance, unmanned driving, and robot vision. Background Art
[0002] Object tracking is an important research direction in the field of computer vision, and its goal is to continuously locate and track a specified object in a video sequence. In recent years, with the development of deep learning technology, significant progress has been made in object tracking algorithms. Early object tracking methods were mainly based on traditional algorithms such as correlation filtering and particle filtering, which were computationally efficient but lacked robustness. Subsequently, deep learning-based tracking algorithms gradually became the mainstream. For example, the Siamese network series of algorithms learn the similarity measure of object appearance features through the Siamese network structure, realizing end-to-end object tracking.
[0003] However, existing visual tracking methods still face many challenges when dealing with complex scenarios. First, drastic changes in object appearance and occlusion interference can lead to tracking failure. Since the object may undergo appearance changes such as deformation, rotation, and scale change during movement, or be partially or completely occluded by other objects, existing methods are difficult to maintain stable tracking performance. Second, complex background interference and similar object interference are also important factors affecting tracking accuracy. When there are interfering objects similar to the target in the scene, it is easy to cause the tracker to drift or switch IDs. In addition, imaging factors such as lighting changes, motion blur, and low resolution will also reduce the discriminability of feature representation and increase the tracking difficulty.
[0004] In long-term tracking scenarios, the target may reappear after completely leaving the field of view, which poses higher requirements for the re-identification ability of tracking algorithms. At the same time, in practical applications, it is often necessary to track multiple targets simultaneously. The mutual occlusion and interaction between targets make the multi-target tracking problem more challenging. In addition, excessive computational resource consumption is another important challenge restricting the practical application of tracking algorithms. Existing deep learning-based tracking methods generally adopt complex network structures and heavy feature extraction processes, which lead to large computational overheads of the algorithms and are difficult to meet the real-time requirements in practical applications. Especially on resource-constrained embedded devices, how to reduce the computational complexity while ensuring the tracking performance is an urgent problem to be solved. At the same time, the storage overhead of large models also poses challenges to practical deployment, and a more lightweight and efficient network structure needs to be designed. These problems seriously restrict the wide application of visual tracking technology in practical scenarios. Summary of the Invention
[0005] The object of the present invention is to provide an intelligent visual tracking method integrating memory distillation and spatio-temporal alignment to solve challenging problems such as large target appearance changes, high motion uncertainty, and long-term occlusion in the prior art. Specifically, the present invention aims to achieve the following technical objectives: First, an adaptive spatio-temporal consistency attention mechanism is proposed. Through the collaborative action of the temporal attention branch and the spatial attention branch, it effectively models the dependencies of the target in the temporal and spatial dimensions, enhancing the adaptability of the tracking algorithm to target motion changes. Second, a feature optimization framework based on memory distillation is designed. A dynamic memory bank structure is introduced to store and manage target features, and the feature representation ability of the teacher network is transferred to the student network through the knowledge distillation mechanism, improving the robustness and reliability of the tracking model in long-term tracking scenarios. Third, a spatio-temporal consistency loss function is constructed, including trajectory smoothing loss, speed matching loss, and feature consistency loss. These losses are jointly optimized through multi-task learning to maintain the coherence of the tracking results in time and space, and improve the smoothness and accuracy of the tracking trajectory. Fourth, a lightweight backbone network and model compression technology are adopted to reduce the computational complexity without significantly sacrificing performance, realizing real-time inference of the tracking algorithm. In summary, the present invention strives to start from multiple aspects such as spatio-temporal modeling, memory optimization, consistency constraint, and model acceleration to comprehensively improve the performance and efficiency of visual tracking, providing strong support for the application and deployment of intelligent visual systems.
[0006] To achieve the above objectives, the present invention adopts the following technical solutions:
[0007] S1: Construct a feature extraction module based on a shared backbone network, and the feature extraction module includes:
[0008] S1.1: First, perform preprocessing on the input target template frame and search frame by dividing them into image patches. Specifically, divide the image into a sequence of non-overlapping image patches of size 16×16 pixels. For the template frame, its image patch sequence is represented as where H r and W r are the height and width of the template frame respectively. For the search frame, its image patch sequence is represented as where H s and W s are the height and width of the search frame respectively.
[0009] S1.2: Next, extract features from the divided image patch sequences through a shared backbone network. For each image patch of the template frame obtain the corresponding feature representation through the feature extraction function φ(·) where i = 1, 2,..., N z . Similarly, for each image patch of the search frame obtain the feature representation where j = 1, 2,..., N x , and the feature dimension C is uniformly set to 256.
[0010] S1.3: Finally, construct a multi-level feature pyramid structure to obtain multi-scale feature representations. This structure contains 5 feature layers from shallow to deep, and each feature layer performs channel dimensionality reduction through 1×1 convolution operations to obtain feature maps with different spatial resolutions. The specific feature compression ratios from shallow to deep are
[0011] S2: Construct an Adaptive Spatio-Temporal Consistency Attention (ASTCA) module, and the ASTCA module includes:
[0012] S2.1: Design parallel temporal attention branch and spatial attention branch. This module uses the multi-head attention mechanism for feature modeling, with the input feature dimension D in = C = 256, the number of attention heads h = 8, and the feature dimension d of each attention head = D in / h = 32.
[0013] S2.2: Calculation process of the temporal attention branch. This branch is mainly used to model the cross-frame temporal dependency relationship. First, generate the query matrix q , key matrix k and value matrix v respectively through the learnable linear transformation matrices W , key matrix and value matrix where X prevRepresents the features of the previous frame. Then, calculate the attention weights Finally, obtain the temporal features
[0014] S2.3: Spatial attention branch calculation process. This branch is mainly used to capture the target-background association within the frame. Similarly, generate the query matrix through a linear transformation Key matrix And value matrix Calculate the spatial attention weights Obtain the spatial features
[0015] S2.4: Feature dynamic fusion process. To adaptively fuse the temporal and spatial features, a weight learning mechanism based on a multi-layer perceptron (MLP) is designed. Concatenate the temporal feature F t And the spatial feature F s And input them into the MLP network. The MLP adopts a three-layer structure, and the hidden layer dimensions are 512, 256, and 2 in sequence. Obtain the fusion weights through the softmax function The final fused feature is obtained through weighted summation: F = w 1 F t + w 2 F s , where w 1 + w 2 = 1, ensuring weight normalization.
[0016] S3: Construct a spatio-temporal memory distillation network (STMDN), and the STMDN includes:
[0017] S3.1: Design of the dynamic memory bank structure. The present invention designs a feature storage and management mechanism based on a dynamic memory bank. The memory bank consists of K memory slots, where K = 64 is the total number of memory slots. Each memory slot is used to store and maintain the appearance features and motion feature information of the target. Mathematically, the memory bank can be represented as a matrix Where Represents the C-dimensional feature vector maintained by the k-th memory slot. The memory bank adopts a dynamic update mechanism, and determines the update strategy by calculating the similarity between the current frame features and the historical memory slots in real time. Specifically, for the current frame features First, calculate its cosine similarity s k = cos(f t , m k ). Then, select the most similar memory slot for update according to the similarity score, and the update formula is Where α ∈ [0, 1] is the dynamic update rate.
[0018] S3.2: Implementation of the gated read-write mechanism. This mechanism consists of two core operations: reading and writing. In the reading operation, first, the memory content is linearly transformed through a learnable weight matrix and a bias vector , and then a read gate signal is generated through the sigmoid activation function σ(·). Then, a candidate content is generated using the weight matrix and the bias vector , where tanh(·) is the hyperbolic tangent activation function. The final reading result is obtained through the element-wise product of the gated unit and the candidate content, realizing the selective reading of the memory content. The writing operation adopts a similar gated mechanism. First, the write gate signal is calculated, where and are learnable parameters. Then, a candidate update content is generated, where and are learnable parameters. Finally, the memory content is updated through the gated update mechanism , where ⊙ represents the Hadamard product (element-wise product).
[0019] S3.3: Design of the knowledge distillation mechanism. The present invention designs a knowledge transfer mechanism based on feature distillation to guide the student network to learn better feature representations through the teacher network. Specifically, the teacher network and the student network respectively extract features from the input image to generate feature maps , where C is the number of feature channels, and H and W respectively represent the spatial height and width of the feature map. To make the knowledge transfer process smoother, a temperature coefficient T is introduced in the distillation process for feature softening. In the present invention, T = 4 is set. For the softened feature map, the L 2 distance between the output features of the teacher network and the student network is calculated to construct the distillation loss function: where and respectively represent the feature values of the teacher network and the student network at the position (c, h, w). The present invention introduces an attention guidance mechanism in the distillation process to guide the student network to focus on important feature channels by calculating the channel attention weights of the teacher network feature map.
[0020] S4: Construct a spatio-temporal consistency (STC) loss function, and the STC loss function includes:
[0021] S4.1: Trajectory smoothing loss. To ensure the smoothness of the target trajectory, the present invention designs a trajectory smoothing loss function based on the L1 norm. The target position is represented by the bounding box parameter vector b t =[x t , y t , w t , h t , where (x t , y t ) is the target center coordinate, and (w t , h t ) are the target width and height. The trajectory smoothing loss function is defined as the L1 distance between the bounding box parameters of adjacent frames: L smooth =||b t -b t-1 || 1 .
[0022] S4.2: Velocity matching loss. To constrain the consistency of the target motion, the present invention introduces a velocity matching loss function based on the L2 norm. The instantaneous velocity of the target at time t is defined as the difference in the bounding box parameters of adjacent frames: v t =(b t -b t-1 ) / Δt, where Δt is the inter-frame time interval. The velocity matching loss function is defined as the L2 distance between the velocity vectors of adjacent frames: L vel =||v t -v t-1 || 2 .
[0023] S4.3: Feature consistency loss. First, perform L2 normalization on the feature vector: f t =f t / ‖f t ‖ 2 . Then calculate the cosine similarity of the normalized feature vectors of adjacent frames. The feature consistency loss function is defined as: L feat =1 - cos(f t , f t-1 ). This loss function can constrain the gradual change of the target appearance feature and improve the stability of tracking.
[0024] S4.4: Joint optimization process. The present invention adopts a multi-task learning framework to jointly optimize the classification loss, regression loss, and spatio-temporal consistency loss. The classification loss uses the binary cross-entropy function: L cls =BCE(p, y), where p is the predicted probability and y is the true label. The regression loss uses the L1 distance function: L reg =||b - b gt || 1 , where b gtare the true bounding box parameters. The spatio-temporal consistency loss is the weighted sum of the above three loss terms: L stc = L smooth + L vel + L feat . The final combined loss function is defined as: L = L cls + λ 1 L reg + λ 2 L stc , where the weight coefficients λ 1 = 5 and λ 2 = 2 are used to balance the contributions of each loss term.
[0025] Compared with the prior art, the present invention has significant advantages in terms of tracking accuracy, long-term tracking performance, temporal consistency, and computational efficiency.
[0026] First, the present invention innovatively combines an adaptive spatio-temporal attention mechanism and a memory distillation network, effectively improving the accuracy of the tracking algorithm. The adaptive spatio-temporal attention mechanism can dynamically adjust the attention weights in the time and space dimensions according to the target motion state, enabling the algorithm to adaptively focus on the key regions and key frames of the target, thereby more accurately capturing the motion changes of the target. The memory distillation network guides the student network to learn more robust and discriminative feature representations through the teacher network, further enhancing the adaptability of the tracking algorithm to target appearance changes. The combination of the two has achieved a significant improvement in accuracy on the standard tracking benchmark dataset.
[0027] Second, the present invention introduces a feature management mechanism based on a dynamic memory bank, significantly improving the performance of the tracking algorithm in long-term occlusion scenarios. Traditional tracking algorithms usually rely only on the current frame or short-term memory to update the target appearance model, resulting in the easy loss of the target when long-term occlusion occurs. However, through the dynamic memory bank mechanism of the present invention, it can adaptively store and update diverse appearance features of the target, forming a long-term memory ability. When occlusion occurs, the historical appearance features stored in the memory bank can provide key clues for target re-identification, enabling the tracker to quickly re-locate the target when the target reappears. Therefore, the long-term tracking success rate of the present invention is significantly higher than that of existing methods.
[0028] Thirdly, the present invention designs a spatio-temporal consistency loss function, which significantly improves the smoothness and continuity of the tracking trajectory. Existing tracking algorithms usually process each frame independently, ignoring the temporal correlation between adjacent frames, resulting in jitter and incoherence problems in the generated tracking trajectories. The present invention establishes a smoothness penalty mechanism between the prediction results of adjacent frames by introducing spatio-temporal consistency constraints. On the one hand, by minimizing the L1 distance of the bounding box parameters of adjacent frames, the generated target trajectory becomes smoother; on the other hand, by minimizing the L2 distance of the target speed of adjacent frames, the predicted target motion becomes more continuous. Quantitative evaluations on standard tracking datasets show that both the trajectory smoothness and the target motion prediction accuracy of the present invention are significantly better than existing methods.
[0029] Finally, the present invention adopts knowledge distillation technology to compress the model size, significantly improving the computational efficiency while maintaining the tracking performance. Specifically, by introducing a teacher-student network architecture, a large-scale pre-trained teacher network is used to guide the training process of a small student network. During the distillation process, the knowledge of the teacher network is transferred to the student network in the form of soft labels, guiding the student network to learn more compact and discriminative feature representations. At the same time, through channel recalibration of the feature map and non-local attention mechanism, the key knowledge of the teacher network is further refined. Experimental results show that the student network after knowledge distillation can still achieve similar tracking performance to the teacher network with a significant reduction in the number of parameters, demonstrating the effectiveness of knowledge distillation technology in tracking model compression. The tracking speed of the present invention meets the real-time requirement and can meet the actual application needs.
[0030] In summary, by integrating memory distillation and spatio-temporal alignment technologies, the present invention has made significant improvements in multiple aspects such as tracking accuracy, long-term tracking ability, temporal consistency, and computational efficiency, representing an important advancement in intelligent visual tracking technology and providing an effective solution for stable, accurate, and real-time object tracking in complex scenarios. Brief Description of the Drawings
[0031] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0032] Figure 1 is the overall flowchart of the embodiment of the present invention;
[0033] Figure 2 is the structural schematic diagram of the overall network model of the embodiment of the present invention;
[0034] Figure 3 is the structural schematic diagram of the adaptive spatio-temporal consistency attention module of the embodiment of the present invention;
[0035] Figure 4Schematic diagram of the spatio-temporal memory distillation network according to the embodiment of the present invention;
[0036] Figure 5 Line chart of the overall performance comparison between the embodiment of the present invention and existing tracking algorithms on the LaSOT tracking dataset;
[0037] Figure 6 Radar chart of the attribute-based comparison between the embodiment of the present invention and existing tracking algorithms on the LaSOT tracking dataset;
[0038] Figure 7 Schematic diagram of the tracking results of several typical sequences in the LaSOT tracking dataset according to the embodiment of the present invention. Detailed implementation manners
[0039] Embodiment 1: Overall network architecture
[0040] As Figure 1 shown, the intelligent visual tracking method integrating memory distillation and spatio-temporal alignment proposed by the present invention includes the following main steps:
[0041] 1. Feature extraction stage:
[0042] First, the input target template frame and search frame are subjected to feature extraction through a shared ViT-Base backbone network. As Figure 2 shown, this network first divides the input image into image patches of size 16×16, and then performs feature extraction through 12 layers of Transformer encoders to generate multi-scale feature maps. For the template frame, a feature sequence is generated. For the search frame, a feature sequence
[0043] 2. Spatio-temporal attention modeling stage:
[0044] As Figure 3 shown, the Adaptive Spatio-Temporal Consistency Attention (ASTCA) module contains two parallel branches. The temporal attention branch models the temporal dependence relationship by calculating the attention weights between the features of the current frame and the features of the historical frames. The spatial attention branch captures the spatial dependence relationship by calculating the attention weights inside the current frame. The output features of the two branches are adaptively fused through learnable weight coefficients w 1 and w 2 : F = w 1 F t + w 2 F s .
[0045] 3. Memory distillation optimization stage:
[0046] As Figure 4 shown, the Spatiotemporal Memory Distillation Network (STMDN) consists of the following key components. The dynamic memory bank maintains K = 64 memory slots, and each memory slot stores the appearance and motion features of the target. The gated read-write mechanism controls the access and update of the memory through the read gate r = σ(W r m + b r ) and the write gate w = σ(W w m + b w ). During the knowledge distillation process, the teacher network (ResNet-50) guides the student network (MobileNetV2) to learn discriminative feature representations.
[0047] Example 2: Tracking Performance Evaluation
[0048] The present invention has carried out a comprehensive performance evaluation on the LaSOT long-term tracking dataset. As Figure 5 shown, the tracking algorithm of the present invention has achieved significant improvements in two key metrics: Precision and Success Rate. Specifically, the precision of the present invention reaches 0.823, which is a 6% improvement compared to the existing best-performing tracking algorithm; in terms of the success rate, the present invention reaches a level of 0.727, which is a 2.8% improvement compared to the existing best method. These quantitative results fully demonstrate the superior performance of the present invention in long-term tracking tasks.
[0049] From Figure 5 the performance comparison curves, it can be seen that the tracking algorithm of the present invention maintains an obvious leading advantage under different overlap thresholds. Especially in the high overlap threshold range (such as IoU > 0.6), the precision curve and success rate curve of the present invention are significantly higher than those of other algorithms, indicating that the present invention can generate more accurate and complete target bounding boxes. This benefits from the innovative designs of the present invention in feature representation, spatiotemporal modeling, and memory optimization, enabling the tracker to better adapt to target appearance changes and cope with complex tracking environments.
[0050] Generally speaking, as one of the most challenging long-term tracking datasets currently, LaSOT covers a large number of real-scene videos, posing a severe test on the robustness and generalization ability of tracking algorithms. The excellent results achieved by the present invention on this dataset fully illustrate its great potential in practical applications, providing an efficient and reliable new idea for solving long-term tracking problems.
[0051] As Figure 6 shown, the present invention shows strong robustness under various challenging attributes. Specifically:
[0052] In the full occlusion scenario, thanks to the dynamic memory mechanism, the tracking success rate of the present invention reaches 0.646, which is attributed to the target history features stored in the memory, which enables the tracker to quickly relocate when the target reappears.
[0053] For the deformation attribute, the tracking success rate of the present invention is 0.743, which is attributed to the adaptive spatiotemporal attention mechanism that can dynamically focus on the key areas of the target and effectively adapt to the non-rigid deformation of the target.
[0054] In the fast motion scenario, the tracking success rate of the present invention reaches 0.607, which is attributed to the constraint of spatiotemporal consistency loss, which makes the generated tracking trajectory smoother and more coherent, thereby improving the tracking accuracy of the moving target.
[0055] For the background clutter problem, the tracking success rate of the present invention is 0.675, which is attributed to the memory distillation learning that integrates spatiotemporal information to enhance the tracker's ability to distinguish between the target and the background.
[0056] In terms of scale variation, the tracking success rate of the present invention reaches 0.726, which is attributed to the extraction and fusion strategy of multi-scale features that enables the tracker to accurately estimate the scale variation of the target.
[0057] In summary, the present invention has achieved significant performance improvements under various tracking challenges, showing excellent robustness and generalization capabilities. These experimental results fully demonstrate the effectiveness and advancement of the present invention in solving the challenges of complex tracking environments.
[0058] Example 3: Typical scenario analysis
[0059] like Figure 7 As shown, the present invention exhibits excellent performance on a number of representative tracking sequences:
[0060] Occlusion scenario: When the target is blocked by other objects, the historical features stored in the memory bank can quickly resume tracking when the target reappears
[0061] Deformation scenario: Adaptive spatiotemporal attention mechanism can dynamically focus on the key areas of the target and adapt to the deformation changes of the target
[0062] Fast motion scenes: The constraint of spatiotemporal consistency loss makes the tracking trajectory smoother and more coherent.
[0063] It should be noted that in this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0064] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An intelligent visual tracking method integrating memory distillation and spatiotemporal alignment, characterized in that: The following steps are involved: S1: Construct a feature extraction module based on a shared backbone network, the feature extraction module includes: S1.1: First, the input target template frame and search frame are preprocessed by image block division. Specifically, the image is divided into a sequence of non-overlapping image blocks of size 16×16 pixels. For the template frame, its image block sequence is represented as in H r and W r are the height and width of the template frame respectively. For the search frame, its image block sequence is represented as in H s and W s are the height and width of the search frame, respectively. S1.2: Next, the shared backbone network is used to extract features from the divided image block sequence. The corresponding feature representation is obtained through the feature extraction function φ(·) where i = 1, 2, ..., N z Similarly, for each image block of the search frame Get feature representation where j = 1, 2, ..., N x ,The feature dimension C is uniformly set to 256. S1.3: Finally, a multi-level feature pyramid structure is constructed to obtain multi-scale feature representation. The structure contains 5 feature layers from shallow to deep. Each feature layer performs channel dimension reduction through 1×1 convolution operation to obtain feature maps with different spatial resolutions. The specific feature compression ratios from shallow to deep are: S2: Construct an adaptive spatiotemporal consistency attention (ASTCA) module, the ASTCA module includes: S2.1: Design parallel temporal attention branches and spatial attention branches. This module uses a multi-head attention mechanism for feature modeling, with an input feature dimension D in = C = 256, the number of attention heads h = 8, and the feature dimension of each attention head d = D in / h=32. S2.2: Temporal attention branch calculation process. This branch is mainly used to model the temporal dependencies across frames. First, through the learnable linear transformation matrix W q , W k and W v Generate query matrices separately Key Matrix Sum Matrix Where X prev represents the features of the previous frame. Then, the attention weight is calculated Finally, the time characteristics S2.3: Spatial attention branch calculation process. This branch is mainly used to capture the object-background association within the frame. Similarly, the query matrix is generated by linear transformation Key Matrix Sum Matrix Calculating spatial attention weights Get spatial features S2.4: Dynamic feature fusion process. In order to adaptively fuse temporal and spatial features, a weight learning mechanism based on multi-layer perceptron (MLP) is designed. t and spatial feature F s After concatenation, the data is input into the MLP network. The MLP adopts a three-layer structure, and the hidden layer dimensions are 512, 256, and 2 respectively. The fusion weight is obtained through the softmax function The final fusion feature is obtained by weighted summation: F = w1F t +w2F s , where w1+w2=1, ensuring weight normalization. S3: Construct a spatiotemporal memory distillation network (STMDN), wherein the STMDN includes: S3.1: Dynamic memory library structure design. The present invention designs a feature storage and management mechanism based on a dynamic memory library. The memory library consists of K memory slots, where K = 64 is the total number of memory slots. Each memory slot is used to store and maintain the surface features and motion feature information of the target. From a mathematical point of view, the memory library can be represented as a matrix in represents the C-dimensional feature vector maintained by the kth memory slot. The memory library adopts a dynamic update mechanism, which determines the update strategy by calculating the similarity between the current frame feature and the historical memory slot in real time. Specifically, for the current frame feature First, calculate the cosine similarity s with all memory slots k =cos(f t , m k ). Then, according to the similarity score, the closest memory slot is selected for update. The update formula is Where α∈[0, 1] is the dynamic update rate. S3.2: Implementation of gated read and write mechanism. This mechanism includes two core operations: read and write. In the read operation, firstly, the learnable weight matrix and the bias vector Memory content Perform a linear transformation and then generate a read gating signal through the sigmoid activation function σ(·) Next, use the weight matrix and the bias vector Generate candidate content Where tanh(·) is the hyperbolic tangent activation function. The final reading result is the element-by-element product of the gate unit and the candidate content. The write operation uses a similar gating mechanism. First, the write gating signal is calculated. in and is a learnable parameter. Then generate candidate update content in and is a learnable parameter. Finally, through the gated update mechanism Complete the update of the memory content, where ⊙ represents the Hadamard product (element-wise product). S3.3: Design of knowledge distillation mechanism. This paper designs a knowledge transfer mechanism based on feature distillation, which guides the student network to learn better feature representation through the teacher network. Specifically, the teacher network and the student network extract features from the input image and generate feature maps. Where C is the number of feature channels, H and W represent the spatial height and width of the feature map, respectively. In order to make the knowledge transfer process smoother, a temperature coefficient T is introduced in the distillation process to soften the features. In the present invention, T=4 is set. For the softened feature map, the distillation loss function is constructed by calculating the L2 distance between the output features of the teacher network and the student network: in and Respectively represent the characteristic values of the teacher network and the student network at the position (c, h, w). The present invention introduces an attention guidance mechanism in the distillation process, by calculating the channel attention weight of the teacher network feature map To guide the student network to focus on important feature channels. S4: Construct a spatiotemporal consistency (STC) loss function, wherein the STC loss function includes: S4.1: Trajectory smoothing loss. In order to ensure the smoothness of the target trajectory, the present invention designs a trajectory smoothing loss function based on the L1 norm. The target position is calculated by the bounding box parameter vector b t =[x t ,y t , w t ,h t ] means, where (x t ,y t ) is the target center coordinate, (w t ,h t ) are the target width and height. The trajectory smoothness loss function is defined as the L1 distance of the bounding box parameters of adjacent frames: L smooth =||b t -b t-1 ||1. S4.2: Speed matching loss. In order to constrain the consistency of the target motion, the present invention introduces a speed matching loss function based on the L2 norm. The instantaneous speed of the target at time t is defined as the difference of the bounding box parameters of adjacent frames: v t =(b t -b t-1 ) / Δt, where Δt is the time interval between frames. The velocity matching loss function is defined as the L2 distance of the velocity vectors of adjacent frames: L vel =||v t -v t-1 ||2. S4.3: Feature consistency loss. First, the feature vector is L2 normalized: f t =f t / ||f t ||2. Then the cosine similarity of the normalized feature vectors of adjacent frames is calculated, and the feature consistency loss function is defined as: L feat =1-cos(f t , f t-1 ). This loss function can constrain the gradient of the target’s apparent features and improve the stability of tracking. S4.4: Joint optimization process. The present invention adopts a multi-task learning framework to jointly optimize the classification loss, regression loss and spatiotemporal consistency loss. The classification loss uses a binary cross entropy function: L cls =BCE(p, y), where p is the predicted probability and y is the true label. The regression loss uses the L1 distance function: L reg =||bb gt ||1, where b gt is the true bounding box parameter. The spatiotemporal consistency loss is the weighted sum of the above three loss terms: L stc =L smooth +L vel +L feat The final joint loss function is defined as: L = L cls +λ1L reg +λ2L stc , where the weight coefficients λ1=5 and λ2=2 are used to balance the contribution of each loss term.
2. The intelligent visual tracking method integrating memory distillation and spatiotemporal alignment according to claim 1, characterized in that: The shared backbone network adopts the Vision Transformer (ViT-Base) architecture as the feature extraction backbone. The network receives an RGB image of size 224×224×3 as input, and first divides the input image into a sequence of image blocks of size 16×16. Each image block is mapped to a feature space of D=768 dimensions through a linear projection layer. The network contains 12 standard Transformer encoder layers, each of which contains 12 attention heads for modeling long-range dependencies between image blocks. In order to maintain position information, the network introduces a learnable absolute position encoding, which is optimized by back propagation. During the feature extraction process, the network generates multi-scale feature maps layer by layer, and its downsampling rate is from shallow to deep. The corresponding number of feature channels is {64, 128, 256, 512, 1024}. This multi-scale feature representation can effectively capture the semantic information of different levels of the target and provide rich feature representation for subsequent target positioning and tracking.
3. The intelligent visual tracking method integrating memory distillation and spatiotemporal alignment according to claim 1, characterized in that: The calculation process of the multi-head attention mechanism includes the following steps: First, for the input feature matrix Where N represents the sequence length and D represents the feature dimension. It is divided into h attention heads for parallel calculation. The feature dimension of each attention head is d = D / h. Then, through three learnable linear transformation matrices and Generate query matrix, key matrix and value matrix respectively. Specifically, query matrix Q = XW Q , key matrix K = XW K , value matrix V = XW V Then, calculate the attention weight matrix where divided by This is to alleviate the vanishing gradient problem. Multiply the attention weight matrix by the value matrix to get the attention output H = AV. Finally, concatenate the outputs of the h attention heads and output the projection matrix Perform linear transformation to obtain the final multi-head attention output MultiHead(X)=Concat(H1,...,H h )W O This multi-head parallel attention mechanism enables the model to simultaneously focus on different representation subspaces of the input features, thereby improving the effect of feature extraction.
4. The intelligent visual tracking method integrating memory distillation and spatiotemporal alignment according to claim 1, characterized in that: The memory bank update strategy adopts a dynamic update mechanism, which specifically includes the following technical implementation details: The memory bank is updated in real time frame by frame, and the memory capacity is fixed at K = 64 memory slots. The memory slot replacement strategy adopts the least frequent replacement algorithm (LFU), that is, when a memory slot needs to be replaced, the memory slot with the lowest frequency of use is selected for replacement. The memory slots are initialized randomly using a Gaussian distribution with a mean of 0 and a standard deviation of 0.01, that is, m k ~N(0, 0.01). The memory reading process includes two methods: one is the soft addressing method based on the attention mechanism, which calculates the attention weights of the query vector and all memory slots. to achieve soft addressing; the other is hard addressing based on the nearest neighbor, by calculating the cosine similarity s k =cos(q,m k ) Select the memory slot with the highest similarity. Memory writing adopts a two-stage update mechanism: first, incremental update is performed, and the new memory content is obtained by weighted combination of historical information and current information, that is, Among them, α = 0.9 is the historical information retention rate; then selective update is performed, and the update operation is performed only when the similarity between the new and old memory contents exceeds the threshold, that is, when The memory content is updated when the similarity threshold α=0.
7.
5. The intelligent visual tracking method integrating memory distillation and spatiotemporal alignment according to claim 1, characterized in that: The knowledge distillation process uses a teacher-student network architecture for knowledge transfer. The teacher network uses ResNet-50 as the backbone network. The teacher network contains three feature layers {C3, C4, C5}, and the corresponding number of feature channels is {512, 1024, 2048} respectively. The student network uses a lightweight MobileNetV2 as the backbone network, which also contains three feature layers {C3, C4, C5}, and the corresponding number of feature channels is {96, 160, 320} respectively. In order to achieve effective knowledge transfer between the teacher network and the student network, the present invention designs a multi-level distillation strategy: first, the features are aligned through channels through a 1×1 convolution operation, and the high-dimensional features of the teacher network are mapped to the same feature space as the student network; secondly, the Squeeze-and-Excitation (SE) module is introduced to recalibrate the feature channels. This module first compresses the spatial dimensions through global average pooling to obtain channel descriptors. Then through the two-layer fully connected network F se (z) = σ(W2δ(W1z)) to learn the correlation between channels, where and is the learnable weight, r = 16 is the dimensionality reduction ratio, δ and σ represent the ReLU and Sigmoid activation functions respectively; finally, the non-local attention module is used to capture the spatial dependency, which calculates the correlation f(x i , x j )=θ(x i ) T φ(x j ) to establish long-range dependencies, where θ(·) and φ(·) are feature transformation functions implemented as 1×1 convolutions.
6. The intelligent visual tracking method integrating memory distillation and spatiotemporal alignment according to claim 1, characterized in that: The prediction head network structure includes two parallel sub-networks: the classification branch and the regression branch. The classification branch adopts a three-layer fully connected network structure, and the input feature dimension is D in =256, the dimensions of the two hidden layers in the middle are both 256, and the output dimension is D out =2. Specifically, the first fully connected layer takes the input feature Mapped to 256-dimensional latent space: h1 = ReLU (W1x + b1), where is the weight matrix, is the bias vector, and ReLU(·) is the rectified linear unit activation function. The calculation process of the second fully connected layer is: h2=ReLU(W2h1+b2), where and is a learnable parameter. The last fully connected layer outputs the binary classification probability: p = softmax(W3h2+b3), where and is a learnable parameter. To prevent overfitting, a dropout layer is added after each hidden layer, and the dropout rate is set to 0.
1. The regression branch uses a similar three-layer fully connected network structure, with an input dimension of 256, two middle hidden layer dimensions of 256, and a final output dimension of 4, corresponding to the center coordinates (x, y) and size (w, h) of the target bounding box. The calculation process of the regression branch can be expressed as: h1 = ReLU (W1x + b1), h2 = ReLU (W2h1 + b2), b = W3h2 + b3, where is the weight matrix, is the bias vector. Similarly, a dropout layer is added after each hidden layer for regularization, and the dropout rate is set to 0.
1.
7. The intelligent visual tracking method integrating memory distillation and spatiotemporal alignment according to claim 1, characterized in that: The training strategy adopts an end-to-end joint optimization method. In terms of optimizer configuration, AdamW optimizer is selected to update model parameters, where the basic learning rate is set to α = 1×10 -4 , weight decay coefficient λ = 1 × 10 -4 Used to suppress overfitting. During the training process, a mini-batch training method with a batch size of 16 was used, and a total of 50 rounds of iterative training were performed. In order to make the training process more stable, a dynamic learning rate scheduling strategy was designed: first, a 5-round linear warm-up process was performed, and the learning rate gradually increased from 0 to the basic learning rate; After preheating, the cosine annealing strategy is used to dynamically adjust the learning rate. Its mathematical expression is: where α min =1×10 -6 is the minimum learning rate, α max is the basic learning rate, t is the current iteration number, and T is the total number of rounds. In terms of data enhancement, a variety of image transformation strategies are designed: the input image is flipped horizontally and vertically with a probability of 0.5; random rotation is performed in the range of [-30°, 30°]; the scale transformation adopts a random scaling ratio of [0.8, 1.2]; the adjustment range of brightness and contrast is set to [0.8, 1.2]; and Gaussian blur enhancement is introduced, and its standard deviation σ is randomly sampled in the range of [0.1, 2.0].
Citation Information
Cited By
Multi-target tracking method and system based on non-appearance chain trajectory correlation model
CN120912642A
Landslide identification method based on multilevel heterogeneous knowledge distillation
CN121074679A
Ultrasonic image discrimination perception pre-training method based on cooperative training framework
CN121213578A
Target tracking method and system based on trajectory perception
CN121330013A
Motion sensing multi-target tracking method based on query transfer
CN121459000A