Time sequence alignment visual language tracking method and system
Through the temporal semantic enhancement module and the cross-modal feature fusion module, combined with the constancy learning balance strategy, the problems of insufficient static description and insufficient multimodal integration of visual language tracking technology in dynamic environments are solved, achieving more efficient target tracking and stronger environmental adaptability.
Patent Information
- Application Number
- CN202510797526.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-23
AI Technical Summary
Existing visual language tracking technology has problems in dynamic environments, such as insufficient guidance from static language descriptions, insufficient integration of multimodal features, poor dynamic adaptability, and insufficient perceptual consistency, which leads to degraded target tracking performance.
A temporal semantic enhancement module is used to dynamically update the language tag weights, a cross-modal feature fusion module is used to integrate visual and language features, and a constancy learning balance strategy is implemented to improve the learning consistency of the model in different perception tasks.
It significantly improves the accuracy and stability of target tracking, enhances the robustness and adaptability of the system in complex environments, simplifies the computational complexity, and improves the efficiency and accuracy of multimodal tracking.
Smart Images

Figure CN120689634A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of combining computer vision with natural language processing, particularly the development and application of multimodal tracking systems. Specifically, this invention relates to visual language tracking technology, focusing on the efficient tracking of objects in dynamic video sequences using natural language descriptions. This technology has broad applications in fields such as intelligent surveillance, virtual reality, augmented reality, human-computer interaction, and autonomous navigation, providing a new solution for more accurately capturing and identifying objects in videos. Background Art
[0002] With the rapid development of computer vision and natural language processing technologies, visual language tracking technology has gradually become an important research direction in many fields, such as intelligent monitoring, autonomous driving, virtual reality, and human-computer interaction. However, current existing technologies still face many technical problems and shortcomings in the process of effectively integrating visual information with language descriptions:
[0003] 1. Limitations of Static Language Descriptions: Existing visual language tracking methods mostly rely on static language descriptions to guide dynamic visual sequences. This approach fails to provide continuous, real-time contextual information, resulting in a significant degradation in tracking performance when the target moves rapidly or is occluded.
[0004] 2. Insufficient cross-modal information integration: Although some recent research has attempted to combine visual data and language descriptions, existing methods often simply concatenate the two or use traditional feature fusion techniques. This approach fails to fully explore the deep relationship between vision and language, thus affecting the effective integration, relevance, and accuracy of features.
[0005] 3. Lack of dynamic adaptability: Existing technologies are mostly based on static features and model training, making them difficult to adapt to real-time changes in target state. In real applications, target features such as appearance, shape, and color may change over time, and existing models are unable to adjust the corresponding language descriptions in a timely manner, resulting in unstable and inaccurate tracking results.
[0006] 4. Perceptual consistency issues: Visual tracking tasks often require processing multiple perceptual features (such as shape, color, and size). Existing technologies are inadequate in learning and adjusting between these features. The lack of effective methods to address loss fluctuations between different tasks affects the overall stability and accuracy of the model.
[0007] In summary, existing technologies still have significant deficiencies in terms of effective integration of vision and language, dynamic adaptability, and multi-task learning. Therefore, a new vision-language tracking method is urgently needed to address these key issues and improve target tracking performance in real-time dynamic environments. Summary of the Invention
[0008] To overcome the shortcomings of the above-mentioned existing technologies, the present invention provides a temporally aligned visual language tracking method. This tracking method aims to improve the performance of object tracking in dynamic scenes by effectively integrating visual data and natural language descriptions. Specifically, the present invention mainly includes the following key technical features:
[0009] 1. Temporal Semantic Enhancement Module:
[0010] This module constructs a spatiotemporal semantic alignment matrix. By dynamically updating language tags, it transforms static language descriptions into dynamic tag weights that are aligned with the time of the tracking sequence. This allows for continuous provision of relevant language guidance as the target changes, effectively improving tracking accuracy and stability.
[0011] 2. Cross-modal feature fusion module:
[0012] This module integrates spatial features from the spatiotemporal semantic alignment matrix and effectively combines templates and search tags through weighted vector calculation. This innovative feature fusion method allows visual information and language descriptions to be more closely integrated, thereby improving the efficiency and accuracy of multimodal tracking.
[0013] 3. Constant learning balance strategy:
[0014] This strategy monitors loss fluctuations and dynamically adjusts the model's learning focus to improve the consistency of the model's learning across different perceptual tasks (such as shape, color, and size). This approach ensures that the model is more robust and adaptable when facing complex and changing scenarios.
[0015] Through the above method, the present invention can effectively solve the problems of insufficient guidance, insufficient integration of multimodal features, and poor model adaptability caused by static language descriptions in the existing technology. The temporal alignment visual language tracking method proposed in this invention not only improves the performance of the tracking method in dynamic scenes, but also significantly enhances the reliability and accuracy of the system when dealing with complex situations such as occlusion and environmental interference, thereby providing a new and effective solution for the development of the field of visual language tracking. Specifically, it includes the following steps:
[0016] Step 1: Obtain visual data and corresponding natural language description, where the visual data includes a template frame and a search frame;
[0017] Step 2: Perform feature extraction on the input data to extract basic visual markup features and natural language markup features;
[0018] Step 3: Build a spatiotemporal semantic alignment matrix through the temporal semantic enhancement module and dynamically update the natural language tag features so that the spatiotemporal semantic alignment features of the next frame contain time information.
[0019] Step 4: Use the cross-modal feature fusion module to extract the spatial information in the spatiotemporal semantic alignment matrix of the template frame and the search frame and fuse them to form an enhanced feature set with contextual information;
[0020] Step 5: Construct a loss function to train the network constructed in steps 2 to 4.
[0021] In step 6, the trained network is used to generate an enhanced feature set, and the tracking method is applied to generate the bounding box of the target.
[0022] Furthermore, the natural language description in step 1 provides contextual information about the target to be tracked.
[0023] Furthermore, the visual encoder Transformer and the language representation model RoBERTa are used to perform feature extraction on the input data to extract visual tag features and natural language tag features.
[0024] Furthermore, the specific implementation of step 3 is as follows:
[0025] (1) Constructing spatiotemporal semantic alignment matrix: using template marking features , Search Tag Features and word embedding tag features To enhance their deep semantics, M, N and L represent the number of template frames, search frames and natural language tags respectively, and D represents the feature dimension. Each tag feature in each tag feature group is subjected to Hadamard product calculation to generate two matrices containing tags: , called the spatiotemporal semantic alignment matrix; Include OK Column markers, Include OK Column markers;
[0026] (2) Language tag weight calculation: Pearson correlation coefficient is used to calculate the correlation between natural language tag features, i.e., word embedding tag features and visual tag features, to generate the weight of word embedding tag features. ;
[0027] (3) Dynamic update: After calculating the weight of the word embedding tag feature, the weight update is used ; According to the target's movement direction and speed in the video stream, the weight of the natural language tag feature is adjusted in real time, Updated The calculation formula is as follows:
[0028]
[0029] Similarly, the Pearson correlation coefficient was used to calculate and Update weights between , Depend on The vertical pixel level average is obtained and the renew for :
[0030]
[0031] Finally, the Pearson correlation coefficient was used to calculate and The updated weights ,Depend on Synthetic update marker It combines template frames, search frames, and natural language description information:
[0032]
[0033] Each natural language tag feature is Appropriate weights are assigned to them.
[0034] Furthermore, the weight calculation formula of the word embedding tag feature in (2) is as follows:
[0035]
[0036] in, Represents the i-th Calculated and The weight between , is the Pearson correlation coefficient, is the number of markers; , yes Any one of Any tag in should generate an appropriate weight for ,and Depend on The pixel-level average in the vertical direction is obtained.
[0037] Furthermore, after the temporal semantic enhancement in step 3, the spatiotemporal semantic alignment matrix and The temporal information is contained in the cross-modal feature fusion module, and the cross-modal feature fusion module is responsible for using and The spatial information in the image is combined to form an enhanced feature set with contextual information. and Specifically, With the template tag feature Asking Hadamard to get ; With the search tag feature Asking Hadamard to get ; Depend on The average pixel value in the horizontal direction is obtained. Depend on The average pixel value in the horizontal direction is obtained. and Perform a cross-correlation operation and output an estimated distribution of the target position.
[0038] Furthermore, in step 5, the loss function constructed includes: regression loss GIoU loss, classification loss cross entropy loss and feature alignment loss;
[0039] Feature alignment loss The formalization is as follows:
[0040]
[0041] in represents the cross entropy, Indicates a label, and denote natural language tags and visual tag groups, respectively. Compute the cosine similarity matrix between the two marker groups.
[0042] Furthermore, in step 5, a loss accumulator is designed to store and update the loss during mini-batch training. The formula for calculating the loss accumulator is as follows:
[0043]
[0044] in is the attenuation factor, Indicates time The overall loss of .
[0045] Furthermore, in step 5, using the current loss and time The difference between the average losses in the loss accumulator As an indicator, the exponential moving average model EMA is used to effectively monitor the loss:
[0046]
[0047] Setting a hyperparameter As a threshold, it is used to judge whether a given sample shows imbalance in the perceptual constancy task. If Exceeding this threshold , then the sample is classified as an outlier and masked, and its corresponding loss value is set to zero, effectively excluding the sample from the mini-batch.
[0048] The present invention also provides a temporally aligned visual language tracking system, comprising:
[0049] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the time-aligned visual language tracking method as described in the above technical solution.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] 1. Improve tracking accuracy:
[0052] By introducing a temporal semantic enhancement module, this invention effectively aligns dynamic language markers with visual features, enabling the system to respond to changes in target state in real time. This dynamic update mechanism significantly improves target tracking accuracy, especially in dynamic scenes and rapidly changing environments.
[0053] 2. Enhanced multimodal feature fusion:
[0054] Through a cross-modal feature fusion module, the present invention achieves a deep integration of visual information and language descriptions. Compared with traditional simple splicing or feature fusion methods, TAVLT can effectively extract and integrate spatiotemporal semantic information, improving the overall performance of multimodal tracking and allowing the model to better understand and process complex visual and language inputs.
[0055] 3. Improve model robustness:
[0056] The implementation of a constancy learning balancing strategy enables the system to demonstrate enhanced stability across diverse perceptual characteristics, such as shape, color, and size constancy. By dynamically adjusting the learning strategy, the model can effectively cope with challenging scenarios such as occlusion and lighting changes, ensuring the high reliability of the tracking system in real-world applications.
[0057] 4. Flexible adaptation to complex scenarios:
[0058] The design of this invention enables the time-aligned visual language tracking method to flexibly adapt to various complex dynamic scenarios. This capability not only improves the applicability of the system, but also provides more reliable technical support for various application scenarios such as intelligent monitoring, robot navigation, and human-computer interaction.
[0059] 5. Simplify computational complexity:
[0060] Through a modular design that independently processes templates, detection, and language features, TAVLT simplifies the overall computational process while fully utilizing computing resources. This optimization not only makes the invention competitive in performance but also significantly improves operational efficiency.
[0061] In summary, the temporally aligned visual language tracking method of the present invention not only effectively overcomes many defects of the existing technology, but also significantly improves the accuracy, reliability and adaptability of target tracking, providing an innovative solution for the development of multimodal tracking technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 This is a flow chart of an embodiment of the present invention.
[0063] Figure 2 This is a diagram of the constancy learning balance strategy of the present invention.
[0064] Figure 3 Schematic diagram comparing the input data in an embodiment of the present invention with traditional concatenated learning; (a) Traditional visual-language tracking frameworks concatenate visual and language data and feed them into the model simultaneously. (b) In contrast, the framework proposed in this invention feeds the template image, detection image, and language specifications into the model separately. While (a) enables the computation of long-range dependencies between multimodal data through concatenation, it also imposes a significant computational burden. In contrast, the framework proposed in this invention (b) integrates multimodal information and dynamically adjusts the weights of the labels by learning an alignment matrix. DETAILED DESCRIPTION
[0065] The present invention will be further described below with reference to the accompanying drawings and examples.
[0066] The temporal alignment visual language tracking method of the present invention flexibly copes with the target tracking problem in dynamic scenes through a series of precisely constructed modules. Figure 1 The specific embodiments of the present invention are further described.
[0067] In this embodiment, the present invention uses a complex video surveillance scene as an example. The target in this scene moves, changes in scale, and changes in posture, and there are challenges such as occlusion and lighting changes. The process of temporally aligned visual language tracking can be performed by the following steps:
[0068] An embodiment of the present invention provides a time-aligned visual language tracking method, comprising the following steps:
[0069] Step 1. Input data
[0070] Receives input visual data and a corresponding natural language description. The visual data includes template frames and search frames, while the natural language description provides contextual information about the target to be tracked. For example, "The boy sitting at the edge of the boat."
[0071] Step 2: Initial feature extraction
[0072] Use the Transformer visual encoder and language representation model RoBERTa to perform feature extraction on the input, extracting basic visual and natural language features (see Figure 1 "Overview" in the middle left): Template markup , search tags and word embedding tags M, N and L represent the number of template frames, search frames and natural language tags respectively, and D represents the feature dimension.
[0073] Step 3: Temporal Semantic Enhancement Module
[0074] Through the temporal semantics reinforcement module, the system constructs a spatiotemporal semantic alignment matrix, dynamically updates language tags, and adjusts their weights based on the current state of the target. The core of this step is to convert static language descriptions into dynamic tag weights that are aligned with the tracking sequence time. This process includes the following sub-steps:
[0075] (1) Constructing spatiotemporal semantic alignment matrix: using 、 and To enhance their deep semantics, each tag in each tag group undergoes a Hadamard product calculation to generate two matrices containing the tags: , called the spatiotemporal semantic alignment matrix. Include OK Column markers, Include OK Column markers. It combines template features and language embedding, and It integrates search features and language embedding.
[0076] (2) Language tag weight calculation: Use the Pearson correlation coefficient to calculate the correlation between language tags and visual features, generate weights, and use the weights of word embedding tags Calculation example:
[0077]
[0078] in, Represents the i-th Mark calculated and The weight between , is the Pearson correlation coefficient, is the number of markers. , yes Any one of Any tag in should generate an appropriate weight for .and Depend on Perform vertical pixel-level averaging Similarly, the subsequent weight updates, such as arrive The weight of Depend on Perform vertical pixel-level averaging After the weight is updated, it becomes 。 )、 arrive The weights will also be calculated in the same way (see Figure 1 The whole process can be regarded as the temporal semantic reinforcement module in Updated to .
[0079] The optimized language tag weights are updated through a time window algorithm, which can consider the language tag weights of the previous moment and the visual features of the current moment at each moment, thereby achieving smooth weight transfer and reducing instability caused by visual interference.
[0080] (3) Dynamic update: After calculating the weight, use the weight to update the mark According to the target's movement direction and speed in the video stream, the weight of the language tag is adjusted in real time (see Figure 1 The middle "Transformer encoder" section). Updated The calculation formula is as follows:
[0081]
[0082] Similarly, the Pearson correlation coefficient can also be used to calculate and Update weights between . And use renew for :
[0083]
[0084] Finally, the present invention uses Pearson correlation coefficient to calculate and The updated weights .Depend on Synthetic update marker Combines templates, search and language information:
[0085]
[0086] Each language tag is in The appropriate weights are assigned to In the next frame, it will be used as the word embedding tag in this frame, that is, , start the work of the next frame. After the dynamic language features are updated, they will affect the fusion of the next frame. and , instead of calculating the Hadamard product, calculate the spatiotemporal semantic alignment matrix of the next frame. To intuitively illustrate the effect of language tags after dynamic update, please refer to Figure 1 In this example, a scene with the descriptive text "a boy sitting by the boat" is presented. In the carefully designed temporal semantic reinforcement module, key tags like "boy" are given the highest weight, while "edge" and "boat" gradually decrease in importance and their weight as auxiliary tags decreases accordingly. The presence of "sitting" negatively impacts the boy's current state, so its weight as a distractor is extremely low. The remaining words are treated as background tags, with minimal impact on the current frame.
[0087] The temporal semantic enhancement module also has an adaptive adjustment mechanism that allows it to autonomously adjust the effective duration and update frequency of language tags in different scenarios according to the target's motion characteristics and environmental changes, so as to improve the flexibility and adaptability of tracking.
[0088] Through the above improvements, the temporal semantic enhancement module is further enhanced in its ability to respond to dynamic visual information and its connection with language descriptions, thereby effectively solving the problems of inaccurate and instability in tracking caused by static descriptions in the existing technology.
[0089] Step 4: Cross-modal feature fusion module
[0090] After the temporal semantic enhancement in step 3, the spatiotemporal semantic alignment matrix and The module has the timing information, and this module is responsible for using and The spatial information in the is used to perform multimodal fusion. The specific process is as follows:
[0091] (1) Feature aggregation: By calculating weighted vectors, adjacent visual features are combined with corresponding language tags to form an enhanced feature set with contextual information. and Specifically, With template features Asking Hadamard to get ; With search features Asking Hadamard to get . Depend on Find the average pixel value in the horizontal direction (indicated as ) to obtain, Depend on Calculated by averaging the pixels in the horizontal direction. and Perform a cross-correlation operation to output an estimated distribution of target locations. Later, we will propose enhanced alignment between modalities to monitor and evaluate the quality of this distribution. For details, see step 5. This establishes the matching between the template and the search feature.
[0092] (2) Generating tracking features: The obtained enhanced feature set is used as the input for subsequent target tracking to improve the robustness and accuracy of the features (see Figure 1 The middle "Transformer encoder" section).
[0093] Step 5: Constant learning of balance strategy
[0094] During training, the constant learning balance strategy ensures that the model is consistent when facing different perceptual tasks. By monitoring the loss fluctuations and dynamically adjusting the learning focus as needed, the stability of the model in multi-task processing is improved (see Figure 2 By calculating the loss rate, the constancy learning balancing strategy can set a threshold to effectively balance the model's understanding of different constancy types during training, ultimately improving tracking accuracy.
[0095] To avoid evaluating the loss for all samples at each time step, the present invention develops a loss accumulator to store and update the loss during mini-batch training. A key reason for storing the loss in mini-batches is that the loss typically decreases steadily throughout the training process. Therefore, this approach helps in calculating the average loss using controlled weights. The formula for calculating the loss accumulator is as follows:
[0096]
[0097] in is the attenuation factor, Indicates time The overall loss when the feature is aligned. It is usually composed of the common regression loss GIoU loss and classification loss cross entropy loss. In order to optimize the relationship between different modal features, the present invention additionally designs feature alignment loss Specifically, the alignment between modalities is enhanced by comparing the features of various modalities, thereby ensuring a strong semantic integration of visual and language features. The goal is to promote effective fusion by encouraging similarity between visual and language features with the same target. The formalization is as follows:
[0098]
[0099] in represents the cross entropy, Represents the manually annotated label, that is, the true value, and denote the language and visual marker groups, respectively, which are and / The smallest unit of composition. Compute the cosine similarity matrix between the two marker groups.
[0100] The present invention uses the exponential moving average model (EMA) to effectively monitor the loss, thereby avoiding excessive computational requirements. When the loss outliers fluctuate greatly, it indicates that the model's performance on the perceptual constancy task is unbalanced. Therefore, the present invention uses the current loss and the time step The difference between the average losses in the loss accumulator As an indicator to evaluate the balance of the model in this task:
[0101]
[0102] along with As the value of increases, the possibility of abnormal fluctuations in loss during training will also increase. The present invention sets a hyperparameter As a threshold, it is used to judge whether a given sample shows imbalance in the perceptual constancy task. Exceeding this threshold , then the sample is classified as an outlier and masked, and its corresponding loss value is set to zero, effectively excluding the sample from the mini-batch. This paper conceptualizes this exclusion as a regularization technique to balance different perceptual constancy tasks during visual language model training. Specifically, it reallocates resources from shape constancy tasks to more challenging color and size constancy tasks.
[0103] Step 6: Target tracking
[0104] The system then applies tracking methods using the processed, enhanced feature set to generate a bounding box for the target. For example, a regression model can be used to predict the target's next location, outputting precise tracking results. Finally, the system visualizes the tracking results and feeds this information to the user interface in real time, providing the operator with a dynamic monitoring view and ensuring timely capture of changes in the video scene.
[0105] Table 1 shows the results of our proposed method on the LaSOT and TNL2K datasets, both of which consist of detailed natural language descriptions rich in semantic information and long videos. As can be seen from the table, our method significantly outperforms other trackers. This demonstrates its remarkable robustness in scenarios sensitive to historical information, such as aspect ratio changes, full occlusion, scale changes, and out-of-view conditions. This robustness is attributed to the framework's ability to enhance the temporal density of language descriptions. This capability, through the temporal semantic enhancement module, optimizes semantic alignment to minimize the temporal discrepancy between visual and language data. In other words, the temporal semantic enhancement module can simultaneously process and update cues from multiple sources, integrating rich contextual information. This integration enhances the model's understanding of complex scenes. Our method enhances the integration of visual and language modalities through a dynamic label adjustment mechanism using the Pearson correlation coefficient, ensuring more accurate tracking in diverse environments. A cross-modal feature fusion module and a spatiotemporal semantic alignment matrix further refine this integration, while a feature alignment loss promotes semantic coherence. Furthermore, a constancy learning balancing strategy enhances the model's adaptability to diverse perception tasks.
[0106] Table 1 Simulation experiments of the present invention on two data sets
[0107] UVLTrack comes from the literature Unifying visual and vision-language tracking viacontrastive learning. Proceedings of the AAAI Conference on ArtificialIntelligence. Vol. 38. No. 5. 2024.
[0108] OSDT comes from the document One-stream stepwise decreasing for vision-languagetracking. IEEE Transactions on Circuits and Systems for Video Technology (2024).
[0109] ATT comes from the literature Consistencies are all you need for semi-supervisedvision-language tracking. Proceedings of the 32nd ACM InternationalConference on Multimedia. 2024.
[0110] Through the above-described implementation, the temporally aligned visual language tracking method of the present invention resolves the contradiction between static descriptions and dynamic environments in the prior art by effectively integrating visual and language information. The various modules work together to enable the system to dynamically adapt to complex scenarios, significantly improving tracking accuracy and performance. This example demonstrates an effective solution to a technical problem and provides innovative reference for future visual language tracking systems.
[0111] On the other hand, an embodiment of the present invention also provides a temporally aligned visual language tracking system, including:
[0112] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the time-aligned visual language tracking method as described in the above technical solution.
[0113] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A temporally aligned visual language tracking method, characterized in that: The steps include: Step 1: Obtain visual data and corresponding natural language description, where the visual data includes a template frame and a search frame; Step 2: Perform feature extraction on the input data to extract basic visual markup features and natural language markup features; Step 3: Build a spatiotemporal semantic alignment matrix through the temporal semantic enhancement module and dynamically update the natural language tag features so that the spatiotemporal semantic alignment features of the next frame contain time information. Step 4: Use the cross-modal feature fusion module to extract the spatial information in the spatiotemporal semantic alignment matrix of the template frame and the search frame and fuse them to form an enhanced feature set with contextual information; Step 5: Construct a loss function to train the network constructed in steps 2 to 4. In step 6, the trained network is used to generate an enhanced feature set, and the tracking method is applied to generate the bounding box of the target.
2. The temporally aligned visual language tracking method according to claim 1, wherein: The natural language description in step 1 provides contextual information about the target to be tracked.
3. The temporally aligned visual language tracking method according to claim 1, wherein: The visual encoder Transformer and the language representation model RoBERTa are used to perform feature extraction on the input data to extract visual tag features and natural language tag features.
4. The temporally aligned visual language tracking method according to claim 1, wherein: The specific implementation of step 3 is as follows: (1) Constructing spatiotemporal semantic alignment matrix: using template marking features , Search Tag Features and word embedding tag features To enhance their deep semantics, M, N and L represent the number of template frames, search frames and natural language tags respectively, and D represents the feature dimension. Each tag feature in each tag feature group is subjected to Hadamard product calculation to generate two matrices containing tags: , called the spatiotemporal semantic alignment matrix; Include OK Column markers, Include OK Column markers; (2) Language tag weight calculation: Pearson correlation coefficient is used to calculate the correlation between natural language tag features, i.e., word embedding tag features and visual tag features, to generate the weight of word embedding tag features. ; (3) Dynamic update: After calculating the weight of the word embedding tag feature, the weight update is used ; According to the target's movement direction and speed in the video stream, the weight of the natural language tag feature is adjusted in real time, Updated The calculation formula is as follows: ; Similarly, the Pearson correlation coefficient was used to calculate and Update weights between , Depend on The vertical pixel level average is obtained and the renew for : ; Finally, the Pearson correlation coefficient was used to calculate and The updated weights ,Depend on Synthetic update marker It combines template frames, search frames, and natural language description information: ; Each natural language tag feature is Appropriate weights are assigned to them.
5. The temporally aligned visual language tracking method according to claim 4, wherein: The weight calculation formula of the word embedding tag feature in (2) is as follows: ; in, Represents the i-th Calculated and The weight between , is the Pearson correlation coefficient, is the number of markers; , yes Any one of Any tag in should generate an appropriate weight for ,and Depend on The pixel-level average in the vertical direction is obtained.
6. The temporally aligned visual language tracking method according to claim 4, wherein: After the temporal semantic enhancement in step 3, the spatiotemporal semantic alignment matrix and The temporal information is contained in the cross-modal feature fusion module, and the cross-modal feature fusion module is responsible for using and The spatial information in the image is combined to form an enhanced feature set with contextual information. and Specifically, With the template tag feature Asking Hadamard to get ; With the search tag feature Asking Hadamard to get ; Depend on The average pixel value in the horizontal direction is obtained. Depend on The average pixel value in the horizontal direction is obtained. and Perform a cross-correlation operation and output an estimated distribution of the target position.
7. The temporally aligned visual language tracking method according to claim 1, wherein: In step 5, the loss functions constructed include: regression loss GIoU loss, classification loss cross entropy loss and feature alignment loss; Feature alignment loss The formalization is as follows: ; in represents the cross entropy, Indicates a label, and denote natural language tags and visual tag groups, respectively. Compute the cosine similarity matrix between the two marker groups.
8. The temporally aligned visual language tracking method according to claim 1, wherein: In step 5, a loss accumulator is also designed to store and update the loss during mini-batch training. The formula for calculating the loss accumulator is as follows: ; in is the attenuation factor, Indicates time The overall loss of .
9. The temporally aligned visual language tracking method according to claim 8, wherein: In step 5, using the current loss and time The difference between the average losses in the loss accumulator As an indicator, the exponential moving average model EMA is used to effectively monitor the loss: ; Setting a hyperparameter As a threshold, it is used to judge whether a given sample shows imbalance in the perceptual constancy task. If Exceeding this threshold , then the sample is classified as an outlier and masked, and its corresponding loss value is set to zero, effectively excluding the sample from the mini-batch.
10. A temporally aligned visual language tracking system, characterized in that include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the temporal alignment visual language tracking method as described in any one of claims 1 to 9.
Citation Information
Cited By
Traffic target tracking method and system based on language updating and memory modeling
CN121921342A