Description multi-target tracking method based on language decoupling and fine-grained multi-modal feature alignment

By decoupling language description into local description and motion state, and combining a static semantic enhancement module and a motion perception alignment module, the problems of insufficient semantic understanding and loose feature alignment in multi-target tracking are solved, and accurate target tracking in complex environments is achieved.

CN120976266APending Publication Date: 2025-11-18XIAMEN UNIV +3
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511117198.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing multi-target tracking methods struggle to meet the specific target tracking needs of natural language instructions in complex environments. In particular, in dynamic real-world scenarios, the diversity of target features, occlusion, and environmental interference lead to insufficient semantic understanding capabilities. Furthermore, visual-linguistic features lack fine-grained alignment at the regional level, and the integration of motion state with visual temporal features is not tight.

Method used

By decoupling language description into local description and motion state, and combining static semantic enhancement module and motion perception alignment module, fine-grained multimodal feature alignment is achieved, thereby enhancing language understanding and target tracking capabilities.

Benefits of technology

It enhances the model's ability to perform refined parsing of complex semantics, improves the accuracy of single-frame target detection and the robustness of cross-frame association, and enables accurate tracking in dynamic backgrounds and multi-target interaction scenarios, outperforming existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976266A_ABST
    Figure CN120976266A_ABST
Patent Text Reader

Abstract

The invention discloses a reference multi-target tracking method based on language decoupling and fine-grained multi-modal feature alignment, and relates to a computer vision technology. The method comprises the following steps: A, giving a training data set containing a video sequence and language description; b, inputting the video sequence in the step A into a backbone network to extract visual features, and inputting language description into a language model to extract text features; and C, performing multi-modal alignment and fusion through a cross attention mechanism according to the visual features and the language features extracted in the step B. And D, decoupling the language features extracted in the step B into local description and a motion state. And E, inputting the refined features extracted in the step C and the local description extracted in the step D into a static semantic enhancement module to extract target information. And F, associating the current frame target obtained in the step E with the existing trajectory by using a Hungary matching algorithm. G, inputting the matched target features in the step F and the motion state in the step D into a motion perception alignment module to enhance the target recognition capability; the tracking performance of the method is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to computer vision technology, specifically to a method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment. Background Technology

[0002] In recent years, multi-object tracking has become an important research direction in computer vision, showing broad application prospects in key areas such as autonomous driving, video surveillance, and smart cities. However, the complexity and challenges of multi-object tracking tasks cannot be ignored, especially in dynamically changing real-world scenarios. Factors such as the diversity of targets, occlusion problems, and environmental interference significantly increase the difficulty of the task. Traditional multi-object tracking methods typically rely on detecting all targets present in the scene and achieving target tracking by associating all detected targets. However, this reliance severely restricts its interactive capabilities, especially when tracking specific targets based on natural language instructions in complex environments, where traditional methods often fail to meet practical needs.

[0003] With the rapid development of vision-language models, denotational multi-object tracking has made significant progress in the field of computer vision. These models, through the alignment of visual and linguistic features, demonstrate excellent ability to track specified target classes, overcoming the interactive limitations of traditional tracking methods. However, applying vision-language models to multi-object tracking tasks still faces many challenges. On the one hand, denotational multi-object tracking not only requires accurate target localization but also the use of effective association strategies to link targets into complete trajectories across consecutive frames. On the other hand, the complexity of video data, such as lighting variations, occlusion, and viewpoint switching, further increases the task difficulty. Furthermore, the diversity of target features in real-world scenes and the complexity of linguistic expressions place higher demands on the robustness of the models. Especially when dealing with complex linguistic descriptions, accurately understanding and decomposing multiple semantic elements (such as appearance features, spatial relationships, and motion states) in the description becomes a key challenge. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies, such as insufficient semantic understanding of complex language descriptions, lack of fine-grained alignment of visual-linguistic features at the region level, and insufficient integration of motion state and visual temporal features. This invention provides a referential multi-target tracking method based on language decoupling and fine-grained multimodal feature alignment. By decomposing linguistic expressions into local descriptions and motion states, language understanding is enhanced, achieving accurate referential multi-target tracking.

[0005] To achieve the above-mentioned objectives, the present invention provides the following technical solutions.

[0006] This invention provides a method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment, comprising the following steps:

[0007] A. Given a training dataset containing video sequences and language descriptions;

[0008] B. Input the video sequence from step A into the backbone network to extract visual features, and simultaneously input the language description into the language model to extract language features;

[0009] C. Based on the visual and linguistic features extracted in step B, a multimodal alignment and fusion process is performed using a cross-attention mechanism to obtain refined multimodal features;

[0010] D. Decouple the language features extracted in step B into local descriptions and motion states;

[0011] E. Input the multimodal features extracted in step C and the local descriptions extracted in step D into the static semantic enhancement module to extract target information;

[0012] F. Associate the target features of the current frame obtained in step E with the existing trajectory using the Hungarian matching algorithm to obtain the matched target trajectory;

[0013] G. Input the target features stored in the matched target trajectory in step F and the motion state in step D into the motion perception alignment module to enhance the target recognition capability.

[0014] In step A, the given training dataset containing video sequences and language descriptions can specifically be as follows: Given a denotational multi-object tracking video dataset, this dataset contains multiple video sequences, each video consisting of several consecutive frames. It consists of, and is accompanied by, natural language descriptions corresponding to the targets in the video frames. These language descriptions include fine-grained information such as the target's motion state, appearance features, and spatial location.

[0015] In step B, the process of inputting the video sequence from step A into the backbone network to extract visual features, and simultaneously inputting the language description into the language model to extract text features, can be specifically described as follows: inputting the video sequence from step A into a pre-trained ResNet50 backbone network to extract image-level features, and obtaining the features of the current frame. The language description is then input into a pre-trained RoBERTa language model to extract sentence features. Among them, statement features Includes sentence features Features of words .

[0016] In step C, the visual features and linguistic features extracted in step B are aligned and fused using a cross-attention mechanism for multimodal processing. Specifically, this process involves: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] Sentence features Multimodal fusion and alignment are performed using a cross-attention mechanism, defined as follows:

[0017]

[0018] in The image features representing language enhancement are then used as input to the Transformer encoder for refinement; through a cross-attention mechanism, the model can dynamically capture the correlation between language descriptions and visual features, thereby enhancing the semantic consistency of the target representation.

[0019] In step D, the decoupling of the language features extracted in step B into local descriptions and motion states can specifically involve: decoupling the word features extracted in step B into local descriptions and motion states. A binary mask is obtained through word attribute analysis. ; An element with a value of 1 indicates that the word corresponding to that position is a description of motion state, while a value of 0 indicates that the word corresponding to that position is a local description; then through... Decoupling a sentence into a motion description and a local description is defined as follows:

[0020]

[0021] in Indicates the characteristics of motion state. These represent local descriptive features, which serve as feature inputs for the subsequent static semantic enhancement module and motion-aware alignment module, respectively.

[0022] In step E, the refined features extracted in step C and the local description extracted in step D are input into the static semantic enhancement module to extract target information. The specific steps may be as follows:

[0023] Local descriptive features are input into the static semantic enhancement module, and the defined target query is first initialized using the global semantic enhancement module. And obtain enhanced features Next, Enhancement is achieved through bidirectional attention to the local description input, the process of which is defined as follows:

[0024] in and These represent enhanced image and language features, respectively. Next, we will... Image features with language enhancement The process of interactively extracting target information is defined as follows:

[0025]

[0026] in This represents the combination of self-attention and cross-attention mechanisms. This represents the extracted target features;

[0027] use Updated, further enhancing visual-linguistic feature alignment to align target features with local linguistic descriptions, the process is defined as follows:

[0028]

[0029] in, This indicates similarity calculation. Indicates channel dimension, This represents the image features enhanced by local description. The output is then processed by a multilayer perceptron to map the target features into spatial coordinates.

[0030] In step F, the step of associating the current frame target obtained in step E with the existing trajectory using the Hungarian matching algorithm can be specifically described as follows: using the Hungarian matching algorithm to match the target features obtained in step E with the existing target trajectory. The process of performing associations and generating association prediction results is defined as follows:

[0031] ,

[0032] in, The features representing the matching between the current frame and the target trajectory are used to obtain the target trajectory of the current frame, ensuring the continuity and stability of the tracking process.

[0033] In step G, the specific steps for inputting the matched target features from step F and the motion state from step D into the motion perception alignment module can be as follows:

[0034] The matched target features in step F Compared with the motion state in step D The input is fed into the motion-aware alignment module, where motion state descriptions are used to enhance target recognition capabilities. A cross-attention mechanism is employed to combine existing trajectory information with the current frame target update, thereby enhancing the spatiotemporal consistency of tracking. First, the motion state is used... Aligning and enhancing the motion properties of targets is a process defined as follows:

[0035]

[0036] in This indicates target features that enhance motion attributes. The temperature coefficient, with a value of 0.5, represents the degree of influence of the motion state description. The target in the current frame is updated using the constructed trajectory features to enhance the target's spatiotemporal perception capability. After linear layer and layer normalization processing, the spatial position of the target is predicted again to obtain the spatial position of the referred target. The motion perception alignment module significantly improves the tracking robustness of the model in complex scenes by introducing motion state description and trajectory information.

[0037] To improve the accuracy of language-guided object recognition, this invention introduces a static semantic enhancement module. Through a hierarchical multimodal feature interaction enhancement mechanism, a strong alignment relationship at the region level is established between vision and language, thereby generating a more discriminative object representation.

[0038] This invention proposes a motion-aware alignment module that combines the dynamic components of linguistic description with visual temporal features to explicitly align object queries with motion representations, thereby achieving accurate object trajectory prediction across frames and enhancing adaptability to complex motion patterns. Therefore, this invention can accurately detect and generate trajectories for specific targets in real-world scenarios, achieving competitive performance in challenging tracking scenarios.

[0039] Compared with the prior art, the present invention has the following outstanding technical effects:

[0040] 1. This invention decouples language description into local description (static features such as appearance and space) and motion state (dynamic features), thereby achieving refined parsing of complex semantics and improving the model's understanding accuracy of natural language instructions;

[0041] 2. The static semantic enhancement module of this invention establishes strong regional alignment between vision and language, generates more discriminative target representations, and significantly improves the accuracy of single-frame target detection.

[0042] 3. The motion-aware alignment module of the present invention combines the motion state in language with visual temporal features, enhances the adaptability to complex motion patterns, reduces trajectory breaks caused by occlusion, changes in lighting, etc., and improves the robustness of cross-frame association.

[0043] 4. In challenging scenarios such as dynamic backgrounds and multi-target interactions, this invention achieves accurate tracking of specified targets through fine-grained multimodal feature fusion, outperforming existing mainstream methods. Attached Figure Description

[0044] Figure 1 This is an overall flowchart of an embodiment of the present invention.

[0045] Figure 2 This is a flowchart for the static semantic enhancement module. Detailed Implementation

[0046] The method of the present invention will be described in detail below with reference to the accompanying drawings and embodiments. This embodiment is implemented under the premise of the technical solution of the present invention, and provides implementation methods and specific operation processes. However, the protection scope of the present invention is not limited to the following embodiments.

[0047] like Figure 1 As shown, the implementation of this embodiment of the invention includes the following steps:

[0048] A. Given a denoted multi-object tracking video dataset containing multiple video sequences, each video consisting of several consecutive frames. It consists of, and is accompanied by, natural language descriptions corresponding to the targets in the video frames. These linguistic descriptions include fine-grained information such as the target's motion state, appearance features, and spatial location.

[0049] B. Input the video sequence from step A into a pre-trained ResNet50 backbone network to extract image-level features and obtain the features of the current frame. The language description is then input into a pre-trained RoBERTa language model to extract sentence features. The sentence features output by the RoBERTa model Includes sentence features Features of words The ResNet50 outputs a 2048×H×W feature map (H and W are the spatial dimensions after downsampling) in the conv5 stage. Global average pooling yields 2048-dimensional global features for the image frame, while retaining the 1024×H×W feature map output in the conv4 stage for subsequent region-level feature alignment. RoBERTa segments and encodes the input text, outputting word embedding features of dimension [seq_len, 768] (seq_len is the length of the segmented sentence, and 768 is the word vector dimension). Sentence-level global semantic features are obtained through the vectors corresponding to the CLS tokens, while retaining all word vector features. This achieves hierarchical output of sentence features + word features, providing a fine-grained semantic foundation for language decoupling to local descriptions and motion states.

[0050] C. Extracted image features Sentence features Multimodal fusion and alignment are performed using a cross-attention mechanism, defined as follows:

[0051]

[0052] in The image features representing language enhancement are then used as input to the Transformer encoder for refinement. Through a cross-attention mechanism, the model can dynamically capture the correlation between language descriptions and visual features, thereby improving the cross-modal consistency and semantic expressiveness of feature representations.

[0053] D. Extract the word features from step B. A binary mask is obtained through word attribute analysis. . An element with a value of 1 indicates that the word corresponding to that position is a description of motion state, while a value of 0 indicates that the word corresponding to that position is a local description. Then, through... Word features Performing positional multiplication decouples a sentence into a motion description and a local description. This process is defined as follows:

[0054]

[0055] in Indicates the characteristics of motion state. These represent local descriptive features, which serve as feature inputs for the subsequent static semantic enhancement module and motion-aware alignment module, respectively. Explicit modeling of sentence content structure through language decoupling operations helps improve the ability of subsequent modules to model the differences in semantic types, further enhancing multimodal collaborative expression.

[0056] E. Describing local features The input is fed into the static semantic enhancement module, and the flowchart of the static semantic enhancement module is as follows: Figure 2 As shown, this approach comprises two main branches: linguistic features and visual features. First, the local descriptive features of the input are encoded by mining semantic associations through bidirectional attention. Next, spatial attention is used to highlight the target region in the image features for encoding. Then, cross-attention is used to guide visual feature selection based on linguistic semantics, and visual regions feed back into linguistic semantics, achieving cross-modal feature alignment. Next, multimodal features are fused through interactive attention, and target query is optimized by self-attention. Finally, a fully connected layer outputs the target's spatial coordinates and matching confidence, achieving accurate target detection guided by linguistic semantics, establishing strong alignment at the visual-linguistic region level, and improving the accuracy of referential object recognition. The specific steps are as follows:

[0057] First, the global semantic enhancement module is used to extract sentence features. The target query is defined as the initialization definition of the value vector and key vector. And obtain enhanced features Next, we will... Enhancement is achieved through bidirectional attention to the local description input, the process of which is defined as follows:

[0058] in, and These represent enhanced image and language features, respectively; Image features with language enhancement The process of interactively extracting target information is defined as follows:

[0059]

[0060] in, This represents the combination of self-attention and cross-attention mechanisms. This represents the extracted target features.

[0061] use Update target features To further enhance visual-linguistic feature alignment and ensure consistency between target features and local linguistic descriptions, the process is defined as follows:

[0062]

[0063] in, This indicates similarity calculation. Indicates channel dimension, This represents the image features enhanced by local description, and the next output is... The target features are mapped to spatial coordinates and matching prediction values ​​by a multilayer perceptron, where the matching prediction value indicates whether the detected target matches the linguistic description.

[0064] F. Use the Hungarian matching algorithm to match the target features obtained in step E with the existing target trajectory. The process of performing associations and generating association prediction results is defined as follows:

[0065] ,

[0066] in, This represents the features that match the target trajectory in the current frame. The target trajectory in the current frame is obtained using the method described above, ensuring the continuity and stability of the tracking process.

[0067] G. Match the target features from step F. Compared with the motion state in step D The input is fed into the motion-aware alignment module, where motion state descriptions are used to enhance target recognition capabilities. A cross-attention mechanism is employed to combine existing trajectory information with the current frame target update, thereby improving the spatiotemporal consistency of tracking. First, the motion state is used... Aligning and enhancing the motion properties of targets is a process defined as follows:

[0068]

[0069] in, This indicates target features that enhance motion attributes. The temperature coefficient, representing the degree of influence of the motion state description, is set to 0.5. Next, the constructed trajectory features are used to update the target in the current frame, enhancing the target's spatiotemporal awareness. After linear layer and layer normalization processing, the target's spatial position is predicted again, resulting in the referred target's spatial position. The motion-aware alignment module significantly improves the model's tracking robustness in complex scenes by introducing motion state descriptions and trajectory information.

[0070] To verify the performance of the present invention, the performance of different methods on multiple metrics (HOTA, DetA, AssA, DetPr, AssPr, etc.) was compared. The comparative experimental results on the Refer-KITTI dataset are shown in Table 1.

[0071] Table 1

[0072]

[0073] As shown in Table 1, the proposed method (DKGTranck) exhibits significant advantages in the multi-target tracking task. In key metrics such as HOTA (Hybrid Detection and Association Performance), DetA (Detection Accuracy), and AssA (Association Accuracy), the proposed method (DKGTranck) outperforms other comparative methods. For example, the HOTA score reaches 52.01, far exceeding classic methods such as FairMOT (28.46) and ByteTrack (25.29). Experiments demonstrate that the proposed method is superior in its comprehensive tracking capability, combining detection accuracy and trajectory association, and can more accurately identify and associate target trajectories. Compared to other methods that rely on traditional detection-association frameworks (such as DeepSORT and TransRMOT), DKGTranck compensates for its shortcomings in semantic understanding and cross-frame tracking through deep fusion of multimodal features. This invention achieves breakthroughs in detection (DetPr, DetA) and association (AssPr, AssA) metrics by enhancing semantic understanding through language decoupling, strengthening region alignment through static modules, and optimizing temporal association through motion modules. This demonstrates that the method of this invention effectively improves the detection and tracking accuracy of referential targets in complex scenarios, and its performance is significantly better than that of existing mainstream methods.

[0074] Different referential multi-target tracking methods were tested on the Refer-KITTI-V2 dataset. The results of each method on the detection backbone network and multiple tracking performance metrics (HOTA, DetA, AssA, DetPr, AssPr, LocA, etc.) were statistically analyzed to compare the tracking performance of different methods on this dataset. The comparative experimental results on the Refer-KITTI-V2 dataset are shown in Table 2.

[0075] Table 2

[0076]

[0077] As shown in Table 2, the proposed method (DKGTranck) achieves a HOTA (Overall Performance) value of 35.26, significantly outperforming other comparative methods (such as FairMOT at 22.53 and ByteTrack at 24.59). Experiments demonstrate that this invention excels in comprehensive tracking capabilities, combining detection accuracy and trajectory association accuracy. It is well-suited for Refer-KITTI-V2 dataset scenarios (such as road target tracking, including complex dynamic objects like vehicles and pedestrians), enabling more precise target detection and cross-frame association. In terms of detection and association accuracy, the invention achieves the highest scores in DetA (23.04), AssA (54.13), DetPr (36.88), and AssPr (83.85). Combined with the proposed language decoupling, static semantic enhancement, and motion-aware alignment modules, this invention strengthens visual-linguistic feature alignment, improving target detection accuracy and the correctness of trajectory association. The localization accuracy (LocA) of this invention is 91.65, outperforming existing mainstream methods. Experiments show that this invention provides more accurate spatial localization of targets and can precisely match the spatial features of verbal descriptions and visual targets. Through multi-module collaborative optimization, this invention is better suited for language-guided multi-target tracking tasks in road scenarios. It effectively solves problems such as insufficient semantic understanding, coarse feature alignment, and poor motion pattern adaptation. This invention provides a superior technical option for multi-target tracking applications in fields such as autonomous driving and intelligent transportation.

[0078] In summary, the main innovations of this invention are: a. using a vision-language model for target tracking; b. enhancing the model's language understanding ability through language decoupling; c. introducing a static semantic enhancement module and a motion perception alignment module to improve the accuracy and robustness of target tracking from spatial and temporal perspectives, respectively.

[0079] This invention proposes a referential multi-target tracking method based on language decoupling and fine-grained multimodal feature alignment. It decouples complex linguistic information into fine-grained local descriptions and motion states to achieve accurate target tracking. By precisely integrating semantic information into the tracker, the model's understanding of natural language is enhanced, achieving accurate object tracking. A pre-trained language model is used to extract motion states and local descriptions from linguistic expressions. Local descriptions are used to detect potential candidate objects based on static visual features within each frame, while motion states distinguish candidate objects that match the motion. This allows local descriptions and motion states to complement each other, thereby improving the understanding of referential expressions and video content. To accurately capture region-level features and align them with local descriptions, a static semantic enhancement module is proposed. This module enables interaction between sentence embedding and initialization queries to generate queries suitable for tracking. Through a multi-layered attention mechanism, this module achieves efficient fusion of linguistic descriptions and visual features, significantly improving the accuracy of target localization.

[0080] Furthermore, to address the challenge of aligning trajectories with object motion states in the temporal domain, a motion-aware alignment module is proposed. This module leverages motion descriptions from linguistic representations to aid object tagging in understanding temporal information. By aligning motion descriptions with object queries, the gap between visual and linguistic modalities is bridged. This alignment enhances tracking robustness under complex motion patterns and mitigates ambiguity caused by inter-frame appearance changes, thereby achieving a coherent understanding of the target trajectory.

[0081] The method proposed in this invention can achieve efficient and accurate referential target tracking in real-world scenarios, while demonstrating strong adaptability to different scenarios and target categories. Experimental results show that this method significantly outperforms existing mainstream methods in referential multi-target tracking tasks, especially in tracking targets based on complex statements, providing new directions and inspirations for the research and application of referential multi-target tracking.

Claims

1. A method for tracking multiple targets based on language decoupling and fine-grained multimodal feature alignment, characterized in that... Includes the following steps: A. Given a training dataset containing video sequences and language descriptions; B. Input the video sequence from step A into the backbone network to extract visual features, and simultaneously input the language description into the language model to extract language features; C. Based on the visual and linguistic features extracted in step B, a multimodal alignment and fusion process is performed using a cross-attention mechanism to obtain refined multimodal features; D. Decouple the language features extracted in step B into local descriptions and motion states; E. Input the multimodal features extracted in step C and the local descriptions extracted in step D into the static semantic enhancement module to extract target information; F. Associate the target features of the current frame obtained in step E with the existing trajectory using the Hungarian matching algorithm to obtain the matched target trajectory; G. Input the target features stored in the matched target trajectory in step F and the motion state in step D into the motion perception alignment module to enhance the target recognition capability.

2. The method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment as described in claim 1, characterized in that... In step A, the given training dataset containing video sequences and language descriptions specifically involves: providing a denotational multi-object tracking video dataset containing multiple video sequences, each video consisting of several consecutive frames. It consists of, and is accompanied by, natural language descriptions corresponding to the targets in the video frames. These language descriptions include fine-grained information about the target's motion state, appearance features, and spatial location.

3. The method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment as described in claim 1, characterized in that... In step B, the process of inputting the video sequence from step A into the backbone network to extract visual features, and simultaneously inputting the language description into the language model to extract language features, specifically involves: inputting the video sequence from step A into a pre-trained ResNet50 backbone network to extract image-level features, and obtaining the features of the current frame. The language description is then input into a pre-trained RoBERTa language model to extract sentence features. Among them, statement features Includes sentence features Features of words .

4. The method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment as described in claim 1, characterized in that... In step C, the visual features and linguistic features extracted in step B are aligned and fused using a cross-attention mechanism. Specifically, the extracted image features are... Sentence features Multimodal fusion and alignment are performed using a cross-attention mechanism, defined as follows: in The image features representing language enhancement are then used as input to the Transformer encoder for refinement; through a cross-attention mechanism, the model dynamically captures the correlation between language descriptions and visual features to enhance the semantic consistency of the target representation.

5. The method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment as described in claim 1, characterized in that... In step D, the decoupling of the language features extracted in step B into local descriptions and motion states specifically involves: decoupling the word features extracted in step B into local descriptions and motion states. A binary mask is obtained through word attribute analysis. ; An element with a value of 1 indicates that the word corresponding to that position is a description of motion state, while a value of 0 indicates that the word corresponding to that position is a local description; then through... Decoupling a sentence into a motion description and a local description is defined as follows: in Indicates the characteristics of motion state. These represent local descriptive features, which serve as feature inputs for the subsequent static semantic enhancement module and motion-aware alignment module, respectively.

6. The method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment as described in claim 1, characterized in that... In step E, the process of inputting the multimodal features extracted in step C and the local descriptions extracted in step D into the static semantic enhancement module to extract target information involves the following steps: Local descriptive features are input into the static semantic enhancement module, and the defined target query is first initialized using the global semantic enhancement module. And obtain enhanced features Next, Enhancement is achieved through bidirectional attention to the local description input, the process of which is defined as follows: in and These represent enhanced image and language features, respectively. Next, we will... Image features with language enhancement The process of interactively extracting target information is defined as follows: in This represents the combination of self-attention and cross-attention mechanisms. This represents the extracted target features; use Updated, further enhancing visual-linguistic feature alignment to align target features with local linguistic descriptions, the process is defined as follows: in, This indicates similarity calculation. Indicates channel dimension, The image features are represented by local description enhancements; then the output is processed by a multilayer perceptron to map the target features into spatial coordinates.

7. The method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment as described in claim 1, characterized in that... In step F, the step of associating the target features of the current frame obtained in step E with the existing trajectory using the Hungarian matching algorithm specifically involves: using the Hungarian matching algorithm to match the target features obtained in step E with the existing target trajectory. The process of performing associations and generating association prediction results is defined as follows: , in, The features representing the matching between the current frame and the target trajectory are used to obtain the target trajectory of the current frame, thereby ensuring the continuity and stability of the tracking process.

8. The method for pointing multi-target tracking based on language decoupling and fine-grained multimodal feature alignment as described in claim 1, characterized in that... In step G, the specific steps of aligning the target features stored in the matched target trajectory in step F with the motion state input motion perception alignment module in step D are as follows: The matched target features in step F Compared with the motion state in step D The input motion-aware alignment module uses motion state descriptions to enhance target recognition capabilities and utilizes a cross-attention mechanism combined with existing trajectory information to update the target in the current frame, thereby enhancing the spatiotemporal consistency of tracking. First, the motion state is used... Aligning and enhancing the motion properties of targets is a process defined as follows: in Target features that enhance motion attributes The temperature coefficient represents the degree of influence of the control motion state description; the target in the current frame is updated using the constructed trajectory features to enhance the target's spatiotemporal perception capability; after linear layer and layer normalization processing, the spatial position of the target is predicted again to obtain the spatial position of the referred target; The motion-aware alignment module improves the model's tracking robustness in complex scenarios by introducing motion state descriptions and trajectory information.

Citation Information

Cited By

  • Traffic target tracking method and system based on language updating and memory modeling

    CN121921342A