ATS monovision-oriented semantic-driven 3D single-target tracking method
By integrating visual, language and geometric features in monocular video, the existing 3D visual tracking methods are solved inadequate accuracy and robustness in complex scenarios, precise tracking of 3D objects is achieved and system cost is reduced.
Patent Information
- Application Number
- CN202510130217.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-06
AI Technical Summary
Existing 3D visual tracking methods do not perform well in complex scenarios, and it is difficult to integrate visual, language and geometric features, resulting in insufficient tracking accuracy and robustness, and relying on expensive multi-sensor devices.
A 3D single-object tracking method for semantic-driven ATS single vision is proposed. By constructing a large-scale data set, multi-modal features are extracted, vision-language encoder and memory-enhanced tracking decoder are designed, and visual, language and geometric features are integrated to achieve accurate tracking of 3D objects in monocular videos.
Significantly improves the accuracy and robustness of 3D object tracking in monocular videos, reduces dependence on multi-sensor devices, enhances human-like perception of the tracking system, and provides a large-scale data set to support training and evaluation.
Smart Images

Figure CN120107847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of ATS 3D visual tracking, natural language processing and multimodal deep learning, and in particular to a semantically driven 3D single target tracking method under ATS single vision. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technology, the recognition and tracking of 3D objects have become key technologies in the fields of autonomous driving, intelligent monitoring and robot navigation. However, traditional visual target tracking methods mainly rely on information and often have difficulty in achieving sufficient robustness in scenes, such as rapid changes in target appearance, eye-catching, complex lighting changes, etc. At the same time, when tracking targets, humans can combine visual and language clues to achieve more intelligent description and reasoning. Therefore, natural language description Carl target tracking has become a research hotspot. The 3D visual language tracking method can accurately track ATS from monocular RGB video by combining visual, language and geometric features. This method does not rely on expensive multi-sensor equipment, but tries to imitate the human picking method, nothing more. Tracking through natural language descriptions and clues provides new solutions for visual intelligent transportation, unmanned driving, security monitoring and other fields. For example, in unmanned driving scenarios, the system can accurately locate and track the target vehicle through natural language instructions (such as "track the red car in front"), thereby improving the level of interaction and reducing the implementation cost.
[0003] Existing 3D visual tracking methods mainly rely on depth sensors (such as LiDAR) and grid camera systems to obtain object information in 3D space. Although these methods can provide spatial data of coordinates, they have problems in cost, equipment complexity and real-time performance.
[0004] The existing methods have the following shortcomings:
[0005] (1) The existing 2D tracking methods and 3D tracking methods that rely on multiple sensors differ from the natural way humans perceive things. These methods usually rely on data from multiple sensors such as lidar, radar, or depth cameras, which differs from the natural way humans perceive things, which mainly relies on visual and language cues, limiting the application of tracking systems in a wider range of scenarios; (2) Most existing visual-language tracking methods are limited to two-dimensional space and lack the ability to perceive objects in three-dimensional space. This limits their widespread application in the real world, especially in scenarios that require three-dimensional spatial information; (3) Traditional single target tracking (SOT) methods mainly rely on visual features extracted from video frames. These methods often perform poorly when faced with complex scenarios such as target appearance changes, occlusions, and lighting changes; (4) Although existing 3D tracking technologies can provide accurate three-dimensional information, they usually rely on expensive multi-sensor equipment, which not only increases costs but also limits the feasibility and economy of these methods in practical applications; (5) When dealing with ATS multimodal 3D object tracking tasks in monocular videos, existing technologies often cannot effectively integrate visual, language, and geometric features, resulting in insufficient tracking accuracy and robustness; (6) Existing methods have shortcomings in achieving human-like accuracy and robustness, especially in complex and dynamic task environments, where it is difficult to cope with changes in target appearance or the presence of interference, and it is difficult to maintain continuity; (7) Existing visual language tracking technologies are mostly limited to two-dimensional space and fail to effectively simulate the natural perception ability of humans in complex three-dimensional environments, which limits their ability to perceive and track objects in three-dimensional space.
[0006] The present invention aims to solve the challenges faced by existing target tracking (SOT) methods when tracking 3D objects in monocular videos, especially the performance degradation problem in complex scenarios such as target appearance changes, occlusions, and lighting changes. In addition, most of the existing technologies are limited to two-dimensional space, lack the ability to perceive objects in three-dimensional space, and rely on multi-sensor data, which is different from the natural perception method of humans who mainly rely on visual and language clues. The present invention proposes a natural language-driven ATS multimodal 3D target tracking technology Mono3DVLT-MT, which aims to achieve tracking of three-dimensional objects through monocular RGB video and natural language description, thereby being closer to human perception and improving the practicality and economy of the tracking system. Summary of the invention
[0007] In view of this, the present invention provides a semantically driven 3D single target tracking method under ATS single vision.
[0008] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0009] A semantically driven 3D single target tracking method for ATS single vision includes the following steps:
[0010] Step 1: Build a large-scale dataset;
[0011] Step 2: Multimodal feature extraction of visual data and language data;
[0012] Step 3: Designing the Vision-Language Encoder
[0013] Design a language-guided visual encoder and a deep encoder to perform global contextual reasoning with pixel-level attention.
[0014] Step 4: Design a memory-enhanced tracking decoder
[0015] A memory-enhanced tracking decoder is introduced. The decoder continuously optimizes the query by interacting with the TTM, realizing an iterative process to continuously optimize and store query information.
[0016] Step 5: Design the Tracking Head
[0017] Use multiple multi-layer perceptrons (MLPs) to predict the properties of the target object in each frame of the video;
[0018] Step 6: Construct a comprehensive loss function.
[0019] Preferably, in step 1, the large-scale data set of Mono3DVLT is used, which contains 79,158 natural language descriptions, which are generated by ChatGPT and manually optimized based on the V2X-Seq data set.
[0020] Preferably, the step 2 specifically includes the following steps:
[0021] (2a) For visual data, the Swin Transformer model is used to extract multi-scale visual features of each frame in the video; the geometric depth features of each frame are obtained through a lightweight depth predictor;
[0022] (2b) For language data, the pre-trained RoBERTa model is used to extract the features of the input language description; language tokens and multi-scale visual features are extracted through linear layers.
[0023] Preferably, in step 3, the encoder uses multi-scale deformable attention MSDA and multi-head cross attention MHCA to strengthen the connection between visual features and language description;
[0024] The specific steps include:
[0025] Step 3.1: For each 3D object, use its corresponding 2D bounding box to crop the object area from the original RGB video and obtain the corresponding 2D image information;
[0026] Step 3.2: Using the 2D image information, extract the visual features of the 3D object through the pre-trained SwinTransformer network v , and processes the visual features f through a multi-head self-attention mechanism v , obtain fine features f v ', calculate the attention weights between different parts of the image features, capturing the relationship between various regions within the object;
[0027] Step 3.3: For each 3D object, use its corresponding 3D geometric information to construct a text description through the designed fixed template. The text description is input into the pre-trained RoBERTa model, which encodes it into a text embedding vector f t ;
[0028] Step 3.4: Under the guidance of image features, relevant semantic information is extracted from 3D text features and integrated with visual features to form a joint representation of object appearance and geometric attributes to obtain complete object information f a .
[0029] Preferably, in step 3.4, the specific implementation method is as follows:
[0030] Adopt a dual-head attention mechanism to take the image feature f v Considered as query Q, the CLS from 3D text features is used as key K and value V for cross attention calculation, and the formula is as follows:
[0031]
[0032] Compute query-key attention graph A tt , and aggregate the weight information to obtain a visual and 3D text-aware query Q', as follows:
[0033]
[0034] Preferably, the step 4 specifically includes the following steps:
[0035] Step 4.1: For a given description, focus on the relevant parts and apply a bidirectional attention mechanism to achieve a preliminary fusion between language and object features, so that language and object features complement and enhance each other; language features guide the model to focus on the visual aspects of objects related to the description, while the visual features of the object enrich the semantic information of the language description, and language features and object features are used alternately as queries, keys, and values;
[0036] Step 4.2: The fused object and language features are first concatenated and then input into the module for adaptive fusion to capture the interaction between cross-modal features. The input features are first channel-mixed through convolution and then activated using the SiLU function; the processed features are input into the forward and backward modules, which work in parallel to capture contextual information from different directions in the feature sequence.
[0037] Preferably, in step 4.1, the process is as follows:
[0038] O 2 T=MHCA(p t ,f a ,f a ),T 2 O=MHCA(f a ,p t ,p t )
[0039] Among them, Q 2 T represents the feature mapping from vision to language, and multi-head cross attention MHCA is used to transform the language feature P t and the visual features of the object f a Fusion; T 2 O represents the feature mapping from language to vision.
[0040] Preferably, in step 4.2, the process is as follows:
[0041] x input =Concat(MLP(T 2 O),MLP(O 2 T))
[0042] x mixed =SiLU(Conv 1x1 (x input ))
[0043] S forward =SSM(x mixed )s
[0044] S backward =Flip(SSM(Flip(x mixed )))
[0045] Among them, the input feature x input The two feature MLPs (T 2 O) and MLP(O 2 T) The result after fusion; the mixed feature x mixed The SiLU activation function is applied to the 1×1 convolutional layer Conv 1x1The features obtained introduce nonlinear relationships; forward sequence features S forward It is the result of further processing by the adaptive sequence module SSM; the backward sequence feature S backward Using the sequence flip operation Flip and the adaptive sequence module to mixed Processed.
[0046] Preferably, step 5 includes object category, 2D and 3D bounding box sizes, and object orientation and depth.
[0047] Preferably, in step 6, the comprehensive loss function is designed to include 2D loss L2D, 3D loss L3D and depth map loss Ldmap.
[0048] Compared with the prior art, the present invention has achieved the following technical effects:
[0049] (1) This paper proposes an innovative natural language-driven ATS multimodal 3D target tracking method in monocular video, namely Mono3DVLT-MT, which significantly improves the accuracy and robustness of 3D object tracking in monocular video by integrating visual, language and geometric features;
[0050] (2) Compared with the prior art, the present invention not only reduces the reliance on expensive multi-sensor equipment and reduces system costs, but also enhances the human-like perception capability of the tracking system by simulating the natural way humans use visual and language descriptions to perceive and track objects;
[0051] (3) The present invention also constructs a large-scale Mono3DVLT-V2X dataset, which contains 79,158 natural language descriptions, providing rich training and evaluation resources for single target tracking tasks;
[0052] (4) The end-to-end network Mono3DVLT-MT designed by the present invention adopts a memory-enhanced Token TuringMachine (TTM) structure, which optimizes the query process and improves data processing efficiency;
[0053] (5) Through extensive experiments and ablation studies, the present invention shows significant superiority in various benchmark tests, verifying the effectiveness of the proposed method.
[0054] In summary, the present invention not only achieves technological innovation, but also demonstrates significant advantages in practical applications, providing a new solution for the field of monocular video 3D visual language tracking and promoting the development of the field of computer vision. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Generate a flow chart for the data of the present invention;
[0056] Figure 2 is a visual-linguistic encoder diagram of the present invention;
[0057] Figure 3 This is a diagram of a tracking decoder with TTM according to the present invention. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] The present invention discloses a semantically driven 3D single target tracking method under ATS single vision, which is characterized by comprising the following steps:
[0060] Step 1: Build a large-scale dataset
[0061] This paper first proposes a large-scale dataset called Mono3DVLT, which contains 79,158 natural language descriptions for single target tracking;
[0062] These descriptions are generated by ChatGPT and manually optimized based on the V2X-Seq dataset to ensure the accuracy and uniqueness of the descriptions. Figure 1 shown.
[0063] Step 2: Multimodal feature extraction
[0064] (2a) For visual data, the present invention uses the Swin Transformer model to extract multi-scale visual features of each frame in the video; and obtains the geometric (depth) features of each frame through a lightweight depth predictor;
[0065] (2b) For language data, the present invention uses a pre-trained RoBERTa model to extract features of the input language description; and further extracts language tokens and multi-scale visual features through a linear layer.
[0066] Step 3: Designing the Vision-Language Encoder
[0067] The present invention designs a language-guided visual encoder and a deep encoder to perform global context reasoning based on pixel-level attention. The encoder uses multi-scale deformable attention (MSDA) and multi-head criss-cross attention (MHCA) to strengthen the connection between visual features and language descriptions, such as Figure 2 shown.
[0068] Step 3.1: For each 3D object, use its corresponding 2D bounding box to crop the object area from the original RGB video and obtain the corresponding 2D image information;
[0069] Step 3.2: Using the 2D image information, extract the visual features of the 3D object through the pre-trained Swin Transformer network v , and processes the visual features f through a multi-head self-attention mechanism v , obtain fine features f v ', calculate the attention weights between different parts of the image features, effectively capturing the relationship between various regions within the object;
[0070] Step 3.3: For each 3D object, use its corresponding 3D geometric information to construct a text description through the designed fixed template. The text description is input into the pre-trained RoBERTa model, which encodes it into a text embedding vector f t ;
[0071] Step 3.4: Under the guidance of image features, relevant semantic information is selectively extracted from 3D text features and fused with visual features to form a joint representation of object appearance and geometric attributes to obtain complete object information f a ;
[0072] The specific implementation method is as follows:
[0073] Adopt a dual-head attention mechanism to take the image feature f v is regarded as the query (Q), and [CLS] from the 3D text feature is used as the key (K) and value (V) for cross attention calculation, the formula is as follows:
[0074]
[0075] Compute query-key attention graph A tt , and aggregate the weight information to obtain a visual and 3D text-aware query Q', as follows:
[0076]
[0077] Step 4: Design a memory-enhanced tracking decoder
[0078] The present invention introduces a memory-enhanced tracking decoder that uses an improved Token Turing Machine (TTM) architecture; the decoder continuously optimizes queries through interaction with the TTM, implementing an iterative process to continuously optimize and store query information, thereby achieving efficient tracking, such as Figure 3 shown.
[0079] Step 4.1: For a given description, selectively focus on the relevant parts and apply a bidirectional attention mechanism to achieve a preliminary fusion between language and object features, so that language and object features complement and enhance each other. Language features guide the model to focus on the visual aspects of objects related to the description, while the visual features of the object enrich the semantic information of the language description. Language features and object features are used alternately as queries, keys, and values. The process is as follows:
[0080] O 2 T=MHCA(p t ,f a ,f a ),T 2 O=MHCA(f a ,p t ,p t )
[0081] Among them, Q 2 T represents the feature mapping from vision to language, and multi-head cross attention MHCA is used to transform the language feature P t and the visual features of the object f a Fusion; T 2 O represents the feature mapping from language to vision.
[0082] Step 4.2: The fused object and language features are first concatenated and then input into the module for adaptive fusion to effectively capture the interaction between cross-modal features. The input features are first channel-mixed through convolution and then activated using the SiLU function. The processed features are then input into the forward and backward modules, which work in parallel to capture contextual information from different directions in the feature sequence. The process is as follows:
[0083] x input =Concat(MLP(T 2 O),MLP(O 2 T))
[0084] x mixed =SiLU(Conv 1x1 (x input ))
[0085] S forward =SSM(x mixed )s
[0086] S backward =Flip(SSM(Flip(x mixed )))
[0087] Among them, the input feature x input The two feature MLPs (T 2 O) and MLP(O2 T) The result after fusion; the mixed feature x mixed The SiLU activation function is applied to the 1×1 convolutional layer Conv 1x1 The features obtained introduce nonlinear relationships; forward sequence features S forward It is the result of further processing by the adaptive sequence module SSM; the backward sequence feature S backward Using the sequence flip operation Flip and the adaptive sequence module to mixed Processed.
[0088] Step 5: Design the tracking head
[0089] The tracking head of the present invention uses multiple multi-layer perceptrons (MLPs) to predict the properties of the target object in each frame of the video, including object category, 2D and 3D bounding box size, object orientation and depth, etc.
[0090] Step 6: Construct a comprehensive loss function
[0091] A comprehensive loss function including 2D loss (L2D), 3D loss (L3D) and depth map loss (Ldmap) is designed to optimize the tracking performance.
[0092] Through this technical solution, the Mono3DVLT method can effectively handle 3D object tracking tasks in monocular videos and maintain high tracking performance even in complex scenarios such as target appearance changes, occlusions, and lighting changes.
[0093] In addition, the method does not rely on expensive multi-sensor equipment, which improves the practicality and economy of the system. By integrating visual, language and geometric features, the present invention can achieve accurate tracking of 3D objects in videos, providing a new solution for the field of computer vision.
[0094] Example 1
[0095] Simulation conditions
[0096] The present invention conducts experiments using the deep learning framework Pytorch-GPU 2.0.0 on a Linux operating system with an Intel(R) Xeon(R) Gold 6248CPU@2.50GHz hardware configuration and an NVIDIA Tesla V100 GPU as the graphics processor.
[0097] The monocular video dataset used in the experiment is the V2X-Seq dataset, which contains rich vehicle attributes and natural language descriptions.
[0098] The methods compared in the experiment are as follows:
[0099] One is the zero-sample target localization method, which can locate the target from the text description without target annotation. It is mainly used to find the corresponding object in the image through text guidance. It is denoted as ZSGNet in the experiment.
[0100] The other is the dynamic graph attention mechanism for reference expression understanding method. This is a lightweight visual positioning method that uses the dynamic graph attention mechanism to achieve efficient positioning of the target through natural language reference expressions. It is denoted as FAOA in the experiment.
[0101] In addition, there is a recursive subquery construction method. This is a method to improve single-stage visual localization by constructing recursive subqueries, which can decompose complex reference expressions into more tractable subexpressions, thereby improving localization performance. It is denoted as ReSC in the experiment.
[0102] The last one is the Transformer-based visual localization method. By using the Transformer architecture to achieve joint modeling of language and visual features, efficient target localization is achieved through cross-modal interaction. This method emphasizes the deep fusion of language and visual signals. It is denoted as TransVG in the experiment.
[0103] Simulation content
[0104] According to the specific implementation of the present invention, a series of benchmark experiments are designed in the Mono3DVLT task to compare the performance of various methods in target positioning and 3D tracking tasks, including the zero-shot target positioning method ZSGNet, the dynamic graph attention mechanism method FAOA, the recursive subquery construction method ReSC, and the Transformer-based visual positioning method TransVG.
[0105] Table 1: Comparison of Mono3DVLT-MT and baseline methods
[0106]
[0107]
[0108] As shown in the experimental results in Table 1, the Mono3DVLT-MT method proposed in the present invention is superior to the ZSGNet, FAOA, ReSC and TransVG methods in terms of accuracy and robustness of target positioning and 3D tracking. Specifically, it is manifested in the following aspects: The Mono3DVLT-MT method significantly improves the positioning accuracy, and performs particularly well in scenarios with complex language descriptions, target occlusion and lighting changes. Compared with the zero-sample target positioning method ZSGNet, this method is more adaptable to diverse language descriptions. By integrating visual, language and geometric features, the 3D target tracking accuracy of Mono3DVLT-MT in complex scenes is superior to that of the dynamic graph attention mechanism FAOA and the recursive subquery construction method ReSC. Compared with the TransVG method that relies on the Transformer architecture, the method of the present invention shows higher efficiency in processing long sequence videos, while reducing dependence on hardware resources.
[0109] In summary, the Mono3DVLT-MT method successfully improved the overall performance of the task by innovatively combining text-guided 2D positioning and geometry-assisted 3D tracking, achieving higher tracking accuracy and scene adaptability.
[0110] From the experimental results, it can be seen that due to the design of the multimodal feature extractor, visual-language encoder, memory-enhanced tracking decoder and comprehensive loss function adopted by the present invention, it is possible to effectively integrate visual, language and geometric information to achieve accurate tracking of 3D objects in monocular video. This verifies the advancement and practicality of the present invention in the field of 3D visual language tracking of monocular video, and demonstrates its advantages in improving tracking performance and reducing system costs.
[0111] The above description is only a preferred embodiment of the present invention and does not limit the technical scope of the present invention. Therefore, any slight modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A semantically driven 3D single target tracking method for ATS single vision, characterized in that: The steps include: Step 1: Build a large-scale dataset; Step 2: Multimodal feature extraction of visual data and language data; Step 3: Designing the Vision-Language Encoder Design a language-guided visual encoder and a deep encoder to perform global contextual reasoning with pixel-level attention. Step 4: Design a memory-enhanced tracking decoder A memory-enhanced tracking decoder is introduced. The decoder continuously optimizes the query by interacting with the TTM, realizing an iterative process to continuously optimize and store query information. Step 5: Design the Tracking Head Use multiple multi-layer perceptrons (MLPs) to predict the properties of the target object in each frame of the video; Step 6: Construct a comprehensive loss function.
2. According to the semantically driven 3D single target tracking method for ATS single vision according to claim 1, it is characterized in that: In step 1, the large-scale dataset of Mono3DVLT is used, which contains 79,158 natural language descriptions generated by ChatGPT and manually optimized based on the V2X-Seq dataset.
3. According to the semantically driven 3D single target tracking method for ATS single vision according to claim 1, it is characterized in that: The step 2 specifically includes the following steps: (2a) For visual data, the Swin Transformer model is used to extract multi-scale visual features of each frame in the video; the geometric depth features of each frame are obtained through a lightweight depth predictor; (2b) For language data, the pre-trained RoBERTa model is used to extract the features of the input language description; language tokens and multi-scale visual features are extracted through linear layers.
4. The semantically driven 3D single target tracking method for ATS single vision according to claim 1, characterized in that: In step 3, the encoder uses multi-scale deformable attention MSDA and multi-head cross attention MHCA to strengthen the connection between visual features and language description; The specific steps include: Step 3.1: For each 3D object, use its corresponding 2D bounding box to crop the object area from the original RGB video and obtain the corresponding 2D image information; Step 3.2: Using the 2D image information, extract the visual features of the 3D object through the pre-trained SwinTransformer network v , and processes the visual features f through a multi-head self-attention mechanism v , obtain fine features f v ', calculate the attention weights between different parts of the image features, capturing the relationship between various regions within the object; Step 3.3: For each 3D object, use its corresponding 3D geometric information to construct a text description through the designed fixed template. The text description is input into the pre-trained RoBERTa model, which encodes it into a text embedding vector f t ; Step 3.4: Under the guidance of image features, relevant semantic information is extracted from 3D text features and integrated with visual features to form a joint representation of object appearance and geometric attributes to obtain complete object information f a .
5. The semantically driven 3D single target tracking method for ATS single vision according to claim 4 is characterized in that: In step 3.4, the specific implementation method is as follows: Adopt a dual-head attention mechanism to take the image feature f v Considered as query Q, the CLS from 3D text features is used as key K and value V for cross attention calculation, and the formula is as follows: Q=Linear(f v′ ),K,V=Linear(f t cls ) Compute query-key attention graph A tt , and aggregate the weight information to obtain a visual and 3D text-aware query Q', as follows:
6. The semantically driven 3D single target tracking method for ATS single vision according to claim 1, characterized in that: The step 4 specifically includes the following steps: Step 4.1: For a given description, focus on the relevant parts and apply a bidirectional attention mechanism to achieve a preliminary fusion between language and object features, so that language and object features complement and enhance each other; language features guide the model to focus on the visual aspects of objects related to the description, while the visual features of the object enrich the semantic information of the language description, and language features and object features are used alternately as queries, keys, and values; Step 4.2: The fused object and language features are first concatenated and then input into the module for adaptive fusion to capture the interaction between cross-modal features. The input features are first channel-mixed through convolution and then activated using the SiLU function; the processed features are input into the forward and backward modules, which work in parallel to capture contextual information from different directions in the feature sequence.
7. The semantically driven 3D single target tracking method for ATS single vision according to claim 6, characterized in that: In step 4.1, the process is as follows: O2T=MHCA(p t ,f a ,f a ),T2O=MHCA(f a ,p t ,p t ) Among them, Q2T represents the feature mapping from vision to language, and multi-head cross attention MHCA is used to transform the language feature P t and the visual features of the object f a Fusion; T2O represents feature mapping from language to vision.
8. The semantically driven 3D single target tracking method for ATS single vision according to claim 6, characterized in that: In step 4.2, the process is as follows: x input =Concat(MLP(T2O),MLP(O2T)) x mixed =SiLU(Conv 1x1 (x input )) S forward =SSM(x mixed )s S backward =Flip(SSM(Flip(x mixed ))) Among them, the input feature x input It is the result of fusing the two features MLP(T2O) and MLP(O2T) through the Concat operation; the mixed feature x mixed The SiLU activation function is applied to the 1×1 convolutional layer Conv 1x1 The features obtained introduce nonlinear relationships; forward sequence features S forward It is the result of further processing by the adaptive sequence module SSM; the backward sequence feature S backward Using the sequence flip operation Flip and the adaptive sequence module to mixed Processed.
9. The semantically driven 3D single target tracking method for ATS single vision according to claim 1, characterized in that: In step 5, the object category, 2D and 3D bounding box sizes, direction and depth of the object are included.
10. The semantically driven 3D single target tracking method for ATS single vision according to claim 1, characterized in that: In step 6, a comprehensive loss function is designed including 2D loss L2D, 3D loss L3D and depth map loss Ldmap.