Vehicle target tracking method and system based on natural language reference
By constructing a natural language-based vehicle target tracking system based on BERT and CLIP, the problems of accurate target designation and cross-view association in existing technologies are solved. This system achieves accurate positioning and continuous tracking over a wide area in complex scenarios, improving the flexibility and robustness of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-04-03
AI Technical Summary
Existing vehicle target tracking technologies cannot accurately identify specific vehicles in multi-target scenarios and lack cross-view semantic understanding mechanisms, leading to identity continuity problems and trajectory breaks, making it difficult to achieve intelligent, human-like interaction and continuous tracking over a wide area.
By constructing a semantically driven initialization mechanism, a language-enhanced identity maintenance system, and a cross-perspective semantic unification framework, and integrating BERT and CLIP dual-language models, multimodal feature extraction and cross-perspective target association are achieved. Natural language instructions are used for accurate target localization and initialization. Combined with short-term and long-term identity maintenance mechanisms, the differences in semantic mapping under different perspectives are resolved.
It enables precise positioning of vehicle targets in complex scenarios, ensuring identity consistency and trajectory integrity, improving the system's flexibility and robustness, supporting collaborative work of multi-source heterogeneous sensors, and achieving continuous tracking over a wide area.
Smart Images

Figure CN121789149A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle target tracking technology, and specifically refers to a vehicle target tracking method and system based on natural language reference. Background Technology
[0002] In recent years, with the rapid development of computer vision and artificial intelligence technologies, vehicle target tracking has become one of the core technologies in fields such as intelligent transportation systems, autonomous driving, and public safety monitoring. This technology aims to achieve continuous localization and identity preservation of vehicle targets in video sequences, providing crucial data support for tasks such as traffic flow analysis, behavior recognition, and event detection. From early traditional methods based on Kalman filtering and mean shift to the current mainstream deep learning-based target detection and multi-target tracking frameworks, this field has made significant progress. Existing methods typically construct and associate vehicle trajectories by fusing appearance features, motion information, and spatiotemporal context, which to some extent meets the basic needs of practical applications.
[0003] However, through in-depth analysis and practical application verification of existing technologies, it has been found that traditional vehicle tracking methods still have significant limitations: 1. Existing technologies lack precise target designation and interaction capabilities based on natural language: Traditional vehicle tracking systems cannot accurately locate specific vehicles through semantic commands in multi-target scenarios. Operators must manually select targets from numerous detection results using methods such as mouse clicks and bounding boxes. This operation is cumbersome and extremely inaccurate in complex scenarios such as high-density traffic flow and overlapping targets. The system cannot understand high-level semantic commands such as "track the black sedan driving on the far left across the line," limiting its application in law enforcement, security, and other scenarios requiring rapid and accurate target designation, and making it difficult to achieve an intelligent and user-friendly interactive experience.
[0004] 2. Existing methods face challenges in maintaining identity continuity in complex environments: In the short term, the association mechanism relying on vision and motion models is insufficient in the discrimination ability in scenarios such as occlusion and similar appearance, leading to frequent identity switching; In the long term, the lack of an effective semantic memory mechanism makes it impossible to cope with the challenge of re-identification after the target disappears and reappears for a long time, resulting in tracking interruption and trajectory breakage.
[0005] 3. Existing methods are limited to single sensors or homogeneous camera networks and lack cross-view semantic understanding mechanisms: they cannot handle the spatial mapping differences of the same semantic command under different viewpoints. For example, the spatial orientation of "vehicles driving from left to right" is completely different in aerial photography from above, side-view fixed cameras, and front-view vehicle cameras. This leads to confusion in cross-view target reference and failure of association, making it difficult to achieve continuous trajectory tracking in a wide area. There are also problems such as coverage blind spots, uncompensated occlusion, and target loss due to viewpoint differences. Summary of the Invention
[0006] To address the technical problems existing in the prior art, the present invention provides a vehicle target tracking method and system based on natural language reference, the technical solution of which is as follows: On the one hand, a vehicle target tracking method based on natural language reference is provided, which includes: S1. Receive and preprocess video stream input from multiple source sensors and natural language commands from the user to obtain standardized image data stream and key semantic information; S2. Input the standardized image data stream and key semantic information into the semantic reference initialization module. Through cross-modal feature extraction, deep fusion and semantic consistency verification, achieve accurate target localization and initialization, and output the target mask. S3. Input the target mask into a language enhancement and memory-driven full-time identity maintenance system, integrating short-term identity maintenance and long-term recovery mechanisms to ensure identity consistency and trajectory integrity. The full-time identity maintenance system includes a short-term identity maintenance submodule and a long-term identity recovery submodule. The short-term identity maintenance submodule receives the target mask, realizes inter-frame identity association through multimodal feature fusion and optimized matching, and outputs stable tracking target information and all target information. The long-term identity recovery submodule receives the stable tracking target information output by the short-term identity maintenance submodule, constructs a cross-modal semantic memory library by drawing on the idea of contrastive learning, realizes identity recovery after the target disappears and reappears, and outputs long-term recovered target information. S4. Input the multi-view raw image stream from multiple source sensors into the cross-view semantic unification layer, preprocess and semantically align it to provide a unified semantic representation basis for subsequent cross-view target association, and output unified semantic representation and processed multi-view image data. S5. The unified semantic representation and processed multi-view image data output from the cross-view semantic unification layer, all target information output from the short-term identity maintenance submodule, and long-term recovery target information output from the long-term identity recovery submodule are input into the cross-view semantic unification association module. By coordinate system transformation and semantic direction standardization, the semantic mapping differences under different perspectives are resolved, and the accurate association and trajectory fusion of cross-view targets are realized, forming continuous tracking in a wide area.
[0007] On the other hand, a vehicle target tracking system based on natural language reference is provided, the system comprising: The receiving preprocessing module is used to receive and preprocess video stream input from multiple source sensors and natural language instructions from users to obtain standardized image data streams and key semantic information. The semantic reference initialization module is used to input the standardized image data stream and key semantic information into the semantic reference initialization module. Through cross-modal feature extraction, deep fusion and semantic consistency verification, it achieves accurate target localization and initialization and outputs the target mask. A full-time identity maintenance system is used to input the target mask into a language enhancement and memory-driven full-time identity maintenance system, integrating short-term identity maintenance and long-term recovery mechanisms to ensure identity consistency and trajectory integrity. The full-time identity maintenance system includes a short-term identity maintenance submodule and a long-term identity recovery submodule. The short-term identity maintenance submodule receives the target mask, realizes inter-frame identity association through multimodal feature fusion and optimized matching, and outputs stable tracking target information and all target information. The long-term identity recovery submodule receives the stable tracking target information output by the short-term identity maintenance submodule, constructs a cross-modal semantic memory based on the idea of contrastive learning, realizes identity recovery after the target disappears and reappears, and outputs long-term recovered target information. The cross-view semantic unification layer is used to input the raw image streams from multiple sources of sensors, preprocess and semantically align them, provide a unified semantic representation basis for subsequent cross-view target association, and output unified semantic representation and processed multi-view image data. The cross-view semantic unified association module is used to input the unified semantic representation and processed multi-view image data output from the cross-view semantic unified layer, all target information output from the short-term identity maintenance submodule, and long-term recovered target information output from the long-term identity recovery submodule into the cross-view semantic unified association module. By transforming the coordinate system and standardizing the semantic direction, the differences in semantic mapping under different perspectives are resolved, and the accurate association and trajectory fusion of cross-view targets are achieved, forming continuous tracking in a wide area.
[0008] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described vehicle target tracking method based on natural language reference.
[0009] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement the above-described vehicle target tracking method based on natural language reference.
[0010] The beneficial effects of the technical solution provided by this invention include at least the following: 1. The semantic reference initialization module of the present invention achieves deep semantic parsing of user text or voice commands by integrating a reference segmentation model of visual and linguistic modalities. It can accurately locate and initialize targets in complex scenarios, significantly improving the system's flexibility and human-computer interaction efficiency.
[0011] 2. The language enhancement and memory-driven full-time identity maintenance system of the present invention ensures identity consistency through two collaborative sub-modules: short-term and long-term. At the short-term level, language description and segmentation mask are integrated to enhance the ability to identify targets and effectively prevent identity confusion. At the long-term level, a cross-modal semantic memory is established to achieve reliable identity recovery and form a complete "prevention-recovery" mechanism.
[0012] 3. The cross-view semantic unification and target association module for multi-source heterogeneous sensors of this invention breaks through the limitations of traditional single-camera systems, supporting collaborative work of various sensors such as fixed monitoring, vehicle-mounted sensors, and drones. It innovatively constructs a cross-view semantic spatial mapping mechanism, solving the problem of spatial mapping differences for the same semantic command under different views through viewpoint geometric correction and semantic direction standardization. It achieves accurate target identification, precise association, and trajectory stitching across views, significantly improving system robustness and coverage through the complementary advantages of multiple sensors. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of a vehicle target tracking method based on natural language reference provided in an embodiment of the present invention; Figure 2 This is a general block diagram of a vehicle target tracking method based on natural language reference provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the multimodal data input and preprocessing process provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the semantic reference initialization module processing procedure provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the short-term identity maintenance submodule processing procedure provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the identity decision matching process provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the processing procedure of the multidimensional similarity calculation module provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the long-term identity recovery submodule processing procedure provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the cross-perspective semantic unification layer processing procedure provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the cross-perspective semantic unified association module provided in an embodiment of the present invention; Figure 11 This is a block diagram of a vehicle target tracking system based on natural language reference provided in an embodiment of the present invention; Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0015] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0016] This invention provides a vehicle target tracking method based on natural language reference. By constructing a semantically driven initialization mechanism, a language-enhanced identity maintenance system, and a cross-view semantic unification framework, it achieves intelligent processing throughout the entire process, from natural language instruction parsing to multi-view continuous tracking. It integrates the BERT and CLIP dual-language models: BERT is used for deep semantic understanding and feature fusion in the initialization phase, while CLIP is used for semantic verification and semantic feature extraction in identity maintenance. Drawing on the multimodal fusion concept of reference segmentation technology, it receives video stream input from multiple source sensors and natural language instructions from the user. Through the cross-modal attention mechanism of visual-language models such as CLIP, it integrates semantic instructions... The system directly parses the data into a target segmentation mask, enabling interactive initialization without manual annotation. It constructs a language-enhanced and memory-driven end-to-end identity maintenance system. In short-term tracking, it integrates language descriptions, visual appearance (based on a ReID network), and segmentation mask features to prevent identity confusion. In long-term tracking, it draws on contrastive learning to build a cross-modal semantic memory, enabling identity recovery after target disappearance and reappearance. For heterogeneous sensors such as fixed surveillance cameras, vehicle-mounted cameras, and drones, it designs a cross-view semantic unification and target association module. Through coordinate system transformation and semantic direction standardization, it resolves semantic mapping differences under different viewpoints, achieving accurate cross-view target association and trajectory fusion, forming continuous tracking over a wide area.
[0017] To support the implementation of this invention, a multimodal training dataset including semantic annotation is first constructed: Video streams at 25fps and 1920×1080 resolution are collected from heterogeneous devices such as fixed surveillance cameras, vehicle-mounted cameras, and drone aerial photography. Based on this, a three-layer annotation system is established: The first layer annotates vehicle targets with bounding boxes and records their position coordinates (x, y, w, h), tracking ID, and attributes such as vehicle color and model. The second layer is linguistic referential annotation, constructing natural language descriptions for each target, including appearance features, spatial relationships, and driving behavior. For example, a target vehicle can be labeled as: "Red SUV, located in the middle lane, changing lanes to the left," "The large red SUV in the middle," "The red SUV changing lanes to the left in the second lane," etc. Each target is annotated with 3-5 different descriptions, forming a language-visual pairing dataset. The third layer is cross-view ground truth annotation, manually annotating the true correspondence of the same vehicle under different sensor views for vehicle targets appearing simultaneously from multiple views, recording cross-view identity correspondence ground values, and providing supervision signals for training the cross-view association module. Meanwhile, for labeled samples including directional semantic instructions (such as "drive from left to right" and "drive forward"), the corresponding standardized direction vectors are labeled in a unified world coordinate system. This vector represents the standardized driving direction indicated by the command in the world coordinate system. Furthermore, calibration information such as the intrinsic matrix, extrinsic parameters, and distortion parameters of each sensor are collected and recorded, providing basic data for cross-view geometric transformation. Finally, based on the complete labeled data, a complete module consisting of a semantic reference initialization module, a full-time identity maintenance system, a cross-view semantic unification layer, and a cross-view semantic unification association module is trained. The trained model is then used to perform vehicle target tracking based on natural language reference in this embodiment of the invention.
[0018] This invention provides a vehicle target tracking method based on natural language reference. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of this method is shown below. Figure 2 The diagram shown is an overall block diagram of the method. The processing flow may include the following steps:
[0019] S1. Receive and preprocess video stream input from multiple source sensors and natural language commands from the user to obtain standardized image data stream and key semantic information; like Figure 3As shown, in this embodiment of the invention, during runtime, inputs are simultaneously received from multiple sources of sensors, including fixed surveillance cameras, vehicle-mounted cameras, and drone aerial photography, as well as natural language commands input by the user (such as "track that black SUV"). The video stream undergoes preprocessing processes including image enhancement, distortion correction, and resolution normalization, ultimately forming a standardized image data stream; simultaneously, the natural language commands undergo text cleaning and word segmentation, extracting key semantic information (such as vehicle attributes, location, and behavior).
[0020] S2. Input the standardized image data stream and key semantic information into the semantic reference initialization module. Through cross-modal feature extraction, deep fusion and semantic consistency verification, achieve accurate target localization and initialization, and output the target mask. Optionally, such as Figure 4 As shown, the processing procedure of the semantic reference initialization module specifically includes: After the standardized image data stream is input into the visual backbone network, multi-scale visual features are extracted through multi-layer convolution and pooling operations. These correspond to feature representations of different receptive fields: high-resolution features capture fine edge and texture information, medium-resolution features encode local structural features, and low-resolution features represent global semantic information. The key semantic information is first input into a language encoder, where a pre-trained language model BERT converts the text description into a high-dimensional semantic feature vector (which contains multi-dimensional semantic information such as the target vehicle's attributes, location, and behavior). Then, a deep semantic feature extraction module uses a multi-layer Transformer encoder and a self-attention mechanism to capture long-distance dependencies between words, generating a more discriminative deep semantic feature representation. ; In the cross-modal attention fusion stage, a parallel dual-branch architecture is adopted to achieve deep fusion of linguistic and visual features: the first branch uses a cross-attention mechanism to establish semantic correspondence, integrating deep semantic features. As the query vector, visual features serve as the key and value vectors. An attention weight matrix is calculated for each scale i. After softmax normalization, the visual features are weighted and summed to generate fused features, allowing the model to dynamically adjust the attention given to different regions of the image based on the text description. The second branch introduces a multi-head attention enhancement mechanism, capturing semantic associations at different levels from multiple independent representation subspaces in parallel to generate enhanced features. Residual connections preserve the original visual feature information, preventing information loss during deep fusion. Then, semantically aligned feature fusion is performed, concatenating and adaptively weighting the outputs of the two parallel branches. Features at different scales are bilinearly upsampled to align spatial dimensions. Channel concatenation and convolutional dimensionality reduction operations generate unified fused multimodal features. This feature integrates visual appearance, semantic attributes, and spatial structure information, and is then used to generate identity features through a feature extraction network. This serves as a benchmark for maintaining one's identity in the future; The fused multimodal features The data is fed into two parallel output head branches for processing: the pixel-level segmentation head uses a cross-attention mechanism to generate the target mask. The Sigmoid activation function is used to output the probability that each pixel belongs to the target, and the mask quality score is also output. Used to evaluate the reliability of segmentation results; the bounding box regression head generates the target bounding box through a multi-head attention enhancement mechanism. The output is normalized center coordinates and width / height parameters; the results from the two outputs are then combined for semantic consistency verification to comprehensively evaluate the initialization quality: first, the CLIP model is used to extract the visual features of the target mask region and the linguistic features of the user's text description, and the cosine similarity between the two is calculated as a semantic similarity index. The algorithm evaluates the degree of matching between the segmentation results and semantic instructions; it also calculates the overlap between the target mask and the bounding box as a spatial IoU metric. The spatial consistency of the results from the two output heads is evaluated, and the mask quality score of the pixel-level segmentation head output is used. This score reflects the connectivity, boundary smoothness, and confidence distribution of the segmentation results. The overall confidence score is calculated through weighted fusion. (weight) (Typically 0.4, 0.3, 0.3) Then, a threshold judgment is performed: when... Confirm successful initialization and output. , and Establish a baseline identity profile for the target. Set a preset threshold (usually 0.7); otherwise, reject the current result and ask the user to re-describe or provide a clearer reference, thereby ensuring high reliability of initialization.
[0021] S3. Input the target mask into a language enhancement and memory-driven full-time identity maintenance system, integrating short-term identity maintenance and long-term recovery mechanisms to ensure identity consistency and trajectory integrity. The full-time identity maintenance system includes a short-term identity maintenance submodule and a long-term identity recovery submodule. The short-term identity maintenance submodule receives the target mask, realizes inter-frame identity association through multimodal feature fusion and optimized matching, and outputs stable tracking target information and all target information. The long-term identity recovery submodule receives the stable tracking target information output by the short-term identity maintenance submodule, constructs a cross-modal semantic memory library by drawing on the idea of contrastive learning, realizes identity recovery after the target disappears and reappears, and outputs long-term recovered target information. Optionally, such as Figure 5 As shown, the processing procedure of the short-term identity maintenance submodule specifically includes: For the candidate target detected in frame t, the multimodal feature extraction stage is first entered: based on the target mask, three types of complementary features are extracted in parallel to construct a robust target representation, making full use of the synergistic effect of visual, semantic and geometric information. Among them, the depth visual features are extracted by the ReID network in the mask area to capture the visual information of the target and form a high-dimensional visual feature vector; the language semantic features are extracted by the CLIP model to maintain semantic consistency with the user's reference and ensure that the tracked target always conforms to the constraints of the language description; the segmentation mask features are extracted based on the target mask to extract geometric shape and contour information (including the target's aspect ratio, area, boundary curvature and other spatial structural characteristics) to enhance the robustness to target deformation and occlusion. The three types of features characterize the target characteristics from different dimensions and lay the foundation for subsequent fusion.
[0022] Then, the feature fusion enhancement stage is performed: a feature alignment network is used to project the three complementary features into a unified embedding space, eliminating the distribution differences between different feature representations and ensuring the comparability of features of each modality; then, through the attention enhancement mechanism, the contribution of each feature is dynamically calculated using a gated attention module: first, the three aligned features are concatenated and input into a multilayer perceptron, and the interdependencies between features are extracted through nonlinear transformation; then, a sigmoid activation function is used to generate three gate weight values, each weight corresponding to the importance score of a feature class. The gate mechanism can adaptively adjust the contribution of each feature according to the characteristics of the current scene (such as increasing the weight of visual features when the illumination changes, enhancing the weight of geometric features when there is severe occlusion, and strengthening the weight of semantic features when there is multi-target confusion, etc.). The fusion representation generates a comprehensive representation vector by weighted summing of each feature, which not only retains the unique advantages of each modality, but also achieves effective integration of information and improves the discriminative ability of target representation; Based on the fused representation, the multi-dimensional similarity calculation module is invoked to perform a matching evaluation between each candidate target in frame t and each tracked target in frame (t-1), including: All candidate targets in frame t are paired with all tracked targets in frame (t-1). Each target pair is sequentially input into the multi-dimensional similarity calculation module, which comprehensively considers multiple dimensions: appearance similarity, semantic similarity, and spatial IoU similarity (appearance similarity measures the visual similarity of targets by comparing the fused visual feature vectors; semantic similarity is based on CLIP semantic features to ensure that the target is consistent with the initial text description; spatial IoU similarity uses motion models such as Kalman filtering to predict the position of the target in the current frame and calculates the IoU between the predicted box and the detection box as a spatial constraint). A comprehensive similarity score is output for each target pair. The scores of all target pairs are organized into an M×N dimensional similarity matrix S according to the index relationship between candidate targets and tracked targets, where M is the number of candidate targets, N is the number of tracked targets, and S(i,j) represents the matching score between the i-th candidate target and the j-th tracked target, providing a quantitative basis for subsequent global optimal matching. Then it enters the identity matching and decision-making stage, such as Figure 6 As shown, the Hungarian algorithm is used to solve the similarity matrix S, achieving globally optimal target association and avoiding local suboptimal solutions that may be caused by greedy strategies. This includes: The similarity matrix S is converted into a cost matrix C, and mapped using the cost conversion formula C(i,j)=1-S(i,j). The Hungarian algorithm is used for global optimal matching. By adjusting the cost matrix C, a unique candidate target is found for each tracked target, and the total matching cost is minimized. Then, the confidence of the matching results is verified: if the similarity of the matching results is higher than the preset threshold (usually set to 0.6), the identity matching is confirmed to be successful, the candidate target is associated with the corresponding tracking target, and the feature representation, bounding box position and motion state of the target are updated; if the similarity of the matching results is lower than the preset threshold, the candidate target is determined to be a newly appearing target in the scene, a new tracking ID is assigned to it and the tracking state is initialized. After the identity matching decision is completed, the target information is distributed in two ways according to the target status: the stable tracking target information is output to the long-term identity recovery module for identity preservation and recovery; all target information is output to the cross-view semantic unified association module to support multi-view collaborative perception.
[0023] Optionally, such as Figure 7 As shown, the multidimensional similarity calculation module is applied in the short-term identity maintenance submodule, the long-term identity recovery submodule, and the cross-perspective semantic unified association module. It uses "target pairs" as input units, with each target pair including a complete feature set of the candidate target and the historical target. Through a structured parallel computing process, it ultimately outputs a comprehensive similarity score. The processing includes: After receiving the target pair, the two sets of input features are preprocessed and aligned to ensure that all features (visual, semantic, spatial, etc.) are in a unified representation space, laying the foundation for subsequent parallel comparison. Then, the parallel multidimensional computation stage begins, simultaneously generating similarity scores for multiple core dimensions: appearance similarity is based on the cosine similarity of visual features to measure appearance consistency; semantic similarity utilizes the Gaussian kernel similarity of linguistic features to evaluate the semantic matching degree with the initial description; spatial IoU similarity calculates the overlap between the predicted and detected positions using the intersection-union ratio (IoU); and contextual similarity integrates spatiotemporal environmental features to capture the association between behavioral patterns and scene context. Different sub-modules calculate all or part of these four similarities: the short-term identity maintenance sub-module calculates the first three similarities; and the long-term identity recovery sub-module and the cross-perspective semantic unified association module calculate all four similarities. Then, these independently calculated scalar scores are combined into a unified similarity vector, which is then fed into a normalization layer for processing to eliminate the dimensional differences between the dimensions and ensure that they are comparable on the same scale. For the normalized similarity vector, two strategies are provided for fusion: the first is a fixed weight strategy, which predefines a set of weight coefficients according to the specific task (such as short-term tracking, long-term recovery or cross-view association) and obtains a comprehensive score by weighted summation; the second is a more flexible adaptive weight strategy, which introduces a lightweight attention network. This network takes the current similarity vector as input, dynamically learns and outputs the optimal weights for each dimension, and realizes intelligent adjustment of the evaluation focus according to the specific matching context. Through the above process, a comprehensive similarity score is output for each input target pair. When processing multiple target pairs, these scores are organized into a similarity matrix for subsequent matching decisions.
[0024] Optionally, such as Figure 8 As shown, the processing procedure of the long-term identity recovery submodule specifically includes: A cross-modal semantic memory is constructed to establish a complete identity profile for the target. This memory adopts a multi-level feature storage architecture, and maintains four types of core features simultaneously for each tracked target: visual features (extracting visual representations such as appearance details, color and texture of the target), linguistic description features (encoding the semantic attributes of the target through natural language description, such as high-level semantics such as category, color, size, etc.), spatial features (capturing the spatial position relationship of the target in the scene), and contextual features (capturing the environmental context information of the target in the scene). This multi-modal feature fusion strategy ensures the robustness of identity representation, enabling identity recognition from multiple complementary dimensions and effectively dealing with the failure of a single feature dimension. To maintain the timeliness and stability of the memory features, an exponential moving average (EMA) update mechanism is introduced to dynamically maintain the stored features. After the short-term identity maintenance submodule successfully matches the target, four core features of the target in the current frame are extracted: visual features, linguistic description features, spatial features, and contextual features, as new features. For each type of feature, the exponential moving average (EMA) update mechanism is applied to update the new features of the current frame with the historical features of that type of feature stored in the memory. For the fusion update, for the k-th type of feature, a weighted update is performed using the momentum parameter α (usually set to 0.9), and the update formula is as follows: When a new candidate target is detected, it is paired one by one with all historical targets in the memory that have been lost, constructing multiple candidate target-historical target pairs. Using the continuously updated features of the memory, each target pair is input into the multidimensional similarity calculation module for parallel evaluation. Independent scores for appearance similarity, semantic similarity, spatial IoU similarity, and contextual similarity are calculated simultaneously. Then, through comprehensive similarity weighted calculation, an adaptive weight strategy is used to fuse the scores of each dimension into a unified comprehensive confidence score. The candidate matching pair with the highest comprehensive confidence score is selected for fusion judgment. The reliability of the matching is evaluated by combining the distribution characteristics of the scores of each dimension. Then, a rigorous spatiotemporal consistency verification is performed, including checking whether the target disappearance time is within a preset reasonable range and whether the reappearance location conforms to the motion model prediction based on the historical trajectory. This verification process ensures the rationality of the recovery results in the temporal and spatial dimensions. For all verified recovered targets, their long-term recovered target information, including the recovered tracking ID, current position, updated feature vector, and confidence score, is output to the cross-view semantic unified association module to participate in multi-view trajectory fusion. At the same time, the state information of the corresponding target in the memory bank is updated and synchronized to the short-term identity maintenance submodule to ensure continuous tracking in subsequent frames. For candidate targets that fail verification, a new tracking ID will be assigned to them and the corresponding memory bank entries will be initialized to establish a brand-new identity profile.
[0025] S4. Input the multi-view raw image stream from multiple source sensors into the cross-view semantic unification layer, preprocess and semantically align it to provide a unified semantic representation basis for subsequent cross-view target association, and output unified semantic representation and processed multi-view image data. Optionally, such as Figure 9 As shown, the processing procedure of the cross-perspective semantic unification layer specifically includes: The multi-view raw image stream first enters the preprocessing module to perform image enhancement and distortion correction operations to ensure that images from different sensors have a consistent quality benchmark; Then, based on the intrinsic and extrinsic parameters obtained from the calibration of each sensor, a unified coordinate system transformation is performed, and the obtained intrinsic parameter matrix is used. and external references Establish a world coordinate system as a unified reference for sensors. The transformation relationship between image coordinates and world coordinates is as follows: This transformation maps image observations from different perspectives to a unified world coordinate system, laying a geometric foundation for subsequent multi-perspective semantic understanding; After completing coordinate system one, the cross-view semantic alignment network learns the mapping relationship between visual features observed from different viewpoints and a unified semantic space, ensuring that the same target captured by multiple sensors maintains consistent semantic expression. The network adopts a multi-head attention mechanism to process visual features from various viewpoints simultaneously and capture common semantic representations between different viewpoints. For directional descriptions in user semantic commands (such as "from left to right" or "drive forward"), a standardized direction vector is defined in a unified coordinate system. This requires learning the mapping relationship from the textual semantic description to the standardized direction vector in the world coordinate system. Therefore, a contrastive learning approach is used to train the semantic alignment network, enabling it to map semantically identical directional descriptions from different perspectives to the same standardized direction vector space. This applies to user text commands. and the corresponding true normalized direction vector A contrastive learning loss function is constructed to make the mapping result of the text description close to the true direction and far away from other incorrect directions: in Represents text instructions After the mapping results of the semantic alignment network, To achieve true standardization, For other candidate directions, Represents cosine similarity. The temperature parameter ensures the consistency of spatial mapping of semantic instructions under different perspectives, so that the directional description can be accurately understood as a unified directional meaning in the perspectives of fixed surveillance cameras, vehicle cameras and drones. Then, the adaptive weight allocation module dynamically calculates the fusion weights based on the observation quality of each sensor, and comprehensively considers factors such as image clarity, target visibility, and viewpoint suitability to perform weighted fusion of semantic features from different viewpoints. The final output is a weighted fusion of unified semantic representation and processed multi-view image data, providing a consistent reference benchmark for subsequent cross-view association.
[0026] S5. The unified semantic representation and processed multi-view image data output from the cross-view semantic unification layer, all target information output from the short-term identity maintenance submodule, and long-term recovery target information output from the long-term identity recovery submodule are input into the cross-view semantic unification association module. By coordinate system transformation and semantic direction standardization, the semantic mapping differences under different perspectives are resolved, and the accurate association and trajectory fusion of cross-view targets are realized, forming continuous tracking in a wide area.
[0027] Optionally, such as Figure 10 As shown, the processing procedure of the cross-perspective semantic unified association module specifically includes: By utilizing the output of the cross-view semantic unification layer, the comparability of target observations from different perspectives is ensured in a unified semantic space. For short-term tracking targets, long-term recovery targets, and candidate targets detected in the unified semantic representation from different perspectives, cross-view candidate target pairing is constructed. For each cross-viewpoint target pair, the multi-dimensional similarity calculation module is invoked for evaluation. Based on the aligned features and coordinate information of the unified semantic layer, a comprehensive similarity score containing multiple dimensions is calculated: appearance similarity compares the visual feature similarity of the targets in multi-view observations; semantic similarity is based on feature similarity in the unified semantic space to ensure semantic consistency of the targets; spatial IoU similarity calculates the overlap between the predicted position and the detected position through intersection-union ratio; and contextual similarity integrates spatiotemporal environmental features to capture the association between behavioral patterns and scene context. Then, the multi-constraint comprehensive scoring module weights and fuses these multi-dimensional similarity measures to generate a comprehensive confidence score and establishes a multi-view association graph, where nodes represent target detection under all views and the weight of the edges is the comprehensive scoring result. Then, the global optimization trajectory stitching module uses graph optimization algorithms (such as maximum weight matching or confidence propagation) to solve the global optimal association, ensuring that the observations of the same vehicle under different perspectives are correctly associated with the same global ID, and spatiotemporally stitching the successfully matched cross-perspective target trajectories to form a global continuous trajectory; Then, the trajectory fusion module generates a unified representation including a global ID, a complete spatiotemporal trajectory sequence, and tracking confidence, providing downstream applications with continuous vehicle tracking results over a wide area.
[0028] like Figure 11 As shown, this embodiment of the invention also provides a vehicle target tracking system based on natural language reference, the system comprising: The receiving preprocessing module 1110 is used to receive and preprocess video stream input from multiple source sensors and natural language instructions from users to obtain standardized image data streams and key semantic information. The semantic reference initialization module 1120 is used to input the standardized image data stream and key semantic information into the semantic reference initialization module, and achieve accurate target localization and initialization through cross-modal feature extraction, deep fusion and semantic consistency verification, and output the target mask; The all-time identity maintenance system 1130 is used to input the target mask into a language enhancement and memory-driven all-time identity maintenance system, integrating short-term identity maintenance and long-term recovery mechanisms to ensure identity consistency and trajectory integrity. The all-time identity maintenance system includes a short-term identity maintenance submodule and a long-term identity recovery submodule. The short-term identity maintenance submodule receives the target mask, realizes inter-frame identity association through multimodal feature fusion and optimized matching, and outputs stable tracking target information and all target information. The long-term identity recovery submodule receives the stable tracking target information output by the short-term identity maintenance submodule, constructs a cross-modal semantic memory based on the idea of contrastive learning, realizes identity recovery after the target disappears and reappears, and outputs long-term recovered target information. The cross-view semantic unification layer 1140 is used to input the raw image streams from multiple sources of sensors, preprocess and semantically align them, provide a unified semantic representation basis for subsequent cross-view target association, and output unified semantic representation and processed multi-view image data. The cross-view semantic unified association module 1150 is used to input the unified semantic representation and processed multi-view image data output from the cross-view semantic unified layer, all target information output from the short-term identity maintenance submodule, and long-term recovered target information output from the long-term identity recovery submodule into the cross-view semantic unified association module. By transforming the coordinate system and standardizing the semantic direction, the differences in semantic mapping under different perspectives are resolved, and the accurate association and trajectory fusion of cross-view targets are achieved, forming continuous tracking in a wide area.
[0029] The vehicle target tracking system based on natural language reference provided in this embodiment of the invention has a functional structure that corresponds to the vehicle target tracking method based on natural language reference provided in this embodiment of the invention, and will not be described again here.
[0030] Figure 12 This is a schematic diagram of the structure of an electronic device 1200 provided in an embodiment of the present invention. The electronic device 1200 may vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 1201 and one or more memories 1202. The memory 1202 stores at least one instruction, which is loaded and executed by the processor 1201 to implement the steps of the above-described vehicle target tracking method based on natural language reference.
[0031] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the aforementioned vehicle target tracking method based on natural language reference. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0032] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0033] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A vehicle target tracking method based on natural language reference, characterized in that, The method includes: S1. Receive and preprocess video stream input from multiple source sensors and natural language commands from the user to obtain standardized image data stream and key semantic information; S2. Input the standardized image data stream and key semantic information into the semantic reference initialization module. Through cross-modal feature extraction, deep fusion and semantic consistency verification, achieve accurate target localization and initialization, and output the target mask. S3. Input the target mask into a language enhancement and memory-driven full-time identity maintenance system, integrating short-term identity maintenance and long-term recovery mechanisms to ensure identity consistency and trajectory integrity. The full-time identity maintenance system includes a short-term identity maintenance submodule and a long-term identity recovery submodule. The short-term identity maintenance submodule receives the target mask, realizes inter-frame identity association through multimodal feature fusion and optimized matching, and outputs stable tracking target information and all target information. The long-term identity recovery submodule receives the stable tracking target information output by the short-term identity maintenance submodule, constructs a cross-modal semantic memory library by drawing on the idea of contrastive learning, realizes identity recovery after the target disappears and reappears, and outputs long-term recovered target information. S4. Input the multi-view raw image stream from multiple source sensors into the cross-view semantic unification layer, preprocess and semantically align it to provide a unified semantic representation basis for subsequent cross-view target association, and output unified semantic representation and processed multi-view image data. S5. The unified semantic representation and processed multi-view image data output from the cross-view semantic unification layer, all target information output from the short-term identity maintenance submodule, and long-term recovery target information output from the long-term identity recovery submodule are input into the cross-view semantic unification association module. By coordinate system transformation and semantic direction standardization, the semantic mapping differences under different perspectives are resolved, and the accurate association and trajectory fusion of cross-view targets are realized, forming continuous tracking in a wide area.
2. The method according to claim 1, characterized in that, The processing procedure of the semantic reference initialization module specifically includes: After the standardized image data stream is input into the visual backbone network, multi-scale visual features are extracted through multi-layer convolution and pooling operations. These correspond to feature representations of different receptive fields: high-resolution features capture fine edge and texture information, medium-resolution features encode local structural features, and low-resolution features represent global semantic information. The key semantic information is first input into a language encoder, where the pre-trained language model BERT converts the text description into a high-dimensional semantic feature vector. Then, a deep semantic feature extraction module uses a multi-layer Transformer encoder and a self-attention mechanism to capture long-distance dependencies between words, generating a more discriminative deep semantic feature representation. ; In the cross-modal attention fusion stage, a parallel dual-branch architecture is adopted to achieve deep fusion of linguistic and visual features: the first branch uses a cross-attention mechanism to establish semantic correspondence, integrating deep semantic features. As the query vector, visual features serve as the key and value vectors. An attention weight matrix is calculated for each scale i. After softmax normalization, the visual features are weighted and summed to generate fused features, allowing the model to dynamically adjust the attention given to different regions of the image based on the text description. The second branch introduces a multi-head attention enhancement mechanism, capturing semantic associations at different levels from multiple independent representation subspaces in parallel to generate enhanced features. Residual connections preserve the original visual feature information, preventing information loss during deep fusion. Then, semantically aligned feature fusion is performed, concatenating and adaptively weighting the outputs of the two parallel branches. Features at different scales are bilinearly upsampled to align spatial dimensions. Channel concatenation and convolutional dimensionality reduction operations generate unified fused multimodal features. This feature integrates visual appearance, semantic attributes, and spatial structure information, and is then used to generate identity features through a feature extraction network. This serves as a benchmark for maintaining one's identity in the future; The fused multimodal features The data is fed into two parallel output head branches for processing: the pixel-level segmentation head uses a cross-attention mechanism to generate the target mask. The Sigmoid activation function is used to output the probability that each pixel belongs to the target, and the mask quality score is also output. Used to evaluate the reliability of segmentation results; the bounding box regression head generates the target bounding box through a multi-head attention enhancement mechanism. The output is normalized center coordinates and width / height parameters; the results from the two outputs are then combined for semantic consistency verification to comprehensively evaluate the initialization quality: first, the CLIP model is used to extract the visual features of the target mask region and the linguistic features of the user's text description, and the cosine similarity between the two is calculated as a semantic similarity index. The algorithm evaluates the degree of matching between the segmentation results and semantic instructions; it also calculates the overlap between the target mask and the bounding box as a spatial IoU metric. The spatial consistency of the results from the two output heads is evaluated, and the mask quality score of the pixel-level segmentation head output is used. This score reflects the connectivity, boundary smoothness, and confidence distribution of the segmentation results. The overall confidence score is calculated through weighted fusion. Then, a threshold judgment is performed: when... Confirm successful initialization and output. , and Establish a baseline identity profile for the target. A preset threshold is set; otherwise, the current result is rejected and the user is asked to re-describe or provide a clearer reference, thereby ensuring high reliability of initialization.
3. The method according to claim 1, characterized in that, The processing procedure of the short-term identity maintenance submodule specifically includes: For the candidate target detected in frame t, the multimodal feature extraction stage is first entered: based on the target mask, three types of complementary features are extracted in parallel to construct a robust target representation, making full use of the synergistic effect of visual, semantic and geometric information. Among them, the depth visual features are extracted by the ReID network in the mask area to capture the visual information of the target and form a high-dimensional visual feature vector; the language semantic features are extracted by the CLIP model to maintain semantic consistency with the user's reference and ensure that the tracked target always conforms to the constraints of the language description; the segmentation mask features are extracted based on the target mask to extract geometric shape and contour information, enhancing the robustness to target deformation and occlusion. The three types of features characterize the target characteristics from different dimensions, laying the foundation for subsequent fusion. Then, the feature fusion enhancement stage is performed: a feature alignment network is used to project the three complementary features into a unified embedding space, eliminating the distribution differences between different feature representations and ensuring the comparability of features of each modality; then, through the attention enhancement mechanism, the contribution of each feature is dynamically calculated using a gated attention module: first, the three aligned features are concatenated and input into a multilayer perceptron, and the interdependencies between features are extracted through nonlinear transformation; then, a sigmoid activation function is used to generate three gate weight values, each weight corresponding to the importance score of a feature class. The gate mechanism can adaptively adjust the contribution of each feature according to the characteristics of the current scene. The fusion representation generates a comprehensive representation vector by weighted summing of each feature, which not only retains the unique advantages of each modality, but also achieves effective information integration and improves the discriminative ability of the target representation; Based on the fused representation, the multi-dimensional similarity calculation module is invoked to perform a matching evaluation between each candidate target in frame t and each tracked target in frame (t-1), including: All candidate targets in frame t are iterated and paired with all tracked targets in frame (t-1). Each target pair is sequentially input into the multidimensional similarity calculation module, which comprehensively considers multiple dimensions: appearance similarity, semantic similarity, and spatial IoU similarity. A comprehensive similarity score is output for each target pair. The scores of all target pairs are organized into an M×N dimensional similarity matrix S according to the index relationship between candidate targets and tracked targets, where M is the number of candidate targets, N is the number of tracked targets, and S(i,j) represents the matching score between the i-th candidate target and the j-th tracked target, providing a quantitative basis for subsequent global optimal matching. Then, the identity matching decision-making stage begins, using the Hungarian algorithm to solve the similarity matrix S, achieving the globally optimal target association and avoiding local suboptimal solutions that might result from greedy strategies. This includes: The similarity matrix S is converted into a cost matrix C, and mapped using the cost conversion formula C(i,j)=1-S(i,j). The Hungarian algorithm is used for global optimal matching. By adjusting the cost matrix C, a unique candidate target is found for each tracked target, and the total matching cost is minimized. Then, the confidence of the matching results is verified: if the similarity of the matching results is higher than the preset threshold, the identity matching is confirmed to be successful, the candidate target is associated with the corresponding tracking target, and the feature representation, bounding box position and motion state of the target are updated; if the similarity of the matching results is lower than the preset threshold, the candidate target is determined to be a newly appearing target in the scene, a new tracking ID is assigned to it and the tracking state is initialized. After the identity matching decision is completed, the target information is distributed in two ways according to the target status: the stable tracking target information is output to the long-term identity recovery module for identity preservation and recovery; all target information is output to the cross-view semantic unified association module to support multi-view collaborative perception.
4. The method according to claim 3, characterized in that, The multidimensional similarity calculation module is applied in the short-term identity maintenance submodule, the long-term identity recovery submodule, and the cross-perspective semantic unified association module. It uses "target pairs" as input units, with each target pair including a complete feature set of the candidate target and historical targets. Through a structured parallel computation process, it ultimately outputs a comprehensive similarity score. The processing includes: After receiving the target pair, the two sets of input features are preprocessed and aligned to ensure that all features are in a unified representation space, laying the foundation for subsequent parallel comparison. Then, the parallel multidimensional computation stage begins, simultaneously generating similarity scores for multiple core dimensions: appearance similarity is based on the cosine similarity of visual features to measure appearance consistency; semantic similarity utilizes the Gaussian kernel similarity of linguistic features to evaluate the semantic matching degree with the initial description; spatial IoU similarity calculates the overlap between the predicted and detected positions using the intersection-union ratio (IoU); and contextual similarity integrates spatiotemporal environmental features to capture the association between behavioral patterns and scene context. Different sub-modules calculate all or part of these four similarities: the short-term identity maintenance sub-module calculates the first three similarities; and the long-term identity recovery sub-module and the cross-perspective semantic unified association module calculate all four similarities. Then, these independently calculated scalar scores are combined into a unified similarity vector, which is then fed into a normalization layer for processing to eliminate the dimensional differences between the dimensions and ensure that they are comparable on the same scale. For the normalized similarity vector, two strategies are provided for fusion: one is a fixed weight strategy, which predefines a set of weight coefficients according to the specific task and obtains a comprehensive score by weighted summation; the other is a more flexible adaptive weight strategy, which introduces a lightweight attention network. This network takes the current similarity vector as input, dynamically learns and outputs the optimal weights for each dimension, and realizes intelligent adjustment of the evaluation focus according to the specific matching context. Through the above process, a comprehensive similarity score is output for each input target pair. When processing multiple target pairs, these scores are organized into a similarity matrix for subsequent matching decisions.
5. The method according to claim 1, characterized in that, The processing procedure of the long-term identity recovery submodule specifically includes: A cross-modal semantic memory is constructed to establish a complete identity profile for the target. This memory adopts a multi-level feature storage architecture and maintains four types of core features simultaneously for each tracked target: visual features, language description features, spatial features, and contextual environment features. This multi-modal feature fusion strategy ensures the robustness of identity representation and can recognize identities from multiple complementary dimensions, effectively dealing with the failure of a single feature dimension. To maintain the timeliness and stability of the memory features, an exponential moving average (EMA) update mechanism is introduced to dynamically maintain the stored features. After the short-term identity maintenance submodule successfully matches the target, four core features of the target in the current frame are extracted: visual features, linguistic description features, spatial features, and contextual features, as new features. For each type of feature, the exponential moving average (EMA) update mechanism is applied to update the new features of the current frame with the historical features of that type of feature stored in the memory. For the fusion update, the momentum parameter α is used for weighted updating of the k-th feature, and the update formula is as follows: When a new candidate target is detected, it is paired one by one with all historical targets in the memory that have been lost, constructing multiple candidate target-historical target pairs. Using the continuously updated features of the memory, each target pair is input into the multidimensional similarity calculation module for parallel evaluation. Independent scores for appearance similarity, semantic similarity, spatial IoU similarity, and contextual similarity are calculated simultaneously. Then, through comprehensive similarity weighted calculation, an adaptive weight strategy is used to fuse the scores of each dimension into a unified comprehensive confidence score. The candidate matching pair with the highest comprehensive confidence score is selected for fusion judgment. The reliability of the matching is evaluated by combining the distribution characteristics of the scores of each dimension. Then, a rigorous spatiotemporal consistency verification is performed, including checking whether the target disappearance time is within a preset reasonable range and whether the reappearance location conforms to the motion model prediction based on the historical trajectory. This verification process ensures the rationality of the recovery results in the temporal and spatial dimensions. For all verified recovered targets, their long-term recovered target information, including the recovered tracking ID, current position, updated feature vector, and confidence score, is output to the cross-view semantic unified association module to participate in multi-view trajectory fusion. At the same time, the state information of the corresponding target in the memory bank is updated and synchronized to the short-term identity maintenance submodule to ensure continuous tracking in subsequent frames. For candidate targets that fail verification, a new tracking ID will be assigned to them and the corresponding memory bank entries will be initialized to establish a brand-new identity profile.
6. The method according to claim 1, characterized in that, The processing procedure of the cross-perspective semantic unification layer specifically includes: The multi-view raw image stream first enters the preprocessing module to perform image enhancement and distortion correction operations to ensure that images from different sensors have a consistent quality benchmark; Then, based on the intrinsic and extrinsic parameters obtained from the calibration of each sensor, a unified coordinate system transformation is performed, and the obtained intrinsic parameter matrix is used. and external references Establish a world coordinate system as a unified reference for sensors. The transformation relationship between image coordinates and world coordinates is as follows: This transformation maps image observations from different perspectives to a unified world coordinate system, laying a geometric foundation for subsequent multi-perspective semantic understanding; After completing coordinate system one, the cross-view semantic alignment network learns the mapping relationship between visual features observed from different viewpoints and a unified semantic space, ensuring that the same target captured by multiple sensors maintains consistent semantic expression. The network adopts a multi-head attention mechanism to process visual features from various viewpoints simultaneously and capture common semantic representations between different viewpoints. For directional descriptions in user semantic commands, a standardized direction vector is defined in a unified coordinate system. This requires learning the mapping relationship from textual semantic descriptions to standardized direction vectors in the world coordinate system. Therefore, a contrastive learning approach is used to train a semantic alignment network, enabling it to map semantically identical directional descriptions from different perspectives to the same standardized direction vector space for user text commands. and the corresponding true normalized direction vector A contrastive learning loss function is constructed to make the mapping result of the text description close to the true direction and far away from other incorrect directions: in Represents text instructions After the mapping results of the semantic alignment network, To achieve true standardization, For other candidate directions, Represents cosine similarity. The temperature parameter ensures the consistency of spatial mapping of semantic instructions under different perspectives, so that the directional description can be accurately understood as a unified directional meaning in the perspectives of fixed surveillance cameras, vehicle cameras and drones. Then, the adaptive weight allocation module dynamically calculates the fusion weights based on the observation quality of each sensor, and comprehensively considers factors such as image clarity, target visibility, and viewpoint suitability to perform weighted fusion of semantic features from different viewpoints. The final output is a weighted fusion of unified semantic representation and processed multi-view image data, providing a consistent reference benchmark for subsequent cross-view association.
7. The method according to claim 1, characterized in that, The processing procedure of the cross-perspective semantic unified association module specifically includes: By utilizing the output of the cross-view semantic unification layer, the comparability of target observations from different perspectives is ensured in a unified semantic space. For short-term tracking targets, long-term recovery targets, and candidate targets detected in the unified semantic representation from different perspectives, cross-view candidate target pairing is constructed. For each cross-viewpoint target pair, the multi-dimensional similarity calculation module is invoked for evaluation. Based on the aligned features and coordinate information of the unified semantic layer, a comprehensive similarity score containing multiple dimensions is calculated: appearance similarity compares the visual feature similarity of the targets in multi-view observations; semantic similarity is based on feature similarity in the unified semantic space to ensure semantic consistency of the targets; spatial IoU similarity calculates the overlap between the predicted position and the detected position through intersection-union ratio; and contextual similarity integrates spatiotemporal environmental features to capture the association between behavioral patterns and scene context. Then, the multi-constraint comprehensive scoring module weights and fuses these multi-dimensional similarity measures to generate a comprehensive confidence score and establishes a multi-view association graph, where nodes represent target detection under all views and the weight of the edges is the comprehensive scoring result. Then, the global optimization trajectory stitching module uses a graph optimization algorithm to solve the global optimal association, ensuring that the observations of the same vehicle under different perspectives are correctly associated with the same global ID, and spatiotemporally stitches the successfully matched cross-perspective target trajectories to form a global continuous trajectory. Then, the trajectory fusion module generates a unified representation including a global ID, a complete spatiotemporal trajectory sequence, and tracking confidence, providing downstream applications with continuous vehicle tracking results over a wide area.
8. A vehicle target tracking system based on natural language reference, characterized in that, The system includes: The receiving preprocessing module is used to receive and preprocess video stream input from multiple source sensors and natural language instructions from users to obtain standardized image data streams and key semantic information. The semantic reference initialization module is used to input the standardized image data stream and key semantic information into the semantic reference initialization module. Through cross-modal feature extraction, deep fusion and semantic consistency verification, it achieves accurate target localization and initialization and outputs the target mask. A full-time identity maintenance system is used to input the target mask into a language enhancement and memory-driven full-time identity maintenance system, integrating short-term identity maintenance and long-term recovery mechanisms to ensure identity consistency and trajectory integrity. The full-time identity maintenance system includes a short-term identity maintenance submodule and a long-term identity recovery submodule. The short-term identity maintenance submodule receives the target mask, realizes inter-frame identity association through multimodal feature fusion and optimized matching, and outputs stable tracking target information and all target information. The long-term identity recovery submodule receives the stable tracking target information output by the short-term identity maintenance submodule, constructs a cross-modal semantic memory based on the idea of contrastive learning, realizes identity recovery after the target disappears and reappears, and outputs long-term recovered target information. The cross-view semantic unification layer is used to input the raw image streams from multiple sources of sensors, preprocess and semantically align them, provide a unified semantic representation basis for subsequent cross-view target association, and output unified semantic representation and processed multi-view image data. The cross-view semantic unified association module is used to input the unified semantic representation and processed multi-view image data output from the cross-view semantic unified layer, all target information output from the short-term identity maintenance submodule, and long-term recovered target information output from the long-term identity recovery submodule into the cross-view semantic unified association module. By transforming the coordinate system and standardizing the semantic direction, the differences in semantic mapping under different perspectives are resolved, and the accurate association and trajectory fusion of cross-view targets are achieved, forming continuous tracking in a wide area.
9. An electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, characterized in that, The at least one instruction is loaded and executed by the processor to implement the vehicle target tracking method based on natural language reference as described in any one of claims 1-7.
10. A computer-readable storage medium storing at least one instruction, characterized in that, The at least one instruction is loaded and executed by the processor to implement the vehicle target tracking method based on natural language reference as described in any one of claims 1-7.
Citation Information
Cited By
Multi-person posture fusion and interaction display system based on multi-path image data
CN122049145A