Multi-target tracking method based on dynamic semantic focus migration
By constructing a semantic memory pool and a temporal transfer matrix, and combining deformable convolution and graph neural networks, the visual feature alignment and tracking strategies are dynamically adjusted, solving the problems of semantic drift and fusion instability in existing methods, and achieving robust target tracking in complex scenes.
Patent Information
- Application Number
- CN202511329276.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-01-02
AI Technical Summary
Existing target tracking methods struggle to effectively handle high-level semantic instructions in complex scenarios, lack modeling of dynamic changes in semantic concerns, leading to semantic drift, unstable fusion, and misjudgment of tracking paths, especially under conditions of visual occlusion, appearance changes, or semantic drift.
By constructing a semantic memory pool and a temporal transfer matrix, deformable convolution is used to achieve nonlinear spatial alignment of visual features. A spatiotemporal attention heatmap is generated by combining the semantic focus transfer rule matrix. A semantic focus consistency index is introduced to evaluate the fusion stability. Target association is optimized through graph neural networks, and the tracking strategy is dynamically adjusted to improve system stability.
It enhances the semantic interpretability and tracking foresight of the multi-target tracking system, improves the understanding and spatial alignment of natural language targets, ensures robust target pointing ability in complex scenarios, and achieves accurate matching and identity preservation among multiple targets.
Smart Images

Figure CN121259040A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent vehicle target tracking technology, and relates to a multi-target tracking method that integrates natural language semantics and visual perception information. Background Technology
[0002] With the development of intelligent driving and human-computer interaction technologies, how to achieve flexible control of perception systems through natural language descriptions has become an important direction for the development of multi-target tracking systems. Most existing target tracking methods rely on visual information for physical feature matching, which makes it difficult to effectively process the high-order semantic instructions contained in linguistic information. In particular, in complex scenarios, problems such as semantic ambiguity, attention drift, and target ambiguity exist.
[0003] Existing research has attempted to introduce visual-language fusion mechanisms to improve the accuracy and interactivity of target selection. However, most methods remain at the static fusion level, lacking the ability to model the dynamic migration of semantic attention as it evolves with language expression, and failing to capture the temporal migration trend of user command focus. Furthermore, current fusion mechanisms also have significant shortcomings in semantic consistency assessment, dynamic change adaptation, and feature compensation, failing to address representation instability and tracking degradation caused by visual occlusion, appearance changes, or semantic drift. In complex multi-target scenarios, there is a lack of structured modeling and strategy selection mechanisms for semantic focus scheduling, severely limiting their stability, interpretability, and practical application effectiveness. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a multi-target tracking method based on dynamic semantic focus migration, so as to solve the problems of lack of effective modeling and utilization of dynamic changes in semantic focus in semantic modeling, cross-modal fusion and tracking decision-making, which leads to semantic drift, unstable fusion and misjudgment of tracking path.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A multi-target tracking method based on dynamic semantic focus shift, the method comprising:
[0007] Image sequences and user natural language commands are collected separately, and visual and semantic features are extracted separately. The semantic features are stored in a semantic memory pool, and candidate bounding boxes are generated using YOLOX. Unsupervised semantic clustering is performed on the semantic memory pool to extract semantic focus prototype vectors, and a temporal transfer matrix is constructed. A semantic focus transfer pattern matrix is constructed based on the semantic focus prototype vectors and the temporal transfer matrix.
[0008] Nonlinear spatial alignment of visual features is achieved through deformable convolution under semantic guidance, and a spatiotemporal attention heatmap that integrates spatial distribution and temporal trend is generated by combining the semantic focus migration law matrix; a semantic focus consistency index is introduced to evaluate the fusion stability. When the index exceeds the set threshold, spatial relocation and pose compensation based on depth estimation and SO(3) Lie group interpolation are triggered to complete the reconstruction of semantic features and output multimodal fusion features that combine vision and semantics.
[0009] Based on the semantic focus migration pattern matrix and multimodal fusion features, the position of the candidate target in the next frame is predicted. Then, according to the prediction result of the target position and the heat map response, the semantic-visual template of the target is dynamically updated. A semantic graph structure with candidate targets as nodes and focus associations as edges is constructed, and the association between targets is dynamically updated using a graph neural network. The node features are the updated semantic-visual templates. The candidate state is evaluated by semantic perturbation and robustness scoring, and the tracking strategy is dynamically adjusted to output stable and consistent multi-target tracking results.
[0010] Furthermore, the acquired semantic features are stored in a semantic memory pool, which is then populated with historical semantic features. Unsupervised semantic clustering is performed on the semantic memory pool using spectral clustering to obtain K semantic focus prototype vectors Q. The average activation level of each semantic focus category in historical natural language instructions is calculated to obtain the average response intensity vector. Used to measure the importance of different categories;
[0011] The frequency of semantic focus category shifts between two consecutive time steps is statistically analyzed, and a time transition matrix T is constructed.
[0012] By fusing the average response intensity vector and the temporal transfer matrix, a semantic focus transfer pattern matrix M is constructed to quantify the joint transfer trend between different semantic focus categories in the temporal and spatial dimensions, as shown in the following equation:
[0013]
[0014] Where, diag(·) represents mapping a vector to a diagonal matrix; λ∈[0,1] represents the balance factor for the fusion of space and time.
[0015] Furthermore, nonlinear spatial alignment of visual features is achieved through deformable convolution under semantic guidance, thereby aligning semantic features F... l Mapping to visual feature F v With the same spatial resolution, a spatialized language guidance map L is obtained; for each candidate target in the current frame t... Given the number of candidate targets in the current frame t, and using L as a guide, analyze the feature vectors corresponding to the candidate targets. Perform deformable convolution operations to achieve non-linear spatial alignment of visual features under semantic guidance. After alignment, the feature vectors of candidate targets are written back to the global visual feature map through remapping to obtain the visual alignment feature map F. a .
[0016] Furthermore, a spatiotemporal attention heatmap is generated by combining the semantic focus transfer pattern matrix with spatial distribution and temporal trend, and the average response intensity vector is then analyzed based on the semantic focus transfer pattern matrix M. Perform a weighted update to obtain the temporally enhanced attention distribution vector. Utilizing candidate targets The corresponding semantic category, from the semantic focus prototype set Extract the corresponding prototype vector from the data and combine it with... The weights of the visual alignment feature map F a Weighted similarity calculations are performed on candidate regions to generate a spatiotemporal attention heatmap A that highlights the candidate target regions most relevant to semantics. f ;
[0017] Spatiotemporal attention heatmap A f The attention intensity at each pixel location is a weight, and the visual alignment feature map F is... a The corresponding positional feature vectors are weighted to construct a semantic diffusion feature map F. d Then calculate the semantic diffusion feature map F. d Visually aligned feature map F a The pixel-level differences between them are used to define the Semantic Focus Consistency Index (SPI) as follows:
[0018]
[0019] Where H and W are the height and width of the spatial dimension, respectively, and i and j represent the pixel coordinate indices of the image in the spatial dimension. The channel vector representing the spatial position (i,j) in the visual alignment feature map. Let ||·||2 represent the channel vector at spatial location (i,j) in the semantic diffusion feature map, and ||·||2 represent the Euclidean distance.
[0020] When the SPI exceeds the set threshold, spatial relocation and pose compensation based on depth estimation and SO(3) Lie group interpolation are triggered to complete the reconstruction of semantic features and output multimodal fusion features that combine vision and semantics.
[0021] Furthermore, when the SPI exceeds a set threshold, a depth map D is generated from the input image I using the depth estimation network DPT; then, it is combined with the camera intrinsic parameter matrix K. c Projecting pixel coordinates (x, y) and their depth values into pseudo-point cloud coordinates p in 3D space.(x,y) From the spatiotemporal attention heatmap A f A point set is formed by selecting high-response regions. And obtain the target point set from the historical semantic state. Establish a source-target pairing relationship, where N is the number of points; then, use singular value decomposition to perform rigid body transformation estimation on the two point sets, and solve for the rotation matrix R∈SO(3) and the translation vector ι to make the transformed source point p n Align the corresponding target point p′ n This minimizes the sum of squared registration errors.
[0022] In the Lie group space SO(3), spherical linear interpolation is performed on the rotation matrix R to generate an intermediate interpolation rotation matrix R. q The transformed 3D point cloud is reprojected into the image space, and the feature map F is visually aligned. a Mid-sampling yields the visual feature map F after semantic-guided relocalization. r :
[0023]
[0024] Here, Reproject(·) represents projecting 3D points back to the 2D image coordinate system using the camera model, sampling visual features from the visually aligned feature map at the corresponding positions, and generating a new feature map F. r .
[0025] Furthermore, the output of multimodal fusion features combining visual and semantic elements includes: if the SPI value exceeds a set threshold, obtaining the visual feature map F through spatial relocalization and pose compensation processing. r If the SPI value does not exceed the set threshold, the visual feature map F extracted from the input image I will be used. v As visual feature map F r ;
[0026] The current semantic feature F l Through the linear mapping matrix W l Projected onto the visual channel space, a multi-head cross-attention mechanism is introduced, using semantic features F l For the query vector, the visual feature map F r Channel-selective enhancement is performed on the key and value vectors to generate a fused multimodal feature map:
[0027] F f =CrossAttn(F r W l F l )
[0028] Where CrossAttn(·) represents the multi-head cross-attention mechanism; Ff This represents a multimodal fusion feature map.
[0029] Furthermore, based on the semantic focus migration pattern matrix and multimodal fusion features, the position of the candidate target in the next frame is predicted, and based on the multimodal fusion feature map F... f The feature vector of the candidate target in the current frame With semantic focus prototype vector set Calculate the similarity to obtain the semantic relevance vector 'a' for each target. κ Combining the semantic focus migration pattern matrix M, a location prediction module is constructed, using the semantic relevance vector a. κ For semantic prior, the potential regions of interest in the next frame are modeled. This location prediction module uses a Transformer-based spatiotemporal feature encoder to fuse the multimodal feature map F of the current frame. f By jointly encoding with the semantic focus migration pattern matrix M, the spatial position of each candidate target in the next frame is directly output. With confidence distribution
[0030] Based on the predicted location and the multimodal fusion feature map of the current frame, combined with the spatiotemporal attention heatmap A f High-confidence, strong semantic response regions of each candidate target are selected as new semantic-visual templates. Semantic-visual template from the previous frame Weighted fusion is performed, with the weights determined by the semantic focus response intensity of high-confidence, strong semantic response regions and the semantic focus consistency index (SPI).
[0031] Furthermore, a semantic graph structure is constructed with candidate targets as nodes and focus associations as edges, where each node has a semantic focus category. The set of semantic-visual templates and semantic focus prototype vectors of candidate target κ Calculate the cosine similarity, and then determine the category by selecting the semantic focus prototype vector corresponding to the one with the largest cosine similarity; the edge weights between any two nodes u and o in the graph structure. Simultaneously considering semantic transfer strength and visual appearance similarity:
[0032]
[0033] Where α and β are adjustable fusion coefficients, α+β=1, and M is the semantic focus transfer pattern matrix. Indicates the semantic focus category of frame t. arrive The prior transfer strength is given by cos(·,·), which represents the cosine similarity operator. and Let u and o represent the semantic-visual templates of nodes u and o in frame t, respectively. Cosine similarity is used to measure the consistency of the targets corresponding to the two nodes in the semantic and appearance feature space.
[0034] The constructed graph structure is input into a graph neural network. Through multi-layer semantic information propagation and node state updates, it optimizes multi-objective data association and matching results. During the propagation process of the graph neural network, the features of each node are initialized as follows: Then use edge weights The message passing between adjacent nodes is controlled by the following update rules:
[0035]
[0036] in, This represents the feature vector of node u in the l-th layer; Represents the set of neighbors of node u; Indicates the set of neighbors Normalize the weights of all neighbors to ensure that the weights of all neighbors are positive and their sum is 1; W (l) Let be the learnable weight matrix of the l-th layer, projecting the neighbor features into the new feature space; σ(·) is the non-linear activation function (ReLU); after propagation through L layers, the updated node representation is obtained.
[0037] After completing the feature update for the current frame, due to the updated semantic-visual template It integrates the target's own features with neighborhood context information, possessing both discriminative power and temporal consistency; therefore, it is used as a trajectory feature set, along with the trajectory feature set of the previous frame. Similarity calculations are performed to construct a cross-frame correlation matrix. Then, using a Hungarian algorithm matching strategy, each candidate target in the current frame is assigned a corresponding trajectory from the previous frame, achieving cross-frame correlation of targets and updating tracking results. The resulting matching results constitute the set of tracked targets for the current frame.
[0038] Furthermore, a set of perturbation simulation mechanisms is introduced during the candidate target tracking process. By setting occlusion areas and focus shift perturbations, candidate observation states of candidate targets are constructed; combined with the spatiotemporal attention heatmap A of the current frame. f With fusion feature F f For each candidate target state, a confidence assessment is performed: from A f The average response value of the candidate target region is obtained to measure the semantic attention intensity of the candidate target. Then, the semantic-visual templates of the candidate target in the current frame and in multiple historical frames are compared to obtain its feature consistency score, which is used to measure the stability of the candidate target in the time dimension.
[0039] Based on this, a robustness score is obtained by weighted fusion of semantic attention intensity and feature consistency. The higher the score, the higher the credibility of the target under occlusion, appearance drift, or semantic interference. The system further executes strategy scheduling based on the robustness score results. The tracking strategy adapts to different robustness scores by dynamically adjusting the template update frequency and feature weights. When the score is high, the update frequency is increased and the adaptability to appearance changes is enhanced. When the score is low, the update frequency is reduced and more historical information is retained to avoid drift. The scheduling strategy dynamically allocates computing resources according to the target task priority and environmental complexity. In high-priority or complex scenarios, spatial relocation and high-precision matching modules are called to ensure stability. In low-priority or simple scenarios, lightweight tracking paths are used to improve real-time performance.
[0040] Finally, based on the results of robustness verification and policy scheduling, the optimal target state is selected and stable and consistent multi-target tracking results are output.
[0041] The beneficial effects of this invention are as follows:
[0042] 1) This invention proposes a multi-target tracking method based on dynamic semantic focus migration. By constructing a hierarchical semantic memory pool and introducing a semantic evolution prediction mechanism, it effectively models the migration law of semantic focus in the time and space dimensions driven by language input, enhances the guiding role of semantic information in multi-target tracking, and improves the semantic interpretability and tracking foresight of the system.
[0043] 2) Compared with existing target detection or tracking methods based solely on images or language, this invention designs a dynamic guidance mechanism that integrates visual attention and language expression, enabling the system to maintain robust target pointing capabilities even when facing complex scenarios such as semantic ambiguity, occlusion, and blurriness, significantly improving the understanding and spatial alignment capabilities of natural language targets.
[0044] 3) In the data association stage, the semantic focus migration pattern matrix is combined with the visual similarity between templates to construct a multi-target association graph and use graph neural networks (GNN) to perform multi-layer information propagation and node state updates to achieve accurate matching and identity preservation among multiple targets; at the same time, the candidate states are verified and the strategy is scheduled by combining semantic perturbation simulation and robust scoring mechanism to ensure the continuity and stability of the association results.
[0045] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0047] Figure 1 This is a schematic diagram of a multi-target tracking method based on dynamic semantic focus migration provided in an embodiment of the present invention;
[0048] Figure 2 Diagram showing the installation location of the image acquisition equipment; Detailed Implementation
[0049] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0050] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0051] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0052] This invention provides a multi-target tracking method based on dynamic semantic focus migration. This method integrates image and language information. First, visual and language features are extracted through a dual-path coding structure. Candidate target boxes are obtained using YOLOX, and the language features are stored in a historical semantic memory pool. Combining the clustering results of multi-frame semantic states, spatial response weights, and temporal migration trends, a semantic evolution prediction function is proposed to construct a semantic focus migration law matrix, quantifying the joint evolution mode of targets in the spatiotemporal dimensions and providing accurate semantic priors. Second, a spatial alignment and dynamic guidance fusion network (DGSF-Net) is constructed based on this matrix to generate a spatiotemporal attention heatmap, guiding the nonlinear alignment of visual and semantic features, and evaluating the fusion stability with the semantic focus consistency index (SPI). When the SPI exceeds the threshold, spatial relocation and pose compensation based on depth estimation and SO(3) interpolation are triggered to maintain the robustness of features in complex scenes. Finally, a multi-target association graph is constructed by combining the fusion features and the regularity matrix with visual similarity. Graph neural networks (GNNs) are used to optimize data association and identity preservation. Semantic perturbation and robust scoring mechanisms are used for policy scheduling to dynamically adjust the update frequency and computing resources, ensuring that the output of semantically consistent, trajectory-continuous and stable tracking results is achieved under long-term and multi-interference conditions.
[0053] like Figure 1 The image shows a multi-target tracking method provided by an embodiment of the present invention. The method is as follows:
[0054] I. Image sequences are acquired via an in-vehicle camera, and user commands are obtained by combining them with natural language input. A dual-channel encoding architecture is designed to extract visual and semantic features separately. Then, the YOLOX object detector is used to obtain candidate bounding boxes, and a historical semantic state memory pool is constructed to store and manage historical semantic data of the targets. Combined with the current language input, the dynamic migration trend of the user's attention semantic target (i.e., semantic focus) in the spatial and temporal dimensions is modeled based on the historical semantic sequence.
[0055] Specifically, by proposing a semantic evolution prediction function, the distribution and change relationship of semantic focus at different times are modeled, and the spatial distribution weight and temporal migration intensity are integrated to generate a semantic focus migration law matrix for target tracking guidance.
[0056] 101: Place the vehicle-mounted camera on the top front of the intelligent connected vehicle to capture images of the front of the vehicle, and simultaneously input manual commands into the system.
[0057] 102: For image data First, basic visual features are extracted using ResNet50, and then combined with a Feature Pyramid Network (FPN) to obtain global and local multi-scale fused semantic features, which are uniformly represented as... Where C, H, and W represent the number of channels, height, and width, respectively. Then, the YOLOX object detector is used to generate a set of target candidates for the current frame t in the image space. and its corresponding instance feature vector Used for subsequent feature alignment and fusion. d represents the number of candidates for this frame. v Dimensions representing visual features.
[0058] 103: For each user language command input Where t represents the t-th frame, its word-level semantic features are extracted using the RoBERTa language model. The semantic features of the current language instructions are obtained through average pooling. After encoding, the semantic feature is stored in the semantic memory pool. (Initialized to empty), and then the semantic memory pool is filled with historical semantic features. This is used for subsequent clustering and temporal modeling of the semantic representation. Here, L is the length of the input instruction, d is the dimension of each word vector, and i represents the i-th instruction.
[0059] 104: A semantic memory pool storing multiple historical semantic vectors As input, model its evolution over time.
[0060] First, unsupervised semantic clustering is performed using spectral clustering to obtain K semantic focus categories, and a corresponding semantic focus prototype vector is generated for each category. Here, the semantic focus category represents the different semantic clusters obtained by clustering, and the semantic focus prototype vector is the centralized representation of the category in the feature space; the two correspond one-to-one. Based on this, the average activation level of each semantic focus category in historical instructions is further calculated: specifically, for each historical instruction, its response value in each semantic focus category is calculated, and the response values of the same category are accumulated and averaged over time to obtain the average response intensity vector. This is used to measure the importance of different categories. Simultaneously, the frequency of semantic focus category shifts between two consecutive historical instructions is statistically analyzed to construct a temporal transition matrix. This approach aims to characterize the evolution of semantic focus over time. Finally, the average response intensity vector is fused with the time transition matrix to construct a semantic focus transition pattern matrix. The following formula is used to quantify the joint transfer trend between different semantic focus categories in the temporal and spatial dimensions:
[0061]
[0062] Where, diag(·) represents mapping the vector to a diagonal matrix so that it can be fused with T; λ∈[0,1] represents the balance factor for the fusion of space and time.
[0063] II. Based on the constructed semantic focus migration rule matrix M, a spatial alignment and dynamic guidance fusion network (DGSF-Net) is constructed.
[0064] First, nonlinear spatial alignment of visual feature maps guided by semantic vectors is achieved through deformable convolution DCNv3. Then, a spatiotemporal attention heatmap guiding target tracking is generated using matrix M. Next, a semantic focus consistency index (SPI) is designed to measure the alignment consistency between visual and linguistic features. When SPI > 0.3, a progressive reconstruction process is triggered. Pose compensation and dynamic relocalization of target features are performed through depth estimation and rotation interpolation in the SO(3) Lie group space, thereby dynamically adjusting the feature representation of the target. Finally, a dynamic fusion representation guided by semantic focus migration is output.
[0065] Specifically:
[0066] 201: First, the semantic features of the encoded language Mapping to visual feature maps With the same spatial resolution, a spatialized language guidance graph L is obtained. For each candidate target instance... Using L as a guide, its feature vector Deformable convolution (DCNv3) is performed to achieve non-linear spatial alignment of visual features under semantic guidance. After alignment, the feature vectors of candidate targets are written back to the global visual feature map through remapping to obtain the visually aligned feature map.
[0067] 202: Based on the semantic focus transfer pattern matrix M, combined with the semantic focus prototype set With the average response intensity vector For each candidate target Construct a semantically guided spatiotemporal attention heatmap. First, based on the semantic focus transfer pattern matrix M, the average response intensity vector... Perform a weighted update to obtain the temporally enhanced attention distribution vector. Subsequently, using candidate targets The corresponding semantic category, from the semantic focus prototype set Extract the corresponding semantic focus prototype vector and combine it with... The weights of the visual alignment feature map F a Weighted similarity calculations are performed on candidate regions within the dataset. The resulting spatiotemporal attention heatmap... It can highlight the candidate target regions that are most relevant to the semantics.
[0068] 203: Spatiotemporal attention heatmap A f The attention intensity at each pixel location is a weight, and the visual alignment feature map F is... a The corresponding positional feature vectors are weighted to construct a diffusion-weighted feature map. Then, the semantic diffusion feature map F is calculated. d Alignment with the original image F a The pixel-level differences between them are defined by the Semantic Focus Consistency Index (SPI) as follows:
[0069]
[0070] Here, SPI stands for Semantic Focus Consistency Index, which measures the pixel-level difference between the semantically diffused feature map and the original semantically aligned feature map. A smaller value indicates stronger semantic stability. When SPI > 0.3 (an empirical threshold, set to 0.3 here), the reconstruction process is triggered. H and W represent the height and width of the spatial dimension, respectively, and i and j represent the pixel coordinate indices of the image in the spatial dimension. Let i and j represent the channel vectors at spatial positions (i,j), and ||·||2 represent the Euclidean distance.
[0071] 204: If the SPI value exceeds a set threshold, spatial relocation will be triggered. Specifically, the depth estimation network (DPT) is first used to re-evaluate the image. Generate depth map Then combine with the camera intrinsic parameter matrix Project the pixel coordinates (x, y) and their depth values into pseudo-point cloud coordinates in 3D space. From the spatiotemporal attention heatmap A f A point set is formed by selecting high-response regions. And obtain the target point set from the historical semantic state. Establish source-target pairing relationship; then use singular value decomposition (SVD) to perform rigid body transformation estimation on the two point sets, by solving the rotation matrix R∈SO(3) and translation vector. Make the transformed source point p n Align it with its corresponding target point p′ as much as possible n This minimizes the sum of squared registration errors. Simultaneously, to improve temporal continuity and spatial smoothness, spherical linear interpolation (SLERP) is performed on the rotation matrix in the Lie group space SO(3) to generate an intermediate interpolation rotation matrix R. qHere, q represents the q-th intermediate rotation matrix generated during the interpolation process. Because spherical linear interpolation generates multiple intermediate rotation matrices during the transition from the initial state to the target state of the rotation matrix, it enhances temporal continuity and spatial smoothness. q is used to distinguish between different intermediate rotation matrices. Finally, the transformed 3D point cloud is reprojected into the image space, and the feature map F is visually aligned. a Mid-sampling yields the visual feature map after semantic-guided relocalization. As shown in equation (3):
[0072]
[0073] Reproject(·) means using the camera model to project 3D points back to the 2D image coordinate system, and sampling the visual features in the original feature map at the corresponding positions to generate a new feature map.
[0074] 205: Obtaining the semantically guided visual feature map F r Subsequently, regardless of whether spatial relocation steps are performed, it is necessary to deeply integrate it with the semantics of the current frame in order to construct a multimodal joint representation with spatial consistency and semantic interpretability.
[0075] Specifically, visual feature maps As the main input, if the SPI exceeds a set threshold, it indicates that the current visual features and semantics have a low matching degree, and there are problems such as spatial misalignment between semantics and vision. At this time, the relocalization step can effectively correct the spatial position of the visual features, so that the visual features are better aligned with the region of interest of the semantics. Therefore, this feature map is the relocalization result F. r If the threshold is not exceeded, it indicates that the current visual features and semantics already have good consistency, and there is no need for complex spatial relocalization operations. The semantic alignment feature F can be used directly. a At the same time, the current language semantic representation will be... Through linear mapping matrix The image is projected onto the visual channel space. Then, a multi-head cross-attention mechanism is introduced, using semantic representation as the query and visual feature maps as the key and value, to perform channel-selective enhancement, generating a fused multimodal feature map, as shown in Equation (4):
[0076]
[0077] Among them, W l This indicates that the semantic representation F l The linear projection matrix mapped to the same spatial dimension as the visual channel is a learnable linear mapping matrix obtained through training; CrossAttn(·) represents the multi-head cross attention mechanism; This is represented as the final semantically guided multimodal fusion feature.
[0078] III. Based on the semantic focus transfer pattern matrix M, and combined with multimodal fusion features F f The system first predicts the target's position in the next frame. Then, based on the prediction and heatmap response, it dynamically updates the target's semantic-visual template (i.e., a cross-modal feature representation that integrates target appearance features and semantic focus information, used for identity preservation and data association in multi-frame tracking). Next, it constructs a semantic graph structure with candidate targets as nodes and focus associations as edges, and uses a graph neural network (GNN) to dynamically update the associations between targets. Finally, through robustness evaluation and focus intensity-guided policy scheduling, it outputs stable and consistent multi-target tracking results.
[0079] Specifically:
[0080] 301: Based on multimodal fusion feature map F f The feature vector of the candidate target in the current frame With semantic focus prototype vector set Calculate the similarity to obtain the semantic relevance vector of each candidate target κ. Combining the semantic focus migration pattern matrix M, a location prediction module is constructed, using the semantic relevance vector a. κ For semantic priors, potential regions of interest in the next frame are modeled. This module employs a Transformer-based spatiotemporal feature encoder to fuse the multimodal features F of the current frame. f By jointly encoding with the semantic focus migration pattern matrix M, the spatial position of each candidate target in the next frame is directly output. With confidence distribution This allows the prediction to remain stable and accurate even in complex scenarios such as occlusion and changes in appearance.
[0081] 302: After obtaining the predicted target location, a multi-target semantic-visual template dynamic update mechanism driven by attention heatmaps is constructed. Specifically, based on the predicted location and the multimodal fusion features of the current frame, combined with the spatiotemporal attention heatmap A... f High-confidence, strong semantic response regions of each target κ are selected as new templates. With the previous frame template (Initializing the first frame as empty) weighted fusion is performed, with weights determined by the average response intensity. The semantic focus consistency index (SPI) is used to determine the weight of new templates. A lower SPI (higher semantic stability) results in a higher weight for new templates, while a higher SPI retains more historical template information to ensure template robustness. This dynamically updates the appearance and semantic representation of each target, mitigating appearance drift and semantic ambiguity. The updated template set is then used. This will serve as the initial node feature input for the subsequent graph neural network association module.
[0082] 303: Constructing a graph structure based on semantic focus migration patterns to achieve multi-objective data association. This involves updating the obtained template set. As node features, a multi-target association graph is constructed. For frame t, each node in the graph represents a candidate target, whose semantic focus category... The set of semantic-visual templates and semantic focus prototype vectors of candidate target κ Calculate the cosine similarity, and then select the semantic focus prototype vector corresponding to the one with the highest cosine similarity to determine the category. The edge weights between any two nodes u and o in the graph... Simultaneously considering semantic transfer strength and visual appearance similarity, the definition is as shown in equation (5):
[0083]
[0084] Where α and β are adjustable fusion coefficients, satisfying α+β=1, and M is the semantic focus transfer pattern matrix. Indicates the semantic focus category of frame t. arrive The prior transfer strength is given by cos(·,·), which represents the cosine similarity operator. and Let u and o represent the semantic-visual templates of nodes u and o in frame t, respectively. The cosine similarity is used to measure the consistency between the semantic and appearance feature spaces of the targets corresponding to the two nodes. (Edge weight) The two pieces of information are combined and used as a weighting factor in the information propagation process in the graph structure.
[0085] The constructed graph-structured input graph neural network (GNN) module optimizes multi-objective data association and matching results through multi-layer semantic information propagation and node state updates. Specifically, during the graph neural network propagation process, the features of each node are initialized as follows: Then use edge weights The message passing between adjacent nodes is controlled by the following update rules:
[0086]
[0087] in, This represents the feature vector of node u in the l-th layer; Represents the set of neighbors of node u; Indicates the set of neighbors Normalize the weights of all neighbors to ensure that the weights of all neighbors are positive and their sum is 1; W (l)Let be the learnable weight matrix of the l-th layer, used to project neighbor features into the new feature space; σ(·) is the non-linear activation function (ReLU). After propagation through L layers, the updated node representation is obtained. The updated semantic-visual template not only retains its own semantic-visual information, but also incorporates the contextual relationships of neighboring targets.
[0088] After completing the feature update for the current frame, due to the updated semantic-visual template It integrates the target's own features with neighborhood context information, possessing both discriminative power and temporal consistency; therefore, it is used as a trajectory feature set, along with the trajectory feature set of the previous frame. Similarity calculations are performed to construct a cross-frame association matrix. Then, using a Hungarian algorithm matching strategy, each candidate target in the current frame is assigned a corresponding trajectory from the previous frame, achieving cross-frame association of targets and updating tracking results. The resulting matching results constitute the set of tracking targets for the current frame, providing input for subsequent robustness verification and scheduling.
[0089] 304: A semantically guided robustness verification and policy scheduling module is constructed to verify the tracking template. During the tracking process, a set of perturbation simulation mechanisms are introduced, constructing candidate observation states of the target by setting occlusion regions and focus shift perturbations. This is combined with the spatiotemporal attention heatmap A of the current frame. f With fusion feature F f For each candidate target state, a confidence assessment is performed: from A f The average response value of the candidate target region is obtained to measure the semantic attention intensity of the candidate target. Next, the semantic-visual templates of the candidate target in the current frame and multiple historical frames are compared for similarity to obtain its feature consistency score, which measures the stability of the candidate target over time. Based on this, the semantic attention intensity and feature consistency are weighted and fused to obtain a robustness score. A higher score indicates that the target still has high credibility under occlusion, appearance drift, or semantic interference. Further policy scheduling is performed based on the robustness score results. The tracking strategy dynamically adjusts the template update frequency and feature weights to adapt to different robustness scores. When the score is high, the update frequency is increased to enhance adaptability to appearance changes; when the score is low, the update frequency is decreased and more historical information is retained to avoid drift. The scheduling strategy dynamically allocates computing resources based on the target task priority and environmental complexity. In high-priority or complex scenarios, spatial relocation and high-precision matching modules are invoked to ensure stability; in low-priority or simple scenarios, lightweight tracking paths are used to improve real-time performance. Finally, the optimal target state is selected based on the robustness verification and policy scheduling results, and stable and consistent multi-target tracking results are output.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A multi-target tracking method based on dynamic semantic focus shift, characterized in that, The method includes: Image sequences and user natural language commands are collected separately, and visual and semantic features are extracted separately. The semantic features are stored in a semantic memory pool, and candidate bounding boxes are generated using YOLOX. Unsupervised semantic clustering is performed on the semantic memory pool to extract semantic focus prototype vectors, and a temporal transfer matrix is constructed. A semantic focus transfer pattern matrix is constructed based on the semantic focus prototype vectors and the temporal transfer matrix. Nonlinear spatial alignment of visual features is achieved through deformable convolution under semantic guidance, and a spatiotemporal attention heatmap that integrates spatial distribution and temporal trend is generated by combining the semantic focus migration law matrix; a semantic focus consistency index is introduced to evaluate the fusion stability. When the index exceeds the set threshold, spatial relocation and pose compensation based on depth estimation and SO(3) Lie group interpolation are triggered to complete the reconstruction of semantic features and output multimodal fusion features that combine vision and semantics. Based on the semantic focus migration pattern matrix and multimodal fusion features, the position of the candidate target in the next frame is predicted. Then, according to the prediction result of the target position and the heat map response, the semantic-visual template of the target is dynamically updated. A semantic graph structure with candidate targets as nodes and focus associations as edges is constructed, and the association between targets is dynamically updated using a graph neural network. The node features are the updated semantic-visual templates. The candidate state is evaluated by semantic perturbation and robustness scoring, and the tracking strategy is dynamically adjusted to output stable and consistent multi-target tracking results.
2. The method according to claim 1, characterized in that, The acquired semantic features are stored in a semantic memory pool, which is then populated with historical semantic features. Unsupervised semantic clustering is performed on the semantic memory pool using spectral clustering to obtain K semantic focus prototype vectors Q. k ; The average response intensity vector is obtained by statistically analyzing the average activation level of each semantic focus category in historical natural language instructions. Used to measure the importance of different categories; The frequency of semantic focus category shifts between two consecutive time steps is statistically analyzed, and a time transition matrix T is constructed. By fusing the average response intensity vector and the temporal transfer matrix, a semantic focus transfer pattern matrix M is constructed to quantify the joint transfer trend between different semantic focus categories in the temporal and spatial dimensions, as shown in the following equation: Where, diag(·) represents mapping a vector to a diagonal matrix; λ∈[0,1] represents the balance factor for the fusion of space and time.
3. The method according to claim 2, characterized in that, Nonlinear spatial alignment of visual features is achieved through deformable convolution under semantic guidance, and semantic features F are then aligned. l Mapping to visual feature F v With the same spatial resolution, a spatialized language guidance map L is obtained; for each candidate target in the current frame t... Given the number of candidate targets in the current frame t, and using L as a guide, analyze the feature vectors corresponding to the candidate targets. Perform deformable convolution operations to achieve non-linear spatial alignment of visual features under semantic guidance. After alignment, the feature vectors of candidate targets are written back to the global visual feature map through remapping to obtain the visual alignment feature map F. a .
4. The method according to claim 3, characterized in that, By combining the semantic focus transfer pattern matrix, a spatiotemporal attention heatmap is generated that integrates spatial distribution and temporal trend. Based on the semantic focus transfer pattern matrix M, the average response intensity vector is... Perform a weighted update to obtain the temporally enhanced attention distribution vector. Utilizing candidate targets The corresponding semantic category, from the semantic focus prototype set Extract the corresponding prototype vector from the data and combine it with... The weights of the visual alignment feature map F a Weighted similarity calculations are performed on candidate regions to generate a spatiotemporal attention heatmap A that highlights the candidate target regions most relevant to semantics. f ; Spatiotemporal attention heatmap A f The attention intensity at each pixel location is a weight, and the visual alignment feature map F is... a The corresponding positional feature vectors are weighted to construct a semantic diffusion feature map F. d Then calculate the semantic diffusion feature map F. d Visually aligned feature map F a The pixel-level differences between them are used to define the Semantic Focus Consistency Index (SPI) as follows: Where H and W are the height and width of the spatial dimension, respectively, and i and j represent the pixel coordinate indices of the image in the spatial dimension. The channel vector representing the spatial position (i,j) in the visual alignment feature map. Let ||·||2 represent the channel vector at spatial location (i,j) in the semantic diffusion feature map, and ||·||2 represent the Euclidean distance. When the SPI exceeds the set threshold, spatial relocation and pose compensation based on depth estimation and SO(3) Lie group interpolation are triggered to complete the reconstruction of semantic features and output multimodal fusion features that combine vision and semantics.
5. The method according to claim 4, characterized in that, When the SPI exceeds a set threshold, a depth map D is generated from the input image I using the depth estimation network DPT; then, it is combined with the camera intrinsic parameter matrix K. c Projecting pixel coordinates (x, y) and their depth values into pseudo-point cloud coordinates p in 3D space. (x,y) From the spatiotemporal attention heatmap A f A point set is formed by selecting high-response regions. And obtain the target point set from the historical semantic state. Establish a source-target pairing relationship, where N is the number of points; then, use singular value decomposition to perform rigid body transformation estimation on the two point sets, and solve for the rotation matrix R∈SO(3) and the translation vector ι to make the transformed source point p n Align the corresponding target point p′ n This minimizes the sum of squared registration errors. In the Lie group space SO(3), spherical linear interpolation is performed on the rotation matrix R to generate an intermediate interpolation rotation matrix R. q The transformed 3D point cloud is reprojected into the image space, and the feature map F is visually aligned. a Mid-sampling yields the visual feature map F after semantic-guided relocalization. r : Here, Reproject(·) represents projecting 3D points back to the 2D image coordinate system using the camera model, sampling visual features from the visually aligned feature map at the corresponding positions, and generating a new feature map F. r .
6. The method according to claim 4 or 5, characterized in that, The output of multimodal fusion features combining visual and semantic information includes: if the SPI value exceeds a set threshold, a visual feature map F is obtained through spatial relocalization and pose compensation. r If the SPI value does not exceed the set threshold, the visual feature map F extracted from the input image I will be used. v As visual feature map F r ; The current semantic feature F l Through the linear mapping matrix W l Projected onto the visual channel space, a multi-head cross-attention mechanism is introduced, using semantic features F l For the query vector, the visual feature map F r Channel-selective enhancement is performed on the key and value vectors to generate a fused multimodal feature map: F f =CrossAttn(F r ,W l F l ) Where CrossAttn(·) represents the multi-head cross-attention mechanism; F f This represents a multimodal fusion feature map.
7. The method according to claim 6, characterized in that, Predicting the position of candidate targets in the next frame based on the semantic focus migration pattern matrix and multimodal fusion features, and based on the multimodal fusion feature map F. f The feature vector of the candidate target in the current frame With semantic focus prototype vector set Calculate the similarity to obtain the semantic relevance vector a for each candidate target κ. κ Combining the semantic focus migration pattern matrix M, a location prediction module is constructed, using the semantic relevance vector a. κ For semantic prior, the potential regions of interest in the next frame are modeled. This location prediction module uses a Transformer-based spatiotemporal feature encoder to fuse the multimodal feature map F of the current frame. f By jointly encoding with the semantic focus migration pattern matrix M, the spatial position of each candidate target in the next frame is directly output. With confidence distribution Based on the predicted location and the multimodal fusion feature map of the current frame, combined with the spatiotemporal attention heatmap A f High-confidence, strong semantic response regions of each candidate target are selected as new semantic-visual templates. Semantic-visual template from the previous frame Weighted fusion is performed, with the weights determined by the semantic focus response intensity of high-confidence, strong semantic response regions and the semantic focus consistency index (SPI).
8. The method according to claim 7, characterized in that, Construct a semantic graph structure with candidate targets as nodes and focus associations as edges, where each node has a semantic focus category. The set of semantic-visual templates and semantic focus prototype vectors of candidate target κ Calculate the cosine similarity, and then select the semantic focus prototype vector Q corresponding to the one with the largest cosine similarity. k The category is determined; the edge weight between any two nodes u and o in the graph structure. Simultaneously considering semantic transfer strength and visual appearance similarity: Where α and β are adjustable fusion coefficients, α+β=1, and M is the semantic focus transfer pattern matrix. Indicates the semantic focus category of frame t. arrive The prior transfer strength is given by cos(·,·), which represents the cosine similarity operator. and Let u and o represent the semantic-visual templates of nodes u and o in frame t, respectively. Cosine similarity is used to measure the consistency of the targets corresponding to the two nodes in the semantic and appearance feature space. The constructed graph structure is input into a graph neural network. Through multi-layer semantic information propagation and node state updates, it optimizes multi-objective data association and matching results. During the propagation process of the graph neural network, the features of each node are initialized as follows: Then use edge weights The message passing between adjacent nodes is controlled by the following update rules: in, This represents the feature vector of node u in the l-th layer; Represents the set of neighbors of node u; Indicates the set of neighbors Normalize the weights of all neighbors to ensure that the weights of all neighbors are positive and their sum is 1; W (l) Let be the learnable weight matrix of the l-th layer, projecting the neighbor features into the new feature space; σ(·) is the non-linear activation function (ReLU); after propagation through L layers, the updated node representation is obtained. After completing the feature update for the current frame, the updated semantic-visual template is obtained. It integrates the target's own features with neighborhood context information, possessing both discriminative power and temporal consistency; therefore, it is used as a trajectory feature set, along with the trajectory feature set of the previous frame. Similarity calculations are performed to construct a cross-frame association matrix. Then, using the Hungarian algorithm matching strategy, each candidate target in the current frame is assigned a corresponding trajectory from the previous frame, thereby realizing cross-frame association of targets and updating tracking results. The resulting matching results constitute the set of tracking targets in the current frame.
9. The method according to claim 8, characterized in that, A set of perturbation simulation mechanisms is introduced during candidate target tracking. By setting occlusion areas and focus shift perturbations, candidate observation states of candidate targets are constructed. This is combined with the spatiotemporal attention heatmap A of the current frame. f With fusion feature F f For each candidate target state, a confidence assessment is performed: from A f The average response value of the candidate target region is obtained to measure the semantic attention intensity of the candidate target. Then, the semantic-visual templates of the candidate target in the current frame and in multiple historical frames are compared to obtain the feature consistency score, which is used to measure the stability of the candidate target in the time dimension. Subsequently, a robustness score is obtained by weighted fusion of semantic attention intensity and feature consistency. Based on the robustness score, a strategy scheduling process is further executed. The tracking strategy dynamically adjusts the template update frequency and feature weights to adapt to different robustness scores. When the score is high, the update frequency is increased to enhance adaptability to appearance changes; when the score is low, the update frequency is reduced and more historical information is retained to avoid drift. The scheduling strategy dynamically allocates computational resources based on the target task priority and environmental complexity. In high-priority or complex scenarios, spatial relocation and high-precision matching modules are invoked to ensure stability; in low-priority or simple scenarios, lightweight tracking paths are used to improve real-time performance. Finally, based on the robustness verification and strategy scheduling results, the optimal target state is selected, and stable and consistent multi-target tracking results are output.
Citation Information
Cited By
Intelligent road roller parameter adjusting method and system based on neural network
CN121502307A
Target tracking method and device based on multi-view fusion, equipment and medium
CN122023836A
Target tracking methods, devices, equipment, and media based on multi-view fusion
CN122023836B