Infrared image single target tracking method based on hyperbolic-Euclidean space feature modeling

By introducing hyperbolic-European spatial feature modeling and hyperbolic Transformer network, combining Gaussian model and full convolutional network, the problem of hierarchical feature extraction of infrared image single-target tracking in extreme environments is solved, and the target tracking effect with high precision and strong robustness is achieved.

CN120388046APending Publication Date: 2025-07-29KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510584176.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing infrared image single-target tracking method is difficult to accurately extract hierarchical structural features in extreme environments, resulting in insufficient tracking accuracy and robustness, especially in bad weather conditions such as night and rainy days.

Method used

Using a hyperbolic-European spatial feature modeling method, spatiotemporal visual features and hierarchical structure features are extracted through hyperbolic Transformer network and VisionMamba network, hierarchical trees are constructed and dynamically decoded, and weighted fusion and full convolutional network are used for target positioning to improve feature capture capabilities and anti-interference.

Benefits of technology

It significantly improves the accuracy and robustness of single-target tracking of infrared images, can effectively deal with target deformation, occlusion and rapid movement in complex environments, and provides high-precision tracking solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388046A_ABST
    Figure CN120388046A_ABST
Patent Text Reader

Abstract

The invention relates to an infrared image single target tracking method based on hyperbolic-Euclidean space feature modeling, and belongs to the field of computer vision. The method comprises the following steps: firstly, through a parallel VisionMama network and a hyperbolic Transform network, obtaining a space-time visual feature and a hierarchical structure feature of a search frame; secondly, constructing a hierarchical tree based on a Gaussian model, dynamically decoding the hierarchical structure features of the search frame by adopting an attention mechanism, and optimizing a semantic clustering process through KL regularization constraint to obtain enhanced hierarchical structure features; then, fusing the space-time visual features of the search frames and the enhanced hierarchical structure features by using a weighted fusion technology; and finally, reconstructing the fused features, inputting the reconstructed features into a full convolutional network, and positioning a target by adopting weighted focus loss and multi-target regression loss. According to the method, reliable hierarchical feature extraction can be realized, the adaptability of the model to target deformation, shielding and rapid movement is enhanced, and a high-precision and high-robustness solution is provided for infrared single target tracking in an extreme environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a single-object tracking method for infrared images based on hyperbolic-Euclidean space feature modeling, and belongs to the field of computer vision. Background Art

[0002] Single-object tracking is a core task in computer vision, aiming to real-time track the motion trajectory of a specific target from a video sequence, and is widely applied in fields such as military reconnaissance, border patrol, disaster search and rescue, environmental monitoring, etc. However, under extreme weather conditions such as night, rain, fog, etc., visual target tracking based on visible light (RGB) modality images faces severe challenges or even fails. Thermal infrared (TIR) target tracking generates images by detecting the difference in thermal radiation intensity between objects through a mounted thermal infrared sensor, and can achieve target tracking in complete darkness or bad weather. However, thermal infrared images usually lack obvious features such as textures and edges, making it difficult for algorithms to accurately identify dynamic targets. In addition, infrared images are also vulnerable to noise interference in complex environments, further affecting the tracking accuracy. Therefore, how to accurately extract highly discriminative features in infrared images to effectively cope with various environmental interferences has become a key issue in improving the performance of infrared single-object tracking.

[0003] Most of the existing infrared single-object tracking methods focus on enhancing the saliency of the target or improving template matching, and have made remarkable progress. However, the existing saliency enhancement methods usually only rely on low-level image features and do not fully consider the hierarchical structure features of the target. Template matching methods mostly rely on prior information of the target, and when the target morphology or perspective changes, the effect of template matching drops significantly. In addition, most of the existing deep learning methods rely on feature modeling in Euclidean space, and this modeling method is difficult to effectively extract the strong hierarchical structure or complex geometric relationship of the data. Hyperbolic geometry, as a non-Euclidean geometry model, can better adapt to the complex non-linear and hierarchical structures in infrared images. By introducing feature modeling in hyperbolic space, the hidden hierarchical structure features in the image can be deeply mined when capturing target information, thereby enhancing the tracking performance of the target. Therefore, developing an infrared single-object tracking technology based on hyperbolic feature modeling can provide a more effective solution on the basis of the existing technology to cope with the target tracking challenges in complex environments. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a single-target tracking method for infrared images based on hyperbolic-Euclidean space feature modeling, aiming to improve the hierarchical feature capture ability of infrared target tracking methods. Through the geometric feature projection and cross-attention fusion mechanism of the hyperbolic Transformer, the problem of distortion of hierarchical relationships in the traditional Euclidean space is effectively solved, the target structural representation ability is enhanced, and thus the tracking accuracy and anti-interference ability of the model in complex environments are improved.

[0005] The technical solution of the present invention is: a single-target tracking method for infrared images based on hyperbolic-Euclidean space feature modeling, and the specific steps are as follows:

[0006] Step1: Construct a backbone network, including parallel VisionMamba and hyperbolic Transformer networks. Cut and concatenate the target template frame image and the search frame image to obtain the input of the backbone network. Pass the obtained input through VisionMamba to obtain spatio-temporal visual features, and pass through the hyperbolic Transformer to obtain hierarchical structure features. Decompose the obtained spatio-temporal visual features and hierarchical structure features to obtain the spatio-temporal visual features and hierarchical structure features of the search frame;

[0007] Step2: Build a hierarchical tree based on the Gaussian model, use the attention mechanism to dynamically decode the hierarchical structure features of the search frame, and optimize the semantic clustering process through KL regularization constraints to obtain enhanced hierarchical structure features;

[0008] Step3: Use the weighted fusion technology to associate the spatio-temporal visual features of the search frame with the enhanced hierarchical structure features, strengthen the structural representation of the target, and obtain the fused features;

[0009] Step4: Reshape the fused features into a feature map with a 3D structure, and input it into the fully convolutional network to locate the target through the weighted focal loss and multi-object regression loss.

[0010] The specific content of Step1 is as follows:

[0011] Step1.1: Input image blocks and search area blocks H Z and W Z respectively represent the height and width of the template frame image, H S and W S respectively represent the height and width of the search frame image. Divide and flatten the images to obtain the block sequences and where p×p is the resolution of each block, N Z =H Z W Z / p 2 and NS = H S W S / p 2 are the number of blocks in the template and the search area respectively;

[0012] In the spatio-temporal visual feature branch, a trainable linear projection layer with parameter E is used to project Z p and S p into a D-dimensional hidden space. At the same time, the learnable D-dimensional position embedding P Z and P X are added to the block embeddings of the template and the search area respectively to generate the final template token embedding and the search area token embedding as follows:

[0013]

[0014] Subsequently, the template image patches and the search area patches are concatenated as as the input to the backbone network;

[0015] Step1.2: In the spatio-temporal visual feature extraction branch, VisionMamba is used for global feature extraction and long-range dependence modeling. The input is normalized and linearly projected at the l-th layer of VisionMamba:

[0016]

[0017] where Norm represents normalization and Linear represents linear projection processing, represents the output of the (l - 1)-th layer of VisionMamba, and the intermediate quantities are denoted as v and q. Then, a bidirectional state space model (SSM) is used to process v. In both directions, first, v is convolved one-dimensionally:

[0018] v' o = SiLU(Conv1d(v))

[0019] where Conv1d represents one-dimensional convolution and SiLU represents the activation function, and v' o represents the result after convolving v one-dimensionally. Then, v' o is linearly projected to generate the input matrix B o , the output matrix C o and the time scale parameter Δ o :

[0020] B o = (Linear(v' o ))

[0021] C o =(Linear(v' o ))

[0022]

[0023] wherein, represents a predefined parameter;

[0024] Initialize matrix A using the HiPPO matrix o , and use the zero-order hold technique and the time scale parameter Δ o , to process A o and B o to obtain the discretized A o and B o , denoted as and Subsequently, calculate the intermediate quantities y' forward and y' backward , add them to obtain the output token sequence, which is the output of the l-th layer of VisionMamba

[0025] y' forward = SSM forward (v' o ) ⊙ SiLU(q)

[0026] y' backward = SSM backward (v' o ) ⊙ SiLU(q)

[0027]

[0028] wherein, SSM forward and SSM backward respectively represent the forward state space model and the backward state space model, representing different scanning directions of SSM. The calculation process of SSM is expressed as:

[0029]

[0030] y t = C o h t

[0031] wherein, x t represents the input of the current SSM, h t-1 represents the previous state, h t represents the current state, y tIndicates the output. After being processed by multiple layers of VisionMamba, the spatio-temporal visual feature map is finally obtained. Slice the spatio-temporal visual features to leave the spatio-temporal visual features of the search frame.

[0032] Step1.3: In the hierarchical feature extraction branch parallel to the spatio-temporal visual feature extraction branch, use the hyperbolic vision Transformer to represent the structured features of the target. The hyperbolic vision Transformer introduces hyperbolic position encoding, maps the position encoding to the hyperbolic space through a learnable curvature parameter c, and adds the hyperbolic position encoding to the input block features. Finally, concatenate the image blocks and input them into the backbone network:

[0033]

[0034] Among them, is the initial position encoding, represents the hyperbolic vector multiplication, E pos represents the hyperbolic position encoding. Subsequently, use the Poincaré ball to project onto the hyperbolic space to obtain Add the hyperbolic position encoding to the features of the input block:

[0035]

[0036] Among them, represents the sequence of image blocks after adding the hyperbolic position encoding, represents matrix addition. Use the hyperbolic linear layer to calculate the query Q, key K, and value V:

[0037]

[0038] Among them, represents the Möbius matrix-vector multiplication, W Q , W K , W V is the weight matrix, b Q , b K , b V is the bias term. Calculate the hyperbolic distance d D (Q, K):

[0039]

[0040] Among them, β is a constant for numerical stability. Calculate the attention score using the hyperbolic distance and normalize it through the softmax function to obtain the attention weight:

[0041]

[0042] Among them, α h is a specific scaling factor, and A b,h,i,j represents the attention score calculated using this scaling factor. α b,h,i,j represents the attention weight obtained by normalizing the score through the softmax function. Each attention head aggregates the value vector according to the attention weight to obtain the final output:

[0043]

[0044] Among them, O b,h,i represents the weight of a single attention head, and W O and b O are the weight matrix and bias term specific to O b,h,i . O b,i represents the joint weight of all attention heads. To stabilize the residual connection in the hyperbolic space, a learnable scaling parameter δ is introduced:

[0045]

[0046] Among them, F zs is the input, and O is the output of the attention layer. represents the scaling factor at this stage, and F′ zs represents the output of this layer;

[0047] After being processed by multiple layers of hyperbolic vision transformers, the final output is the structural feature map Slice the structural features to leave the hierarchical features of the search frame

[0048] Specifically, Step 2 is as follows:

[0049] Step 2.1: Model the visual structure using a probability hierarchical tree. Each node represents semantic features with a Gaussian distribution, where the mean vector represents the semantic center and the covariance matrix describes the semantic range;

[0050] Construct a tree structure through recursive mixture Gaussian modeling. Each initial layer node is modeled as an independent Gaussian distribution with a mean vector and a diagonal covariance matrix The upper-layer nodes are the mixture distribution of their child nodes. At the same time, to prevent the variance of the distribution from degenerating to zero, a KL divergence regularization term is introduced to constrain the top-layer node distribution to be consistent with the unit Gaussian prior Finally, the predefined hierarchical tree C l is obtained;

[0051] Step 2.2: Given the predefined hierarchical tree C l and the structural features of the search frame Dynamically decode the hierarchical features through a hierarchical decomposition module, where the dynamic decoding is completed by two stacked Transformer decoders. The decomposition process is as follows:

[0052] At the l-th layer, the hierarchical decomposition module takes C l as the query, as the key and value, and aggregates the semantically related hierarchical features into the nearest semantic clusters through the self-attention mechanism. Each decomposed hierarchical node contains multiple hierarchical feature representations, and the final representation of the node is obtained by averaging multiple hierarchical feature representations. The above process is repeated on all L layers to obtain the L-level hierarchy T S ;

[0053] Step2.3: Map T S and back to the Euclidean space to obtain the hierarchical tree T in the Euclidean space E and the search frame structure features

[0054] Step2.4: Use the attention mechanism with T E as the query, as the key and value for hierarchical feature decomposition, enabling the model to learn the logical structure of the visual hierarchy and finally obtaining the enhanced search frame hierarchical structure features

[0055] The specific content of Step3 is as follows:

[0056] Perform weighted fusion on and through a learnable weight parameter. The fusion process is given by the following formula:

[0057]

[0058] where α is a learnable parameter used to calculate a weight for each channel and is normalized through the Sigmoid activation function to limit the value of α within the range of [0,1]. represents the fused features.

[0059] The specific content of Step4 is as follows:

[0060] Step4.1: Reshape the weighted-fused search region feature sequence into a feature map with a 3D structure and input the reshaped feature map into a fully convolutional network, which consists of multiple stacked convolutional layers and outputs specific task results. The task results include the target classification score map, local offset, and normalized bounding box size. During the inference process, select the position with the highest score in the target classification score map as the position of the target;

[0061] Step 4.2: Adopt weighted focal loss L cls Optimize the target classification score map, and use L1 loss L L1 and generalized intersection over union loss L IoU to optimize the regression of the bounding box. The total loss function L track is expressed as:

[0062] L track = L cls + λ IoU L IoU + λ L1 L L1

[0063] where λ IoU and λ L1 are regularization parameters used to balance the weights of different loss terms.

[0064] The beneficial effects of the present invention are as follows: Hyperbolic geometry is introduced into the field of infrared single-object tracking and co-designed with the VisionMamba network, breaking through the bottleneck of the traditional Euclidean space in characterizing the hierarchical structural features of infrared targets, and significantly improving the tracking accuracy and robustness in complex interference scenarios; Based on the hierarchical feature projection of hyperbolic Transformer and the Gaussian hierarchical tree dynamic decoding mechanism, it can adaptively capture the hierarchical semantic information of the target, alleviating the tracking drift problem caused by fuzzy features of infrared low-texture targets; Through the geometric constraints of hyperbolic position encoding and Möbius operations, reliable hierarchical feature extraction is achieved; Combining the joint optimization strategy of weighted focal loss and multi-object regression loss, the adaptability of the model to target deformation, occlusion, and fast movement is enhanced; The integrated application of these technologies provides a high-precision and strong-robustness technical solution for infrared single-object tracking in extreme environments. Description of the Drawings

[0065] Figure 1 is the flowchart of the steps of the present invention;

[0066] Figure 2 is the visualization effect diagram of the tracking results of the present invention on the PTB-TIR dataset. Detailed Embodiments

[0067] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described below in conjunction with the drawings and specific embodiments, but the protection scope of the present invention is not limited to the described scope.

[0068] Embodiment 1: As Figure 1 shown, an infrared image single-object tracking method based on hyperbolic-Euclidean space feature modeling, the specific steps are as follows:

[0069] Step1: Construct the backbone network, which includes parallel VisionMamba and hyperbolic Transformer networks. Cut and concatenate the target template frame image Z and the search frame image S to obtain the input of the backbone network. Pass the obtained input through VisionMamba to get spatio-temporal visual features, and through the hyperbolic Transformer to get hierarchical features. Decompose the obtained spatio-temporal visual features and hierarchical features to obtain the spatio-temporal visual features and hierarchical features of the search frame.

[0070] Step1.1: Prepare the open-source dataset PTB-TIR. Read the corresponding image pairs according to the index and convert them into PyTorch tensors. Input the image pairs into the backbone network for preprocessing: the input template area size is processed to 128×128 pixels, and the search area is processed to 256×256 pixels. The template image and the search area are divided into blocks with p = 16 and flattened into a block sequence and where N Z = 64, N S = 256.

[0071] Step1.2: In the spatio-temporal visual feature extraction branch, for the block sequences Z p and S p , generate initial features through linear projection (hidden dimension D = 384), embed learnable one-dimensional positional encoding, and concatenate the template and search area features as and input it into the backbone network for visual feature extraction. The backbone network has 24 layers, and each layer contains a bidirectional SSM module. At each layer of the backbone network, after input normalization, linear projection generates a one-dimensional convolutional kernel size of 4, and initialize the state matrix A using the HiPPO matrix o . The output of the bidirectional SSM is connected through a residual connection, and finally the spatio-temporal visual features For subsequent processing convenience, slice the spatio-temporal visual features to leave the spatio-temporal visual features of the search frame

[0072] Step1.3: In the hierarchical feature extraction branch, for the block sequences Z p and S p , generate initial features through Poincaré hyperbolic linear projection (hidden dimension D = 384), with an initial curvature parameter c = 1, embed learnable one-dimensional hyperbolic positional encoding, and concatenate the template and search area features as and input it into the backbone network for visual feature extraction. The backbone network has 12 layers, and each layer contains a hyperbolic Transformer module. Inside each hyperbolic Transformer module, use hyperbolic linear layers to calculate the query Q, key K, and value V:

[0073]

[0074] Among them, represents the Möbius matrix-vector multiplication, is the weight matrix, is the bias term. Calculate the hyperbolic distance d D (Q, K):

[0075]

[0076] Among them, β is a constant for numerical stability. In this invention, β = 1e-15. The attention score is calculated using the hyperbolic distance and normalized by the softmax function to obtain the attention weights:

[0077]

[0078] Among them, α h is a specific scaling factor. In this invention, α h = 0.1, A b,h,i,j represents the attention score calculated using this scaling factor, and α b,h,i,j represents the attention weights obtained by normalizing the scores through the softmax function. Each attention head aggregates the value vectors according to the attention weights to obtain the final output:

[0079]

[0080] Among them, O b,h,i represents the weight of a single attention head, W O and b O are the weight matrix and bias term specific to O b,h,i , and O b,i represents the joint weight of all attention heads. To stabilize the residual connection in the hyperbolic space, a learnable scaling parameter δ is introduced:

[0081]

[0082] Among them, F zs is the input, and O is the output of the attention layer. represents the scaling factor at this stage, and F' zs represents the output of this layer. After passing the output of this layer through the activation function, residual connection, and hyperbolic layer normalization, the final output is obtained.

[0083] After being processed by multiple hyperbolic vision Transformers, the final output is the structural feature map Slice the structural features to leave the hierarchical features of the search frame

[0084] Step 2: Construct a hierarchical tree based on the Gaussian model, adopt the attention mechanism to dynamically decode the hierarchical structure features of the search frame, and optimize the semantic clustering process through KL regularization constraints to obtain the enhanced hierarchical structure features.

[0085] Step 2.1: Let be the set of N1 initial hierarchical nodes. Parameterize each node as a normal distribution with mean vector and diagonal covariance matrix

[0086]

[0087] where is the probability distribution of node I represents the variance, and are randomly initialized, and the reparameterization trick is used for stable sampling. Then, by conditioning each level on the previous level, the remaining (L - 1) levels of the hierarchical tree are obtained. For each level l, construct a set of N l nodes Specifically, the k-th node in the l-th level is approximated by the MoG of its corresponding two child nodes and as:

[0088]

[0089] By conditioning the higher-level nodes on the previous lower-level nodes in sequence, the hierarchical tree is obtained. At the same time, introduce the KL divergence regularization term, denoted as to constrain the similarity between the top-level node distribution and the unit Gaussian prior :

[0090]

[0091] Step 2.2: After obtaining the predefined hierarchical tree T and the search frame structure features the decomposition module composed of two Transformer decoders decomposes G l as the query, as the key and value, so that the semantically related visual features are clustered into the closest semantic clusters:

[0092]

[0093] The decomposed visual hierarchical nodes are composed of 2 l-1consists of visual representations. For each visual hierarchy node The representations are averaged, and the same decomposition process is performed on layer L to obtain the hierarchy of layer L

[0094] Step2.3: Map T S and back to the Euclidean space to obtain the hierarchical tree T E of the Euclidean space and the search frame structure features

[0095] Step2.4: Hierarchical feature dynamic decoding, using the Euclidean space to perform hierarchical encoding on the visual hierarchy structure. Through a decomposition module composed of two Transformer decoders, T E is used as the query, is used as the key and value for decomposition, enabling the model to learn the logical structure of the visual hierarchy and finally obtaining the enhanced search frame hierarchical structure features thus enhancing the model's hierarchical understanding of the global visual representation.

[0096] Step3: Use the weighted fusion technique to associate the spatio-temporal visual features of the search frame with the enhanced hierarchical structure features, strengthen the structured representation of the target, and obtain the fused features.

[0097] Specifically, through a learnable weight parameter pair and are weighted and fused, and the fusion process is given by the following formula:

[0098]

[0099] where α is a learnable parameter used to calculate a weight for each channel and is normalized through the Sigmoid activation function to limit the value of α within the range of [0,1], represents the fused features.

[0100] Step4: Reshape the fused features into a feature map with a 3D structure and input it into a fully convolutional network to localize the target through weighted focal loss and multi-objective regression loss.

[0101] Step4.1: Reshape the weighted fused search region feature sequence into a feature map with a 3D structure, and then input it into a fully convolutional network composed of L stacked Conv-BN-ReLU layers. The output of the FCN contains the target classification score map local offsets and the normalized bounding box sizes Select the position with the highest score in the target classification score map as the position of the target, i.e., (x d , y d ) = argmax (x,y) Cls, (x d , y d ) represents the coordinates of the candidate box with the highest score. The final target box is:

[0102] (x, y, w, h) = (x d + Bia(0, x d , y d ), y d + Bia(1, x d , y d ), Box(0, x d , y d ), Box(1, x d , y d ))

[0103] Among them, (x, y) represents the upper left corner coordinates of the target box, and w and h represent the width and height of the target box respectively.

[0104] Step 4.2: During training, use the weighted focal loss L cls to optimize the target classification score map, and use the L1 loss L L1 and the generalized intersection over union loss L IoU to optimize the regression of the bounding box. The total loss function L track is expressed as:

[0105] L track = L cls + λ IoU L IoU + λ L1 L L1

[0106] Among them, λ IoU = 2 and λ L1 = 5 are regularization parameters used to balance the weights of different loss terms.

[0107] Next, based on the specific implementation records, the present invention will be further described by means of experiments.

[0108] Experimental data: The present invention uses the PTB-TIR dataset. PTB-TIR was published in 2019 and is a comprehensive thermal infrared pedestrian tracking dataset, including 60 carefully annotated thermal infrared tracking sequences from different devices, scenes, and shooting times, including more than 30,000 video frames.

[0109] Experimental Setup: This invention is implemented based on the Pytorch 2.1 framework, with CUDA version 11.8. Model training was carried out using 6 NVIDIA 4090 GPUs, and testing was performed using a single NVIDIA 4060 TI GPU. AdamW was used as the model training optimizer, and the cosine annealing strategy in OneCycleLR was adopted to dynamically adjust the learning rate. The training is divided into two stages. The first stage is carried out on four visible light datasets, GOT-10K, LaSOT, COCO, and TrackingNet. The total number of epochs in the first stage is set to 300, the initial learning rate is 0.0004, and the learning rate is reduced to 0.0001 after 240 epochs. The second stage of training is carried out on the thermal infrared dataset LSOTB-TIR. The total number of epochs in the second stage is set to 60, the initial learning rate is 0.0004, and the learning rate is reduced to 0.0001 after 42 epochs. The settings of the training set follow the official settings. This invention uses the area under the curve (AUC) and precision (Precision) as evaluation metrics, where AUC is the main evaluation metric.

[0110] For comparison, this embodiment also selected classic models ECO_stir (2019, IEEE TIP), AiATrack (2022, ECCV), SimTrack (2022, ECCV), SAOT (2021, ICCV), CSWinTT (2022, CVPR), and STARK (2021, ICCV) as comparison models.

[0111] Experimental Results: Through the above steps, the PTB-TIR dataset was experimentally verified, and the predicted experimental results are shown in Table 1. The original images, ground truths, and predicted results of some datasets are visualized as Figure 2 shown.

[0112] Table 1 Prediction Effect of PTB-TIR Dataset

[0113]

[0114]

[0115] The optimal values of each metric are shown in bold. As can be seen from Table 1, the AUC and Precision of the predicted values of this invention on the PTB-TIR dataset both reach the optimal, indicating that this invention can better model the features of infrared images compared with the prior art, thereby improving the prediction accuracy. Figure 2 The visualization shows the prediction effect diagram of this invention, more intuitively reflecting the excellent performance of this invention. For example, Figure 2 the first column shows that when the target is occluded, the method of this invention can still track the target well. Figure 2The second and fourth rows show the performance of the present invention when similar targets appear, and the third row indicates the adaptability of the present invention to the rapid movement of the target.

[0116] The specific embodiments of the present invention have been described in detail in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. An infrared image single-target tracking method based on hyperbolic-Euclidean space feature modeling, characterized in that, Specifically, it includes the following steps: Step1: Construct a backbone network, which includes parallel VisionMamba and hyperbolic Transformer networks. Cut and concatenate the target template frame image and the search frame image to obtain the input of the backbone network. Pass the obtained input through VisionMamba to get spatio-temporal visual features, and pass it through the hyperbolic Transformer to get hierarchical structure features. Decompose the obtained spatio-temporal visual features and hierarchical structure features to obtain the spatio-temporal visual features and hierarchical structure features of the search frame; Step2: Construct a hierarchical tree based on the Gaussian model. Use the attention mechanism to dynamically decode the hierarchical structure features of the search frame, and optimize the semantic clustering process through KL regularization constraint to obtain enhanced hierarchical structure features; Step3: Use the weighted fusion technology to associate the spatio-temporal visual features of the search frame and the enhanced hierarchical structure features, strengthen the structured representation of the target, and obtain the fused features; Step4: Reshape the fused features into a feature map with a 3D structure and input it into a fully convolutional network to locate the target through weighted focal loss and multi-object regression loss.

2. The single-object tracking method for infrared images based on hyperbolic-Euclidean space feature modeling according to claim 1, wherein, The specific content of Step1 is as follows: Step1.1: Input the image block and the search area block H Z and W Z represent the height and width of the template frame image respectively. H S and W S represent the height and width of the search frame image respectively. The images are blocked and flattened to obtain the block sequences and where p×p is the resolution of each block, and N Z = H Z W Z / p 2 and N S = H S W S / p 2 are the number of blocks of the template and the search area respectively; In the spatio-temporal visual feature branch, Z is projected into a D-dimensional hidden space using a trainable linear projection layer with parameter E p and S p Meanwhile, the learnable D-dimensional position embeddings P Z and P X are added to the patch embeddings of the template and the search region respectively to generate the final template token embedding and the search region token embedding as follows: Subsequently, the template image block and the search area block are concatenated into as the input of the backbone network; Step1.2: In the spatio-temporal visual feature extraction branch, use VisionMamba for global feature extraction and long-range dependence modeling. Normalize and linearly project the input at the l-th layer of VisionMamba: Among them, Norm represents normalization, and Linear represents linear projection processing. represents the output of the (l-1)th layer of VisionMamba, and the obtained intermediate quantities are represented by v and q. Then, the bidirectional state space model is used to process v. In two directions, first, a one-dimensional convolution is performed on v: v' o = SiLU(Conv1d(v)) Among them, Conv1d represents one-dimensional convolution, SiLU represents an activation function, and v' o represents the result after performing one-dimensional convolution on v. Then, for v' o a linear projection is performed to generate the input matrix B o , the output matrix C o and the time scale parameter Δ o : B o = (Linear(v' o )) C o = (Linear(v' o )) Among them, represents a predefined parameter; Initialize matrix A using the HiPPO matrix o , use the zero-order hold technique and the time scale parameter Δ o , for A o and B o perform processing to obtain the discretized A o and B o , denoted as and Subsequently, calculate the intermediate quantities y' forward and y' backward , add them together to obtain the output token sequence, which is the output of the l-th layer of VisionMamba y' forward = SSM forward (v' o ) ⊙ SiLU(q) y' backward = SSM backward (v' o )⊙SiLU(q) Among them, SSM forward and SSM backward respectively represent the forward state space model and the backward state space model, representing different scanning directions of SSM. The calculation process of SSM is expressed as: y t = C o h t Among them, x t represents the input of the current SSM, h t-1 represents the previous state, h t represents the current state, y t represents the output. After being processed by multiple layers of VisionMamba, the spatio-temporal visual feature map is finally obtained Slice the spatio-temporal visual features to leave the spatio-temporal visual features of the search frame Step1.3: In the hierarchical structure feature extraction branch parallel to the spatio-temporal visual feature extraction branch, use the hyperbolic vision Transformer to represent the structured features of the target. The hyperbolic vision Transformer introduces hyperbolic position encoding, and maps the position encoding to the hyperbolic space through the learnable curvature parameter c: Among them, is the initial position encoding, represents hyperbolic vector multiplication, E pos represents hyperbolic position encoding. Subsequently, using the Poincaré ball to project to the hyperbolic space to obtain Add the hyperbolic position encoding to the features of the input block: Among them, represents the sequence of image patches after adding hyperbolic positional encoding, represents matrix addition, and uses a hyperbolic linear layer to calculate the query Q, key K, and value V: Among them, represents the Möbius matrix-vector multiplication, and W Q , W K , W V is the weight matrix, and b Q , b K , b V is the bias term, and calculates the hyperbolic distance d D (Q, K): Among them, β is a constant for numerical stability. Calculate the attention score using the hyperbolic distance and normalize it through the softmax function to obtain the attention weight: Among them, α h is a specific scaling factor, and A b,h,i,j represents the attention score calculated using this scaling factor. α b,h,i,j represents the attention weight obtained by normalizing the scores through the softmax function. Each attention head aggregates the value vectors according to the attention weights to obtain the final output: Among them, O b,h,i represents the weight of a single attention head, W O and b O are the weight matrix and bias term specific to O b,h,i , O b,i represents the combined weight of all attention heads. To stabilize the residual connection in the hyperbolic space, a learnable scaling parameter δ is introduced: Among them, F zs is the input, O is the output of the attention layer, represents the scaling factor of this stage, and F′ zs represents the output of this layer; After being processed by multiple layers of hyperbolic vision transformers, the final output is a structural feature map Slice the structural features to leave the hierarchical features of the search frame 3. A single-object infrared image tracking method based on hyperbolic-Euclidean space feature modeling according to claim 1, characterized in that The specific content of Step2 is as follows: Step2.1: Use probabilistic hierarchical tree to model the visual structure. Each node represents semantic features with a Gaussian distribution, where the mean vector represents the semantic center and the covariance matrix describes the semantic range; Construct a tree structure through recursive Gaussian mixture modeling. Each initial layer node is modeled as an independent Gaussian distribution with a mean vector and a diagonal covariance matrix The upper layer nodes are the mixture distributions of their child nodes. At the same time, a KL divergence regularization term is introduced to constrain the top layer node distribution to be consistent with the unit Gaussian prior to obtain the predefined hierarchical tree C l ; Step 2.2: Given a predefined hierarchical tree C l and search frame structure features Dynamically decode the hierarchical features through a hierarchical decomposition module, which is completed by two stacked Transformer decoders. The decomposition process is as follows: At the l-th layer, the hierarchical decomposition module takes C l as a query, as keys and values, and aggregates semantically related hierarchical features into the nearest semantic clusters through the self-attention mechanism. Each decomposed hierarchical node contains multiple hierarchical feature representations, and the final representation of the node is obtained by averaging multiple hierarchical feature representations. The above process is repeated on all L layers to obtain the L-level hierarchy T S ; Step 2.3: Map T S and back to the Euclidean space to obtain the hierarchical tree T E of the Euclidean space and the search frame structure feature Step2.4: Use the attention mechanism with T E as the query, as the key and value for hierarchical feature decomposition, enabling the model to learn the logical structure of the visual hierarchy and finally obtaining the enhanced search frame hierarchical structure features 4. The single-target tracking method for infrared images based on hyperbolic-Euclidean space feature modeling according to claim 1, characterized in that The specific content of Step3 is as follows: Weighted fusion is performed on and through a learnable weight parameter, and the fusion process is given by the following formula: Among them, α is a learnable parameter used to calculate a weight for each channel and normalize it through the Sigmoid activation function, restricting the value of α within the range of [0, 1]. represents the fused features.

5. A single-target tracking method for infrared images based on hyperbolic-Euclidean space feature modeling according to claim 1, characterized in that The specific content of Step4 is as follows: Step4.1: Reshape the weighted fused search region feature sequence into a 3D-structured feature map and input the reshaped feature map into a fully convolutional network, which consists of multiple stacked convolutional layers and outputs specific task results. The task results include the target classification score map, local offset, and normalized bounding box size. During the inference process, select the position with the highest score in the target classification score map as the position of the target; Step 4.2: Adopt the weighted focal loss L cls to optimize the target classification score map, and use the L1 loss L L1 and the generalized intersection over union loss L IoU to optimize the regression of the bounding box. The total loss function L track is expressed as: L track = L cls + λ IoU L IoU + λ L1 L L1 Among them, λ IoU and λ L1 are regularization parameters used to balance the weights of different loss terms.

Citation Information

Cited By

  • Visual Transform-based dynamic screening medical image target tracking method and device

    CN120823216A

  • A dynamic screening medical image target tracking method and device based on visual Transformer

    CN120823216B