A hyperspectral target tracking method based on dual-stream visual cues
Through the dual-stream visual cueing method, representative bands are selected and spectral and spatiotemporal information are integrated to solve the band redundancy and data scarcity problems in hyperspectral target tracking, and improve the robustness and target tracking success rate.
Patent Information
- Application Number
- CN202410192605.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-21
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-02-21
AI Technical Summary
Existing hyperspectral target tracking technologies suffer from the problems of hyperspectral band redundancy, scarce training data, and insufficient utilization of temporal information, which lead to interference and insufficient robustness in the tracking process.
A hyperspectral target tracking method based on dual-stream visual cues is adopted. Representative bands are selected through the band selection module. The spectral and spatiotemporal information are fused with the dual-stream visual cue. A multi-layer cross-correlation cue layer is designed to enhance the representation ability of the basic modality.
It effectively solves the problems of hyperspectral band redundancy and scarcity of training data, improves the robustness of hyperspectral target tracking and its ability to adapt to complex scenes, and enhances the success rate of target recognition and tracking.
Smart Images

Figure CN117994282B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a hyperspectral target tracking method based on dual-stream visual cues, belonging to hyperspectral target tracking technology. Background Art
[0002] Visual object tracking is a key task in computer vision, widely used in fields such as video surveillance and autonomous driving. Advances in hyperspectral imaging technology, using hyperspectral snapshot sensors, provide detailed compositional information about scene content compared to traditional RGB imagery. The stability of multi-band spectral features in hyperspectral imagery enables hyperspectral object tracking to effectively address challenges such as object deformation and occlusion. This enhances the ability of visual object tracking systems to distinguish objects from complex backgrounds, improves inter-object recognition, and increases target tracking success rates.
[0003] Early material-based hyperspectral trackers used handcrafted feature extractors to capture material information from hyperspectral videos and then employed background-aware correlation filters for target tracking; recent hyperspectral tracking research has employed deep learning features. However, compared to RGB images, hyperspectral images suffer from channel mismatch and spectral redundancy, so the current mainstream approach is to convert hyperspectral images into three-channel images using various dimensionality reduction methods, followed by comprehensive fine-tuning of the parameters of the RGB-based tracker. Despite improvements in hyperspectral tracking technology, current hyperspectral tracking solutions still have significant limitations:
[0004] 1. Hyperspectral band redundancy: Hyperspectral images contain more bands than RGB images, and there is repeated information, which leads to interference and inefficiency in the tracking process.
[0005] 2. Limited hyperspectral training data: Due to the limitations of existing sensor technology, there is a lack of high-quality hyperspectral data, which hinders the effective transfer training of RGB networks and affects the robustness of hyperspectral target tracking.
[0006] 3. Insufficient utilization of temporal information: Traditional hyperspectral trackers often ignore key temporal changes in hyperspectral videos, affecting their ability to handle complex scenes such as rapid object changes and occlusions. Summary of the Invention
[0007] Purpose of the Invention: Hyperspectral images, with their rich spectral information, are beneficial for target tracking in various application scenarios. To leverage the powerful representation capabilities of large-scale pre-trained RGB trackers, the mainstream approach to hyperspectral image tracking involves comprehensive fine-tuning of RGB tracker-based parameters. However, due to redundant downstream hyperspectral band information and scarcity of training data, this approach, while effective, is suboptimal. Furthermore, existing hyperspectral trackers make very limited use of temporal information. To alleviate these issues, this paper proposes a unified hyperspectral target tracking method using spectral-spatiotemporal multimodal dual-stream visual cues.
[0008] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is:
[0009] A hyperspectral target tracking method based on dual-stream visual cues is proposed. The network model used to implement this method is called the HDSP model. The training and tracking process of the HDSP model includes the following steps:
[0010] (1) The t-th frame hyperspectral image and the t-1th frame hyperspectral image As the input of HDSP model; get H through band selection module BSM t The three representative bands are combined into the pseudo color image of the tth frame H t and H t-1 The false color image of frame t is obtained by CIE color matching functions CMFs respectively and the t-1th frame false color image F t 、P t 、F t-1 The corresponding tokens are generated by cropping, dicing and embedding and flattening into the latent space. The first generated token is recorded as the initial token, which is represented by
[0011] (2) Input into the dual-stream visual prompter MDVP, MDVP consists of an initial cross-correlation prompt layer ICPL and multiple subsequent cross-correlation prompt layers CPL; ICPL consists of a motion feature fusion module MFFM and two prompt generation modules PGM, the two prompt generation modules PGM are respectively recorded as the first layer PGM-1 and the first layer PGM-2; the subsequent cross-correlation prompt layer CPL is recorded as the lth layer CPL, the lth layer CPL consists of two prompt generation modules PGM, the two prompt generation modules PGM are respectively recorded as the lth layer PGM-1 and the lth layer PGM-2, l = 2, 3, ..., L; for the convenience of introduction, the present invention is based on the MDVP The generated token information is called basic information flow, based on The generated token information is called spatiotemporal hint stream, based on The generated token information is called the spectral cue stream;
[0012] In MDVP, MFFM will and Fusion to obtain the initial spatiotemporal prompt flow Will and Input into the first layer PGM-1 to get the first layer of spatiotemporal prompt flow Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow The first layer of spatiotemporal prompt stream output by ICPL and the first layer of spectral hint flow As the input of the second layer CPL;
[0013] In the first layer of CPL, and Input into the first layer PGM-1 to obtain the first layer spatiotemporal prompt stream Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow
[0014] (3) Construct an L-layer classification encoder, After element-by-element addition, it is input into the lth layer of the classification encoder to obtain the lth layer basic information flow Recorded as For the L-th layer basic information flow of the L-th layer output of the classification encoder First, the linear layer is used to reduce the channel dimension, and then the input is sent to the prediction head for classification and bounding box regression to achieve target tracking;
[0015] (4) Using classification loss and bounding box regression loss Calculate total loss Utilize total loss Train the HDSP model and finally obtain a trained HDSP model;
[0016] (5) Use the trained HDSP model to track the target in the video frame to obtain the target tracking result.
[0017] Preferably, in step (1), F is calculated based on the following steps: t 、P t 、F t-1 :
[0018] (11) Calculate the pseudo-color image P of the tth frame t , including the following steps:
[0019] (111) The t-th frame hyperspectral image H t Split into a sequence of N bands, the i-th band is recorded as i=1,2,…,N; based on Image entropy Evaluate Information content:
[0020]
[0021] in: Indicates band The probability density function of the band The pixel intensity 255 histogram is normalized to obtain As x, with i as y, Mapped into a two-dimensional coordinate system;
[0022] (112) Use the image entropy of each band to perform density clustering on the bands, select each band as the cluster center in turn, and calculate the local density value and minimum distance
[0023] Local density value Depends on the direct distance between the bands, based on the triangle inequality, the local density value Defined as:
[0024]
[0025] in: express and The Euclidean distance between them; when x<0, When x≥0, d r is the cluster radius, H t The entropy of the N band images is determined by one third of the difference between the maximum and minimum values, and the divisor 3 is the same as the number of channels in RGB;
[0026] According to the band Bands with higher entropy than another image The local density value calculation band Image entropy distance band The minimum distance of image entropy
[0027]
[0028] Minimum distance through Characterization band With band relevance;
[0029] (113) Mapped to the two-dimensional coordinate system, in order to maintain the consistency of the three channels of RGB and the basic information flow, select The highest three cluster centers are used as representative bands, and the three representative bands are combined into the pseudo-color image P of the tth frame. t ;
[0030] (12) Calculate the false color image F of the t-th frame and the t-1-th frame t 、F t-1 , including the following steps:
[0031] (121) For a hyperspectral image H containing H×W pixels and N bands t 、H t-1 , use color matching function CMFs to transform the hyperspectral image H t 、H t-1 Convert to CIEXYZ color space, expressed as:
[0032]
[0033]
[0034] in: Represents color matching function CMFs, I t , I t-1 is the hyperspectral image H t 、H t-1 False color image rendered in the CIEXYZ color space after color matching functions (CMFs); in this case, the 10-deg XYZCMFs function converted by CIE (2006) was used with a step size of 0.1 nm. The default CIEXYZ color space is RGB.
[0035] (122) The Monge-Kantorovitch linear color mapping transformation method based on sample color migration is used to enhance the color closeness between the false color image and the RGB image. t , I t-1 Denoted as F t 、F t-1 ;
[0036] (13) To F t 、P t 、F t-1The corresponding tokens are generated by cropping, dicing and embedding and flattening into the latent space. The first generated token is recorded as the initial token, which is represented by
[0037] Preferably, in step (2), spectral cue stream and spatiotemporal cue stream are generated by fusing multimodal features using MDVP, and the spectral cue stream and spatiotemporal cue stream are inserted into the classification encoder; MDVP is composed of an initial cross-correlation cue layer ICPL and multiple subsequent cross-correlation cue layers CPL;
[0038] (21)ICPL consists of a motion feature fusion module MFFM and two prompt generation modules PGM. The two prompt generation modules PGM are respectively denoted as the first layer PGM-1 and the first layer PGM-2; Input into ICPL to get the first layer of spatiotemporal prompt flow First layer spectral hint flow Expressed as:
[0039]
[0040] (22) In the motion feature fusion module MFFM, first calculate The absolute difference between them is the frame difference Then use two convolutional layers plus a Sigmoid activation function to obtain the initial spatiotemporal hint flow
[0041] (23) and Input into the first layer PGM-1 to get the first layer of spatiotemporal prompt flow Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow
[0042] (24)ICPL output and As the input of the second layer CPL, the output of the l-1 layer CPL and As the input of the l-th layer CPL, it is expressed as
[0043] Preferably, in step (2), the prompt generation module PGM includes three parts: a channel attention module CAM, a spatial attention module SAM and a cross fusion module CFM;
[0044] For the channel attention module CAM, the input token is represented as First, follow the H k ×W kPerform average pooling on the dimension and max pooling Then and The input is fed into a 1×1 convolutional layer for channel dimensionality reduction, and then nonlinear enhancement is performed through a ReLU layer. Then another 1×1 convolutional layer is used and a Sigmoid activation function is added to project the original dimension back to obtain a weight map with the same dimension as the input token C. Finally, the obtained Z channel Multiply by the input token C to get the channel attention token T C =C×Z channel ;
[0045] For the spatial attention module SAM, the input token is represented as First, average pooling is performed along the D dimension and max pooling Then along the D dimension and Connect them and input them into a (H k -1)×(W k -1) Convolutional layer and add Sigmoid activation function to get weight mapping Finally, the obtained Z spatial Multiply by the input token S to get the spatial attention token T S =S×Z spatial ;
[0046] For the cross-fusion module CFM, the prompt token T is obtained by cross-fusion of input token T1 and input token T2. F , T1 passes through a channel attention module CAM and a spatial attention module SAM in turn to obtain T2 is obtained after passing through a channel attention module CAM and a spatial attention module SAM in sequence. The sum of T1' and T2' is obtained as T'=T1'+T2', and T' is input into two different network branches for operation to obtain and right and Summing to get For T1', T2' and Perform cross-fusion to get prompt token
[0047] Preferably, in step (3), when constructing the L-layer classification encoder, all network parameters related to the false color image (i.e., the basic mode, generally the basic mode is the RGB mode, but in this case there is no RGB mode image, so the false color image generated based on the hyperspectral image is used as the basic mode) are frozen, including the network parameters of the block embedding, feature extraction, feature interaction and prediction head.
[0048] Preferably, in step (4), when using the loss function to train the classification encoder, the HDSP model is first initialized using the parameters in the baseline model OSTrack, and then the cross entropy loss is used as the classification loss L cls , using IoU loss as the bounding box regression loss Then according to the classification loss and bounding box regression loss Calculate total loss λ iou 、 is the regularization parameter, L1 represents the mean absolute error, and finally the total loss is used Tracking training is performed on the HDSP model. During training, the backbone network of the HDSP model is frozen, and only the prompt part of the HDSP model is trained, and finally a trained model is obtained. The backbone network of the HDSP model includes F t and F t-1 The block embedding part, classification encoder, prediction head, and the prompt part of the HDSP model include P t The cut embedded part, MDVP.
[0049] Preferably, in step (5), the trained HDSP model is used to track the target in the video frame, and the following requirements are met: ① the input is two adjacent hyperspectral images in the video frame; ② the Siamese architecture is used for target tracking, and the first frame of the hyperspectral image input is initialized as the template frame, and the subsequent frames of hyperspectral images are used as search frames; during the target tracking process, the online template is not used and the template frame is not updated; ③ the maximum value of the final classification score map is used as the center position of the tracked target, and network regression is performed to determine the position of the bounding box.
[0050] Beneficial effects: The hyperspectral target tracking method based on dual-stream visual cues provided by the present invention has the following advantages over the existing technology: 1. The present invention adopts the technology of cue learning, which can be trained under the premise of freezing most parameters, solving the problems of lack of hyperspectral data sets and difficulty in migrating RGB trackers; 2. The band selection module based on density clustering designed by the present invention retains the bands with minimum correlation and maximum information as spectral cue information, which is conducive to removing band redundancy; 3. The dual-stream visual information prompter in the present invention converts multimodal input into a single modality through the designed multi-layer cross-correlation prompt layer, effectively utilizes spectral information and spatiotemporal information, and helps to enhance the representation ability of the basic modality; 4. The dual-stream visual information prompter in the present invention, combined with the attention mechanism, adaptively enhances and fuses the features of spectral information modality and spatiotemporal information modality through multi-layer cross-correlation prompt layer, and adds them to the basic modality, which can efficiently adapt to downstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 Flow chart for the implementation of the method of the present invention;
[0052] Figure 2 Schematic diagram of the HDSP framework of the method of the present invention;
[0053] Figure 3 This is an example diagram of the operation of the BSM of the present invention;
[0054] Figure 4 This is a structural block diagram of the PGM of the present invention. DETAILED DESCRIPTION
[0055] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0056] With its rich spectral information, hyperspectral images are beneficial for target tracking in various application scenarios. To leverage the powerful representation capabilities of large-scale pre-trained RGB trackers, the mainstream approach to hyperspectral image tracking is to fully fine-tune the parameters of RGB-based trackers. However, due to the redundancy of downstream hyperspectral image band information and the scarcity of training data, this approach, while effective, is not optimal. Furthermore, existing hyperspectral trackers make very limited use of temporal information. To alleviate these issues, this paper proposes a unified spectral-spatiotemporal multimodal dual-stream visual cueing method for hyperspectral target tracking. First, a density clustering-based band selection module (BSM) is designed to retain the bands with the minimum correlation and the maximum information content as spectral cue information to remove redundancy. Then, the generated band information and temporal information are used as spectral-spatiotemporal multimodal cues, and a dual-stream visual information cue is proposed. Through the designed multi-layer cross-correlation cue layer, the multimodal input is converted into a single modality, enhancing the representation capability of the basic modality of hyperspectral target tracking.
[0057] like Figure 1 The following is the implementation flow chart of this scheme. Initial preprocessing is performed on the two input hyperspectral image frames using color matching functions (CMFs) and a band selection module (BSM). These images are then cropped and sliced before being fed into an embedding layer to obtain corresponding base and cue tokens. These tokens enter the dual-stream visual cue generator (MDVP), where effective visual cues are generated through a designed initial cross-correlation cue layer (ICPL) and multiple subsequent cross-correlation cue layers (CPL). The cue generation modules (PGMs) within each cross-correlation cue layer enhance and fuse the base, spectral, and spatiotemporal information streams. The visual (spectral-spatiotemporal) cue stream output from the cross-correlation cue layer is then element-wise added to the base information stream and fed into a multi-layer stacked classification encoder for feature extraction and interaction. Finally, the video frames are fed into a trained network model for tracking, yielding the tracking results. This scheme leverages spectral and spatiotemporal modal cue information to enhance the expressive power of the base modality and fully leverage the advantages of cue learning. Each step is described below.
[0058] Step S01: Initial preprocessing of adjacent frame hyperspectral images is performed through color matching functions CMFs and band selection modules BSM.
[0059] like Figure 2 As shown, the t-th frame hyperspectral image and the t-1th frame hyperspectral image As the input of HDSP model; get H through band selection module BSM t The three representative bands are combined into the pseudo color image of the tth frame H t and H t-1 The false color image of frame t is obtained by CIE color matching functions CMFs respectively and the t-1th frame false color image F t 、P t 、F t-1 The corresponding tokens are generated by cropping, dicing and embedding and flattening into the latent space. The first generated token is recorded as the initial token, which is represented by The specific steps include:
[0060] Step 11: Calculate the pseudo-color image P of the tth frame t , including the following steps:
[0061] Step 111: Transform the t-th frame hyperspectral image H t Split into a sequence of N bands, the i-th band is recorded as i=1,2,…,N; based on Image entropy Evaluate Information content:
[0062]
[0063] in: Indicates band The probability density function of the band The pixel intensity 255 histogram is normalized to obtain As x, with i as y, Mapped into a two-dimensional coordinate system; Figure 3 As shown in (a), taking the first frame of the ball video in the test set as an example, the image entropy map of each band of the first frame is calculated, and the superscript numbers represent the image entropy.
[0064] Step 112: Use the image entropy of each band to perform density clustering on the bands, select each band as the cluster center in turn, and calculate the local density value and minimum distance
[0065] Local density value Depends on the direct distance between the bands, based on the triangle inequality, the local density value Defined as:
[0066]
[0067] in: express and The Euclidean distance between them; when x<0, When x≥0, d r is the cluster radius, H t The entropy of the N band images is determined by one third of the difference between the maximum and minimum values, and the divisor 3 is the same as the number of channels in RGB; Figure 3 As shown in (b), the entropy of the fourth band is centered and the entropy of d r The circle is drawn with a radius of 0.01, and the points represented by the squares inside the circle belong to the same cluster as the points in the fourth band.
[0068] According to the band Bands with higher entropy than another image The local density value calculation band Image entropy distance band The minimum distance of image entropy
[0069]
[0070] Minimum distance through Characterization band With band relevance; such as Figure 3 As shown in (c), the connecting line in the figure represents the distance between the 4th band and the band with higher local density than it.
[0071] Step 113: Mapped into a two-dimensional coordinate system, such as Figure 3 As shown in (d), in order to maintain the consistency of the three channels of RGB and the basic information flow, we choose The highest three cluster centers are used as representative bands, and the three representative bands are combined into the pseudo-color image P of the tth frame. t , is input into the network as spectral prompt information;
[0072] Step 12: Calculate the false color image F of the t-th frame and the t-1-th frame t 、F t-1 , including the following steps:
[0073] (121) For a hyperspectral image H containing H×W pixels and N bands t 、H t-1 , use color matching function CMFs to transform the hyperspectral image H t 、H t-1 Convert to CIEXYZ color space, expressed as:
[0074]
[0075]
[0076] in: Represents color matching function CMFs, I t , I t-1 is the hyperspectral image H t 、H t-1 False color image rendered in the CIEXYZ color space after color matching functions (CMFs); in this case, the 10-deg XYZCMFs function converted by CIE (2006) was used with a step size of 0.1 nm. The default CIEXYZ color space is RGB.
[0077] Step 122: Using the Monge-Kantorovitch linear color mapping transformation method based on sample color migration, the color closeness between the false color image and the RGB image is enhanced, and the enhanced I t , I t-1 Denoted as F t 、F t-1 ;
[0078] Step 13: F t 、P t 、F t-1 The corresponding tokens are generated by cropping, dicing and embedding and flattening into the latent space. The first generated token is recorded as the initial token, which is represented by
[0079] Step S02: Utilize the dual-stream visual prompter MDVP to fuse multimodal features to generate spectral prompts and spatiotemporal prompts.
[0080] A dual-stream visual prompter MDVP is designed, and MDVP is used to fuse multimodal features to generate spectral prompt stream and spatiotemporal prompt stream, which are then inserted into the classification encoder. Figure 2 As shown in Figure 1, MDVP consists of an initial cross-correlation cue layer (ICPL) and multiple subsequent cross-correlation cue layers (CPL). ICPL consists of a motion feature fusion module (MFFM) (MFFM is used to capture the spatiotemporal characteristics of hyperspectral targets) and two cue generation modules (PGM), which are denoted as the first layer PGM-1 and the first layer PGM-2. The subsequent cross-correlation cue layer (CPL) is denoted as the lth layer CPL. The lth layer CPL consists of two cue generation modules (PGM), which are denoted as the lth layer PGM-1 and the lth layer PGM-2, where l = 2, 3, …, L. The implementation process of MDVP includes the following steps:
[0081] Step 211: Input into ICPL to get the first layer of spatiotemporal prompt flow First layer spectral hint flow Expressed as:
[0082]
[0083] Step 212: In the motion feature fusion module MFFM, first calculate The absolute difference between them is the frame difference Then use two convolutional layers plus a Sigmoid activation function to obtain the initial spatiotemporal hint flow
[0084] Step 213: and Input into the first layer PGM-1 to get the first layer of spatiotemporal prompt flow Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow
[0085] Step 214: Output ICPL and As the input of the second layer CPL, the output of the l-1 layer CPL and As the input of the l-th layer CPL, it is expressed as:
[0086]
[0087] Step 215: In the first layer of CPL, and Input into the first layer PGM-1 to obtain the first layer spatiotemporal prompt stream Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow
[0088] For the sake of convenience, this case will be based on MDVP The generated token information is called basic information flow, based on The generated token information is called spatiotemporal hint stream, based on The generated token information is called spectral hint stream.
[0089] Step S03: The basic information stream, the spectral information stream, and the spatiotemporal information stream are enhanced and integrated through the prompt generation module PGM module of each cross-correlation prompt layer.
[0090] like Figure 4 As shown in Figure 3, the design hint generation module PGM includes three parts: channel attention module CAM, spatial attention module SAM and cross fusion module CFM.
[0091] For the channel attention module CAM, the input token is represented as First, follow the H k ×W k Perform average pooling on the dimension and max pooling Then and The input is fed into a 1×1 convolutional layer for channel dimensionality reduction, and then nonlinear enhancement is performed through a ReLU layer. Then another 1×1 convolutional layer is used and a Sigmoid activation function is added to project the original dimension back to obtain a weight map with the same dimension as the input token C. Finally, the obtained Z channel Multiply by the input token C to get the channel attention token T C =C×Z channel .
[0092] For the spatial attention module SAM, the input token is represented as First, average pooling is performed along the D dimension and max pooling Then along the D dimension and Connect them and input them into a (H k -1)×(W k -1) Convolutional layer and add Sigmoid activation function to get weight mapping Finally, the obtained Z spatial Multiply by the input token S to get the spatial attention token T S =S×Z spatial .
[0093] For the cross-fusion module CFM, the prompt token T is obtained by cross-fusion of input token T1 and input token T2. F , T1 passes through a channel attention module CAM and a spatial attention module SAM in turn to obtain T2 is obtained after passing through a channel attention module CAM and a spatial attention module SAM in sequence. The sum of T1' and T2' is obtained as T'=T1'+T2', and T' is input into two different network branches for operation to obtain and right and Summing to get For T1', T2' and Perform cross-fusion to get prompt token
[0094] Step S04: Construct an L-layer classification encoder.
[0095] Construct an L-layer classification encoder and After element-by-element addition, it is input into the lth layer of the classification encoder to obtain the lth layer basic information flow Recorded as For the L-th layer basic information flow of the L-th layer output of the classification encoder First, channel dimensionality reduction is performed through a linear layer, and then input into the prediction head for classification and bounding box regression to achieve target tracking.
[0096] When constructing the L-layer classification encoder, all network parameters related to the false color image (i.e., the base modality, generally the RGB modality, but in this case, there is no RGB modality, so a false color image generated based on the hyperspectral image is used as the base modality) are frozen, including the network parameters of the tile embedding, feature extraction, feature interaction, and prediction head.
[0097] Step S05: Freeze the backbone network and train the classification encoder using the loss function.
[0098] When using the loss function to train the classification encoder, the HDSP model is first initialized using the parameters in the baseline model OSTrack, and then the cross entropy loss is used as the classification loss L cls , using IoU loss as the bounding box regression loss Then according to the classification loss and bounding box regression loss Calculate total loss λ iou 、 is the regularization parameter, L1 represents the mean absolute error, and finally the total loss is used Tracking training of HDSP model.
[0099] When training the classification encoder, the backbone network of the HDSP model is frozen, and only the prompt part of the HDSP model is trained, and finally a trained model is obtained; the backbone network of the HDSP model includes F t and F t-1 The block embedding part, classification encoder, prediction head, and the prompt part of the HDSP model include P t The block embedding part, MDVP. Using classification loss and bounding box regression loss Calculate total loss Utilize total loss The HDSP model is trained to finally obtain a trained HDSP model.
[0100] Step S06: Use the trained HDSP model to track the target in the video frame to obtain the target tracking result.
[0101] The trained HDSP model is used to track targets in video frames. The following requirements are met: ① The input is two adjacent hyperspectral images in the video frame; ② The Siamese architecture is used for target tracking. The first hyperspectral image frame is initialized as the template frame, and the subsequent hyperspectral images are used as search frames. During the target tracking process, no online template is used and the template frame is not updated; ③ The maximum value of the final classification score map is used as the center position of the tracked target, and network regression is performed to determine the position of the bounding box.
[0102] A hyperspectral target tracking framework based on dual-stream visual cues is proposed to realize the above-mentioned hyperspectral target tracking. The framework takes OSTrack as the baseline model and mainly includes a band selection module (BSM), a dual-stream visual cue (MDVP), an encoder, and a prediction head. The prediction head includes a commonly used classification head and a commonly used regression head.
[0103] The HDSP model is based on OSTrack and its input is the adjacent hyperspectral image H. t 、H t-1 . H is obtained by color matching function CMFs t 、H t-1 The corresponding false color image F t 、F t-1 , through the pseudo color image P t Specifically, first calculate H t Image entropy of each band, The image entropy is Then, As x, with band sequence i as y Mapped into a two-dimensional coordinate system. Next, use the image entropy of each band to perform density clustering on the bands, select each band as the cluster center in turn, and calculate the local density value and minimum distance Finally, Mapped to a two-dimensional coordinate system, select The highest three cluster centers are used as representative bands, and the three representative bands are combined into the pseudo-color image P of the tth frame. t , which is input into the network as spectral prompt information. After preprocessing, three images F t 、P t 、F t-1 ,The three images are then tiled and embedded and flattened into the latent space to generate the corresponding tokens.
[0104] MDVP is used to fuse multimodal features to generate spectral hint stream and spatiotemporal hint stream, and then insert the spectral hint stream and spatiotemporal hint stream into the classification encoder; MDVP consists of an initial cross-correlation hint layer ICPL and multiple subsequent cross-correlation hint layers CPL. Initial Token Enter ICPL to get spectrum prompt information Time and space prompt information Expressed as In order to capture the spatiotemporal characteristics of hyperspectral targets, a motion feature fusion module MFFM is introduced in ICPL. In MFFM, we first calculate The absolute difference between them is the frame difference Then, two convolutional layers plus a Sigmoid activation function are used to obtain the initial spatiotemporal cue information. Then, and Input into the first layer PGM-1 to get the first layer of spatiotemporal prompt flow Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow The spectral prompt information and spatiotemporal prompt information output by the previous layer of CPL are used as the input of the next layer of CPL. The structure of CPL is similar to that of ICPL, but CPL only uses two PGMs to fuse the spectral prompt information. Time and space prompt information and basic information To generate the spectral hint flow of layer l and spatiotemporal cue flow The definition of the first layer CPL is: L is the number of layers of the MDVP.
[0105] The PGM module of the cross-correlation prompt layer is used to enhance and fuse the basic information flow, spectral information flow, and spatiotemporal information flow, including the channel attention module CAM, the spatial attention module SAM, and the cross fusion module CFM. For the channel attention module CAM, the input token is represented as First, follow the H k ×W k Perform average pooling on the dimension and max pooling Then and The input is fed into a 1×1 convolutional layer for channel dimensionality reduction, and then nonlinear enhancement is performed through a ReLU layer. Then another 1×1 convolutional layer is used and a Sigmoid activation function is added to project the original dimension back to obtain a weight map with the same dimension as the input token C. Finally, the obtained Z channel Multiply by the input token C to get the channel attention token T C =C×Z channel For the spatial attention module SAM, the input token is represented as First, average pooling is performed along the D dimension and max pooling Then along the D dimension and Connect them and input them into a (H k -1)×(W k -1) Convolutional layer and add Sigmoid activation function to get weight mapping Finally, the obtained Z spatial Multiply by the input token S to get the spatial attention token T S =S×Z spatial For the cross-fusion module CFM, the prompt token T is obtained by cross-fusion of input token T1 and input token T2. F , T1 passes through a channel attention module CAM and a spatial attention module SAM in turn to obtain T2 is obtained after passing through a channel attention module CAM and a spatial attention module SAM in sequence. The sum of T1' and T2' is obtained as T'=T1'+T2', and T' is input into two different network branches for operation to obtain and right and Summing to get For T1', T2' and Perform cross-fusion to get prompt token
[0106] Next, After element-by-element addition, it is input into the lth layer of the classification encoder to obtain the lth layer basic information flow The results of the final classification encoder layer are processed through a linear layer for channel dimensionality reduction before being fed into the prediction head for classification and bounding box regression to achieve object tracking. During the tracking process, all network parameters related to the false color image are frozen, including those for tile embedding, feature extraction, feature interaction, and the prediction head.
[0107] According to the classification loss L cls , regression loss L iou Calculate total loss is the regularization parameter, L1 represents the mean absolute error, and finally the total loss is used The HDSP model is trained for tracking. During training, the HDSP model's backbone network is frozen, and only the hint portion of the HDSP model is trained, ultimately yielding a trained model. Finally, the final classification score map is calculated, and the maximum value of the classification score map is used as the center position of the tracked target. Network regression is then performed to determine the position of the bounding box.
[0108] To address the limitations of hyperspectral tracking solutions, this paper employs the concept of cue learning, freezing the base model and focusing on learning visual cues from spectral and spatiotemporal modalities, thereby improving the efficiency of hyperspectral target representation and tracking. In recent years, models employing cue learning have emerged in the multimodal field. However, these models utilize only single-stream cue information and do not consider spatiotemporal information, making them unsuitable for multi-band hyperspectral data. To address this issue, this paper proposes an end-to-end dual-stream cue-based hyperspectral target tracking framework that incorporates auxiliary modal inputs of spectral and spatiotemporal information. Spectral information provides signatures of different substances, helping to distinguish targets from complex backgrounds. Spatiotemporal information characterizes changes in object position, enabling the tracking system to quickly and dynamically adapt to changes in the target's appearance. Experiments on a hyperspectral video tracking dataset demonstrate that this approach achieves superior performance compared to popular RGB and hyperspectral trackers.
[0109] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solutions obtained by equivalent replacement or equivalent transformation fall within the scope of protection of the present invention.
Claims
1. A hyperspectral target tracking method based on dual-stream visual cues, characterized by: The network model used to implement this method is called the HDSP model. The training and tracking process of the HDSP model includes the following steps: (1) The t-th frame hyperspectral image and the t-1th frame hyperspectral image As the input of HDSP model; get H through band selection module BSM t The three representative bands are combined into the pseudo color image of the tth frame H t and H t-1 The false color image of frame t is obtained by CIE color matching functions CMFs respectively and the t-1th frame false color image F t 、P t 、F t-1 The corresponding tokens are generated by cropping, dicing and embedding and flattening into the latent space. The first generated token is recorded as the initial token, which is represented by (2) Input into the dual-stream visual prompter MDVP, MDVP consists of an initial cross-correlation prompt layer ICPL and multiple subsequent cross-correlation prompt layers CPL; ICPL consists of a motion feature fusion module MFFM and two prompt generation modules PGM, the two prompt generation modules PGM are respectively recorded as the first layer PGM-1 and the first layer PGM-2; the subsequent cross-correlation prompt layer CPL is recorded as the l-th layer CPL, the l-th layer CPL consists of two prompt generation modules PGM, the two prompt generation modules PGM are respectively recorded as the l-th layer PGM-1 and the l-th layer PGM-2, l = 2, 3, ..., L; the MDVP based on The generated token information is called basic information flow, based on The generated token information is called spatiotemporal hint stream, based on The generated token information is called the spectral cue stream; In MDVP, MFFM will and Fusion to obtain the initial spatiotemporal prompt flow Will and Input into the first layer PGM-1 to get the first layer of spatiotemporal prompt flow Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow The first layer of spatiotemporal prompt stream output by ICPL and the first layer of spectral hint flow As the input of the second layer CPL; In the first layer of CPL, and Input into the first layer PGM-1 to obtain the first layer spatiotemporal prompt stream Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow (3) Construct an L-layer classification encoder, After element-by-element addition, it is input into the lth layer of the classification encoder to obtain the lth layer basic information flow Recorded as For the L-th layer basic information flow of the L-th layer output of the classification encoder First, the linear layer is used to reduce the channel dimension, and then the input is sent to the prediction head for classification and bounding box regression to achieve target tracking; (4) Using classification loss and bounding box regression loss Calculate total loss Utilize total loss Train the HDSP model and finally obtain a trained HDSP model; (5) Use the trained HDSP model to track the target in the video frame to obtain the target tracking result.
2. The hyperspectral target tracking method based on dual-stream visual cues according to claim 1 is characterized in that: In step (1), F is calculated based on the following steps: t 、P t 、F t-1 : (11) Calculate the pseudo-color image P of the tth frame t , including the following steps: (111) The t-th frame hyperspectral image H t Split into a sequence of N bands, the i-th band is recorded as based on Image entropy Evaluate Information content: in: Indicates band The probability density function of As x, with i as y, Mapped into a two-dimensional coordinate system; (112) Use the image entropy of each band to perform density clustering on the bands, select each band as the cluster center in turn, and calculate the local density value and minimum distance Local density value Depends on the direct distance between the bands, based on the triangle inequality, the local density value Defined as: in: express and The Euclidean distance between them; when x<0, When x≥0, d r is the cluster radius, H t The entropy of the N band images is determined by one third of the difference between the maximum and minimum values, and the divisor 3 is the same as the number of channels in RGB; According to the band The band with higher entropy than the other image The local density value calculation band Image entropy distance band The minimum distance of image entropy Minimum distance through Characterization band With band relevance; (113) Mapped to the two-dimensional coordinate system, in order to maintain the consistency of the three channels of RGB and the basic information flow, select The highest three cluster centers are used as representative bands, and the three representative bands are combined into the pseudo-color image P of the tth frame. t ; (12) Calculate the false color image F of the t-th frame and the t-1-th frame t 、F t-1 , including the following steps: (121) For a hyperspectral image H containing H×W pixels and N bands t 、H t-1 , use color matching function CMFs to transform the hyperspectral image H t 、H t-1 Convert to CIEXYZ color space, expressed as: in: Represents color matching function CMFs, I t , I t-1 is the hyperspectral image H t 、H t-1 False color image presented in CIEXYZ color space after color matching function CMFs; (122) The Monge-Kantorovitch linear color mapping transformation method based on sample color migration is used to enhance the color closeness between the false color image and the RGB image. t , I t-1 Denoted as F t 、F t-1 ; (13) To F t 、P t 、F t-1 The corresponding tokens are generated by cropping, dicing and embedding and flattening into the latent space. The first generated token is recorded as the initial token, which is represented by 3. The hyperspectral target tracking method based on dual-stream visual cues according to claim 1 is characterized in that: In the step (2), spectral cue stream and spatiotemporal cue stream are generated by fusing multimodal features using MDVP, and the spectral cue stream and spatiotemporal cue stream are inserted into the classification encoder; MDVP is composed of an initial cross-correlation cue layer ICPL and multiple subsequent cross-correlation cue layers CPL; (21)ICPL consists of a motion feature fusion module MFFM and two prompt generation modules PGM. The two prompt generation modules PGM are respectively denoted as the first layer PGM-1 and the first layer PGM-2; Input into ICPL to get the first layer of spatiotemporal prompt flow First layer spectral hint flow Expressed as: (22) In the motion feature fusion module MFFM, first calculate The absolute difference between them is the frame difference Then use two convolutional layers plus a Sigmoid activation function to obtain the initial spatiotemporal hint flow (23) and Input into the first layer PGM-1 to get the first layer of spatiotemporal prompt flow Will and Input into the first layer PGM-2 to get the first layer spectral prompt flow (24)ICPL output and As the input of the second layer CPL, the output of the l-1 layer CPL and As the input of the l-th layer CPL, it is expressed as 4. The hyperspectral target tracking method based on dual-stream visual cues according to claim 1 is characterized in that: In step (2), the prompt generation module PGM includes three parts: channel attention module CAM, spatial attention module SAM and cross fusion module CFM; For the channel attention module CAM, the input token is represented as First, follow the H k ×W k Perform average pooling on the dimension and max pooling Then and The input is fed into a 1×1 convolutional layer for channel dimensionality reduction, and then nonlinear enhancement is performed through a ReLU layer. Then another 1×1 convolutional layer is used and a Sigmoid activation function is added to project the original dimension back to obtain a weight map with the same dimension as the input token C. Finally, the obtained Z channel Multiply by the input token C to get the channel attention token T C =C×Z channel ; For the spatial attention module SAM, the input token is represented as First, average pooling is performed along the D dimension and max pooling Then along the D dimension and Connect them and input them into a (H k -1)×(W k -1) Convolutional layer and add Sigmoid activation function to get weight mapping Finally, the obtained Z spatial Multiply by the input token S to get the spatial attention token T S =S×Z spatial ; For the cross-fusion module CFM, the prompt token T is obtained by cross-fusion of input token T1 and input token T2. F , T1 passes through a channel attention module CAM and a spatial attention module SAM in turn to obtain T2 is obtained after passing through a channel attention module CAM and a spatial attention module SAM in sequence. The sum of T1' and T2' is obtained as T'=T1'+T2', and T' is input into two different network branches for operation to obtain and right and Summing to get For T1', T2' and Perform cross-fusion to get prompt token 5. The hyperspectral target tracking method based on dual-stream visual cues according to claim 1 is characterized in that: In the step (3), when constructing the L-layer classification encoder, all network parameters related to the false color image are frozen, including the network parameters of the tile embedding, feature extraction, feature interaction and prediction head.
6. The hyperspectral target tracking method based on dual-stream visual cues according to claim 1, characterized in that: In step (4), when using the loss function to train the classification encoder, the HDSP model is first initialized using the parameters in the baseline model OSTrack, and then the cross entropy loss is used as the classification loss L cls , using IoU loss as the bounding box regression loss Then according to the classification loss and bounding box regression loss Calculate total loss λ iou 、 is the regularization parameter, L1 represents the mean absolute error, and finally the total loss is used Tracking training is performed on the HDSP model. During training, the backbone network of the HDSP model is frozen, and only the prompt part of the HDSP model is trained, and finally a trained model is obtained. The backbone network of the HDSP model includes F t and F t-1 The block embedding part, classification encoder, prediction head, and the prompt part of the HDSP model include P t The cut embedded part, MDVP.
7. The hyperspectral target tracking method based on dual-stream visual cues according to claim 1, characterized in that: In step (5), the trained HDSP model is used to track the target in the video frame, and the following requirements are met:
1. The input is two adjacent hyperspectral images in the video frame; ② Using the Siamese architecture for target tracking, the first frame of the input hyperspectral image is initialized as the template frame, and the subsequent frames of hyperspectral images are used as search frames; During the target tracking process, no online template is used and the template frame is not updated; ③ The maximum value of the final classification score map is used as the center position of the tracked target, and network regression is performed to determine the position of the bounding box.
Citation Information
Patent Citations
Unsupervised hyperspectral video target tracking method based on spatial-spectral feature fusion
CN112766102A
Wireless LAN scholastic tracking system
US6633223B1