Small target tracking method for hyperspectral video

By calculating the band contribution scores of the hyperspectral frame and the initial target template, recombining them into a false color image, and using the twin network for feature extraction and template update, the problem of low accuracy of the hyperspectral target tracker in small target tracking is solved, and high-precision tracking and occlusion processing of small targets are achieved.

CN120747756AActive Publication Date: 2025-10-03NORTHWESTERN POLYTECHNICAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511224274.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-03
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing hyperspectral target trackers have low accuracy in tracking small targets and cannot effectively deal with complex background interference and occlusion, making tracking tasks difficult.

Method used

By calculating the band contribution scores of the hyperspectral frame and the initial target template, a false color image is recombined and distorted. The twin network is used to extract features and update the template to solve the difficulties in extracting small target features and occlusion problems.

Benefits of technology

It improves the tracking accuracy and robustness of small targets, can effectively cope with complex backgrounds and occlusions, and improves the perception ability and tracking accuracy of small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747756A_ABST
    Figure CN120747756A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of target tracking, and discloses a hyperspectral video-oriented small target tracking method, which comprises the following steps of: respectively calculating contribution degree scores of hyperspectral frames with consistent waveband information and all wavebands of an initial target template; forming a plurality of false color hyperspectral images and a plurality of false color initial target images; distortion is carried out to obtain a distorted hyperspectral image and a distorted initial target image, and a region of interest is obtained; features of the distorted hyperspectral image and the distorted initial target image are extracted respectively and fused to obtain a feature map, then regression is carried out in a region of interest in the feature map by using a prediction head to obtain a rectangular frame where a small target is located, and if the rectangular frame where the small target is located is not obtained in the region of interest, the small target is located; if yes, performing global search on the hyperspectral frame to obtain a rectangular frame where the small target is located, namely a tracking result of the small target; according to the invention, through band recombination and template updating, the problems that small target features are difficult to extract and are seriously influenced by shielding and other scenes are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target tracking, and in particular relates to a small target tracking method for hyperspectral video. Background Art

[0002] Object tracking is a fundamental task in computer vision, and small object tracking is a challenging and widely used research area within this field, with applications in fields such as remote sensing image processing and autonomous driving. In many real-world scenarios, the performance of many excellent trackers for small object tracking remains low due to factors such as their small size and complex background interference.

[0003] With the rapid development of hyperspectral sensor technology, hyperspectral video (HSV) generated using hyperspectral imaging can highlight target material information in specific spectral channels. The simultaneous spatial, spectral, and temporal information recorded by HSV is particularly useful for tracking small targets. However, due to the limited resolution of small targets and their susceptibility to interference from complex backgrounds, research and development of hyperspectral small target trackers is essential and urgent. While a large number of hyperspectral object trackers (HOTs) have been developed, mainstream HOTs focus on improving tracking accuracy or speed to meet the requirements for fast and accurate target location and tracking. However, the design and optimization of these HOTs often focus on conventionally sized targets, such as large objects or those with distinct features. These conventional targets typically occupy a large area in an image or video frame and have prominent features, making them easier for tracking algorithms to capture and identify. Small targets in complex scenes may be affected by factors such as rapid motion and occlusion, resulting in a variable appearance and limited features, which poses a greater challenge to tracking. In addition, due to the limited size information of small targets, the existing HOT often cannot obtain enough effective information when extracting features, and it is difficult to effectively deal with challenges such as small targets being occluded, resulting in the inability to accurately capture the motion process of small targets. Summary of the Invention

[0004] The purpose of the present invention is to provide a small target tracking method for hyperspectral video to solve the problem that the existing technology cannot accurately track small targets.

[0005] The present invention adopts the following technical solution: a small target tracking method for hyperspectral video, comprising: Step 1: Calculate the contribution scores of each band of the hyperspectral frame and the initial target template with consistent band information; the initial target template is predefined and contains the typical features of the small target tracked in the hyperspectral frame; Step 2: Sort the hyperspectral frame bands according to their contribution scores, and reassemble the three adjacent bands in the hyperspectral frame to form multiple false-color hyperspectral images; similarly, sort the initial target template bands according to their contribution scores, and reassemble the three adjacent bands in the initial target template to form multiple false-color initial target images; Step 3: The small target area in the false color hyperspectral image and the small target area in the false color initial target image are distorted to obtain a distorted hyperspectral image and a distorted initial target image respectively; and the bounding rectangle of the highlighted area in the saliency map of the distorted hyperspectral image is recorded as the region of interest; Step 4: Use the backbone network of the twin network tracker to extract the features of the distorted hyperspectral image and the distorted initial target image respectively and fuse them to obtain a feature map. Then use the prediction head to regress the region of interest in the feature map to obtain the rectangular box where the small target is located, which is the tracking result of the small target.

[0006] The beneficial effects of the present invention are: When tracking small targets in HSV, the present invention can effectively improve the tracking accuracy of small targets by distorting the small target area. Since small targets are small in size, they only occupy a small number of pixels in the image, and their feature information is relatively limited and easily submerged in background noise. Taking into full consideration their spatial-spectral aliasing correlation, the resolution of the small target area is adaptively amplified to enhance the perception ability of small targets, and the target template is updated in a timely manner to improve the tracking accuracy of small targets. The present invention solves the problems of difficulty in extracting small target features and serious influence of scenes such as occlusion through band reorganization and template updating. When tracking small targets, the hyperspectral frame band reorganization can screen out bands with significant differences between small targets and backgrounds, which helps the feature extraction network extract sufficient and efficient discriminant feature information. The template update judges whether the existing template is occluded, so as to reduce problems such as reduced tracking accuracy and tracking trajectory drift caused by occlusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 Comparison results of the accuracy and success rate of the tracking method of the present invention and the existing hyperspectral target tracker on the MSSOT dataset; Figure 2 Comparison results of the accuracy and success rate of the tracking method of the present invention and the visible light tracker on the MSSOT dataset; Figure 3 The trade-off between tracking accuracy and tracking speed of the tracking method of the present invention and the existing hyperspectral target tracker is compared under the MSSOT dataset. Figure 3 (a) is the trade-off between tracking accuracy and tracking speed; Figure 3 (b) is the trade-off between tracking success rate and tracking speed; Figure 4 Visual tracking results of the tracking method of the present invention and the existing hyperspectral target tracker on four video sequences under the MSSOT dataset; Figure 4 (a) is the tracking result of the video sequence named doublecar-9 in the MSSOT dataset; Figure 4 (b) is the tracking result of the video sequence named human-1 in the MSSOT dataset; Figure 4 (c) is the tracking result of the video sequence named motorcycle-6 in the MSSOT dataset; Figure 4 (d) is the tracking result of the video sequence named double-9 in the MSSOT dataset; Figure 5 Visual tracking results of the tracking method of the present invention and the visible light tracker on four video sequences under the MSSOT dataset; Figure 5 (a) is the tracking result of the video sequence named airplane-2 in the MSSOT dataset; Figure 5 (b) is the tracking result of the video sequence named basketball-0 in the MSSOT dataset; Figure 5 (c) is the tracking result of the video sequence named electriccar-6 in the MSSOT dataset; Figure 5 (d) is the tracking result of the video sequence named boat-13 in the MSSOT dataset. DETAILED DESCRIPTION

[0008] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0009] The present invention discloses a small target tracking method for hyperspectral video, which includes four steps.

[0010] Step 1: Calculate the contribution scores of each band of the hyperspectral frame and the initial target template with consistent band information respectively; the initial target template is predefined and contains the typical features of the small target tracked in the hyperspectral frame.

[0011] Among them, a small target refers to a target whose pixel area is much smaller than the size of the video frame (image). Based on the absolute scale definition, it is generally considered that when the target pixel area is less than 32×32 pixels, the target is a small target.

[0012] Assuming that the hyperspectral frame has B bands, the hyperspectral frame input at time t can be , The i-th band in is , , Each band of is composed of W row pixels and H column pixels, that is, .

[0013] Then, given the initial target template The band information is consistent with the hyperspectral frame. The i-th band in is , Each band of is composed of V row pixels and U column pixels, that is, , 1 <V<W,1<U<H。

[0014] Hyperspectral frame For example, calculate The normalized importance score of each band and the normalized score of the difference between the target and the background in each band are added and averaged to get the contribution score of each band.

[0015] Among them, the method for obtaining the normalized importance score of each band is as follows: the initial spectral features of the hyperspectral frame and the initial target template are extracted in sequence through convolution, batch normalization and ReLU activation function, and then the deep spectral features are extracted with the initial spectral features as input. Then, the attention weight of each band is calculated using global average pooling, fully connected layer and softmax, and then the deep spectral features are combined with the corresponding attention weights to obtain the importance score vector and importance ranking matrix, and then the importance normalized score is obtained.

[0016] The band attention mechanism is used to obtain deep spectral features , and by combining the depth spectral characteristics of each band Multiply by the attention weight corresponding to the band , and obtain the enhanced spectral features , and then get the importance score vector of the band , and generate an importance ranking matrix to obtain the normalized importance score of each band.

[0017] First, the hyperspectral frame is calculated by formulas (1) and (2) respectively. and the initial target template Extract the initial spectral features, specifically: (1) (2) in, Hyperspectral frame The initial spectral characteristics of Initial target template The initial spectral characteristics of is the initial spectral feature of the i-th band in the hyperspectral frame, is the initial spectral feature of the i-th band in the initial target template; is the ReLU activation function, is the batch normalization operation, is a convolution operation, and represents a convolution operation with a convolution kernel size of a×a and a step size of b. The convolution kernel is .

[0018] Extract the deep spectrum features. The initial spectrum features obtained in the previous step are calculated using formulas (3) and (4) to complete the effective extraction of the deep spectrum features. Specifically: (3) (4) in, For hyperspectral frames The extracted deep spectral features, For the initial target template Extracted deep spectral features.

[0019] Calculate the attention weight. The attention weight of each band can be calculated using the deep spectrum characteristics of each band. Specifically: (5) (6) in, Hyperspectral frame The attention weight of the i-th band in , ;and Initial target template The attention weight of the i-th band in , ; For use function, is the fully connected layer (Fully Connected Layer), is the global average pooling operation (Global AveragePooling), is the deep spectral feature of the i-th band in the hyperspectral frame, is the deep spectral feature of the i-th band in the initial target template.

[0020] Computing hyperspectral frames The importance score vector of the hyperspectral frame is converted into Deep spectral characteristics of each band in Multiply by the corresponding attention weight , and further obtain the enhanced spectral characteristics, which is recorded as ; According to formula (8), using residual network With hierarchical feature extraction, the final feature map is obtained through the fully connected layers FC1 and FC2, and the global average pooling GAP operation is performed on the final feature map to obtain the hyperspectral frame The global characteristics of each band in the final use The function normalizes the output to a probability distribution to obtain the importance score vector of each band, which is recorded as .

[0021] (7) (8) in, Hyperspectral frame The spectral characteristics after the enhancement of the i-th band, is the importance score vector of the i-th band in the hyperspectral frame. The importance score vectors of each band are arranged in order to form the importance score matrix of all bands. .

[0022] Calculate the initial target template according to formula (9) and formula (10) Importance score vector, steps and methods to calculate hyperspectral frames The importance score vectors are the same.

[0023] (9) (10) in, Initial target template The spectral characteristics after the enhancement of the i-th band, is the importance score vector of the i-th band in the initial target template. The importance score vectors of each band are arranged in order to form the importance score matrix of all bands. .

[0024] in, and There are two fully connected layers. Reduce the feature dimension from B to ,in is the scaling factor, Map the feature dimension back to B, is a residual connection network, is the matrix multiplication operation.

[0025] According to formula (11), the importance score matrix The importance ranking matrix is ​​obtained by matrix multiplication with its transpose ,matrix The main diagonal values ​​are the hyperspectral frames The importance normalized score , specifically: (11) in, Hyperspectral frame Importance ranking matrix of all bands.

[0026] According to formula (12), the importance score matrix The importance ranking matrix is ​​obtained by matrix multiplication with its transpose ,matrix The main diagonal value of is the initial target template The importance normalized score , specifically: (12) in, Initial target template Importance ranking matrix of all bands.

[0027] The method for obtaining the normalized score of the target and background difference is: Firstly, the pinwheel convolution is used to extract the initial spectral features of the hyperspectral frame and the initial target template and batch normalization is performed. Secondly, the multi-scale feature extraction branch is used to realize the complementary splicing of feature maps. Then, the fully connected layer is used to compress and restore the channel dimension, and the global average pooling is combined to obtain the normalized difference score between the target and the background.

[0028] The target-background separation mechanism is mainly used to measure the difference between small targets and background in each band of the hyperspectral frame, and give higher scores to the bands that can highlight the small target area. Among them, the target-background separation mechanism extracts the hyperspectral frame through formula (13) and formula (14) respectively. and the initial target template The initial spectral characteristics of: (13) (14) in, Hyperspectral frame The initial spectral characteristics of Initial target template The initial spectral characteristics of is the initial spectral feature of the i-th band of the hyperspectral frame, is the initial spectral feature of the i-th band of the initial target template, This is the windmill convolution.

[0029] The obtained initial spectral features are complementary spliced ​​through the multi-scale feature extraction branch, in which the hyperspectral frame The calculation method of the initial spectral characteristics of the initial target template is as follows: The calculation method of the initial spectral characteristics is formula (16): (15) (16) in, For hyperspectral frames The concatenated feature map, For the initial target template The concatenated feature map, For splicing operation.

[0030] Calculate the hyperspectral frame using formula (17) The normalized score of the target-background difference , use formula (18) to calculate the initial target template The normalized score of the target-background difference , first reduce the number of channels to 1 / B of the original through the first fully connected layer, denoted as , and then go through and ReLU activation function and then pass through the second fully connected layer Restore the number of channels to their original size, recorded as , and finally through The normalized score of the target and background difference is obtained after the function : (17) (18) in, Hyperspectral frame The normalized score of the target and background difference, Initial target template The normalized score of the target-background difference.

[0031] Step 2: Sort the hyperspectral frame bands according to their contribution scores, and reassemble the three adjacent bands in the hyperspectral frame to form multiple false-color hyperspectral images; similarly, sort the initial target template bands according to their contribution scores, and reassemble the three adjacent bands in the initial target template to form multiple false-color initial target images.

[0032] Calculate band contribution scores to hyperspectral frames For example, the importance of each band is normalized to score Normalized score of target-background difference Add them and then use Normalize to get the final contribution score of each band , calculated according to formula (19), and for the initial target template Calculate according to formula (20): (19) (20) in, Hyperspectral frame The contribution score of the i-th band, is the importance normalized score of the i-th band of the hyperspectral frame, is the normalized score of the target-background difference in the i-th band of the hyperspectral frame, For addition operation, use The function output is normalized to the value range [0,1].

[0033] in, Initial target template The contribution score of the i-th band, is the normalized importance score of the i-th band of the initial target template, is the normalized score of the difference between the target and the background in the i-th band of the initial target template.

[0034] Hyperspectral frame For example, all the bands in the frame are re-ranked from high to low according to the band contribution scores, and the ranking is recorded as , synthesize false color hyperspectral images according to every three adjacent bands of the newly sorted bands, and finally obtain N false color hyperspectral images. Then for the initial target template Then we get N false color initial target images.

[0035] Bands with higher contribution scores are key bands and are considered to have a favorable impact on subsequent tracking, while bands with lower contribution scores are considered to have an unfavorable impact on subsequent tracking. All bands in the image are reordered in descending order according to the contribution scores to obtain a new band order, and N false color hyperspectral images are obtained by recombining three adjacent bands according to the new band order; the initial target template is Recombination operation is performed on the hyperspectral frame The operation is the same as

[0036] Since the hyperspectral frame With the initial target template The band information in the hyperspectral frame is consistent, so the band contribution score ranking Band contribution score ranking of the initial target template are consistent, that is, the band contribution score rankings of the two are the same, therefore, the N false color hyperspectral images after the hyperspectral frame band recombination N false color initial target images recombined with the initial target template In the figure, when N false color hyperspectral images and N false color initial target images are arranged in order from 1 to N, the three band information contained in each false color hyperspectral image and the false color initial target image are consistent.

[0037] Every three adjacent bands can form a false color hyperspectral image , Every three adjacent bands can form a false color initial target image , so that a total of N false color hyperspectral images and N false color initial target images can be obtained, where When B is an integer multiple of 3, N is the quotient of B divided by 3; otherwise, N is the result of rounding down the quotient of B divided by 3, that is, .

[0038] Step 3: The small target area in the false color hyperspectral image and the small target area in the false color initial target image are distorted to obtain a distorted hyperspectral image and a distorted initial target image respectively; and the bounding rectangle of the highlighted area in the saliency map of the distorted hyperspectral image is recorded as the region of interest.

[0039] Specifically, the distortion is to elastically enlarge the small target area of ​​the input false color hyperspectral image and the false color initial target image while keeping the resolution unchanged. The initial target template is for the small target tracked in the hyperspectral frame. The small target area is obtained by calculating the area in the hyperspectral frame with high similarity to the target in the initial target template. The small target area is the area with a high probability of containing a small target.

[0040] The input image is distorted using the saliency map with the forward mapping, and the output is the distorted image, specifically: (twenty one) in, is the mapping relationship of image distortion, and are the input and output coordinates for reading the image.

[0041] By transforming the false color hyperspectral image and the false color initial target image Resample pixel values ​​and output a distorted hyperspectral image and distort the initial target image , specifically: (twenty two) Image distortion is achieved by backward mapping, where at each output pixel grid position, the inverse mapping is calculated. To find the corresponding input pixel coordinates, use bilinear interpolation to calculate the color depth of adjacent input pixel grid points, and assign the calculated color depth to the output pixel, specifically: (twenty three) The complete process of small target area resolution enlargement operation It can be written as: (twenty four) in, is a non-linear function that returns the bounding box coordinates of the predicted detection, The information will be used to regress to the target position in the original space label.

[0042] Step 4: Use the backbone network of the twin network tracker to extract the features of the distorted hyperspectral image and the distorted initial target image respectively and fuse them to obtain a feature map. Then use the prediction head to regress the rectangular box where the small target is located in the region of interest in the feature map. If the rectangular box where the small target is located is not obtained in the region of interest, a global search is performed on the hyperspectral frame to obtain the rectangular box where the small target is located, which is the tracking result of the small target.

[0043] The present invention adopts a multi-stream Siamese network structure, in which feature extraction and feature fusion as well as target position regression prediction heads all come from a Siamese network tracker pre-trained on a large-scale visible light dataset.

[0044] The present invention uses multiple identical backbone networks to perform feature extraction and fusion operations on the output N distorted initial target images and N distorted hyperspectral images, respectively, to obtain N feature maps. , each set of feature maps is ,use The operation splices these feature maps and feeds them into the prediction head, which reads the position information of the region of interest and performs foreground-background regression, center regression and target box regression operations in the region of interest to obtain the input hyperspectral frame at time t Tracking results , the result will be used for subsequent template updates.

[0045] Template updating includes the occlusion judgment stage and the template update stage. In the occlusion judgment stage, the target prior (π) and the context information of the region of interest are used to generate the input hyperspectral frame at time t. Tracking results The saliency map of the initial target template without occlusion is calculated using the aggregate signature, and the aggregate signature feature saliency map is generated by The saliency map is compared with the aggregate signature feature saliency map to generate a response map, and the threshold judgment method is used to detect whether occlusion occurs.

[0046] In the template update phase, if occlusion is detected, the previous template is used; if no occlusion occurs, the previous template is used. Replace the previous dynamic template as a new one.

[0047] calculate The method of the saliency map is: (25) in, 、 、 is the integrated context information of three different channels of N distorted hyperspectral images, 、 、 is the weight coefficient, and π is the two-dimensional prior related to the tracking target.

[0048] The calculation of the aggregate signature feature saliency map is mainly used to obtain the saliency map of the target without occlusion in the initial target template, and further locate the target. The aggregate signature is gradually smoothed by the Gaussian kernel to obtain the aggregate signature feature saliency map.

[0049] In determining the occurrence of occlusion, the present invention uses a threshold-based method to trigger template update, as shown in formula (26): (26) in, is the aggregate signature feature saliency map response using the initial target template, The saliency map response of is the threshold, is the tracking result of the hyperspectral frame at time t.

[0050] pass The response and the initial target template response generate a response map, which shows the difference between the saliency maps. If the response is far away from the initial target template response, the target is considered to be blocked or out of view. When the response difference exceeds the threshold, the template update mechanism is triggered.

[0051] Example: This example uses the MSSOT dataset to illustrate the implementation process. The MSSOT dataset is constructed using a snapshot spectral camera in 5×5 SFA mode, sampling a spectral band of 680–960 nm, with a frame rate of 25 frames per second and 25 bands. The hyperspectral frames in this dataset have 2045 rows and 1080 columns. An adaptive scaling method is used to reduce the row and column pixels of each band in the input hyperspectral frame to 320 pixels, i.e., W = H = 320. The row and column pixels of each band in the input initial target template are reduced to 128 pixels, i.e., V = U = 128.

[0052] The MSSOT dataset contains 185 training videos and 40 test videos, and includes nine challenge attributes: small target challenge SO, illumination change challenge IV, on-chip rotation challenge OR, shape change challenge SV, occlusion challenge OCC, deformation challenge DEF, motion blur challenge MB, camera motion challenge CM, and background clutter challenge BC.

[0053] The MSSOT data set comes from L. Chen, Y. Zhao, and SG Kong. "SFA-guided mosaictransformer for tracking small objects in snapshot spectral imaging." ISPRSJournal of Photogrammetry and Remote Sensing, vol. 204, pp. 223-236, 2023.

[0054] Hyperspectral frames With the initial target template After input, the band attention mechanism and the target background separation mechanism are respectively used to obtain the normalized importance score of each band and the normalized difference score between the target and the background of each band. The normalized importance score of each band and the normalized difference score between the target and the background of each band are added and averaged to obtain the contribution score of each band, and then the contribution score ranking is obtained. The hyperspectral frame is ranked according to the contribution score ranking of each band. With the initial target template The bands in the list are reordered, and the bands with higher contribution scores are ranked higher, otherwise they are ranked lower.

[0055] The convolution operations used in formulas (1), (2), (7) and (9) In the example, a and b are both set to 1, and the is a convolution kernel of size 1 × 1. The convolution operation used in formulas (3), (4), (15) and (16) is In the example, a is set to 3, b is set to 1, and the is a convolution kernel of size 3×3 .

[0056] After calculation, the hyperspectral frame in the MSSOT dataset HSV is divided into 8 false color hyperspectral images, and the initial target template Divided into 8 false color initial target images.

[0057] The small target areas in the 8 false color hyperspectral images and the small target areas in the 8 false color initial target images are distorted respectively. The distortion operation of the small target area is performed according to formula (24). The prior knowledge of this operation comes from the small target position information in the initial target template. The distorted image is obtained through the distortion operation. The inverse operation of is regressed to the target position in the original spatial label, and the saliency map of the distorted hyperspectral image after resolution amplification of the small target area is calculated by the Gaussian kernel smoothing filter method. The saliency map contains the position information of the small target area with high confidence. A rectangular region of interest is divided according to the length and width of the saliency map. The information of the region of interest will be transmitted to the target position regression prediction head, and the position regression of the small target will focus on this area.

[0058] This embodiment uses a twin feature extraction and fusion backbone network trained on a large-scale visible light dataset and a target position regression prediction head to perform feature extraction and fusion as well as small target position information regression operations. Feature extraction and fusion use the ResNet-50 network, and the target position regression prediction head reads the position information of the region of interest and performs foreground-background regression, center regression, and target box regression operations in the region of interest to obtain small target tracking results.

[0059] According to formula (25), the hyperspectral frame input at time t is obtained Tracking results The significance map of 、 、 are all set to 0.25, tracking the results The saliency map of is compared with the saliency map of the aggregate signature feature to generate a response map, and formula (26) is used to judge Whether the target in is occluded, the threshold When the response difference exceeds the threshold, it is considered that occlusion has occurred, triggering the template update mechanism and continuing to use the previous template. Otherwise, use Replace the previous template.

[0060] This embodiment selects 11 existing HOT and 7 visible light trackers to verify the effectiveness of this embodiment. Among them, the 11 existing HOT are MHT, MFI-HVT, DeepHKCF, BAE-Net, SiamBAG, SiamOHOT, Trans-DAT, MMF-Net, PHTrack, SSTtrack and HDSP. MHT uses hand-crafted features for tracking, while other trackers use deep features for tracking. The manual feature-based tracker is run on a machine equipped with an Intel Core i9-14900HX CPU@2.20 GHz and 32 GB RAM; while the deep feature-based tracker is run on another machine equipped with the same CPU but with an NVIDIA RTX 4060 GPU added.

[0061] The seven visible light trackers selected in this example are SiamCAR, SiamBAN, Stark, TransT, OSTrack, SeqTrack, and AQATrack. Precision and success rate graphs are used to evaluate tracking accuracy. The precision graph shows the accuracy of the trackers with a center position error below a specified 20 pixels and is used to rank the trackers. The success rate graph describes the number of successful frames and the overlap ratio between the predicted bounding box and the true value. The true value exceeds a given threshold, ranging from 0 to 1. This example also uses FPS and FLOPs to compare the real-time performance and computational complexity of the trackers. Specific results are shown in Tables 1 and 2.

[0062] Table 1 AUC comparison results between this example and HOT under different attributes of the MSSOT dataset

[0063] Table 2 AUC comparison results between this embodiment and the visible light tracker under different attributes of the MSSOT dataset

[0064] In Tables 1 and 2, SO stands for small target challenge; IV stands for illumination change challenge; OR stands for on-chip rotation challenge; SV stands for shape change challenge; OCC stands for occlusion challenge; DEF stands for deformation challenge; MB stands for motion blur challenge; CM stands for camera motion challenge; and BC stands for background clutter challenge.

[0065] like Figure 1 and Figure 2 As shown in Figure 2, compared with the existing 11 HOT and 7 visible light trackers, this embodiment has a leading tracking accuracy for small targets, as shown in Figure 2. Figure 3 As shown, this embodiment provides a good balance between tracking accuracy and tracking speed. In addition, according to the results in Tables 1 and 2, this embodiment performs best when dealing with challenging attributes such as occlusion, verifying the effectiveness of the tracking method of this embodiment.

[0066] like Figure 4 and Figure 5 As shown, interference from occlusion and target deformation poses a significant challenge to trackers, preventing most trackers from successfully tracking the target. However, this embodiment effectively eliminates interference such as occlusion and can effectively regress the position of small targets even at very low target resolution. Overall, the method of this embodiment achieves good tracking results on the MSSOT dataset, validating the effectiveness of the present invention.

[0067] Among them, the 11 types of HOT information are: MMF-Net: Z. Li, F. Xiong, J. Zhou, J. Lu, Z. Zhao, and Y. Qian. "Material-guided multiview fusion network for hyperspectral object tracking". IEEE Transactions on Geoscience and Remote Sensing., vol. 62, pp. 1-15, 2024, Art no. 5509415. SSTtrack: Y. Chen, Q. Yuan, Y. Tang, Y. Xiao, J. He, T. Han, Z. Liu, and L. Zhang. "SSTtrack: A unified hyperspectral video tracking framework via modeling spectral-spatial-temporal conditions". Information Fusion., vol.114, p. 102658, 2025. PHTrack: Y. Chen, Y. Tang, X. Su, J. Li, Y. Xiao, J. He, and Q. Yuan."PHTrack: Prompting for hyperspectral video tracking". IEEE Transactions onGeoscience and Remote Sensing., vol. 62, pp. 1-18, 2024, Art no. 5533918. Trans-DAT: Y. Wu, L. Jiao, X. Liu, F. Liu, S. Yang, and L. Li. "Domain adaptation-aware transformer for hyperspectral object tracking". IEEETransactions on Circuits and Systems for Video Technology., vol. 34, pp.8041-8052, 2024. SiamOHOT: C. Sun, X. Wang, Z. Liu, Y. Wan, L. Zhang, and Y. Zhong. "SiamOHOT: A lightweight dual Siamese network for onboard hyperspectral objecttracking via joint spatial-spectral knowledge distillation". IEEETransactions on Geoscience and Remote Sensing., vol. 61, pp. 1-12, 2023, Artno. 5521112. HDSP: R. Yao, L. Zhang, Y. Zhou, H. Zhu, J. Zhao, and Z. Shao. "Hyperspectral object tracking with dual-stream prompt". IEEE Transactions onGeoscience and Remote Sensing., vol. 63, pp. 1-12, 2025, Art no. 5500612. SiamBAG: W. Li, Z. F. Hou, J. Zhou, and R. Tao. "SiamBAG: Bandattention grouping-based Siamese object tracking network for hyperspectralvideos". IEEE Transactions on Geoscience and Remote Sensing., vol. 61, 2023,Art no. 5514712. BAE-Net: Z. Li, F. Xiong, J. Zhou, J. Wang, J. Lu, and Y. Qian. "BAE-Net: A band attention aware ensemble network for hyperspectral objecttracking". in Proceedings of the IEEE International Conference on ImageProcessing. (ICIP), 2020, pp. 2106-2110. MHT: F. Xiong, J. Zhou, and Y. Qian. "Material based object trackingin hyperspectral videos". IEEE Transactions on Image Processing., vol. 29,pp. 3719-3733, 2020. MFI-HVT: Z. Zhang, K. Qian, J. Du, and H. Zhou. "Multi-featuresintegration based hyperspectral videos ttracker". in Proceedings of the 11thWorkshop on Hyperspectral Imaging and Signal Processing: Evolution in RemoteSensing (WHISPERS), 2021, pp. 1-5. DeepHKCF: B. Uzkent, A. Rangnekar, and M. J. Hoffman. "Tracking in aerial hyperspectral videos using deep kernelized correlation filters". IEEE Transactions on Geoscience and Remote Sensing., vol. 57, no. 1, pp. 449-461, 2019. Among them, the information of the 7 visible light trackers is as follows: AQATrack: J. Xie, B. Zhong Z. Mo, S. Zhang, L. Shi, S. Song, and R. Ji. "Autoregressive queries for adaptive tracking with spatio-temporal transformers". in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. (CVPR), 2024, pp. 14572-14581. Stark: B. Yan, H. Peng, J. Fu, D. Wang, and H. Lu. "Learning spatio-temporal transformer for visual tracking". in Proceedings of the International Conference on Computer Vision. (ICCV), 2021, pp. 10428-10437. OSTrack: B. Ye, H. Chang, B. Ma, S. Shan, and X. Chen. "Joint feature learning and relation modeling for tracking: A one-stream framework". in Proceedings of the European Conference on Computer Vision. (ECCV), 2022, pp. 341-357. SeqTrack: X. Chen, H. Peng, D. Wang, H. Lu, and H. Hu. "SeqTrack:Sequence to sequence learning for visual object tracking". in Proceedings ofthe IEEE Conference on Computer Vision and Pattern Recognition. (CVPR), 2023,pp. 14572-14581. TransT: X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu. "Transformer tracking". in Proceedings of the IEEE Conference on ComputerVision and Pattern Recognition. (CVPR), 2021, pp. 8126-8135. SiamCAR: D. Guo, J. Wang, Y. Cui, Z. Wang, and S. Chen. "SiamCAR:Siamese fully convolutional classification and regression for visualtracking". in Proceedings of the IEEE Conference on Computer Vision andPattern Recognition. (CVPR), 2020, pp. 6268-6276. SiamBAN: Z. Chen, B. Zhong, G. Li, S. Zhang, R. Ji, and Jeee. "Siamese box adaptive network for visual tracking". in Proceedings of the IEEEConference on Computer Vision and Pattern Recognition. (CVPR), 2020, pp.6667-6676。

Claims

1. A small target tracking method for hyperspectral video, characterized in that: include: Step 1: Calculate the contribution scores of each band of the hyperspectral frame and the initial target template with consistent band information respectively; The initial target template is predefined and contains typical features of the small target tracked in the hyperspectral frame; Step 2: sorting the hyperspectral frame according to the contribution score of each band, and recombining the three adjacent bands in the hyperspectral frame to form a plurality of false color hyperspectral images; similarly sorting the initial target template according to the contribution score of each band, and recombining the three adjacent bands in the initial target template to form a plurality of false color initial target images; Step 3: distorting the small target area in the false color hyperspectral image and the small target area in the false color initial target image to obtain a distorted hyperspectral image and a distorted initial target image respectively; and recording the circumscribed rectangle of the highlighted area in the saliency map of the distorted hyperspectral image as a region of interest; Step 4: Use the backbone network of the twin network tracker to extract the features of the distorted hyperspectral image and the distorted initial target image respectively and fuse them to obtain a feature map. Then use the prediction head to regress the region of interest in the feature map to obtain the rectangular box where the small target is located, which is the tracking result of the small target.

2. The small target tracking method for hyperspectral video according to claim 1, characterized in that: The contribution score in step 1 is calculated by adding the normalized importance score of each band and the normalized score of the target-background difference, and using softmax normalization to obtain the contribution score of each band.

3. The small target tracking method for hyperspectral video according to claim 2, characterized in that: The method for obtaining the importance normalized score is: The initial spectral features of the hyperspectral frame and the initial target template are extracted in sequence through convolution, batch normalization and ReLU activation function. The deep spectral features are then extracted using the initial spectral features as input. The attention weights of each band are then calculated using global average pooling, fully connected layers and softmax. The deep spectral features are then combined with the corresponding attention weights to obtain the importance score vector and importance ranking matrix, and finally the importance normalized score is obtained.

4. The small target tracking method for hyperspectral video according to claim 2, characterized in that: The method for obtaining the normalized score of the target and background difference is: Firstly, the pinwheel convolution is used to extract the initial spectral features of the hyperspectral frame and the initial target template and batch normalization is performed. Secondly, the complementary splicing of feature maps is realized through the multi-scale feature extraction branch. Then, the fully connected layer is used to compress and restore the channel dimension, and the global average pooling is combined to obtain the normalized difference score between the target and the background.

5. The small target tracking method for hyperspectral video according to claim 1, characterized in that: In step 4, the twin network tracker uses a threshold-based method to trigger template update, and the threshold method is: , in, is the aggregate signature feature saliency map response using the initial target template, for The saliency map response of is the threshold, is the tracking result of the hyperspectral frame at time t.

Citation Information

Patent Citations

  • Small target recognition precision optimization method based on local difference analysis

    CN114462542A

  • Target tracking method for hyperspectral video

    CN119313710A

  • Hyperspectral image band selection method and system based on latent feature fusion

    US20240212307A1