Hyperspectral video target tracking method based on background spectrum capsule characteristics

Through multi-frame spectral curve analysis and capsule network modeling, background spectral capsule features are constructed and background similar routing algorithms are combined to solve the problem of limited performance when the target is similar to the background spectral features of the existing hyperspectral video, and effective distinction between complex backgrounds and robust tracking of targets is achieved.

CN120147422APending Publication Date: 2025-06-13WUXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510257496.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing hyperspectral video target tracking methods are limited in performance when the target is similar to the background spectrum features, and deep learning methods are prone to tracking loss when the lighting conditions fluctuate violently.

Method used

Through multi-frame spectral curve analysis and capsule network modeling, background spectral capsule characteristics are constructed, and target response maps are generated based on background similar routing algorithms, search area positioning is optimized, and spectral feature differences and background context information are fused.

Benefits of technology

Effectively overcome background clutter interference, distinguish target from complex background, and realize robust tracking between continuous frames, improving the accuracy and stability of target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147422A_ABST
    Figure CN120147422A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral video target tracking method based on background spectrum capsule characteristics, and the method comprises the steps: obtaining a spectrum curve of a local region of a first frame of hyperspectral image, segmenting the first frame of hyperspectral image into a target region and a background region, importing the (t-1) th frame, the (t-3) th frame and the (t-5) th frame, obtaining a spectrum curve of each pixel point in a search region, and carrying out the segmentation of the target region and the background region; segmenting background regions of (t-1) th, (t-3) th and (t-5) th frames according to the first frame of target region, sending the background regions, the tth frame and the (t-1) th frame of hyperspectral image into a capsule network to extract background capsule features, and matching the background capsule features with the capsule features of the search region to obtain background spectrum capsule features; and converting the background response diagram generated by the network to obtain a target response diagram, determining the positions of the tth frame and the (t-1) th frame of hyperspectral image prediction frames, obtaining a tracked target, positioning the position of a next frame of search area through the positions of the tth frame and the (t-1) th frame of hyperspectral image prediction frames and a target spectrum curve, and carrying out subsequent tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a hyperspectral video object tracking method based on background spectral capsule features. Background Art

[0002] In the forefront of computer vision, object tracking technology, as an important research direction, has been applied on a large scale in key fields such as intelligent transportation systems, security monitoring, and intelligent robots. The current mainstream tracking methods are mainly designed for visible light video sequences. However, in the case where the spectral features of the object and the background are similar, the performance of traditional visible light tracking models is significantly limited. Compared with the traditional visible light modality, hyperspectral imaging technology can effectively analyze the material feature differences by capturing the spectral response curves of continuous narrow bands, providing a new technical path to solve the above problems, which also promotes the extension of visible light tracking methods to the hyperspectral domain and becomes an important current research trend. The existing tracking method systems mainly cover the following two types of technical paradigms:

[0003] In the tracking methods based on correlation filtering, a typical representative is the context-aware correlation filtering framework. This method extracts the histogram of oriented gradients features and uses the context area around the object as negative samples to participate in the filter training to improve the positioning accuracy. However, in complex scenes, the spatial resolution of the histogram of oriented gradients features is insufficient, and it is difficult to effectively distinguish the object from the background area with similar textures. Especially when the illumination changes suddenly, due to the lack of a systematic sampling strategy for the background area, the uniform suppression weight mechanism it adopts easily leads to random suppression of background clutter by the model, thus causing tracking drift problems.

[0004] For the tracking methods based on deep learning, the fully convolutional Siamese network architecture is a typical implementation scheme. This model constructs a feature matching mechanism between the object template and the search area to locate the object position with the maximum response value. However, such methods have two significant defects: First, they do not effectively fuse the spectral dimension discrimination information, and when the object and the background exhibit highly similar spectral reflection characteristics, their discrimination ability drops sharply; Second, the fixed receptive field design is difficult to adapt to the object scale change, especially when the illumination conditions fluctuate violently, it is easy to cause tracking loss.

[0005] The method proposed by the present invention constructs background spectral capsule features through multi-frame spectral curve analysis and capsule network modeling, combines the background similarity routing algorithm to generate an object response map and optimize the search area positioning. By fusing the spectral feature differences and background context information, it effectively overcomes the interference of background clutter, distinguishes the object from the complex background, and finally realizes robust tracking between consecutive frames based on the dynamic prediction box. Summary of the Invention

[0006] In view of this, the main object of the present invention is to provide a hyperspectral video target tracking method based on background spectral capsule features.

[0007] To achieve the above object, the technical solution of the present invention is realized as follows:

[0008] An embodiment of the present invention provides a hyperspectral video target tracking method based on background spectral capsule features, and the method is as follows:

[0009] Step 1: Load the target position, target box, and target image block of the first hyperspectral image in the hyperspectral image sequence, and load the (t - 1)-th, (t - 3)-th, and (t - 5)-th hyperspectral images in the hyperspectral image sequence; determine the search regions S t-1 , S t-3 , S t-5 ;

[0010] Step 2: Classify the region within the target box of the first hyperspectral image to obtain the target spectral curve C t of the first hyperspectral image, and determine that the global spectral curves of the search regions S t-1 , S t-3 , S t-5 of the (t - 1)-th, (t - 3)-th, and (t - 5)-th hyperspectral images are respectively

[0011] Step 3: Determine that the spectral angle distances of the search regions S , S t-1 , S t-3 , S t-5 are respectively D t-1 , D t-3 , D t-5 ;

[0012] Step 4: Generate background masks for the (t - 1)-th, (t - 3)-th, and (t - 5)-th hyperspectral images through D t-1 , D t-3 , D t-5 , and multiply the (t - 1)-th, (t - 3)-th, and (t - 5)-th hyperspectral images by the background masks of the (t - 1)-th, (t - 3)-th, and (t - 5)-th hyperspectral images to obtain the background regions of the first, (t - 1)-th, (t - 3)-th, and (t - 5)-th hyperspectral images as the background pool;

[0013] Step 5: Determine the capsule feature u t of the current t-th hyperspectral image and the background capsule feature u b of the background pool, and obtain the new capsule feature

[0014] Step 6: Obtain the background spectral capsule feature E of the t-th frame hyperspectral image through the background similarity routing algorithm, and obtain the background response map R of the search area S of the t-th frame through the capsule network decoder t of the background response map R b ;

[0015] Step 7: Determine the target response map R through the background response map R b and determine the target position p through the target response map R t ; t Determine the target position p t ;

[0016] Step 8: Repeat Steps 5 to 7 for the (t - 1)-th frame hyperspectral image to obtain the target position p of the (t - 1)-th frame hyperspectral image t-1 ;

[0017] Step 9: Determine the target spectral curves C t-1 and C t of the (t - 1)-th and t-th frame hyperspectral images after normalization, and determine the spectral curve difference f according to C t-1 and C t ;

[0018] Step 10: Determine the template update U through the spectral curve difference f t→t-1 ;

[0019] Step 11: Determine the predicted position P through the template update U t→t-1 ; pre ;

[0020] Step 12: Load each frame of the hyperspectral image sequence in turn, and repeat Steps 1 to 11 to obtain the target position of each frame of the hyperspectral image, and complete the target tracking of the hyperspectral image sequence

[0021] In the above solution, Step 5 is specifically implemented through the following steps:

[0022] (501) Use the capsule network to extract the capsule feature u t of the current t-th frame hyperspectral image and the background capsule feature u b of the background pool;

[0023] (502) Calculate the similarity t between u b and u as

[0024]

[0025] where represents the i-th background capsule in u b and represents ut The i-th feature capsule in, ||·|| represents the magnitude of the vector, represents the capsule feature u t and the background capsule feature u b The similarity between them ranges from [-1, 1]. A value of 1 means the directions of the two capsules are exactly the same, a value of -1 means the directions are exactly opposite, and 0 means the two capsules are orthogonal;

[0026] (503) Calculate the new capsule feature according to the following formula as

[0027]

[0028] where, represents the new capsule feature The i-th feature capsule in.

[0029] In the above solution, step six is specifically implemented through the following steps:

[0030] (601) Calculate the i-th background capsule according to the following formula and the i-th feature capsule The weight coefficient between them is;

[0031]

[0032] where, when is not 1, it means that the current feature capsule is different from the background capsule and its weight coefficient needs to be reduced;

[0033] (602) Calculate the predicted output between the background capsule feature and the new capsule feature according to the following formula as

[0034]

[0035] where, represents the predicted output of the i-th capsule of the capsule feature to the i-th background capsule in the t-th frame, represents the newly generated feature capsule updated after similarity processing, W tb Contains the transformation of the learning relationship of the i-th feature capsule in the t-th frame to the i-th background capsule;

[0036] (603) Update the initial weight e according to the following formula t|b as

[0037]

[0038] where, v j represents the capsule vector obtained by compressing with the squash function, which is the output of the next layer of capsule j;

[0039] (604) Calculate the coupling coefficient according to the following formula where

[0040]

[0041] where is the normalized weight, i.e., the coupling coefficient, k is the number of initial weights, exp(·) represents the exponential function, and ∑(·) represents the summation operation;

[0042] (605) Calculate the total input o of the next-layer capsule j according to the following formula j

[0043]

[0044] (606) Calculate the capsule vector v according to the following formula j where

[0045]

[0046] where v j represents the output of the lower-layer capsule j, and o j represents the total input of the lower-layer capsule j. When ||o j || → 0, v j → 0; when ||o j || → 0, v j → 1. squash(·) represents the squash function, which can compress the capsule length and keep the vector direction unchanged;

[0047] (607) Integrate the vector v j output by the final layer into capsule features E for expression, and amplify the features E through three deconvolution layers in the capsule network decoder to obtain a background response map R t with the same size as the search region S b of the t-th frame.

[0048] In the above solution, in step nine, determine f according to the following formula

[0049]

[0050] where f represents the spectral curve difference, B represents the number of bands of the hyperspectral image, C t represents the spectral curve of the target region of the t-th frame of the hyperspectral image after normalization, and C t-1 represents the spectral curve of the target region of the (t - 1)-th frame of the hyperspectral image after normalization.

[0051] In the above solution, in step ten, determine U according to the following formula t→t-1 where

[0052]

[0053] Among them, U t→t-1 represents template update, and ε represents the spectral positive number; when f < ε, the objects in the t-th frame and the (t - 1)-th frame are similar, and it is determined as the same object for the next position prediction; when f ≥ ε, the objects in the t-th frame and the (t - 1)-th frame are not similar, and the network is updated.

[0054] In the above solution, in step eleven, P is determined according to the following formula pre is

[0055] P pre = p t + Δp

[0056] = p t + p t - p t-1

[0057] = (x t , y t ) + (x t - x t-1 , y t - y t-1 )

[0058] Among them, P pre represents the predicted position of the (t + 1)-th frame, p t represents the position of the target center in the t-th frame, Δp represents the position difference, p t-1 represents the position of the target center in the (t - 1)-th frame, x t represents the abscissa of the target in the t-th frame, y t represents the ordinate of the target in the t-th frame, x t-1 represents the abscissa of the target in the (t - 1)-th frame, y t-1 represents the ordinate of the target in the (t - 1)-th frame.

[0059] Compared with the prior art, the present invention analyzes multi-frame spectral curves and models with capsule networks, fuses background spectral capsule features and capsule features, combines spectral feature differences and background context information, effectively overcomes background clutter interference, and distinguishes targets from complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flowchart of the present invention;

[0061] Figure 2 is the first-frame hyperspectral image of a book after normalization in an embodiment of the present invention;

[0062] Figure 3In the embodiment of the present invention, the 61st frame of the normalized hyperspectral image of the book of the present invention;

[0063] Figure 4 In the embodiment of the present invention, the 63rd frame of the normalized hyperspectral image of the book of the present invention;

[0064] Figure 5 In the embodiment of the present invention, the 65th frame of the normalized hyperspectral image of the book of the present invention;

[0065] Figure 6 In the embodiment of the present invention, the 66th frame of the normalized hyperspectral image of the book of the present invention;

[0066] Figure 7 In the embodiment of the present invention, the search area S of the 61st frame of the normalized hyperspectral image of the book 61 ;

[0067] Figure 8 In the embodiment of the present invention, the search area S of the 63rd frame of the normalized hyperspectral image of the book 63 ;

[0068] Figure 9 In the embodiment of the present invention, the search area S of the 65th frame of the normalized hyperspectral image of the book 65 ;

[0069] Figure 10 In the embodiment of the present invention, the target spectral curve C of the 1st frame of the normalized hyperspectral image of the book t ;

[0070] Figure 11 In the embodiment of the present invention, the search area S of the 61st frame of the normalized hyperspectral image of the book 61 of the global spectral curve

[0071] Figure 12 In the embodiment of the present invention, the search area S of the 63rd frame of the normalized hyperspectral image of the book 63 of the global spectral curve

[0072] Figure 13 In the embodiment of the present invention, the search area S of the 65th frame of the normalized hyperspectral image of the book 65 of the global spectral curve

[0073] Figure 14 In the embodiment of the present invention, the search area S of the 66th frame of the normalized hyperspectral image of the book 66 of the background spectral capsule feature u b ;

[0074] Figure 15 In the embodiment of the present invention, the search area S of the 66th frame of the normalized book hyperspectral image 66 of the background response map R b ;

[0075] Figure 16 In the embodiment of the present invention, the search area S of the 66th frame of the normalized book hyperspectral image 66 of the target response map R t ;

[0076] Figure 17 In the embodiment of the present invention, the target tracking result map of the 66th frame of the normalized book hyperspectral image;

[0077] Figure 18 In the embodiment of the present invention, the background masks of the 1st, 61st, 63rd, and 65th frames of the normalized book hyperspectral images;

[0078] Figure 19 In the embodiment of the present invention, the background pool composed of the 1st, 61st, 63rd, and 65th frames of the normalized book hyperspectral images. Specific implementation scheme

[0079] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementations described herein are only used to solve the present invention and are not used to limit the present invention.

[0080] The embodiment of the present invention provides a hyperspectral video target tracking method based on background spectral capsule features, as Figure 1 shown, the method is as follows:

[0081] Step 1: Load the target position, target box and target image block of the 1st frame of the hyperspectral image sequence, and load the (t - 1)th, (t - 3)th, and (t - 5)th frames of the hyperspectral image sequence; determine the search areas S t-1 , S t-3 , S t-5 ;

[0082] Specifically, load the target position, target box and target image block of the 1st frame of the hyperspectral image sequence, then load the (t - 1)th, (t - 3)th, and (t - 5)th frames of the hyperspectral image sequence, and perform gray normalization on the 1st, (t - 1)th, (t - 3)th, and (t - 5)th frames of the hyperspectral images to obtain the (t - 1)th, (t - 3)th, and (t - 5)th frames of the normalized hyperspectral images, and determine the search areas S t-1 , S t-3 , St-5 ;

[0083] Among them, the gray-scale ranges of the pixels in the (t-1)-th, (t-3)-th, and (t-5)-th frame hyperspectral images are in the left-closed and right-closed interval from 0 to 1. Here, t-1, t-3, and t-5 represent the frame numbers of the hyperspectral images. Random interpolation processing is performed on the missing sequences, and t is an integer greater than or equal to 6 and less than or equal to 2000;

[0084] Load the target position, target bounding box, and target image patch of the first frame hyperspectral image in the hyperspectral image sequence, then load the 61st, 63rd, and 65th frame hyperspectral images in the hyperspectral image sequence, and perform gray-scale normalization on the 1st, 61st, 63rd, and 65th frame hyperspectral images to obtain the normalized 61st, 63rd, and 65th frame hyperspectral images, and determine the search regions S 61 、S 63 、S 65 ,S t obtained from the target image patch of the previous frame. The size of S t is 1.5 times the size of the target image patch of the previous frame. S t has B bands, and the image resolution is p×q. The image resolution p of S 61 is 105, q is 126, and B is 16. The image resolution p of S 63 is 104, q is 132, and B is 16. The image resolution p of S 65 is 114, q is 133, and B is 16. The total number of frames in the loaded hyperspectral image sequence is 601, and the image resolution is 495×231 pixels. Figure 2 is the normalized first-frame book hyperspectral image in the embodiment of the present invention. Figure 3 is the normalized 61st-frame book hyperspectral image in the embodiment of the present invention. Figure 4 is the normalized 63rd-frame book hyperspectral image in the embodiment of the present invention. Figure 5 is the normalized 65th-frame book hyperspectral image in the embodiment of the present invention. Figure 7 is the search region S of the normalized 61st-frame book hyperspectral image in the embodiment of the present invention. 61 , Figure 8 is the search region S of the normalized 63rd-frame book hyperspectral image in the embodiment of the present invention. 63 , Figure 9 is the search region S of the normalized 65th-frame book hyperspectral image in the embodiment of the present invention. 65 。

[0085] Step 2: Classify the area within the target bounding box of the first-frame hyperspectral image to obtain the target spectral curve C of the first-frame hyperspectral image. t, determine the search regions S of the (t-1)-th, (t-3)-th, and (t-5)-th frame hyperspectral images t-1 , S t-3 , S t-5 The global spectral curves of

[0086] Specifically, classify the target box region in the first-frame hyperspectral image according to a statistical method to obtain the target spectral curve C of the first-frame hyperspectral image t , and the target spectral values in the first to sixteenth bands are {0.2914, 0.3004, 0.3160, 0.3078, 0.3116, 0.2813, 0.3031, 0.2320, 0.2156, 0.2194, 0.2094, 0.2388, 0.2534, 0.2533, 0.2512, 0.3100}.

[0087] Determine the search regions S of the 61st, 63rd, and 65th frame hyperspectral images 61 , S 63 , S 65 The global spectral curves of The global spectral values in the first to sixteenth bands are {0.3065, 0.3144, 0.3077, 0.2961, 0.3042, 0.2697, 0.2723, 0.2286, 0.1930, 0.1884, 0.1962, 0.2236, 0.2147, 0.2146, 0.2443, 0.2889}; The global spectral values in the first to sixteenth bands are {0.3040, 0.3146, 0.3474, 0.3024, 0.3082, 0.2656, 0.2754, 0.2248, 0.1729, 0.1666, 0.1966, 0.2244, 0.2030, 0.2153, 0.2456, 0.2942}; The global spectral values in the first to sixteenth bands are {0.3081, 0.3163, 0.3332, 0.2960, 0.3515, 0.3052, 0.2975, 0.2445, 0.2013, 0.1802, 0.2042, 0.2221, 0.2114, 0.2226, 0.2558, 0.2894}.

[0088] Figure 10 In the embodiment of the present invention, the target spectral curve C of the normalized first-frame book hyperspectral image t , Figure 11 In the embodiment of the present invention, the search region S of the normalized 61st-frame book hyperspectral image61 Global spectral curve Figure 12 In the embodiment of the present invention, it is the search area S of the 63rd frame of the normalized book hyperspectral image 63 Global spectral curve Figure 13 In the embodiment of the present invention, it is the search area S of the 65th frame of the normalized book hyperspectral image 65 Global spectral curve

[0089] Step 3: Through Determine the search area S t-1 、S t-3 、S t-5 The spectral angular distances of are D t-1 、D t-3 、D t-5 ;

[0090] Specifically, through And C t Substitute into the existing spectral angular distance formula to calculate the spectral angular distance between the search area and the target spectral curve

[0091] Exemplarily, the spectral angular distance D 61 of the search area S 61 is a large matrix containing 13,230 values, where each value represents the spectral angular distance between different bands of a specific pixel in the search area S 61 and the target spectral curve C t of the 1st frame of the book hyperspectral image. The spectral angular distance D 63 of the search area S 63 is a large matrix containing 13,728 values, where each value represents the spectral angular distance between different bands of a specific pixel in the search area S 61 and the target spectral curve C t of the 1st frame of the book hyperspectral image. The spectral angular distance D 65 of the search area S 65 is a large matrix containing 15,162 values, where each value represents the spectral angular distance between different bands of a specific pixel in the search area S 61 and the target spectral curve C t of the 1st frame of the book hyperspectral image

[0092] Step 4: Generate a background mask through D t-1 、D t-3 、D t-5 and determine the background regions of the 1st, t-1, t-3, and t-5 frame hyperspectral images as the background pool

[0093] Specifically, by combining with C t substitute into the existing spectral angle distance formula to calculate the spectral angle distance between the search area and the target spectral curve.

[0094] Exemplarily, by D 61 , D 63 , D 65 generate a background mask, and determine the background areas of the 1st, 61st, 63rd, and 65th frame hyperspectral images as the background pool. The experimental training sets the threshold for distinguishing the background and the target as 0.84. Load the 61st frame hyperspectral image of the book sequence. The spectral angle distance of the 12th pixel is 0.6438, which is less than the threshold 0.84, indicating that the background pixel is set to 1. The spectral angle distance of the 47th pixel is 0.9332, which is greater than the threshold 0.84, indicating that the target pixel is set to 0. Obtain the background mask by judging the spectral angle distance of each pixel of the entire hyperspectral image, and multiply the obtained background mask and the hyperspectral image to obtain the hyperspectral image of the pure background. Figure 18 This is the background mask of the 1st, 61st, 63rd, and 65th frame book hyperspectral images after normalization in the embodiment of the present invention. Figure 19 This is the background pool composed of the 1st, 61st, 63rd, and 65th frame book hyperspectral images after normalization in the embodiment of the present invention.

[0095] Step Five: Determine the capsule feature u of the t-th frame hyperspectral image through the capsule network t and the background capsule feature u of the background pool b , and obtain the new capsule feature of the t-th frame hyperspectral image

[0096] Specifically, (501) calculate the similarity t between the capsule feature u b and the background capsule feature u of the background pool as

[0097]

[0098] where, represents the i-th background capsule in the background capsule feature u b , represents the i-th feature capsule in the capsule feature u t , ||·|| represents the magnitude of the vector, represents the similarity between the capsule feature u t and the background capsule feature u b , and its range is [-1, 1]. A value of 1 indicates that the directions of the two capsules are exactly the same, a value of -1 indicates that the directions are exactly opposite, and 0 indicates that the two capsules are orthogonal;

[0099] Exemplarily, the hyperspectral image capsule feature u of the 66th frame after normalization t has a size of 36×36×16×16. Among them, 36×36 represents the spatial size, 16 is the number of channels, and 16 is the number of capsule categories. The background pool is the background capsule feature u b determined by the capsule network, with a size of 36×36×16×16. Among them, 36×36 represents the spatial size, 16 is the number of channels, and 16 is the number of capsule categories. The 7th feature capsule and the 7th background capsule The similarity between them has a value of -1.1432, indicating that these two capsules are not similar. The 9th feature capsule and the 9th background capsule The similarity between them has a value of 0.9345, indicating that these two capsules are similar;

[0100] (502) Calculate the new capsule feature according to the following formula as

[0101]

[0102] where represents the i-th feature capsule in the new capsule feature ;

[0103] Exemplarily, in the present invention the new capsule feature with a size of 36×36×16×16. Among them, 36×36 represents the spatial size, 16 is the number of channels, and 16 is the number of capsule categories. The 7th new capsule feature of the 66th frame hyperspectral image has a value that is the vector of the original 7th feature capsule containing 20736 numerical values. The 9th new capsule feature of the 66th frame hyperspectral image is a vector containing 20736 numerical values.

[0104] Step Six: Obtain the background capsule spectral feature E of the t-th frame hyperspectral image through the background similarity routing algorithm, and obtain the background response map R t of the search area S of the t-th frame through the capsule network decoder b ;

[0105] Specifically, (601) Calculate the weight coefficient between the background capsule feature and the new capsule feature according to the following formula as

[0106]

[0107] where represents the i-th background capsule with the i-th feature capsule The weight coefficient therebetween. When the similarity is not 1, it indicates that the current feature capsule is different from the background capsule and its weight coefficient needs to be reduced;

[0108] Exemplarily, the 7th background capsule and the 7th feature capsule The weight coefficient value is -0.3638. The 9th background capsule and the 9th feature capsule The weight coefficient value is 0.9345;

[0109] (602) Calculate the predicted output between the background capsule feature and the new capsule feature according to the following formula is

[0110]

[0111] wherein, represents the predicted output of the i-th capsule of the capsule feature in the t-th frame to the i-th background capsule, represents the newly generated feature capsule updated after similarity processing, W tb contains the transformation of the learning relationship of the i-th feature capsule in the t-th frame to the i-th background capsule. W tb is learned through the network;

[0112] Exemplarily, the predicted output of the 7th capsule of the capsule feature in the 66th frame to the 7th background capsule is a large matrix, and the predicted output of the 9th capsule of the capsule feature in the 66th frame to the 9th background capsule is a large matrix;

[0113] (603) Update the initial weight e according to the following formula t|b is

[0114]

[0115] wherein, v j represents the capsule vector obtained after being compressed by the squash function, which is the output of the capsule j in the next layer. The larger the value, the stronger the consistency of the capsule, and the initial weight e t|b will also increase; The smaller the value, the weaker the consistency of the capsule, and the initial weight e t|b thus decreases;

[0116] Exemplarily, the initial weight of the predicted output of the 7th capsule of the capsule feature in the 66th frame to the 7th background capsule is The prediction output of the 9th capsule of the capsule feature in the 66th frame for the 9th background capsule of the initial weights

[0117] (604) Calculate the coupling coefficient according to the following formula is

[0118]

[0119] where is the normalized weight, that is, the coupling coefficient, k is the number of initial weights, exp(·) represents the exponential function, and ∑(·) represents the summation operation;

[0120] Exemplarily, the normalized weight value of the 7th capsule is 0.6949, and the normalized weight value of the 9th capsule is 2.5453;

[0121] (605) Calculate the total input o of the next-layer capsule j according to the following formula j

[0122]

[0123] Exemplarily, the total input of the 8th capsule in the next layer of the 7th capsule The total input of the 10th capsule in the next layer of the 9th capsule

[0124] (606) Calculate the capsule vector v according to the following formula j is

[0125]

[0126] where, v j represents the output of the lower-layer capsule j, and o j represents the total input of the lower-layer capsule j. When ||o j || → 0, v j → 0; when ||o j || → 0, v j → 1. squash(·) represents the squash function, which can compress the capsule length and keep the vector direction unchanged;

[0127] (607) Integrate the vector v j output by the final layer into capsule features E for expression, and amplify the features E through three deconvolution layers in the capsule network decoder to obtain a background response map R t of the same size as the search area S b at the t-th frame.

[0128] Exemplarily, the output v of the lower - layer capsule 8 8 has a value of 0.7934, which is integrated into the capsule feature E through the network 8 for expression. E 8 is a vector containing 82944 numerical values. The feature E is magnified through three de - convolution layers in the capsule network decoder 8 to obtain a background response map R with the same size as the search region S of the 63rd frame 63 , which is 114×133 b ; The output v of the lower - layer capsule 10 10 has a value of 0.1056, which is integrated into the capsule feature E through the network 10 for expression. E 10 is a vector containing 82944 numerical values. The feature E is magnified through three de - convolution layers in the capsule network decoder 10 to obtain a background response map R with the same size as the search region S of the 65th frame 65 , which is 104×132 b . Figure 14 In the embodiment of the present invention, the background spectral capsule feature u of the search region S of the 66th - frame book hyperspectral image after normalization 66 . b .

[0129] Step Seven: Determine the target response map R through the background response map R b , and determine the target position p through the target response map R t ; t Specifically, the size of the obtained background spectral capsule feature E is adjusted to 36×36×64, and then these features are magnified through three de - convolution layers, and finally a background response map of 115×132×1 with the same size as the 66th - frame input is obtained t .

[0130] Exemplarily, the target position p of the 66th - frame hyperspectral image

[0131] is (180.5, 47). 66 Figure 15 In the embodiment of the present invention, the background response map R of the search region S of the 66th - frame book hyperspectral image after normalization 66 . b , Figure 16 In the embodiment of the present invention, the target response map R of the search region S of the 66th - frame book hyperspectral image after normalization 66 . t .

[0132] Step Eight: Repeat Steps Five to Seven for the (t - 1)th - frame hyperspectral image to obtain the target position p of the (t - 1)th - frame hyperspectral image t-1 ;​

[0133] Exemplarily, the target position p of the 65th frame hyperspectral image 65 is (180.5, 49.5).

[0134] Step Nine: Determine the target spectral curves C t-1 and C t of the (t - 1)th and tth frame hyperspectral images after normalization, and determine the spectral curve difference f according to C t-1 and C t ; specifically, determine f as

[0135] Specifically, determine f according to the following formula

[0136]

[0137] where f represents the spectral curve difference, B represents the number of bands of the hyperspectral image, with a value of 16, C t represents the spectral curve of the target area of the tth frame hyperspectral image after normalization, and C t-1 represents the spectral curve of the target area of the (t - 1)th frame hyperspectral image after normalization. If this value is closer to 0, it can be considered that C t and C t-1 are more similar in the band interval;

[0138] Exemplarily, the target spectral values of the 65th frame hyperspectral image in the 1st to 16th bands are {0.2974, 0.3014, 0.3155, 0.3178, 0.3122, 0.2813, 0.3029, 0.2316, 0.2156, 0.2294, 0.2114, 0.2354, 0.2578, 0.2333, 0.2502, 0.3140}; the target spectral values of the 66th frame hyperspectral image in the 1st to 16th bands are {0.2972, 0.3014, 0.3155, 0.3177, 0.3122, 0.2813, 0.3028, 0.2316, 0.2156, 0.2290, 0.2114, 0.2354, 0.2578, 0.2333, 0.2502, 0.3140}, and the spectral curve difference f is 0.00000021.

[0139] Step Ten: Determine the template update U t→t-1 ;

[0140] Specifically, determine U according to the following formula t→t-1 as

[0141]

[0142] where U t→t-1It represents template update, and ε represents the positive number of the spectrum. When f < ε, the objects in the t-th frame and the (t - 1)-th frame are similar, and it is determined as the same object for the next position prediction; when f ≥ ε, the objects in the t-th frame and the (t - 1)-th frame are not similar, and the network is updated.

[0143] Exemplarily, the value of ε in the present invention is 10 -6 , and the spectral curve difference f is 0.00000021 which is less than 10 -6 , and the object in the 66th frame of the normalized hyperspectral image of the book is similar to the object in the 65th frame of the hyperspectral image of the book for the next position prediction.

[0144] Step Eleven: Update the template U t→t-1 Determine the predicted position P pre ;

[0145] Specifically, P is determined according to the following formula pre which is

[0146] P pre = p t + Δp

[0147] = p t + p t - p t-1

[0148] = (x t , y t ) + (x t - x t-1 , y t - y t-1 )

[0149] where P pre represents the predicted position of the (t + 1)-th frame, p t represents the position of the target center in the t-th frame, Δp represents the position difference, p t-1 represents the position of the target center in the (t - 1)-th frame, x t represents the abscissa of the target in the t-th frame, y t represents the ordinate of the target in the t-th frame, x t-1 represents the abscissa of the target in the (t - 1)-th frame, y t-1 represents the ordinate of the target in the (t - 1)-th frame.

[0150] Exemplarily, the target position p 65 of the 65th frame of the hyperspectral image is (180.5, 49.5), the target position p 66 of the 66th frame of the hyperspectral image is (180.5, 47), Δp represents the position difference as (0, -2.5), P preIt is indicated that the predicted position of the 67th frame is (180.5, 44.5). Set the predicted position as the center position of the search area for the 67th frame, and expand the size of the target box predicted in the 66th frame by 1.5 times as the search area for the 67th frame, and input it into the network for the next frame tracking.

[0151] Step Twelve: Load each frame of hyperspectral image in the hyperspectral image sequence in turn, repeat Steps One to Eleven to obtain the target position of each frame of hyperspectral image, and complete the target tracking of the hyperspectral image sequence.

[0152] Obtain the predicted position P pre Then, according to the generated search area, reduce the computational amount of the capsule network and perform a series of subsequent tracking operations to obtain the target position.

[0153] Specifically, Figure 14 In the embodiment of the present invention, it is the target tracking result diagram of the 66th frame of the normalized book hyperspectral image. Repeat Steps One to Seventeen to obtain the target positions of 601 frames of hyperspectral images and complete the final target tracking of the hyperspectral image sequence.

[0154] The present invention proposes a new feature called background spectral capsule feature. This feature skillfully integrates the background capsule features in the background pool extracted by the capsule network and the capsule features of the current frame, realizing the deep integration of spatial and spectral information. The background spectral capsule feature not only contains detailed spatial structure information but also deeply reflects the unique properties of the spectral dimension, providing the feature representation ability for hyperspectral video target tracking. Secondly, the present invention proposes a background similarity routing algorithm, which can accurately realize the routing fusion between the background capsule features and the capsule features of the current frame, and then efficiently generate the background spectral capsule feature, which is beneficial to the accurate recognition of background information. In addition, based on the spectral difference result of the target spectral curve, the present invention determines whether the modules are similar and generates a search area for predicting the approximate position of the target in the next frame.

[0155] The above is only a preferred embodiment of the present invention and is not used to limit the protection scope of the present invention.

Claims

1. A hyperspectral video target tracking method based on background spectral capsule features, characterized in that: The method is: Step 1: Load the target position, target frame and target image block of the first frame of the hyperspectral image in the hyperspectral image sequence, and load the t-1, t-3 and t-5 frames of the hyperspectral image in the hyperspectral image sequence; Determine the search area S of the normalized hyperspectral images of the t-1, t-3, and t-5 frames t-1 , S t-3 , S t-5 ; Step 2: Classify the target frame area of ​​the first frame of the hyperspectral image to obtain the target spectral curve C of the first frame of the hyperspectral image t , determine the search area S of the hyperspectral image frames t-1, t-3, and t-5 t-1 , S t-3 , S t-5 The global spectral curves are Step 3: Pass Determine the search area S t-1 , S t-3 , S t-5 The spectral angular distances are D t-1 , D t -3 , D t-5 ; Step 4: Through D t-1 , D t-3 , D t-5 Generate background masks of the t-1, t-3, and t-5 frames of hyperspectral images, and multiply the background masks of the t-1, t-3, and t-5 frames of hyperspectral images by the t-1, t-3, and t-5 frames of hyperspectral images to obtain the background areas of the 1, t-1, t-3, and t-5 frames of hyperspectral images as background pools; Step 5: Determine the capsule feature u of the current t-th frame hyperspectral image through the capsule network t and the background capsule feature u of the background pool b , get the new capsule feature of the t-th frame hyperspectral image Step 6: Obtain the background spectral capsule feature E of the hyperspectral image of the t-th frame through the background similarity routing algorithm, and obtain the search area S of the t-th frame through the capsule network decoder t Background response map R b ; Step 7: Background response map R b Determine the target response map R t , through the target response map R t Determine the target position p t ; Step 8: Repeat steps 5 to 7 for the t-1th frame of hyperspectral image to obtain the target position p of the t-1th frame of hyperspectral image. t-1 ; Step 9: Determine the target spectral curve C of the normalized hyperspectral images of the t-1th frame and the tth frame t-1 and C t , according to C t-1 and C t Determine the spectral curve difference f; Step 10: Determine the template update U by using the spectral curve difference f t→t-1 ; Step 11: Update U through the template t→t-1 Determine the predicted position P pre ; Step 12: load each frame of the hyperspectral image in the hyperspectral image sequence in turn, repeat steps 1 to 11, obtain the target position of each frame of the hyperspectral image, and complete the target tracking of the hyperspectral image sequence.

2. The hyperspectral video target tracking method based on background spectrum capsule features according to claim 1 is characterized in that: The step 5 is specifically implemented by the following steps: (501) Use the capsule network to extract the capsule feature u of the current t-th frame hyperspectral image t and the background capsule feature u of the background pool b ; (502) Calculate u according to the following formula t and u b Similarity for in, Indicates u b The i-th background capsule in Indicates u t The i-th feature capsule in ||·|| represents the size of the vector, Represents capsule feature u t and background capsule feature u b The similarity between them is in the range of [-1,1], where a value of 1 means that the directions of the two capsules are exactly the same, a value of -1 means that the directions are completely opposite, and 0 means that the two capsules are orthogonal; (503) The new capsule characteristics are calculated according to the following formula for in, Indicates new capsule features The i-th feature capsule in .

3. The hyperspectral video target tracking method based on background spectrum capsule features according to claim 1 or 2, characterized in that: The step six is ​​specifically implemented by the following steps: (601) Calculate the i-th background capsule according to the following formula With the i-th feature capsule The weight coefficient between for; Among them, when When it is not 1, it means that the current feature capsule is different from the background capsule and its weight coefficient needs to be reduced; (602) The predicted output between the background capsule feature and the new capsule feature is calculated according to the following formula: in, represents the predicted output of the ith capsule of the capsule feature for the ith background capsule in the tth frame, represents the new feature capsule updated and generated after similar processing, W tb Contains the transformation of the learning relationship between the i-th feature capsule and the i-th background capsule in the t-th frame; (603) Update the initial weight e according to the following formula t|b for Among them, v j represents the capsule vector obtained after compression by the squashing function, which is the output of the next layer of capsule j; (604) The coupling coefficient is calculated according to the following formula for in, is the normalized weight, i.e., the coupling coefficient, k is the number of initial weights, exp(·) represents the exponential function, and Σ(·) represents the summation operation; (605) Calculate the total input o of the next layer capsule j according to the following formula j (606) The capsule vector v is calculated according to the following formula j for Among them, v j represents the output of the lower capsule j, o j represents the total input of the lower capsule j, when ||o j When ||→0, v j →0; when ||o j When ||→0, v j →1, squash(·) represents the squashing function, which can compress the capsule length and keep the vector direction unchanged; (607) The vector v output by the final layer j The network is integrated into a capsule feature E for expression, and the feature E is amplified by three deconvolution layers in the capsule network decoder to obtain the search area S of the tth frame. t Background response map R of the same size b .

4. The hyperspectral video target tracking method based on deep spectral cascade texture features according to claim 3 is characterized in that: In step nine, f is determined according to the following formula: Among them, f represents the difference of spectral curves, B represents the number of bands of hyperspectral images, and C t represents the spectral curve of the target area of ​​the t-th frame hyperspectral image after normalization, C t-1 Represents the spectral curve of the target area of ​​the t-1th frame hyperspectral image after normalization.

5. The hyperspectral video target tracking method based on deep spectral cascade texture features according to claim 4 is characterized in that: In step 10, U is determined according to the following formula: t→t-1 for Among them, U t→t-1 represents template update, ε represents a positive spectral number; when f<ε, the objects in the t-th frame are similar to those in the t-1-th frame, and they are determined to be the same object for the next position prediction; when f≥ε, the objects in the t-th frame are not similar to those in the t-1-th frame, and the network is updated.

6. The hyperspectral video target tracking method based on deep spectral cascade texture features according to claim 5 is characterized in that: In step eleven, P is determined according to the following formula: pre for P pre =p t +Δp =p t +p t -p t-1 (x) t ,and t )+(x t -x t-1 ,and t -and t-1 ) Among them, P pre represents the predicted position of the t+1th frame, p t represents the position of the target center in the tth frame, Δp represents the position difference, and p t-1 represents the position of the target center in the t-1th frame, x t Indicates the horizontal coordinate of the target in the tth frame, y t Indicates the ordinate of the target in the tth frame, x t-1 Indicates the horizontal coordinate of the target in the t-1th frame, y t-1 Represents the vertical coordinate of the target in the t-1th frame.