Hyperspectral Video Object Tracking Method Based on Deep Spectral Cascade Texture Features

By combining the depth spectral cascade texture features and LBP features, the problem that the existing technology is difficult to distinguish between targets and backgrounds under complex backgrounds is solved, and the accuracy and stability of hyperspectral video target tracking is improved.

CN116091545BActive Publication Date: 2025-06-20JIANGXI HANGKE PRECISION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310060920.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-13
Publication Date
2025-06-20
Estimated Expiration
2043-01-13

AI Technical Summary

Technical Problem

Existing target tracking algorithms are difficult to distinguish between targets and backgrounds in complex backgrounds, especially when lighting changes, resulting in reduced tracking accuracy and drift.

Method used

A hyperspectral video target tracking method based on depth spectral cascade texture features is adopted. By combining LBP features and spectral information, the depth average spectral cascade texture features are extracted, the depth features and spectral cascade texture prediction features are fused, and the position prediction mask is generated to suppress background textures.

Benefits of technology

Effectively overcome the impact of light changes, distinguish between targets and backgrounds, improve the accuracy and stability of target tracking, and reduce drift phenomena.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091545B_ABST
    Figure CN116091545B_ABST
Patent Text Reader

Abstract

The present invention discloses a hyperspectral video object tracking method based on depth spectral cascade texture features. The spectral curve of the local area of the t-th frame hyperspectral image is obtained, and the spectral curve of the pixels in the unknown area is subtracted from the spectral curve of the local area. The image is segmented into a target area and a background area. The (t + 1)-th frame is imported, and the spectral curve of each pixel point in the search area is obtained. After dimensionality reduction processing, a one-band image after dimensionality reduction is obtained. The depth feature and texture feature of this image are extracted. According to the image signal-to-noise ratio curve, the weight of each spectral channel is obtained, and it is multiplied by the corresponding texture feature to obtain a spectral cascade feature. The spectral cascade feature is covered with a set mask to obtain a spectral cascade texture prediction feature, and the spectral cascade texture prediction feature is convolved with the depth feature pixel by pixel to obtain a depth-averaged spectral cascade texture feature U k , and the U of the first frame image k is sent to the DCF filter to train a template, and the U of the (t + 1)-th frame image k is sent to the filter template to obtain a response map. According to the distribution strategy, the target scale is estimated, the position of the prediction box is determined, and the tracked target is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a hyperspectral video object tracking method based on deep spectral cascade texture features. Background Art

[0002] As a research hotspot in the field of computer vision, object tracking is widely applied in hot fields such as autonomous driving, public safety, and human-computer interaction.

[0003] Currently, most tracking algorithms are developed based on visible light videos. However, when the color of the object is similar to the background, the existing visible light tracking algorithms cannot achieve satisfactory tracking effects. Hyperspectral images have rich spectral information. Even if the colors are similar, objects of different materials can be easily identified through spectral information. Therefore, extending the algorithms developed based on visible light videos to hyperspectral videos has become a development trend in the field of object tracking. Common object tracking methods generally fall into two categories: correlation filter-based methods and deep learning-based methods.

[0004] For the object tracking algorithm based on filters, a typical algorithm is the context-aware correlation filter tracking algorithm. This algorithm extracts oriented gradient features, incorporates context information into the correlation filter, and uses the target context area as negative samples for the training of the filter to improve tracking accuracy. However, in complex backgrounds, it is difficult to distinguish the target from the background using oriented gradient features. When the illumination changes, the selected context area does not comprehensively sample the background area, and it is accidental to suppress the background. When suppressing background clutter, using the same suppression weight is likely to cause drift.

[0005] For the object tracking algorithm based on deep learning, a typical algorithm is the object tracking algorithm based on a fully convolutional siamese network. This algorithm uses a similarity comparison method to compare the similarity between the target frame and the template image, and gives the position with the maximum similarity, which is the position of the object in the target frame. The algorithm based on deep learning does not utilize spectral information, resulting in the above algorithm being unable to track the object well when the object and the background are extremely similar. This method cannot change the scale of the target area when facing illumination changes and is likely to lose the target.

[0006] The method proposed by the present invention fully considers spectral information and combines it with LBP features, can overcome the influence of illumination changes according to the different spectral features of different objects, effectively distinguish the target from the background, and then effectively complete object tracking. Summary of the Invention

[0007] In view of this, the main object of the present invention is to provide a hyperspectral video object tracking method based on deep spectral cascade texture features.

[0008] To achieve the above object, the technical solution of the present invention is achieved as follows:

[0009] The embodiment of the present invention provides a hyperspectral video target tracking method based on deep spectral cascade texture features, the method comprising:

[0010] Step 1: Load the target position, target frame and target image block of the first frame of the hyperspectral image in the hyperspectral image sequence, then load the t-th frame of the hyperspectral image in the hyperspectral image sequence, and perform grayscale normalization on the t-th frame of the hyperspectral image to obtain the normalized t-th frame of the hyperspectral image, and determine the search area S of the normalized t-th frame of the hyperspectral image. t , where the grayscale range of the pixel of the t-th frame hyperspectral image is a left-closed and right-closed interval from 0 to 1, t represents the number of frames of the hyperspectral image, and t is an integer greater than or equal to 2 and less than or equal to 2000;

[0011] Step 2: Taking the target position of the hyperspectral image of frame t-1 in the hyperspectral image sequence as the center, t A 3×3 pixel image block is intercepted from the top as the local area L of the normalized hyperspectral image of the tth frame t , determine L t The i-th band image The spectral average value, according to L t The i-th band image The spectral average value determines the target local area L of the t-th frame hyperspectral image after normalization t The local spectral curve C l , and obtain C l The local spectral average C in the i-th band li ;

[0012] Step 3: According to S t Subtract L t Get the unknown area U of the normalized hyperspectral image of frame t t , and U t The pixels inside are defined as unknown pixels;

[0013] Step 4: Determine the spectral curve C of the unknown pixel uj , and obtain C uj The gray value C of the jth pixel in the i-th band uij , where j represents the serial number of the unknown pixel;

[0014] Step 5: Pass C uij Minus C li The absolute value of the error ε i,j ;

[0015] Step 6: Determine ε i,j The cumulative sum Q j , compare Qj and the size of the preset error threshold η, to determine U t the j-th unknown pixel U t (j) in the frame belongs to the background pixel or the target pixel; define the set of target pixels as the target region O of the t-th frame hyperspectral image t and define the set of background pixels as the background region J of the t-th frame hyperspectral image t , where η is greater than 0;

[0016] Step Seven: Load the search region S of the (t + 1)-th frame hyperspectral image after normalization t+1 , according to S t+1 obtain the spectral curves C of each pixel point in the search region s ;

[0017] Step Eight: Determine the Euclidean distance D between C si and C li ; j ;

[0018] Step Nine: Determine the deep features E of the one-band grayscale image sequence after dimensionality reduction through the GoogLeNet network;

[0019] Step Ten: Extract the texture features P of S t+1 through the LBP operator and the deep features E;

[0020] Step Eleven: Obtain the i-th band image of S t and determine the image signal-to-noise ratio of According to determine the signal-to-noise ratio spectral value C of the t-th frame hyperspectral image in the i-th band , through C ri obtain the signal-to-noise ratio spectral curve C of the search region ri , according to C r determine the weight w of the i-th spectral channel r ; i ;

[0021] Step Twelve: Determine the spectral cascade texture features R of S i through P and w t+1 ;

[0022] Step Thirteen: Generate the position prediction mask M through O t and J t ;

[0023] Step Fourteen: Determine the spectral cascade texture prediction features Z of S t+1 through R and M;

[0024] Step 15: Divide E into Y groups according to the number of channels, perform per-pixel splitting on the depth features of each group, split them into C convolution kernels, and then perform convolution on Z through the k-th convolution kernel of the v-th group to obtain the k-th depth spectral concatenated texture feature U of the v-th group vk , and average the Y groups of U vk to obtain the depth-averaged spectral concatenated texture feature U k , where Y takes the value of 32, C takes the value of 196, v is an integer greater than or equal to 1 and less than or equal to Y, v represents the sequence number of the depth spectral concatenated feature U vk , k is an integer greater than or equal to 1 and less than or equal to C, and k represents the sequence number of the convolution kernel;

[0025] Step 16: Obtain the U k features extracted from the search area of the first frame of the hyperspectral image, send the U k features extracted from the search area of the first frame into the DCF filter for training to obtain the filter template, and then send the U t+1 features extracted from S k into the filter template to obtain the response map of S t+1 , and determine the target position through the response map of S t+1 ;

[0026] Step 17: Generate the target image patch of the t-th hyperspectral image in the hyperspectral image sequence;

[0027] Step 18: Load each frame of the hyperspectral image sequence in turn, repeat Steps 1 to 17 to obtain the target patches of each frame of the hyperspectral image, and complete the target tracking of the hyperspectral image sequence.

[0028] In the above solution, in Step 4, C is determined according to the following formula uj as

[0029]

[0030] where the range of j is from 1 to p×q, C uij is the spectral value of C uj in the i-th band, is the pixel value of the j-th pixel in the i-th band of the t-th frame in the unknown area.

[0031] In the above solution, Step 6 is specifically implemented through the following steps:

[0032] (601) Set the error threshold η;

[0033] (602) Calculate the error ε according to the following formula i,j and the cumulative sum Q j is

[0034]

[0035] (603) Determine whether the unknown pixel belongs to the target area or the background area according to the following formula;

[0036]

[0037] Among them, when Q j is greater than or equal to η, C uj is the pixel in the background area J t When Q j is less than η, C uj is the pixel in the target area O t inside.

[0038] In the above solution, the eleventh step is specifically implemented through the following steps:

[0039] (1101) Determine according to the following formula as

[0040]

[0041] Among them, represents the target area of the i-th band of the t-th frame of hyperspectral image, represents the background area of the i-th band of the t-th frame of hyperspectral image;

[0042] (1102) Determine the signal-to-noise ratio spectral curve C of the search area according to the following formula r as

[0043]

[0044] Among them, C ri represents the signal-to-noise ratio spectral value of the t-th frame of hyperspectral image in the i-th band;

[0045] (1103) Normalize SNR(S i t ) according to the following formula to obtain the value γ of the normalized image signal-to-noise ratio of the i-th band i as

[0046]

[0047] (1104) To amplify the difference of γ i , perform a non-linear transformation on γ i to obtain the value F after the non-linear transformation of γ i after the non-linear transformation i as

[0048] F i =(γ i ) 2 | i∈{1,...,B}

[0049] (1105) Determine \(w\) according to the following formula i where

[0050]

[0051] In the above solution, the specific steps of step twelve are as follows:

[0052] (1201) Obtain the spectral cascade texture feature \(R_i\) of \(R\) in the \(i\)-th band according to the following formula i where

[0053] \(R_i\) i =\(w\) i \(\times P_i\) i \(\vert\) i∈{1,...,B}

[0054] where \(P_i\) i represents the texture feature of \(P\) in the \(i\)-th band;

[0055] (1202) After \(i\) traverses from 1 to \(B\), obtain the set of all \(R_i\) i which is \(R\)

[0056] \(R = \{R_i\) i \(\vert\) i∈{1,...,B} \}\).

[0057] In the above solution, the specific steps of step thirteen are as follows:

[0058] (1301) Generate a matrix \(M_1\) with the same size as \(S\) t and set the initial value of the matrix to 0;

[0059] (1302) Update \(M_1\) according to the following formula and judge whether the unknown pixel belongs to \(O\) t or \(J\) t . If the unknown pixel belongs to \(O\) t , then assign 1 to the corresponding area position in \(M_1\). If the unknown pixel belongs to \(J\) t , then assign 0.5 to the corresponding area position in \(M_1\). \(M_1\) is

[0060]

[0061] (1303) Generate a position prediction mask \(M\) with the same size as \(M_1\) and assign \(M_1\) to \(M\).

[0062] In the above solution, the specific steps of step fourteen are as follows:

[0063] Multiply the position prediction mask \(M\) by the spectral cascade texture feature \(R\) and obtain the spectral cascade texture prediction feature \(Z\) according to the following formula

[0064] \(Z = M\odot R\)

[0065] Among them, ⊙ represents the dot product operation.

[0066] In the above solution, step 15 is specifically determined through the following steps:

[0067] (1501) Divide E into Y groups according to the number of channels, and perform per-pixel splitting on each group of depth features, splitting them into C convolutional kernels;

[0068] (1502) Obtain U according to the following formula vk where

[0069] U vk = Z * E vk | v∈{1,...,Y}k∈{1,...,C}

[0070] where * represents the convolution operation, and E vk represents the k-th convolutional kernel of the v-th group of depth features of E;

[0071] (1503) Obtain U according to the following formula k where

[0072]

[0073] In the above solution, step 17 is specifically implemented according to the following steps:

[0074] (1701) Determine whether the current frame is the first hyperspectral image in the hyperspectral image sequence. If so, load the target position and target box of the first hyperspectral image in the hyperspectral image sequence. If not, execute the next step;

[0075] (1702) On the (t + 1)-th hyperspectral image in the hyperspectral image sequence, with the target position as the center, intercept the corresponding image patch G through the target box of the t-th hyperspectral image in the hyperspectral image sequence t ;

[0076] (1703) Set the scale pool S according to the following formula p1 where

[0077] Sp1 = {I1, I2,..., I e}

[0078] where I e represents the e-th scale factor, and e represents the serial number of the scale factor;

[0079] (1704) Scale G through I p1 in S e to obtain a set H of target image patches of different sizes on the t-th hyperspectral image in the hyperspectral image sequence t ​t For

[0080] H t ={H X t |H X t =I e ×G t} X∈{1,...,e}

[0081] Wherein, H t represents the X-th image block, and X represents the ordinal number of the image block;

[0082] (1705) Use bilinear interpolation to ensure that the size of each H X t is consistent with G t Extract the U X t feature from all H k and put it into the DCF correlation filter to obtain e response maps, calculate the maximum value of the e response maps, and define the target image block corresponding to the maximum value of the e response maps as H t max , H t max is the target image block of the t-th frame hyperspectral image in the hyperspectral image sequence.

[0083] Compared with the prior art, the present invention fuses the depth feature and the spectral cascade texture prediction feature to obtain the depth average spectral cascade texture feature U k with semantic information, which can overcome the influence of illumination changes according to the different spectral features of different objects and effectively distinguish the target from the background. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 is the flow chart of the present invention;

[0085] Figure 2 is the 50th frame Minions hyperspectral image after normalization in the embodiment of the present invention;

[0086] Figure 3 is the search area S of the 50th frame Minions hyperspectral image after normalization in the embodiment of the present invention 50 ;

[0087] Figure 4 is the target local area L of the 50th frame Minions hyperspectral image after normalization in the embodiment of the present invention 50 of the local spectral curve C l ;

[0088] Figure 5In the embodiment of the present invention, the target region O of the 50th frame of the normalized Minion hyperspectral image 50 ;

[0089] Figure 6 In the embodiment of the present invention, the search region S of the 51st frame of the normalized Minion hyperspectral image 51 The gray-scale image of a single band after dimensionality reduction;

[0090] Figure 7 In the embodiment of the present invention, the search region S of the 51st frame of the normalized Minion hyperspectral image 51 The depth features of the gray-scale image of a single band after dimensionality reduction;

[0091] Figure 8 In the embodiment of the present invention, the search region S of the 51st frame of the normalized hyperspectral image 51 The LBP feature map;

[0092] Figure 9 In the embodiment of the present invention, the search region S of the 50th frame of the normalized Minion hyperspectral image 50 The image signal-to-noise ratio spectral curve;

[0093] Figure 10 In the embodiment of the present invention, the search region S of the 51st frame of the normalized Minion hyperspectral image 51 The spectral cascade texture feature map;

[0094] Figure 11 In the embodiment of the present invention, the search region S of the 51st frame of the normalized Minion hyperspectral image 51 The spectral cascade texture feature map;

[0095] Figure 12 In the embodiment of the present invention, the search region S of the 51st frame of the normalized Minion hyperspectral image 51 The U k Feature map;

[0096] Figure 13 In the embodiment of the present invention, the search region S of the 51st frame of the normalized Minion hyperspectral image 51 The response map obtained by entering the DCF filter;

[0097] Figure 14 In the embodiment of the present invention, the target tracking result map of the 51st frame of the normalized Minion hyperspectral image; Specific implementation scheme

[0098] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to solve the present invention and are not used to limit the present invention.

[0099] An embodiment of the present invention provides a hyperspectral video object tracking method based on deep spectral cascaded texture features, as Figure 1 shown, the method is as follows:

[0100] Step 1: Load the target position, target box and target image patch of the first hyperspectral image in the hyperspectral image sequence, then load the t-th hyperspectral image in the hyperspectral image sequence, and perform gray normalization on the t-th hyperspectral image to obtain the normalized t-th hyperspectral image, and determine the search region S of the normalized t-th hyperspectral image t ;

[0101] Among them, the gray scale range of the pixels of the t-th hyperspectral image is the left-closed and right-closed interval from 0 to 1, t represents the number of frames of the hyperspectral image, and t is an integer greater than or equal to 2 and less than or equal to 2000;

[0102] Specifically, load the 50th hyperspectral image in the hyperspectral image sequence, and perform gray normalization on the 50th hyperspectral image to obtain the normalized 50th hyperspectral image, and determine the search region S of the normalized 50th hyperspectral image 50 , S t is obtained from the target image patch of the previous frame, the size of S t is 1.5 times the size of the target image patch of the previous frame, S t has B bands, the image resolution is p×q, p is 138, q is 127, B is 16, the hyperspectral image sequence has a total of 376 frames, and the image resolution is 423×211 pixels, Figure 2 is the normalized 50th Minion hyperspectral image in the embodiment of the present invention, Figure 3 is the search region S of the normalized 50th Minion hyperspectral image in the embodiment of the present invention 50 .

[0103] Step 2: Take a 3×3 pixel image patch on S t centered on the target position of the (t-1)-th hyperspectral image in the hyperspectral image sequence as the local region L of the normalized t-th hyperspectral image t , calculate the spectral average value of the i-th band image t of L , and determine the target local region L of the normalized t-th hyperspectral image through the spectral average value of the i-th band image t of L ​t The local spectral curve C l , and obtain C l The local spectral average C in the i-th band li ;

[0104] The local spectral curve C is determined according to the following formula l for

[0105]

[0106] Among them, A(.) represents the averaging operation, i represents the serial number of the band, i is an integer greater than or equal to 1 and less than or equal to 16, B represents the total number of bands, B is 16;

[0107] Specifically, the target position of the 49th hyperspectral image in the hyperspectral image sequence is taken as the center, and the center coordinates of the target position of the 49th hyperspectral image in the hyperspectral image sequence are (64, 59). 50 A 3×3 pixel image block is intercepted as the local area L of the 50th frame of the hyperspectral image after normalization. 50 , calculate L 50 The fifth band image L1 50 The spectral average value is 0.4127, through L 50 The spectral average of the 16 band images is used to determine the target local area L of the 50th frame of the hyperspectral image after normalization. 50 The local spectral curve C l , and obtain C l The spectral average values ​​in the 1st to 16th bands are {0.2871, 0.3196, 0.3542, 0.3619, 0.4127, 0.4329, 0.4783, 0.5896, 0.5132, 0.7403, 0.6325, 0.6832, 0.5849, 0.5912, 0.6839, 0.7825} respectively. Figure 4 is the target local area L of the 50th frame of the Minion hyperspectral image 50 The local spectral curve C l , which is a line graph composed of 16 local spectral averages depicted in band order.

[0108] Step 3: Through S t Subtract L t Get the unknown area U of the normalized hyperspectral image of frame t t , and U t The pixels inside are defined as unknown pixels;

[0109] Determine the unknown area U according to the following formula t for

[0110] Ut =S t -L t

[0111] Among them, U t There are i bands and the image resolution is (p×q-9);

[0112] Specifically, the search area S of the 50th frame of the hyperspectral image after normalization is 50 The number of internal pixels is 138×127, i.e., 17526. The local area L of the 50th frame of the hyperspectral image after normalization is 50 The number of internal pixels is 3×3, i.e., 9. The unknown area U of the 50th frame of the hyperspectral image after normalization is 50 The number of unknown pixels is 17517, and the unknown area U t It is a circular area, which is helpful to find the target pixel in the unknown area.

[0113] Step 4: Determine the spectral curve C of the unknown pixel uj , and obtain C uj The gray value C of the jth pixel in the i-th band uij , where j represents the serial number of the unknown pixel;

[0114] Determine C according to the following formula uj for

[0115]

[0116] Among them, j ranges from 1 to p×q, C uij C uj The spectral value in the i-th band, is the jth pixel value of the ith band in the tth frame of the unknown area;

[0117] Specifically, the unknown area U of the 50th frame of the hyperspectral image after normalization is 50 , j ranges from 1 to 17517, U 50 The 32nd pixel value of the first band is 0.3771. When j is 32, U 50 The spectrum curve C of the 32nd unknown pixel u32 The spectral values ​​in the 1st to 16th bands are {0.3771, 0.4232, 0.4572, 0.5123, 0.5783, 0.5863, 0.6114, 0.6235, 0.5874, 0.7840, 0.6455, 0.6845, 0.5874, 0.6324, 0.7145, 0.6874} respectively.

[0118] Step 5: Pass C uij Minus C liThe absolute value of the obtained error ε i,j ;

[0119] The error ε is determined according to the following formula i,j is

[0120] ε i,j = |C uij - C li | i∈{1,...,B},j∈{1,...,p×q-9}

[0121] where |·| represents the absolute value operation;

[0122] Specifically, for the unknown region U of the 50th frame hyperspectral image after normalization 50 , the error ε of the 32nd pixel in the first band 1,32 is 0.090. When j takes 32, the absolute values of the differences between the spectral values of C u32 and C l in the first to the 16th bands are successively, ε 1,32 is 0.090, ε 2,32 is 0.1036, ε 3,32 is 0.1030, ε 4,32 is 0.1504, ε 5,32 is 0.1656, ε 6,32 is 0.1534, ε 7,32 is 0.1331, ε 8,32 is 0.0339, ε 9,32 is 0.0742, ε 10,32 is 0.0437, ε 11,32 is 0.0130, ε 12,32 is 0.0013, ε 13,32 is 0.0025, ε 14,32 is 0.0412, ε 15,32 is 0.0306, ε 16,32 is 0.0951. When j takes 45, the absolute values of the differences between the spectral values of C u45 and C l in the first to the 16th bands are successively, ε 1,45 is 0.096, ε 2,45 is 0.1236, ε 3,45 is 0.1130, ε 4,45 is 0.1504, ε 5,45 is 0.1656, ε 6,45 is 0.1534, ε 7,45 is 0.1331, ε 8,45 is 0.0359, ε 9,45 is 0.0722, ε 10,45 is 0.0457, ε 11,45is 0.0230, ε 12,45 is 0.0023, ε 13,45 is 0.0035, ε 14,45 is 0.0432, ε 15,45 is 0.0356, ε 16,45 It is 0.0911.

[0123] Step 6: Set the error threshold η and calculate ε i,j The cumulative sum Q j , compare Q j and the size of η, judge U t The jth unknown pixel U t (j) Is it a background pixel or a target pixel? The set of target pixels is defined as the target area O of the hyperspectral image of the tth frame. t , the set of background pixels is defined as the background area J of the hyperspectral image of the tth frame t , where η is greater than 0;

[0124] This is achieved through the following steps:

[0125] (101) Setting an error threshold η;

[0126] Specifically, η is 1.5;

[0127] (102) The error ε is calculated according to the following formula i,j The cumulative sum Q j for

[0128]

[0129] Specifically, the unknown area U of the 50th frame of the hyperspectral image after normalization is 50 The cumulative sum of the 32nd pixel error in Its value is 1.2346, and the unknown area U of the 50th frame of hyperspectral image after normalization is 50 The cumulative sum of the 45th pixel error Its value is 1.6107;

[0130] (103) Determine whether the unknown pixel belongs to the target area or the background area according to the following formula;

[0131]

[0132] Among them, U t (j) is U t The jth unknown pixel in the j When C is greater than or equal to η uj is the background area J t Pixels within, when Q j When C is less than η uj The target area Ot Pixels within

[0133] Specifically, the unknown region U of the 50th frame hyperspectral image after normalization 50 In which, C u32 And C l The total error Q between 32 Is 1.2346 which is less than 1.5, so the 32nd pixel U within U 50 Belongs to the target region O 50 (32) of the pixels within 50 In which, C u45 And C l The total error Q between 45 Is 1.6837 which is greater than 1.5, so the 45th pixel U within U 50 Belongs to the background region J 50 (45) of the pixels within 50 In which, Figure 5 In the embodiment of the present invention, the target region O of the 50th frame Minions hyperspectral image after normalization 50 , Since the accurate segmentation of the target and the background is challenging, the tracking method proposed by the present invention retains as many target pixels as possible. As can be clearly seen from Figure 5 , The white part is O 50 , The black part is J 50 , Although the segmentation result of the present invention is not very accurate, the eyes of the toy are well preserved. The effective distinction between the target and the background after such segmentation is the key to accurate tracking.

[0134] Step Seven: Load the search region S of the (t + 1)th frame hyperspectral image after normalization t+1 , According to the following formula, obtain the spectral curve C of each pixel point in the search region through S t+1 Which is s For

[0135]

[0136] Wherein, Represents the jth pixel point in the ith band of the S t+1 Region, C si Represents the spectral value of C s On the ith band, the number of unknown pixels is p'×q', p' represents the number of rows of S t+1 , q' represents the number of columns of S t+1 ;

[0137] Specifically, load the search region S of the 51st frame hyperspectral image after normalization 51 , p' is 139, q' is 128, and the spectral value of the 32nd pixel point in the 1st band is 0.3571, and the search area S of the 51st frame of the hyperspectral image after normalization 51 The spectral values of the 32nd pixel in the 1st to 16th bands are successively {0.3571, 0.4512, 0.4572, 0.5233, 0.5493, 0.5943, 0.6354, 0.6455, 0.5754, 0.7640, 0.6554, 0.6875, 0.5774, 0.6414, 0.7235, 0.6754}

[0138] Step eight, calculate C according to the following formula si and C li The Euclidean distance D j is

[0139]

[0140] where D j is the gray value of the j-th pixel point on S after dimensionality reduction t+1 Traverse all values of j to obtain the one-band gray-scale image of S t+1 The one-band gray-scale image after dimensionality reduction, and process each frame of the hyperspectral image sequence in turn to obtain the one-band gray-scale image sequence after dimensionality reduction;

[0141] Specifically, the search area S of the 51st frame of the hyperspectral image after normalization 51 The spectral values C of the 32nd pixel point in the 1st to 16th bands si and C li The differences are successively {0.0700, 0.1316, 0.1030, 0.1614, 0.1366, 0.1614, 0.1571, 0.0559, 0.0622, 0.0237, 0.0229, 0.0043, -0.0075, 0.0502, 0.0396, -0.1071}, and the squares of the differences are successively {0.0049, 0.0173, 0.0106, 0.0260, 0.0187, 0.0260, 0.0247, 0.0031, 0.0039, 0.0006, 0.0005, 0.0000, 0.0001, 0.0025, 0.0016, 0.0115}. The gray value D of the 32nd pixel point on S after dimensionality reduction 51 is 0.1520. Traverse all values of j to obtain the one-band gray-scale image of S 32 after dimensionality reduction 51 In the example of the present invention, the search area S of the 51st frame of the Minions hyperspectral image after normalization Figure 6 The one-band gray-scale image after dimensionality reduction, and the search area S of the 51st frame of the Minions hyperspectral image after normalization 51 The one-band gray-scale image after dimensionality reduction 51The one-band grayscale image after dimensionality reduction contains 138×127 pixels and one band. Figure 6 In the example of the present invention, the search area S of the 51st frame of the Minions hyperspectral image after normalization 51 The one-band grayscale image after dimensionality reduction.

[0142] Step Nine: Determine the deep feature E of the one-band grayscale image after dimensionality reduction through the GoogLeNet network;

[0143] Specifically, the GoogLeNet network is a prior art. In the present invention, the inception(4c) layer of GoogLeNet is used to extract the deep feature. In the present invention, the size of E is a deep feature of 14×14×512, where 14×14 represents the spatial size and 196 is the number of channels. Figure 7 In the embodiment of the present invention, the search area S of the 51st frame of the Minions hyperspectral image after normalization 51 The deep feature of the one-band grayscale image after dimensionality reduction.

[0144] Step Ten: Use the LBP operator to extract the texture feature map P of the search area S of the (t + 1)th frame of the hyperspectral image after normalization t+1 ;

[0145] Among them, the circular LBP operator is used in the present invention, its radius is one pixel, the number of sampling points is 8, the size of P is 81×88×16, 81×88 is the spatial resolution of P, and 16 is the number of channels of P. Figure 8 In the embodiment of the present invention, the search area S of the 51st frame of the hyperspectral image after normalization 51 The LBP feature map of.

[0146] Step Eleven: Obtain the i-th band image of S t Determine of image signal-to-noise ratio According to Determine the signal-to-noise ratio spectral value C of the t-th frame of the hyperspectral image in the i-th band ri , through C ri Obtain the signal-to-noise ratio spectral curve C of the search area r , according to C r To obtain the weight w of the i-th spectral channel i ;

[0147] Because the target may have only very small differences between adjacent frames, such as position, shape, and illumination. So in this case, we use the prior information of the t-th frame to calculate the weight of each band of the (t + 1)th frame;

[0148] (201) Determine according to the following formula be

[0149]

[0150] Among them, represents the target area of the i-th band of the t-th frame of hyperspectral image, represents the background area of the i-th band of the t-th frame of hyperspectral image;

[0151] (202) Determine the signal-to-noise ratio spectral curve C of the search area according to the following formula r is

[0152]

[0153] Among them, C ri represents the signal-to-noise ratio spectral value of the t-th frame of hyperspectral image in the i-th band;

[0154] Specifically, obtain the 16-band images of the search area S of the 50th frame of hyperspectral image after normalization 50 of, in the search area S of the 50th frame of hyperspectral image after normalization 50 the 1st to 16th bands, are successively {8.62, 5.32, 6.12, 8.51, 6.31, 5.63, 6.45, 7.82, 6.33, 5.65, 7.65, 5.45, 6.65, 8.43, 7.45, 6.32. are successively 4.82, 4.73, 3.48, 2.87, 3.54, 4.21, 4.33, 5.01, 4.21, 3.12, 3.46, 3.21, 4.83, 4.16, 4.32, 3.78}. are successively {0.7884, 0.1247, 0.7586, 1.9652, 0.7825, 0.3373, 0.4896, 0.5609, 0.5036, 0.8109, 1.2110, 0.6978, 0.3768, 1.0264, 0.7245, 0.6720}, step (203) will perform normalization to [0,1] operation on ;

[0155] (203) Perform normalization processing on the according to the following formula to obtain the value γ after normalization of the image signal-to-noise ratio of the i-th band i is

[0156]

[0157] Specifically, the search area S of the 50th frame of hyperspectral image after normalization 50The weights γ1, γ2, γ3, γ4, γ5, γ6, γ7, γ8, γ9, and γ of the SNR channels are 0.0666, 0.0105, 0.0641, 0.01661, 0.0661, 0.0285, 0.0414, 0.0474, 0.0426, and 10 0.0685, respectively, and γ 11 is 0.1024, and γ 12 is 0.0590, and γ 13 is 0.0319, and γ 14 is 0.0868, and γ 15 is 0.0612, and γ 16 is 0.0568. For texture features, the higher the SNR, the more obvious the texture features of the target;

[0158] (204) is to amplify the γ i difference. According to the following formula, γ i is obtained after performing a non - linear transformation on γ i The value F after the non - linear transformation is i as follows

[0159] F i =(γ i ) 2 | i∈{1,...,B}

[0160] Specifically, for the search region S of the 50th - frame hyperspectral image after normalization 50 The weights F1, F2, F3, F4, F5, F6, F7, F8, F9, and F of the amplified high - SNR channels are 0.9956, 0.9999, 0.9959, 0.9724, 0.9956, 0.9992, 0.9983, 0.9978, 0.9982, and F 10 is 0.9953, and F 11 is 0.9895, and F 12 is 0.9965, and F 13 is 0.9990, and F 14 is 0.9925, and F 15 is 0.9962, and F 16 is 0.9968;

[0161] (205) determines w according to the following formula i as follows

[0162]

[0163] Specifically, for the search region S of the 50th - frame hyperspectral image after normalization 50The weights of the 1st to 16th bands are {0.0625, 0.0628, 0.0626, 0.0611, 0.0625, 0.0628, 0.0627, 0.0627, 0.0627, 0.0625, 0.0622, 0.0626, 0.0628, 0.0623, 0.0626, 0.0626} respectively. Figure 9 In the embodiment of the present invention, the search area S of the 50th frame of the Minion hyperspectral image after normalization 50 of the image signal-to-noise ratio spectral curve.

[0164] Step Twelve: Determine the spectral cascade texture feature R of S through P and w i of S t+1 ;

[0165] (301) Obtain the spectral cascade texture feature R of R in the i-th band according to the following formula i as

[0166] R i = W i × P i | i∈{1,...,B}

[0167] where P i represents the texture feature of P in the i-th band;

[0168] (302) After i traverses from 1 to B, obtain all R i which is R

[0169] R = {R i | i∈{1,...,B}}

[0170] Specifically, the size of R in the present invention is 81×88×16, 81×88 is the spatial resolution of R, and 16 is the number of channels of R. Figure 10 In the embodiment of the present invention, the spectral cascade texture feature map of the search area S of the 51st frame of the Minion hyperspectral image after normalization 51 ;

[0171] Step Thirteen: Generate the position prediction mask M through O t and J t , specifically:

[0172] (501) Generate a matrix M1 with the same size as S t , and set the initial value of the matrix to 0;

[0173] (502) Update M1 according to the following formula, and judge whether the unknown pixel belongs to O t or J t , if the unknown pixel belongs to O t, then assign 1 to the corresponding area position in M1. If the unknown pixel belongs to J t , then assign 0.5 to the corresponding area position in M1. M1 is

[0174]

[0175] (503) Generate a position prediction mask M with the same size as M1, and assign M1 to M;

[0176] Specifically, the algorithm proposed in the present invention does not consider special situations such as rapid displacement or drastic scale change. Generally, the target may only have very small differences between adjacent frames, such as position, shape, and illumination. In this case, the target in the t-th frame and the (t + 1)-th frame are very similar in terms of shape and size. Therefore, this mask can be used to predict the approximate position of the target in the next frame of hyperspectral image, and can also be used to suppress excessive background textures.

[0177] Step Fourteen: Determine S through R and M t+1 of the spectral concatenated texture prediction feature Z;

[0178] Multiply the position prediction mask M by the spectral concatenated texture feature R, and obtain the spectral concatenated texture prediction feature Z according to the following formula as

[0179] Z = M ⊙ R

[0180] where ⊙ represents the dot product operation.

[0181] Specifically, in the present invention, the size of Z is 81×88×16, 81×88 is the spatial resolution of Z, and 16 is the number of channels of Z. Figure 11 In the embodiment of the present invention, it is the spectral concatenated texture feature map of the search area S of the 51st frame Minion hyperspectral image after normalization 51 .

[0182] Step Fifteen: As a kind of texture feature, Z still retains the local features of the texture feature while adding spectral information. In order not to destroy this local feature and achieve the effective fusion of spectral information, texture information, and semantic information, divide E into Y groups according to the number of channels, perform per-pixel splitting on each group of depth features, split them into C convolutional kernels, and then perform convolution on Z through the k-th convolutional kernel of the v-th group to obtain the k-th depth spectral concatenated texture feature U of the v-th group vk , and take the average of the Y groups of U vk to obtain the depth-averaged spectral concatenated texture feature U k ;

[0183] (501) Divide E into Y groups according to the number of channels, and perform per-pixel splitting on each group of depth features to split them into C convolutional kernels;

[0184] Specifically, in the present invention, Y takes the value of 32 and C takes the value of 196;

[0185] (502) Obtain U according to the following formula vk is

[0186] U vk = Z * E vk | v∈{1,...,Y}k∈{1,...,C}

[0187] where, * represents the convolution operation, and E vk represents the k-th convolution kernel of E in the v-th group of depth features;

[0188] (503) Obtain U according to the following formula k is

[0189]

[0190] Specifically, in the present invention, the size of U k is 81×88×196, 81×88 is the resolution of U k , and 196 is the number of channels of U k In the embodiment of the present invention, the size of the spectral concatenated texture feature Z of the 51st frame of hyperspectral image is 81×88×16, and the size of the depth feature E is 14×14×512. The depth feature E is divided into 32 groups, with 16 channels in each group. In addition, each group of depth features is divided into 196 convolution kernels, and the size of each convolution kernel is 1x1x16. Finally, each convolution kernel is convolved with Z to obtain U with a size of 81×88×196 k feature, Figure 12 In the embodiment of the present invention, it is the U 51 feature map of the search region S of the 51st frame of normalized Minion hyperspectral image. k

[0191] Step Sixteen: Obtain the U k feature extracted from the search region of the 1st frame of hyperspectral image, send the U k feature extracted from the search region of the 1st frame into the DCF filter for training to obtain a filter template, and then send the U t+1 feature extracted from S k into the filter template to obtain the response map of S t+1 , and determine the target position through the response map of S t+1 ;

[0192] Specifically, Figure 13 in the embodiment of the present invention, it is the response map obtained by the search region S of the 51st frame of normalized Minion hyperspectral image entering the DCF filter. 51

[0193] Step 17: Generate the target image patch of the t-th hyperspectral image in the hyperspectral image sequence;

[0194] Specifically, it is implemented according to the following steps:

[0195] (601) Determine whether the current frame is the first hyperspectral image in the hyperspectral image sequence. If so, load the target position and target box of the first hyperspectral image in the hyperspectral image sequence. If not, proceed to the next step;

[0196] Specifically, this frame is the 50th hyperspectral image in the hyperspectral image sequence, and operation (602) is executed;

[0197] (602) On the (t + 1)-th hyperspectral image in the hyperspectral image sequence, with the target position as the center, intercept the corresponding image patch G through the target box of the t-th hyperspectral image in the hyperspectral image sequence t ;

[0198] Specifically, the corresponding image patch intercepted by the target box of the 50th hyperspectral image in the hyperspectral image sequence is G 50 ;

[0199] (603) Set the scale pool S according to the following formula p1 as

[0200] Sp1 = {I1, I2,..., I e}

[0201] where I e represents the e-th scale factor, and e represents the serial number of the scale factor;

[0202] Specifically, e is 3, and Sp1 is set to {0.99, 1.00, 1.01}, that is, I1 is 0.99, I2 is 1.00, and I3 is 1.01;

[0203] (604) Scale G through I p1 in S e to obtain a set H of target image patches of different sizes on the t-th hyperspectral image in the hyperspectral image sequence t as t H

[0204] = {H t |H X t = I X t × G e} t X∈{1,...,e}

[0205] where H t represents the X-th image patch, and X represents the ordinal number of the image patch;​

[0206] Specifically, for the 50th frame of the normalized Minion hyperspectral image,

[0207] (605) Use bilinear interpolation to ensure that the size of each H X t is the same as that of G t Extract U from all Hs X t features and put them into the DCF correlation filter to obtain e response maps. Calculate the maximum value of the e response maps, and define the target image patch corresponding to the maximum value of the e response maps as H k ; H t max is the target image patch of the t-th frame hyperspectral image in the hyperspectral image sequence; t max Specifically, H1

[0208] = 0.99×G 50 , H2 50 = 1.00×G 50 , H3 50 = 1.01×G 50 , the maximum value of the 3 response maps is H3 50 which is H 50 ; 50 max .

[0209] Step Eighteen: Load each frame of the hyperspectral image sequence in turn, repeat Steps One to Seventeen to obtain the target patches of each frame of the hyperspectral image, and complete the target tracking of the hyperspectral image sequence;

[0210] Specifically, Figure 14 In the embodiment of the present invention, this is the target tracking result map of the 51st frame of the normalized Minion hyperspectral image. Repeat Steps One to Seventeen to obtain the target patches of 376 frames of hyperspectral images, and complete the final target tracking of the hyperspectral image sequence.

[0211] The present invention proposes a new feature called spectral cascade texture feature, which is formed by combining the texture feature extracted by the LBP operator with the signal-to-noise ratio (SNR) spectral curve. As a new feature applicable to hyperspectral video target tracking, this feature contains rich spatial and spectral information;

[0212] Based on the segmentation result of the error threshold, the present invention generates a position mask for predicting the approximate position of the target in the next frame. In addition, the generated mask suppresses the texture features of the background. The present invention also proposes a new depth-averaged spectral cascade texture feature (U k), which is obtained by performing per-pixel convolution on the depth feature and the depth spectral cascaded texture prediction feature, contains rich spatial and spectral information. To combine texture, spectral, and semantic information, the present invention proposes a feature fusion method called per-pixel joint convolution, which can effectively preserve the local features of U k ;

[0213] As described above, only the preferred embodiments of the present invention are given, and they are not used to limit the protection scope of the present invention.

Claims

1. A hyperspectral video object tracking method based on depth spectral cascaded texture features, characterized in that, The method is as follows: Step 1: Load the target position, target bounding box, and target image patch of the first hyperspectral image in the hyperspectral image sequence. Then load the t-th hyperspectral image in the hyperspectral image sequence, and perform gray normalization on the t-th hyperspectral image to obtain the normalized t-th hyperspectral image, and determine the search region S of the normalized t-th hyperspectral image t , where the gray level range of the pixels of the t-th hyperspectral image is the left-closed and right-closed interval from 0 to 1, t represents the number of frames of the hyperspectral image, and t is an integer greater than or equal to 2 and less than or equal to 2000; Step 2: Centering on the target position of the (t - 1)-th frame hyperspectral image in the hyperspectral image sequence, intercept an image block of 3×3 pixels on S t as the local region L of the t-th frame hyperspectral image after normalization t . Determine the i-th band image t of L . According to the spectral average value of the i-th band image t of L , determine the local spectral curve C t of the target local region L of the t-th frame hyperspectral image after normalization l , and obtain the local spectral average value C l of C li at the i-th band; Step 3. According to S t subtract L t to obtain the unknown region U of the t-th frame hyperspectral image after normalization t , and define the pixels within U t as unknown pixels; Step 4. Determine the spectral curve C of the unknown pixel uj , and obtain C uj The gray value C of the j-th pixel in the i-th band uij , where j represents the serial number of the unknown pixel; Step Five: Obtain the error ε by subtracting the absolute value of C uij from C li ; i,j ​ Step 6. Determine ε i,j The cumulative sum Q j Compare Q j with the preset error threshold η, and determine whether the j-th unknown pixel U t in U t (j) belongs to the background pixel or the target pixel; define the set of target pixels as the target region O t of the t-th frame hyperspectral image, and define the set of background pixels as the background region J t of the t-th frame hyperspectral image, where η > 0; Step 7: Load the search region S of the (t + 1)-th frame hyperspectral image after normalization t+1 , and based on S t+1 obtain the spectral curves C of each pixel point in the search region s ; Step 8. Determine C si and C li 's Euclidean distance D j ; C si represents the spectral value of C s at the i-th band; Step Nine: Determine the depth feature E of a sequence of single-band grayscale images after dimensionality reduction through the GoogLeNet network; Step Ten: Extract the texture feature P of S through the LBP operator and the depth feature E t+1 ; Step Eleven: Obtain S t The i-th band image of Determine The image signal-to-noise ratio of According to Determine the signal-to-noise ratio spectral value C of the t-th frame of hyperspectral image in the i-th band ri , through C ri Obtain the signal-to-noise ratio spectral curve C of the search area r , according to C r Determine the weight w of the i-th spectral channel i ; Step Twelve: Determine S i through P and w t+1 for the spectral cascade texture feature R; Step 13: Generate a position prediction mask M through O t and J t ​ Step Fourteen: Determine S through R and M t+1 The spectral cascade texture prediction feature Z of Step 15: Divide E into Y groups according to the number of channels, perform per-pixel splitting on the depth features of each group, split them into C convolution kernels, and then perform convolution on Z through the k-th convolution kernel of the v-th group to obtain the k-th depth spectral concatenated texture feature U of the v-th group vk , and average the Y groups of U vk to obtain the depth-averaged spectral concatenated texture feature U k , where Y takes the value of 32, C takes the value of 196, v is an integer greater than or equal to 1 and less than or equal to Y, and v represents the serial number of the depth spectral concatenated feature U vk , k is an integer greater than or equal to 1 and less than or equal to C, and k represents the serial number of the convolution kernel; Step Sixteen: Obtain the U feature extracted from the search region of the first frame of the hyperspectral image, send the U feature extracted from the search region of the first frame into the DCF filter for training to obtain a filter template, and then send the U feature extracted from S into the filter template to obtain the response map of S, and determine the target position through the response map of S; k feature, send the U k feature extracted from the search region of the first frame into the DCF filter for training to obtain a filter template, and then send the U t+1 feature extracted from S k into the filter template to obtain the response map of S t+1 , and determine the target position through the response map of S t+1 ; Step Seventeen: Generate the target image patch of the t-th frame hyperspectral image in the hyperspectral image sequence; Step Eighteen: Load each frame of the hyperspectral image in the hyperspectral image sequence in turn, repeat Steps One to Seventeen, obtain the target patch of each frame of the hyperspectral image, and complete the target tracking of the hyperspectral image sequence.

2. The hyperspectral video object tracking method based on depth spectral cascaded texture features according to claim 1, characterized in that, In Step 4, determine C according to the following formula uj is where j ranges from 1 to p×q, C uij is C uj the spectral value in the i-th band, is the pixel value of the j-th pixel in the i-th band of the t-th frame in the unknown area, and B represents the total number of bands.

3. The hyperspectral video object tracking method based on depth spectral cascaded texture features according to claim 1 or 2, characterized in that, The specific implementation of Step Six is as follows: (601) Set the error threshold η; (602) Calculate the error ε according to the following formula i,j The cumulative sum Q of j is (603) Determine whether the unknown pixel belongs to the target area or the background area according to the following formula; Among them, when Q j is greater than or equal to η, C uj is the pixel within the background region J t When Q j is less than η, C uj is the pixel within the target region O t B represents the total number of bands.

4. The hyperspectral video object tracking method based on depth spectral cascaded texture features according to claim 3, characterized in that, The specific implementation of Step Eleven is as follows: (1101) Determine according to the following formula be where, A(.) represents the operation of taking the average value, represents the target region of the i-th band of the t-th frame hyperspectral image, represents the background region of the i-th band of the t-th frame hyperspectral image; (1102) Determine the signal-to-noise ratio spectral curve C of the search area according to the following formula r is Among them, C ri represents the signal-to-noise ratio spectral value of the t-th frame hyperspectral image in the i-th band; (1103) Normalize according to the following formula to obtain the normalized value γ of the image signal-to-noise ratio of the i-th band after the normalization process i is (1104) is to amplify the difference of γ i After that, γ i is non-linearly transformed according to the following formula to obtain γ i The value F after non-linear transformation i is F i = (γ i ) 2 | i∈{1,...,B} (1105) Determine w according to the following formula i be 5. The hyperspectral video object tracking method based on deep spectral cascaded texture features according to claim 4, wherein, The specific steps of Step Twelve are as follows: (1201) Obtain the spectral cascade texture feature R of R in the i-th band according to the following formula i is R i = w i × P i | i∈{1,...,B} where P i represents the texture feature of P in the i-th band; After i iterates from 1 to B, all sets of R are obtained i The set is R R = {R i | i∈{1,...,B}}.

6. The hyperspectral video object tracking method based on deep spectral cascaded texture features according to claim 5, wherein, The specific content of Step Thirteen is as follows: (1301) Generate a matrix M1 with the same size as S t and set the initial value of the matrix to 0; (1302) Update M1 according to the following formula, and judge whether each unknown pixel belongs to O t or J t . If the unknown pixel belongs to O t , then assign 1 to the corresponding area position in M1. If the unknown pixel belongs to J t , then assign 0.5 to the corresponding area position in M1. M1 is (1303) Generate a position prediction mask M with the same size as M1, and assign M1 to M.

7. The hyperspectral video object tracking method based on deep spectral cascaded texture features according to claim 6, wherein, The specific content of Step Fourteen is as follows: Multiply the position prediction mask M by the spectral concatenated texture feature R, and obtain the spectral concatenated texture prediction feature Z according to the following formula as Z = M⊙R where ⊙ represents the dot product operation.

8. The hyperspectral video object tracking method based on deep spectral cascaded texture features according to claim 7, wherein, The specific determination of Step Fifteen is through the following steps: (1501) Divide E into Y groups according to the number of channels, perform per-pixel splitting on each group of depth features, and split them into C convolutional kernels; (1502) Obtain U according to the following formula vk where U vk = Z * E vk | v∈{1,...,Y}k∈{1,...,C} Among them, * represents the convolution operation, and E vk represents the k-th convolution kernel of E in the v-th group of depth features; (1503) Obtain U according to the following formula k be 9. The hyperspectral video object tracking method based on deep spectral cascaded texture features according to claim 8, wherein, The specific implementation of Step Seventeen is based on the following steps: (1701) Determine whether the current frame is the first frame of the hyperspectral image sequence. If so, load the target position and target box of the first frame of the hyperspectral image in the hyperspectral image sequence. If not, execute the next step; On the (t + 1)-th frame hyperspectral image in the hyperspectral image sequence, centered at the target position, intercept the corresponding image patch G by the target box of the t-th frame hyperspectral image in the hyperspectral image sequence t ; (1703) Set the scale pool S according to the following formula p1 Let Sp1 = {I1, I2,..., I e} Among them, I e represents the e-th scaling factor, and e represents the serial number of the scaling factor; (1704) Scale G by I in S according to the following formula to obtain a set H of target image patches of different sizes on the t-th frame hyperspectral image in the hyperspectral image sequence p1 where I in e and G in t are used to generate the set H of target image patches of different sizes on the t-th frame hyperspectral image in the hyperspectral image sequence t Let H be t ={H X t |H X t =I e ×G t} X∈{1,...,e} Among them, H t represents the Xth image block, and X represents the ordinal number of the image block; (1705) Use bilinear interpolation to ensure that each H X t has the same size as G t and extract U X t features from all H k and put them into the DCF - related filter to obtain e response maps. Calculate the maximum values of the e response maps, and define the target image patch corresponding to the maximum value of the e response maps as H t max . H t max is the target image patch of the t - th frame hyperspectral image in the hyperspectral image sequence.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on NSCT and SAE

    CN107122733A

  • Hyperspectral target tracking method combining spatial spectrum characteristics and correlation filtering

    CN110276782A