Hyperspectral Target Detection and Recognition Method and System Based on Semantic and Hyperspectral Feature Fusion
By fusing semantic and null spectral features in hyperspectral images, using technologies such as unsupervised anomaly detection, dual-stream convolutional neural network and cubic long and short-time memory network, real-time and accurate detection and classification of targets in hyperspectral images is solved, and the problem of non-real-time target detection and high false alarm rates in the prior art is solved.
Patent Information
- Application Number
- CN202510472960.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Existing hyperspectral image object detection methods cannot achieve real-time object detection and classification, especially targets with strong motion characteristics such as ships and aircraft are difficult to effectively track and identify.
The hyperspectral object detection and recognition method based on the fusion of semantic and null spectral features is adopted, and the unsupervised anomaly detection is performed, and the dual-stream convolutional neural network is performed for fine detection, combined with the cubic long and short-term memory network for dynamic tracking, and the target classification is performed using a support vector machine based on historical spectral training.
The spatial domain object detection at the pixel-by-pixel level of hyperspectral images is realized, the target detection accuracy is improved, the false alarm rate is reduced, and the problem of difficulty in capturing time-sensitive target trajectory in the air is solved.
Smart Images

Figure CN119992238B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of hyperspectral remote sensing image target detection methods and systems, and specifically, to a hyperspectral target detection and recognition method and system based on the fusion of semantic and spatio-spectral features. Background Art
[0002] Hyperspectral remote sensing images acquire geometric, radiometric, and spectral information of the scene, forming three-dimensional image information. Through spectral subdivision, hyperspectral technology obtains the spectral curve of any pixel. Different substances have different characteristic spectral lines, which are the "fingerprints" of the substances themselves. The use of hyperspectral detection technology can achieve the ability to finely identify land, ocean, and air targets.
[0003] However, the amount of data in the entire image scene of hyperspectral images is huge, and it still needs to be sent to the ground and processed manually in large quantities before it can be applied. A large amount of relevant information about targets strongly related to motion characteristics such as ships and airplanes cannot be obtained in real time. Therefore, there is an urgent need to design a hyperspectral target detection and recognition method, system, and electronic device based on the fusion of semantic and spatio-spectral features to perform high-confidence tracking, detection, and classification of the above targets.
[0004] Patent application document CN116958807A discloses a hyperspectral image target detection method based on momentum-free contrast learning and Transformer network, including: First, an encoder and a momentum encoder for spectral feature extraction in hyperspectral target detection tasks are designed based on the encoder of the Transformer. In order to not ignore the local detail information in the spectrum while paying attention to the long-distance dependence and self-similarity of the spectrum, the spectral feature extraction encoder and the momentum encoder focus on the local detail information of the spectrum through the designed overlapping spectral block feature mapping and interactive token feed-forward layer. Secondly, spectral discrimination ability learning is carried out through the method of unsupervised momentum contrast learning, where the queue and the momentum encoder updated slowly in a momentum manner are used to provide a sufficient number of and highly consistent negative sample features to help the model learn better representations. Finally, the detection results obtained through cosine similarity are non-linearly stretched twice through exponential and normalization operations, and power function and normalization operations to suppress the background. However, this patent cannot completely solve the existing technical problems and also cannot meet the requirements of the present invention. Summary of the Invention
[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a hyperspectral target detection and recognition method and system based on the fusion of semantic and spatio-spectral features.
[0006] According to the hyperspectral target detection and recognition method based on the fusion of semantic and spatio-spectral features provided by the present invention, it includes:
[0007] Step 1: Extract abnormal targets in the hyperspectral image based on the unsupervised anomaly detection algorithm, perform rapid and rough target detection through the constrained energy minimization operator, and eliminate false alarms based on the inter-frame motion characteristics comparison;
[0008] Step 2: Use a two-stream convolutional neural network to perform fine detection on the rough detection results, obtain the spatial-spectral feature information of the target, and combine the cubic long short-term memory network to predict the confidence interval range of the target to achieve dynamic tracking;
[0009] Step 3: Use the support vector machine trained based on historical spectra to classify the detected target and determine its category.
[0010] Preferably, the unsupervised anomaly detection algorithm in step 1 adopts the RX detection operator, and the expression is:
[0011]
[0012] in, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image;
[0013] By calculating the energy distribution of the target in the direction of the smallest eigenvalue of the covariance matrix and combining it with the target motion distance threshold between adjacent frames, false alarms are eliminated.
[0014] Preferably, the two-stream convolutional neural network in step 2 includes an upper branch and a lower branch, each branch includes an input, 9 convolutional layers are used in each branch to extract rich spectral information of the input pixels, and a one-dimensional convolutional layer is used to implement the convolution operation;
[0015] In the network, a convolution layer with a kernel step size of 2 is used instead of a pooling layer. To maximize the retention of spectral features, all features extracted by the convolution layer with a kernel step size of 2 are added to the features extracted by the last layer based on different average pooling layers. Then, the final features of each branch are obtained through AVG pooling layer and fully connected layer operations.
[0016] In a two-stream convolutional neural network, is the target prior pixel, is the target pixel, is the background pixel, constructed according to the training samples, and the input of the upper branch is always , the input of the current branch is When , the label of the training sample is 1, and the input of the next branch is When , the label is 0; after multiple convolution operations, pooling operations and a full connection operation, the final features of the two branches are obtained, recorded as and , and then combine the two features:
[0017]
[0018] Finally, the output of the two-stream convolutional neural network is obtained through the last fully connected layer and a Sigmoid function.
[0019] Preferably, the cubic long short-term memory network in step 2 is composed of a space branch, a time branch and an output branch, the input includes the latitude and longitude, speed, acceleration and historical trajectory of the target, and the output is the predicted result of the target position at the next moment; by calculating the target pixel movement speed , and the motion distance of adjacent frames , the confidence interval is limited to the 10-pixel neighborhood of the previous frame position, where V is the target speed, is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging, and f is the video frame rate.
[0020] Preferably, in step 3, the support vector machine adopts a soft margin optimization model, trains a classifier based on historical spectral data, and distinguishes between aircraft, ships and other target categories, and its loss function is:
[0021]
[0022] in, is the weight vector, is the bias term, and are the feature vector and corresponding label of the i-th sample respectively.
[0023] The hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion provided by the present invention comprises:
[0024] Coarse detection module: extracts abnormal targets in hyperspectral images based on unsupervised anomaly detection algorithms, performs rapid coarse target detection through constrained energy minimization operators, and eliminates false alarms based on inter-frame motion characteristic comparisons;
[0025] Precision detection module: Use a two-stream convolutional neural network to perform precision detection on the rough detection results, obtain the spatial-spectral feature information of the target, and combine the cubic long short-term memory network to predict the confidence interval range of the target to achieve dynamic tracking;
[0026] Classification module: Use support vector machines trained based on historical spectra to classify the detected targets and determine their categories.
[0027] Preferably, the unsupervised anomaly detection algorithm in the coarse detection module adopts the RX detection operator, which is expressed as:
[0028]
[0029] Among them, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image;
[0030] By calculating the energy distribution of the target in the direction of the small eigenvalues of the covariance matrix and combining the target motion distance threshold between adjacent frames, false alarms are excluded.
[0031] Preferably, the two-stream convolutional neural network in the fine detection module includes an upper branch and a lower branch. Each branch contains an input. In each branch, 9 convolutional layers are used to extract rich spectral information of the input pixels, and a one-dimensional convolutional layer is used to implement the convolution operation;
[0032] In the network, a convolutional layer with a kernel stride of 2 is used instead of the pooling layer. To maximize the retention of spectral features, based on different average pooling layers, all the features extracted by the convolutional layer with a kernel stride of 2 are added to the features extracted by the last layer, and then the final features of each branch are obtained through the operations of the AVG pooling layer and the fully connected layer;
[0033] In the two-stream convolutional neural network, is the target prior pixel, is the target pixel, is the background pixel, constructed according to the training samples. The input of the upper branch is always , when the input of the lower branch is , the label of the training sample is 1, and when the input of the lower branch is , the label is 0; after multiple convolution operations, pooling operations and one fully connected operation, the final features of the two branches are obtained, denoted as and , and then the two features are combined:
[0034]
[0035] Finally, the output of the two-stream convolutional neural network is obtained through the last fully connected layer and a Sigmoid function.
[0036] Preferably, the cubic long short-term memory network in the fine detection module is composed of a spatial branch, a temporal branch and an output branch. The input includes the longitude, latitude, speed, acceleration and historical trajectory of the target, and the output is the predicted result of the target position at the next moment; by calculating the target pixel motion speed , and the motion distance between adjacent frames , the confidence interval is limited to the 10-pixel neighborhood range of the previous frame position. Among them, V is the target speed, is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging, and f is the video frame rate.
[0037] Preferably, the support vector machine in the classification module adopts a soft margin optimization model, trains a classifier based on historical spectral data to distinguish aircraft, ships, and other target categories, and its loss function is:
[0038]
[0039] where is the weight vector, is the bias term, and are the feature vector and the corresponding label of the i-th sample, respectively.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] The present invention provides a hyperspectral target detection and recognition method based on the fusion of semantic and spatial-spectral features. A semantic segmentation network model is constructed in the spatial domain of the hyperspectral image, an adaptive spatial-spectral joint optimization model is designed, and a spatial-spectral joint cascade detector is constructed, realizing pixel-level spatial domain target detection in the hyperspectral image, improving the target detection accuracy and reducing the false alarm rate, and solving the problem of difficult trajectory capture of airborne time-sensitive targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:
[0043] Figure 1 is a flowchart of a hyperspectral target detection and recognition method based on the fusion of semantic and spatial-spectral features of the present invention;
[0044] Figure 2 is a flowchart of an anomaly detection based on an unsupervised method of the present invention;
[0045] Figure 3a and Figure 3b are a flowchart and a result diagram of an improved constrained energy minimization target detection of the present invention, respectively;
[0046] Figure 4 is a structure diagram of a two-stream convolutional neural network of the present invention;
[0047] Figure 5a and Figure 5b are a flowchart and a result diagram of a target classification and recognition of a support vector machine of the present invention, respectively;
[0048] Figure 6 is a structure diagram of a hyperspectral target detection and recognition system based on the fusion of semantic and spatial-spectral features of the present invention;
[0049] Figure 7 This is a structural diagram of a hyperspectral target detection and recognition electronic device based on the fusion of semantic and spatial-spectral features of the present invention. Specific implementation manner
[0050] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all belong to the protection scope of the present invention.
[0051] Embodiment
[0052] Refer to Figure 1 , this embodiment provides a hyperspectral target detection and recognition method based on the fusion of semantic and spatial-spectral features, mainly including:
[0053] Step 1: Extract abnormal targets in the hyperspectral image through anomaly detection based on unsupervised methods, perform rapid rough detection of the targets through the CEM operator, and exclude false alarms by comparing inter-frame information;
[0054] Step 2: Perform fine detection through a two-stream convolutional neural network, obtain target-related information, and use a cubic long short-term memory network to determine the confidence interval range to track and detect the target;
[0055] Step 3: Classify the detected targets using an SVM trained based on historical spectra to determine which type of target it belongs to.
[0056] Furthermore, refer to Figure 2 , this embodiment provides a flowchart of anomaly detection based on unsupervised methods.
[0057] The anomaly detection method based on unsupervised methods uses the RX detection operator, and its form is:
[0058]
[0059] Among them, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image. In the above formula has the same form as the Mahalanobis distance. Essentially, the RX algorithm can be regarded as the inverse process of principal component analysis. Principal component analysis compresses most of the meaningful image information from the original image feature space into a space with a few uncorrelated principal components as the basis. Obviously, some targets with very low occurrence probabilities in the image (small targets, abnormal targets) will not be included in these principal components. On the contrary, they have a greater probability of appearing in the covariance matrix On the eigenvector direction corresponding to the small eigenvalue. The RX algorithm searches for abnormal targets by calculating value. If there are abnormal targets in the graph, then their corresponding energy will be very small and may correspond to the small eigenvalues of the covariance matrix . And the smaller the eigenvalue, the larger the value, so that the abnormal targets in the image can be effectively detected.
[0060] Using the RX detection operator, targets such as airplanes and ships existing in the hyperspectral image can be detected as abnormal targets. However, other abnormal points in the image often output as noise. Therefore, the abnormal detection realizes the function of rough target detection when the target spectrum is unknown. Further precise detection is needed to eliminate noise and reduce the false alarm rate.
[0061] Furthermore, referring to Figure 3a , this embodiment provides a flowchart of an improved constrained energy minimization target detection.
[0062] Denote as the set of all observed samples, where is an arbitrary sample pixel vector, is the number of pixels, is the number of image bands. Assume that is the target of interest. The purpose of CEM is to design a FIR linear filter to minimize the filtered output energy:
[0063]
[0064] is the covariance matrix, and the solution of the above formula is the CEM operator , that is:
[0065]
[0066] Applying the CEM operator to each pixel in the image will obtain the distribution of the target in the image and realize the detection of the target.
[0067] Assume that V is the target moving speed (m / s), is the flight inclination angle of the airplane, r is the spatial resolution (m) of the video hyperspectral camera imaging, f is the video frame rate (fps), and W is the video width (m), as shown in Table 1.
[0068] According to the above information, the target pixel movement speed reflected on the image plane (pixels / s) can be calculated as:
[0069]
[0070] Adjacent frame target motion distance (in pixels) is:
[0071]
[0072] Time for the target to stay in the field of view (in s) is:
[0073]
[0074] Number of frames for the target to stay in the field of view is:
[0075]
[0076] Table 1. Motion characteristics of aircraft and ships reflected in the image
[0077]
[0078] Through calculation, it can be obtained that the residence time of the target in the field of view is at least 12 s, and the imaging frequency of the hyperspectral video satellite is 5 frames per second. Therefore, the target will appear continuously within at least 60 frames. The distance that the moving target moves between two adjacent frames is conservatively estimated to be less than 20 pixels. Therefore, the target usually appears in the vicinity of the target position in the previous frame (within the 20-pixel neighborhood).
[0079] Therefore, for anomaly detection of two consecutive frames of images, for all targets detected in the previous frame of image, if they do not appear within the range of 20 pixels in the next frame, they can be identified as false alarms.
[0080] Furthermore, referring to Figure 3b , this embodiment provides a result map of an improved constrained energy minimization target detection.
[0081] Furthermore, referring to Figure 4 , this embodiment provides a structure diagram of a two-stream convolutional neural network.
[0082] Before training, a hybrid pixel selection strategy based on sparse representation and classification is proposed to select typical background samples in the hyperspectral image, and then sufficient target samples are generated through some typical background samples and target priors. During the training process, the training samples (positive training samples labeled 1 constructed from target priors and target samples, and negative training samples labeled 0 constructed from target priors and background samples) are input into a carefully designed two-stream convolutional network to learn the discriminative ability. During the testing process, the test samples (composed of target priors and detected pixels) are classified by the well-trained two-stream convolutional network. The output of the network constitutes the final detection result.
[0083] The dual-stream convolutional neural network includes two branches: an up branch and a down branch. Each branch contains an input. In each branch, nine convolutional layers are used to extract rich spectral information of the input pixels. One-dimensional convolutional layers are used to implement the convolutional operations, followed by ReLU layers. Considering that the pooling layer may cause loss of spectral information when spectral dimension reduction is required, convolutional layers with a kernel stride of 2 are used in the network instead of the pooling layer. To maximize the retention of spectral features, based on different average pooling layers, all the features extracted by the convolutional layer with a kernel stride of 2 are added to the features extracted by the last layer. Then the final features of each branch are obtained through AVG pooling layer and fully connected layer operations.
[0084] In the network, is the target prior pixel, is the target pixel, is the background pixel. According to the construction of the previous training samples, the input of the up branch is always , and when the input of the down branch is , the label of the training sample is 1, and when the input of the down branch is , the label is 0. After some convolutional operations, several pooling operations and one fully connected operation, the final features of the two branches are obtained, denoted as 1 and 2; then the two features are combined:
[0085]
[0086] Finally, the output of the network is obtained through the last fully connected layer and a Sigmoid function.
[0087] There are two loss functions in the proposed network. First is the binary cross-entropy loss:
[0088]
[0089] where, is the batch size, is the label of the training sample, is the output of the Sigmoid function. The other is the ICS loss. To improve the effect of target-background separability, the ICS loss is proposed. If the inputs of the two branches are and (label is 1), then they belong to the same class and the distance between them should be minimized; otherwise, they belong to different classes and the distance between them should be maximized. Therefore, the proposed ICS loss is expressed as:
[0090]
[0091] Among them, 1 is the extracted upper branch feature, 2 is the extracted lower branch feature; and are the extracted feature vectors respectively;
[0092] The final loss function is the sum of the ICS loss and the BCE loss, expressed as:
[0093]
[0094] The Cubic Long Short-Term Memory Network is a new structure developed based on the LSTM, consisting of three branches. The spatial branch is used to capture the spatial information of the moving object, the temporal branch is used to process the motion over time, and the output branch is used to combine the previous two branches to generate the predicted frame.
[0095] The spatial branch flows along the z-axis (spatial axis), and the convolution on the z-axis is responsible for capturing and analyzing the moving object. The spatial state is generated by this branch, which carries information about the spatial layout of the moving object.
[0096] The temporal branch flows along the x-axis (time axis), and the convolution aims to obtain and process the motion. The temporal state is generated by this branch, which contains motion information.
[0097] The output branch generates the intermediate or final predicted frame along the y-axis (output axis) according to the predicted motion provided by the temporal branch and the information of the moving object provided by the spatial branch.
[0098] Processing the temporal information and the spatial information separately can obtain better predictions, and this separation can reduce the prediction burden of the network. Stacking multiple CubicLSTM units along the spatial branch and the output branch can form a two-dimensional network. This two-dimensional network can be further constructed into a three-dimensional network (CubicRNN) by evolving along the time axis, and by stacking three layers spatially, the information of the tracking target can be made more prominent and better spatial information can be obtained.
[0099] The motion characteristics of airborne highly maneuverable targets are divided into two categories: "short-term motion characteristics" and "long-term motion characteristics". Among them, the short-term motion characteristics include the longitude, latitude, speed, and acceleration of the target, which may change significantly in a short period. The long-term task characteristics include the historical motion trajectory of the target. Compared with the short-term motion characteristics, the long-term task characteristics are relatively stable during the target's motion process. To make target tracking more accurate, it is necessary to utilize both the short-term motion characteristics and the long-term task characteristics of the target. Therefore, a CubicLSTM network is adopted, and the short-term motion characteristics and the long-term task characteristics of the target are jointly used as the input of the network to predict the target state at the next moment and calculate the confidence range of the target. The definitions of the input and output variables of the adopted CubicLSTM network are as follows:
[0100] Table 2. Input-output relationship
[0101]
[0102] After obtaining the prediction of the target position at the next moment through the CubicLSTM network, further, it is necessary to determine the confidence interval of the target according to the motion characteristics of the target, and within this range, detect the target to obtain the true position of the target, so as to achieve the tracking of the target. Next, the confidence interval range will be analyzed in detail according to the characteristics of the target motion.
[0103] Assume that V is the target speed (m / s), is the orbital inclination angle, r is the spatial resolution (m) of the video hyperspectral camera imaging, f is the video frame rate (fps), and W is the video width (m).
[0104] According to the above information, the target pixel motion speed (pixels / s) reflected on the image plane can be calculated as:
[0105]
[0106] The target motion distance between adjacent frames (pixels) is:
[0107]
[0108] The time (s) for the target to stay in the field of view is:
[0109]
[0110] The number of frames for the target to stay in the field of view is:
[0111]
[0112] It can be calculated that the moving target moves a distance of less than 2.5 pixels between two adjacent frames, and conservatively estimated, it should also be less than 10 pixels. Therefore, the target usually appears in the vicinity area of 10 pixels around the target position in the previous frame. Repeating the target detection within this neighborhood can capture the position of the target in the next frame.
[0113] Furthermore, referring to Figure 5a , this embodiment provides a flowchart for target classification and recognition using a support vector machine.
[0114] Using a hard margin SVM in a linearly inseparable problem will result in classification errors. Therefore, a loss function can be introduced on the basis of maximizing the margin to construct a new optimization problem. Given the input data and the learning objective , the SVM uses a hinge loss function, and the optimization problem of the soft margin SVM is expressed as follows:
[0115]
[0116]
[0117] where is the weight vector, is the bias term, is the regularization parameter, is the number of samples, and are the feature vector and the corresponding label of the i-th sample respectively;
[0118] It can be seen from the above formula that the soft margin SVM is a regularized classifier, and in the formula represents the hinge loss function.
[0119] After the SVM classifier is trained based on historical spectra, it can classify the targets in the refined detection results and distinguish targets such as airplanes and ships.
[0120] Furthermore, referring to Figure 5b , this embodiment provides a target classification and recognition result graph of a support vector machine.
[0121] Furthermore, referring to Figure 6 , this embodiment provides a hyperspectral target detection and recognition system based on the fusion of semantics and spatial-spectral features, mainly including:
[0122] A coarse detection module, which inputs the hyperspectral image to be detected, detects abnormal targets in the hyperspectral image based on the anomaly detection RX operator, performs rapid coarse detection of the target through the CEM operator, and eliminates false alarms by comparing inter-frame information.
[0123] The fine detection module performs fine detection through a two-stream convolutional neural network, obtains target-related information, and uses a cubic long short-term memory network to determine the confidence interval range for tracking and detecting the target.
[0124] The classification module uses an SVM trained based on historical spectra to classify the detected targets and determine which type of target they belong to.
[0125] Further, referring to Figure 7 , this embodiment provides an electronic device, mainly including: at least one memory and at least one processor. Instructions are stored in the at least one memory, and when the instructions are executed by the at least one processor, a hyperspectral target detection and recognition method based on semantic and spatial spectral feature fusion according to an exemplary embodiment of the present disclosure is executed.
[0126] As an example, the electronic device can be a PC computer, a tablet device, a personal digital assistant, a smartphone, or other devices capable of executing the above instructions. Here, the electronic device does not have to be a single electronic device, and can also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device can also be a part of an integrated control system or a system manager, or can be configured as a portable electronic device that interfaces with a local or remote (e.g., via wireless transmission).
[0127] In the electronic device, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. As an example and not a limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0128] The processor can run the instructions or code stored in the memory. Among them, the memory can also store data. The instructions and data can also be sent and received via a network interface device through a network, where the network interface device can adopt any known transmission protocol.
[0129] The memory can be integrated with the processor. For example, RAM or flash memory is arranged within an integrated circuit microprocessor, etc. In addition, the memory can include independent devices, such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory and the processor can be operatively coupled, or can communicate with each other, for example, through I / O ports, network connections, etc., so that the processor can read the files stored in the memory.
[0130] In addition, the electronic device may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device can be connected to each other via a bus and / or a network.
[0131] Those skilled in the art know that, in addition to implementing the systems, devices and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. Therefore, the systems, devices and their respective modules provided by the present invention can be regarded as a kind of hardware components, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware components; the modules for implementing various functions can also be regarded as either software programs for implementing methods or the structures within the hardware components.
[0132] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion, characterized in that: include: Step 1: Extract abnormal targets in the hyperspectral image based on the unsupervised anomaly detection algorithm, perform rapid and rough target detection through the constrained energy minimization operator, and eliminate false alarms based on the inter-frame motion characteristics comparison; Step 2: Use a two-stream convolutional neural network to perform fine detection on the rough detection results, obtain the spatial-spectral feature information of the target, and combine the cubic long short-term memory network to predict the confidence interval range of the target to achieve dynamic tracking; Step 3: Use the support vector machine trained based on historical spectra to classify the detected target and determine its category; The two-stream convolutional neural network in step 2 includes an upper branch and a lower branch, each branch includes an input, 9 convolutional layers are used in each branch to extract rich spectral information of the input pixels, and a one-dimensional convolutional layer is used to implement the convolution operation; In the network, a convolution layer with a kernel step size of 2 is used instead of a pooling layer. To maximize the retention of spectral features, all features extracted by the convolution layer with a kernel step size of 2 are added to the features extracted by the last layer based on different average pooling layers. Then, the final features of each branch are obtained through AVG pooling layer and fully connected layer operations. In a two-stream convolutional neural network, is the target prior pixel, is the target pixel, is the background pixel, constructed according to the training samples, and the input of the upper branch is always , the input of the current branch is When , the label of the training sample is 1, and the input of the next branch is When , the label is 0; after multiple convolution operations, pooling operations and a full connection operation, the final features of the two branches are obtained, recorded as and , and then combine the two features: Finally, the output of the two-stream convolutional neural network is obtained through the last fully connected layer and a Sigmoid function; In step 2, the cubic long short-term memory network is composed of a space branch, a time branch and an output branch. The input includes the latitude and longitude, speed, acceleration and historical trajectory of the target, and the output is the predicted result of the target position at the next moment; by calculating the target pixel movement speed , and the motion distance of adjacent frames , the confidence interval is limited to the 10-pixel neighborhood of the previous frame position, where V is the target speed, is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging, and f is the video frame rate.
2. The hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion according to claim 1 is characterized in that: The unsupervised anomaly detection algorithm in step 1 adopts the RX detection operator, and the expression is: in, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image; By calculating the energy distribution of the target in the direction of the smallest eigenvalue of the covariance matrix and combining it with the target motion distance threshold between adjacent frames, false alarms are eliminated.
3. The hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion according to claim 1 is characterized in that: In step 3, the support vector machine adopts a soft margin optimization model to train a classifier based on historical spectral data to distinguish between aircraft, ships and other target categories, and its loss function is: in, is the weight vector, is the bias term, and are the feature vector and corresponding label of the i-th sample respectively.
4. A hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion, characterized in that: include: Coarse detection module: extracts abnormal targets in hyperspectral images based on unsupervised anomaly detection algorithms, performs rapid coarse target detection through constrained energy minimization operators, and eliminates false alarms based on inter-frame motion characteristic comparisons; Precision detection module: Use a two-stream convolutional neural network to perform precision detection on the rough detection results, obtain the spatial-spectral feature information of the target, and combine the cubic long short-term memory network to predict the confidence interval range of the target to achieve dynamic tracking; Classification module: Use the support vector machine based on historical spectrum training to classify the detected target and determine its category; The two-stream convolutional neural network in the precision detection module includes an upper branch and a lower branch, each branch contains an input, 9 convolutional layers are used in each branch to extract the rich spectral information of the input pixels, and a one-dimensional convolutional layer is used to implement the convolution operation; In the network, a convolution layer with a kernel step size of 2 is used instead of a pooling layer. To maximize the retention of spectral features, all features extracted by the convolution layer with a kernel step size of 2 are added to the features extracted by the last layer based on different average pooling layers. Then, the final features of each branch are obtained through AVG pooling layer and fully connected layer operations. In a two-stream convolutional neural network, is the target prior pixel, is the target pixel, is the background pixel, constructed according to the training samples, and the input of the upper branch is always , the input of the current branch is When , the label of the training sample is 1, and the input of the next branch is When , the label is 0; after multiple convolution operations, pooling operations and a full connection operation, the final features of the two branches are obtained, recorded as and , and then combine the two features: Finally, the output of the two-stream convolutional neural network is obtained through the last fully connected layer and a Sigmoid function; The cubic long short-term memory network in the precision detection module is composed of a space branch, a time branch and an output branch. The input includes the latitude and longitude, speed, acceleration and historical trajectory of the target, and the output is the predicted result of the target position at the next moment; by calculating the target pixel movement speed , and the motion distance of adjacent frames , the confidence interval is limited to the 10-pixel neighborhood of the previous frame position, where V is the target speed, is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging, and f is the video frame rate.
5. The hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion according to claim 4 is characterized in that: The unsupervised anomaly detection algorithm in the coarse detection module adopts the RX detection operator, and the expression is: in, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image; By calculating the energy distribution of the target in the direction of the smallest eigenvalue of the covariance matrix and combining it with the target motion distance threshold between adjacent frames, false alarms are eliminated.
6. The hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion according to claim 4 is characterized in that: The support vector machine in the classification module adopts a soft margin optimization model to train a classifier based on historical spectral data to distinguish between aircraft, ships and other target categories. The loss function is: in, is the weight vector, is the bias term, and are the feature vector and corresponding label of the i-th sample respectively.
Citation Information
Patent Citations
Hyperspectral target detection method based on unsupervised momentum contrast learning
CN116958807A
Hyperspectral image classification method based on combined multi-level spatial spectrum information CNN
CN110084159A
Sensing abnormal data real-time detection method based on artificial intelligence
CN111582298A