Hyperspectral target detection and identification method and system based on semantic and spatial-spectral feature fusion
By fusing semantic and null spectral features in hyperspectral remote sensing images, combining unsupervised anomaly detection, dual-stream convolutional neural network and support vector machine, high-precision and high-efficiency target detection and tracking are achieved, solving the problems of high false alarm rate and difficulty in target trajectory capture in the existing technology.
Patent Information
- Application Number
- CN202510472960.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing hyperspectral remote sensing image object detection method cannot achieve real-time tracking and detection, and the false alarm rate is high, making it difficult to effectively identify targets with strong motion characteristics.
The hyperspectral object detection and recognition method based on the fusion of semantic and null spectral features is adopted, and the unsupervised anomaly detection is performed, and the dual-stream convolutional neural network is used for fine detection, combining the cubic long and short-term memory network to achieve dynamic tracking, and a support vector machine is used for target classification.
The spatial domain object detection at the pixel-by-pixel level of hyperspectral images is realized, the target detection accuracy is improved, the false alarm rate is reduced, and the problem of difficulty in capturing time-sensitive target trajectory in the air is solved.
Smart Images

Figure CN119992238A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of hyperspectral remote sensing image target detection methods and systems, and in particular to a hyperspectral target detection and recognition method and system based on semantic and spatial-spectral feature fusion. Background Art
[0002] Hyperspectral remote sensing images obtain the geometric, radiation, and spectral information of the scene to form three-dimensional image information. Hyperspectral technology obtains the spectral curve of any pixel through spectral segmentation. Different substances have different characteristic spectral lines, which are the "fingerprints" of the substances themselves. Hyperspectral detection technology can achieve the ability to finely identify land, sea, and air targets.
[0003] However, the data volume of the entire image scene of the hyperspectral image is huge, and it still needs to be sent to the ground and processed manually before it can be used. It is impossible to obtain a large amount of relevant information about targets with strong motion characteristics such as ships and aircraft in real time. Therefore, it is urgent to design a hyperspectral target detection and recognition method, system and electronic equipment based on the fusion of semantic and spatial-spectral features to track, detect and classify the above targets with high confidence.
[0004] Patent application document CN116958807A discloses a hyperspectral image target detection method based on momentum-free contrastive learning and Transformer network, including: first, an encoder and a momentum encoder for spectral feature extraction in the hyperspectral target detection task are designed based on the Transformer encoder. In order to pay attention to the long-distance dependence and self-similarity of the spectrum while not ignoring the local detail information in the spectrum, the spectral feature extraction encoder and the momentum encoder focus on the local detail information of the spectrum through the designed overlapping spectral block feature mapping and interactive token feedforward layer. Secondly, the spectral discrimination ability is learned by unsupervised momentum contrastive learning, in which the queue and the momentum encoder slowly updated in momentum are used to provide sufficient and consistent negative sample features to help the model learn better representation. Finally, through the exponential and normalization operations, the power function and normalization operations perform two nonlinear pull-ups on the detection results obtained by cosine similarity to suppress the background. However, this patent cannot completely solve the current technical problems, nor can it meet the needs of the present invention. Summary of the invention
[0005] In view of the defects in the prior art, the object of the present invention is to provide a method and system for hyperspectral target detection and recognition based on the fusion of semantic and spatial-spectral features.
[0006] The hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion provided by the present invention includes: Step 1: Extract abnormal targets in the hyperspectral image based on the unsupervised anomaly detection algorithm, perform rapid and rough target detection through the constrained energy minimization operator, and eliminate false alarms based on the inter-frame motion characteristics comparison; Step 2: Use a two-stream convolutional neural network to perform fine detection on the rough detection results, obtain the spatial-spectral feature information of the target, and combine the cubic long short-term memory network to predict the confidence interval range of the target to achieve dynamic tracking; Step 3: Use the support vector machine trained based on historical spectra to classify the detected target and determine its category.
[0007] Preferably, the unsupervised anomaly detection algorithm in step 1 adopts the RX detection operator, and the expression is:
[0008] in, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image; By calculating the energy distribution of the target in the direction of the smallest eigenvalue of the covariance matrix and combining it with the target motion distance threshold between adjacent frames, false alarms are eliminated.
[0009] Preferably, the two-stream convolutional neural network in step 2 includes an upper branch and a lower branch, each branch includes an input, 9 convolutional layers are used in each branch to extract rich spectral information of the input pixels, and a one-dimensional convolutional layer is used to implement the convolution operation; In the network, a convolution layer with a kernel step size of 2 is used instead of a pooling layer. To maximize the retention of spectral features, all features extracted by the convolution layer with a kernel step size of 2 are added to the features extracted by the last layer based on different average pooling layers. Then, the final features of each branch are obtained through AVG pooling layer and fully connected layer operations. In a two-stream convolutional neural network, is the target prior pixel, is the target pixel, is the background pixel, constructed according to the training samples, and the input of the upper branch is always , the input of the current branch is When , the label of the training sample is 1, and the input of the next branch is When , the label is 0; after multiple convolution operations, pooling operations and a full connection operation, the final features of the two branches are obtained, recorded as and , and then combine the two features:
[0010] Finally, the output of the two-stream convolutional neural network is obtained through the last fully connected layer and a Sigmoid function.
[0011] Preferably, the cubic long short-term memory network in step 2 is composed of a space branch, a time branch and an output branch, the input includes the latitude and longitude, speed, acceleration and historical trajectory of the target, and the output is the predicted result of the target position at the next moment; by calculating the target pixel movement speed , and the motion distance of adjacent frames , the confidence interval is limited to the 10-pixel neighborhood of the previous frame position, where V is the target speed, is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging, and f is the video frame rate.
[0012] Preferably, in step 3, the support vector machine adopts a soft margin optimization model, trains a classifier based on historical spectral data, and distinguishes between aircraft, ships and other target categories, and its loss function is:
[0013] in, is the weight vector, is the bias term, and are the feature vector and corresponding label of the i-th sample respectively.
[0014] The hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion provided by the present invention comprises: Coarse detection module: extracts abnormal targets in hyperspectral images based on unsupervised anomaly detection algorithms, performs rapid coarse target detection through constrained energy minimization operators, and eliminates false alarms based on inter-frame motion characteristic comparisons; Precision detection module: Use a two-stream convolutional neural network to perform precision detection on the rough detection results, obtain the spatial-spectral feature information of the target, and combine the cubic long short-term memory network to predict the confidence interval range of the target to achieve dynamic tracking; Classification module: Use support vector machine based on historical spectrum training to classify the detected target and determine its category.
[0015] Preferably, the unsupervised anomaly detection algorithm in the coarse detection module adopts the RX detection operator, which is expressed as:
[0016] in, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image; By calculating the energy distribution of the target in the direction of the smallest eigenvalue of the covariance matrix and combining it with the target motion distance threshold between adjacent frames, false alarms are eliminated.
[0017] Preferably, the two-stream convolutional neural network in the precision detection module includes an upper branch and a lower branch, each branch includes an input, 9 convolutional layers are used in each branch to extract rich spectral information of the input pixels, and a one-dimensional convolutional layer is used to implement the convolution operation; In the network, a convolution layer with a kernel step size of 2 is used instead of a pooling layer. To maximize the retention of spectral features, all features extracted by the convolution layer with a kernel step size of 2 are added to the features extracted by the last layer based on different average pooling layers. Then, the final features of each branch are obtained through AVG pooling layer and fully connected layer operations. In a two-stream convolutional neural network, is the target prior pixel, is the target pixel, is the background pixel, constructed according to the training samples, and the input of the upper branch is always , the input of the current branch is When , the label of the training sample is 1, and the input of the next branch is When , the label is 0; after multiple convolution operations, pooling operations and a full connection operation, the final features of the two branches are obtained, recorded as and , and then combine the two features:
[0018] Finally, the output of the two-stream convolutional neural network is obtained through the last fully connected layer and a Sigmoid function.
[0019] Preferably, the cubic long short-term memory network in the precision detection module is composed of a space branch, a time branch and an output branch, the input includes the latitude and longitude, speed, acceleration and historical trajectory of the target, and the output is the predicted result of the target position at the next moment; by calculating the target pixel movement speed , and the motion distance of adjacent frames , the confidence interval is limited to the 10-pixel neighborhood of the previous frame position, where V is the target speed, is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging, and f is the video frame rate.
[0020] Preferably, the support vector machine in the classification module adopts a soft margin optimization model, and trains a classifier based on historical spectral data to distinguish between aircraft, ships and other target categories, and its loss function is:
[0021] in, is the weight vector, is the bias term, and are the feature vector and corresponding label of the i-th sample respectively.
[0022] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a hyperspectral target detection and recognition method based on the fusion of semantic and spatial-spectral features, constructs a semantic segmentation network model in the spatial domain of hyperspectral images, designs an adaptive spatial-spectral joint optimization model, and constructs a spatial-spectral joint cascade detector, thereby realizing spatial domain target detection of hyperspectral images at a pixel-by-pixel level, improving target detection accuracy and reducing false alarm rate, and solving the problem of difficulty in capturing the trajectory of time-sensitive targets in the air. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 A flowchart of a hyperspectral target detection and recognition method based on semantic and spatial spectrum feature fusion of the present invention; Figure 2 A flowchart of anomaly detection based on an unsupervised method of the present invention; Figure 3a and Figure 3b They are respectively a flow chart and a result diagram of an improved target detection with constrained energy minimization according to the present invention; Figure 4 A dual-stream convolutional neural network structure diagram of the present invention; Figure 5a and Figure 5b They are respectively a target classification recognition flow chart and a result chart of a support vector machine of the present invention; Figure 6 A structural diagram of a hyperspectral target detection and recognition system based on semantic and spatial spectrum feature fusion according to the present invention; Figure 7 This is a structural diagram of a hyperspectral target detection and recognition electronic device based on the fusion of semantic and spatial-spectral features of the present invention. DETAILED DESCRIPTION
[0024] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0025] Example See also Figure 1 This embodiment provides a hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion, which mainly includes: Step 1: Extract abnormal targets from hyperspectral images based on anomaly detection using unsupervised methods, perform rapid and rough target detection using the CEM operator, and eliminate false alarms by comparing inter-frame information; Step 2: Perform precise detection through a two-stream convolutional neural network to obtain target-related information, and use a cubic long short-term memory network to determine the confidence interval range to track the detected target; Step 3: Use the SVM trained based on historical spectra to classify the detected targets and determine which type of target they belong to.
[0026] Further, see Figure 2 ,This embodiment provides a flowchart of anomaly detection based on unsupervised methods.
[0027] The unsupervised anomaly detection method adopts the RX detection operator, which is in the form of:
[0028] in, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image. It has the same form as the Mahalanobis distance. In essence, the RX algorithm can be regarded as the inverse process of principal component analysis. Principal component analysis compresses most of the meaningful image information from the original image feature space to a space based on a few unrelated principal components. Obviously, some targets (small targets, abnormal targets) with very low probability of appearing in the image will not be included in these principal components. On the contrary, they have a greater probability of appearing in the covariance matrix. The RX algorithm calculates the direction of the eigenvector corresponding to the small eigenvalue of The value of is used to find abnormal targets. If there are abnormal targets in the graph, then its corresponding energy will be small and will likely be different from the covariance matrix The smaller the eigenvalue, the The larger the value, the more effectively abnormal objects in the image can be detected.
[0029] Using the RX detection operator, targets such as aircraft and ships in hyperspectral images can be detected as abnormal targets, but they are often accompanied by other abnormal points in the image as noise output. Therefore, in anomaly detection, the function of rough target detection is realized when the target spectrum is unknown. Further fine detection is required to eliminate noise and reduce the false alarm rate.
[0030] Further, see Figure 3a, this embodiment provides a flowchart of an improved target detection with constrained energy minimization.
[0031] remember is the set of all observed samples, where is any sample pixel vector, is the number of pixels, is the number of bands of the image, assuming is the target of interest. The purpose of CEM is to design a FIR linear filter , so that the filter output energy is minimized:
[0032] is the covariance matrix, and the solution of the above equation is the CEM operator ,Right now:
[0033] Applying the CEM operator to each pixel in the image will yield the target The distribution of the objects in the image is used to detect the target.
[0034] Assume V is the target moving speed (m / s), is the aircraft flight inclination angle, r is the spatial resolution of the video hyperspectral camera imaging (m), f is the video frame rate (fps), and W is the video width (m), as shown in Table 1.
[0035] Based on the above information, the target pixel motion speed reflected on the image plane can be calculated. (pixels / s) is:
[0036] Target movement distance between adjacent frames (pixels) are:
[0037] The time the target stays in the field of view (s) is:
[0038] The number of frames that the target stays in the field of view for:
[0039] Table 1. Motion characteristics of aircraft and ships reflected in the image
[0040] Calculations show that the target stays in the field of view for at least 12 seconds, and the imaging frequency of the hyperspectral video satellite is 5 frames per second, so the target will appear continuously within at least 60 frames. The distance that a moving target moves between two adjacent frames should also be conservatively estimated to be less than 20 pixels. Therefore, the target usually appears in the vicinity of the target position in the previous frame (within a 20-pixel neighborhood).
[0041] Therefore, when performing anomaly detection on two consecutive frames of images, if all targets detected in the previous frame no longer appear within 20 pixels of the next frame, it can be considered a false alarm.
[0042] Further, see Figure 3b , this embodiment provides an improved result graph of target detection with constrained energy minimization.
[0043] Further, see Figure 4 , this embodiment provides a dual-stream convolutional neural network structure diagram.
[0044] Before training, a hybrid pixel selection strategy based on sparse representation and classification is proposed to select typical background samples in hyperspectral images, and then sufficient target samples are generated through some typical background samples and target priors. During training, the training samples (positive training samples with a label of 1 constructed by target priors and target samples, and negative training samples with a label of 0 constructed by target priors and background samples) are input into a well-designed two-stream convolutional network to learn the discriminative ability. During testing, the test samples (consisting of target priors and detection pixels) are classified by the well-trained two-stream convolutional network. The output of the network constitutes the final detection result.
[0045] The two-stream convolutional neural network consists of two branches: an upstream branch and a downstream branch. Each branch contains an input. In each branch, nine convolutional layers are used to extract rich spectral information of the input pixels. The convolution operation is implemented using one-dimensional convolutional layers, which are followed by ReLU layers. Considering that the pooling layer may cause the loss of spectral information when spectral dimension reduction is required, a convolutional layer with a kernel step size of 2 is used instead of the pooling layer in the network. In order to maximize the retention of spectral features, all features extracted by the convolutional layer with a kernel step size of 2 are added to the features extracted by the last layer based on different average pooling layers. The final features of each branch are then obtained through AVG pooling layer and fully connected layer operations.
[0046] In the network, is the target prior pixel, is the target pixel, is the background pixel. According to the previous training sample construction, the input of the upper branch is always , the input of the current branch is When , the label of the training sample is 1, and the input of the next branch is When , the label is 0. After some convolution operations, several pooling operations and one full connection operation, the final features of the two branches are obtained, recorded as 1 and 2; Then combine the two features:
[0047] Finally, the output of the network is obtained through the last fully connected layer and a Sigmoid function.
[0048] There are two loss functions in the proposed network. The first is the binary cross entropy loss:
[0049] in, is the batch size, is the label of the training sample, is the output of the Sigmoid function. The other is the ICS loss. In order to improve the effect of separability between the target and the background, the ICS loss is proposed. If the inputs of the two branches are and (label is 1), they belong to the same class and the distance between them should be minimized; otherwise, they belong to different classes and the distance between them should be maximized. Therefore, the proposed ICS loss is expressed as:
[0050] in, 1 is the extracted upper branch feature, 2 is the extracted lower branch feature; and are the extracted feature vectors respectively; The final loss function It is the sum of ICS loss and BCE loss, expressed as:
[0051] Cubic Long Short-Term Memory Network is a new structure developed based on LSTM, which consists of three branches: a spatial branch for capturing moving objects, a temporal branch for processing motion, and an output branch for combining the first two branches to generate a predicted frame.
[0052] The spatial branch flows along the z-axis (spatial axis), where convolution is responsible for capturing and analyzing moving objects. The spatial state is generated by this branch, which carries information about the spatial layout of moving objects.
[0053] The temporal branch flows along the x-axis (time axis), and the convolution is designed to obtain and process motion. The temporal state is generated by this branch, which contains motion information.
[0054] The output branch generates an intermediate or final prediction frame along the y-axis (output axis) based on the predicted motion provided by the temporal branch and the moving object information provided by the spatial branch.
[0055] Processing temporal information and spatial information separately can obtain better predictions, and this separation can reduce the prediction burden of the network. Stacking multiple CubicLSTM units along the spatial branch and the output branch can form a two-dimensional network. This two-dimensional network can be further constructed into a three-dimensional network (CubicRNN) by evolving along the time axis, and stacking three layers in space can make the information of the tracked target more prominent and obtain better spatial information.
[0056] The motion characteristics of high-maneuverability aerial targets are divided into two categories: "short-term motion characteristics" and "long-term motion characteristics". Among them, short-term motion characteristics include the target's latitude and longitude, speed, and acceleration, etc. These characteristics may change greatly in a short period of time. The long-term task characteristics include the historical motion trajectory of the target. Compared with the short-term motion characteristics, the long-term task characteristics are relatively stable during the target's movement. In order to make the target tracking more accurate, it is necessary to use the short-term motion characteristics and long-term task characteristics of the target at the same time. Therefore, the CubicLSTM network is used, and the short-term motion characteristics and long-term task characteristics of the target are used as the input of the network, so as to predict the target state at the next moment and calculate the confidence range of the target. The input and output variables of the adopted CubicLSTM network are defined as follows: Table 2. Input and output relationships
[0057] After obtaining the prediction of the target position at the next moment through the CubicLSTM network, it is necessary to further determine the target's confidence interval based on the target's motion characteristics, and detect the target within this range to obtain the target's true position, thereby achieving target tracking. The following will analyze the confidence interval range in detail based on the characteristics of the target's motion.
[0058] Assume V is the target speed (m / s), is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging (m), f is the video frame rate (fps), and W is the video width (m).
[0059] Based on the above information, the target pixel motion speed reflected on the image plane can be calculated. (pixels / s) is:
[0060] The target motion distance (pixels) between adjacent frames is:
[0061] The time (s) that the target stays in the field of view is:
[0062] The number of frames that the target stays in the field of view is:
[0063] Through calculation, it can be concluded that the distance moved by the moving target between two adjacent frames is less than 2.5 pixels, and conservatively estimated to be less than 10 pixels. Therefore, the target usually appears in the 10-pixel neighborhood near the target position in the previous frame. Repeating target detection in this neighborhood can capture the target position in the next frame.
[0064] Further, see Figure 5a , this embodiment provides a target classification and recognition flow chart of a support vector machine.
[0065] Using hard margin SVM in linear inseparable problems will result in classification errors, so a new optimization problem can be constructed by introducing a loss function based on maximizing the margin. Given the input data and learning objectives , SVM uses the hinge loss function, and the optimization problem of soft-margin SVM is expressed as follows:
[0066]
[0067] in, is the weight vector, is the bias term, is the regularization parameter, is the sample size, and are the feature vector and corresponding label of the i-th sample respectively; The above formula shows that the soft margin SVM is a Regularized classifier, where represents the hinge loss function.
[0068] After the SVM classifier is trained based on the historical spectra, it can classify the targets of the precision detection results and distinguish between targets such as airplanes and ships.
[0069] Further, see Figure 5b , this embodiment provides a target classification recognition result diagram of a support vector machine.
[0070] Further, see Figure 6 , this embodiment provides a hyperspectral target detection and recognition system based on semantic and spatial spectrum feature fusion, which mainly includes: The coarse detection module inputs the hyperspectral image to be detected, detects abnormal targets in the hyperspectral image based on the anomaly detection RX operator, performs rapid coarse detection of targets through the CEM operator, and compares the inter-frame information to eliminate false alarms.
[0071] The precision detection module performs precision detection through a two-stream convolutional neural network to obtain target-related information, and uses a cubic long short-term memory network to determine the confidence interval range to track the detection target.
[0072] The classification module uses SVM trained based on historical spectra to classify the detected targets and determine which type of target they belong to.
[0073] Further, see Figure 7 This embodiment provides an electronic device, mainly including: at least one memory and at least one processor, wherein the at least one memory stores instructions, and when the instructions are executed by the at least one processor, a hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion according to an exemplary embodiment of the present disclosure is executed.
[0074] As an example, the electronic device may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above instructions. Here, the electronic device is not necessarily a single electronic device, but may also be any device or circuit collection capable of executing the above instructions (or instruction sets) individually or in combination. The electronic device may also be part of an integrated control system or system manager, or may be configured as a portable electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.
[0075] In electronic devices, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0076] The processor can execute instructions or codes stored in the memory, wherein the memory can also store data. Instructions and data can also be sent and received through the network via the network interface device, wherein the network interface device can adopt any known transmission protocol.
[0077] The memory may be integrated with the processor, for example, RAM or flash memory is arranged within an integrated circuit microprocessor or the like. In addition, the memory may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory and the processor may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor can read files stored in the memory.
[0078] In addition, the electronic device may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device may be connected to each other via a bus and / or a network.
[0079] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.
[0080] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion, characterized in that: include: Step 1: Extract abnormal targets in the hyperspectral image based on the unsupervised anomaly detection algorithm, perform rapid and rough target detection through the constrained energy minimization operator, and eliminate false alarms based on the inter-frame motion characteristics comparison; Step 2: Use a two-stream convolutional neural network to perform fine detection on the rough detection results, obtain the spatial-spectral feature information of the target, and combine the cubic long short-term memory network to predict the confidence interval range of the target to achieve dynamic tracking; Step 3: Use the support vector machine trained based on historical spectra to classify the detected target and determine its category.
2. The hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion according to claim 1 is characterized in that: The unsupervised anomaly detection algorithm in step 1 adopts the RX detection operator, and the expression is: in, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image; By calculating the energy distribution of the target in the direction of the smallest eigenvalue of the covariance matrix and combining it with the target motion distance threshold between adjacent frames, false alarms are eliminated.
3. The hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion according to claim 1 is characterized in that: The two-stream convolutional neural network in step 2 includes an upper branch and a lower branch, each branch includes an input, 9 convolutional layers are used in each branch to extract rich spectral information of the input pixels, and a one-dimensional convolutional layer is used to implement the convolution operation; In the network, a convolution layer with a kernel step size of 2 is used instead of a pooling layer. To maximize the retention of spectral features, all features extracted by the convolution layer with a kernel step size of 2 are added to the features extracted by the last layer based on different average pooling layers. Then, the final features of each branch are obtained through AVG pooling layer and fully connected layer operations. In a two-stream convolutional neural network, is the target prior pixel, is the target pixel, is the background pixel, constructed according to the training samples, and the input of the upper branch is always , the input of the current branch is When , the label of the training sample is 1, and the input of the next branch is When , the label is 0; after multiple convolution operations, pooling operations and a full connection operation, the final features of the two branches are obtained, recorded as and , and then combine the two features: Finally, the output of the two-stream convolutional neural network is obtained through the last fully connected layer and a Sigmoid function.
4. The hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion according to claim 1 is characterized in that: In step 2, the cubic long short-term memory network is composed of a space branch, a time branch and an output branch. The input includes the latitude and longitude, speed, acceleration and historical trajectory of the target, and the output is the predicted result of the target position at the next moment; by calculating the target pixel movement speed , and the motion distance of adjacent frames , the confidence interval is limited to the 10-pixel neighborhood of the previous frame position, where V is the target speed, is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging, and f is the video frame rate.
5. The hyperspectral target detection and recognition method based on semantic and spatial-spectral feature fusion according to claim 1 is characterized in that: In step 3, the support vector machine adopts a soft margin optimization model to train a classifier based on historical spectral data to distinguish between aircraft, ships and other target categories, and its loss function is: in, is the weight vector, is the bias term, and are the feature vector and corresponding label of the i-th sample respectively.
6. A hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion, characterized in that: include: Coarse detection module: extracts abnormal targets in hyperspectral images based on unsupervised anomaly detection algorithms, performs rapid coarse target detection through constrained energy minimization operators, and eliminates false alarms based on inter-frame motion characteristic comparisons; Precision detection module: Use a two-stream convolutional neural network to perform precision detection on the rough detection results, obtain the spatial-spectral feature information of the target, and combine the cubic long short-term memory network to predict the confidence interval range of the target to achieve dynamic tracking; Classification module: Use support vector machine based on historical spectrum training to classify the detected target and determine its category.
7. The hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion according to claim 6 is characterized in that: The unsupervised anomaly detection algorithm in the coarse detection module adopts the RX detection operator, and the expression is: in, is any pixel vector in the image, is the sample mean vector, is the sample covariance matrix of the image; By calculating the energy distribution of the target in the direction of the smallest eigenvalue of the covariance matrix and combining it with the target motion distance threshold between adjacent frames, false alarms are eliminated.
8. The hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion according to claim 6 is characterized in that: The two-stream convolutional neural network in the precision detection module includes an upper branch and a lower branch, each branch contains an input, 9 convolutional layers are used in each branch to extract the rich spectral information of the input pixels, and a one-dimensional convolutional layer is used to implement the convolution operation; In the network, a convolution layer with a kernel step size of 2 is used instead of a pooling layer. To maximize the retention of spectral features, all features extracted by the convolution layer with a kernel step size of 2 are added to the features extracted by the last layer based on different average pooling layers. Then, the final features of each branch are obtained through AVG pooling layer and fully connected layer operations. In a two-stream convolutional neural network, is the target prior pixel, is the target pixel, is the background pixel, constructed according to the training samples, and the input of the upper branch is always , the input of the current branch is When , the label of the training sample is 1, and the input of the next branch is When , the label is 0; after multiple convolution operations, pooling operations and a full connection operation, the final features of the two branches are obtained, recorded as and , and then combine the two features: Finally, the output of the two-stream convolutional neural network is obtained through the last fully connected layer and a Sigmoid function.
9. The hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion according to claim 6 is characterized in that: The cubic long short-term memory network in the precision detection module is composed of a space branch, a time branch and an output branch. The input includes the latitude and longitude, speed, acceleration and historical trajectory of the target, and the output is the predicted result of the target position at the next moment; by calculating the target pixel movement speed , and the motion distance of adjacent frames , the confidence interval is limited to the 10-pixel neighborhood of the previous frame position, where V is the target speed, is the orbital inclination, r is the spatial resolution of the video hyperspectral camera imaging, and f is the video frame rate.
10. The hyperspectral target detection and recognition system based on semantic and spatial-spectral feature fusion according to claim 6, characterized in that: The support vector machine in the classification module adopts a soft margin optimization model to train a classifier based on historical spectral data to distinguish between aircraft, ships and other target categories. The loss function is: in, is the weight vector, is the bias term, and are the feature vector and corresponding label of the i-th sample respectively.
Citation Information
Patent Citations
Hyperspectral target detection method based on unsupervised momentum contrast learning
CN116958807A
Hyperspectral image classification method based on combined multi-level spatial spectrum information CNN
CN110084159A
Sensing abnormal data real-time detection method based on artificial intelligence
CN111582298A
Hyperspectral Image Target Detection Method Based on Transfer Learning and Semantic Segmentation
CN114937206A
Hyperspectral abnormal target detection method based on multi-scale analysis and variational auto-encoder
CN118279747A
Cited By
Satellite task planning computing power dynamic distribution system based on heterogeneous multi-core processor
CN120803752A
Hyperspectral image target detection method based on diffusion model and Mangban structure
CN121214170A