A method for generating a saliency map of hyperspectral images based on a semi-supervised neural network

By building a semi-supervised neural network, using twin prediction modules and attention assist modules, combined with fully supervised and weakly supervised data sets, the problem of hyperspectral image data processing is solved, and high-quality significant maps are achieved efficiently generated.

CN114359675BActive Publication Date: 2025-07-18CHONGQING INNOVATION CENTER OF BEIJING INSTITUTE OF TECHNOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210039812.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-07-18
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

The processing of hyperspectral image data is difficult, and the pixel-level labeling cost of neural network models is high, so it is difficult for existing methods to efficiently generate high-quality significant maps.

Method used

A semi-supervised neural network is built, including twin prediction modules and attention-assisted modules, and trains with fully supervised data sets and weakly supervised data sets to generate high-quality significant graphs.

Benefits of technology

This greatly saves pixel-level labeling costs, improves the robustness and detection accuracy of the model, and generates higher quality significant maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359675B_ABST
    Figure CN114359675B_ABST
Patent Text Reader

Abstract

The present invention provides a method for generating a saliency map of a hyperspectral image based on a semi-supervised neural network, comprising: constructing a semi-supervised neural network; the semi-supervised neural network includes a twin prediction module and an attention assistance module; using a fully supervised dataset to pre-train the twin prediction module and the attention assistance module; using the pre-trained attention assistance module to generate a pixel-level mask for a weakly supervised dataset; using the pixel-level mask and the weakly supervised dataset to perform fully supervised training on the pre-trained twin prediction module; using the fully supervised trained twin prediction module to generate a saliency map for the input hyperspectral image. The present invention uses a semi-supervised neural network to extract hyperspectral image features and directly generate a predicted saliency map, greatly saving the pixel-level annotation cost, improving the model robustness and detection accuracy, and being able to generate a saliency map with higher quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more particularly, to a method for generating a hyperspectral image saliency map based on a semi-supervised neural network. Background Art

[0002] Hyperspectral images can simultaneously obtain the spatial information and spectral information of an object, and have dozens or hundreds of spectral bands. Hyperspectral imaging technology has made great progress in terms of spatial resolution, spectral resolution, and portability, and has been widely applied in fields such as ground object remote sensing, precision agriculture, medical diagnosis, and target detection.

[0003] A saliency map simulates the human visual attention mechanism and describes the objects or regions that human eyes focus on in the field of view. The saliency detection task is to let a computer simulate this ability of the human eye, extract the salient objects in the image, and generate a saliency map. In traditional methods, people mainly calculate local or global contrast using shallow features such as color, texture, and brightness to obtain the most salient regions. In recent years, neural network models based on deep learning have been introduced, which can learn deep features and have greatly improved in terms of robustness and detection accuracy.

[0004] Hyperspectral images have the characteristic of integrating spectrum and image. Each pixel has a complete spectral curve, which can better reflect the material characteristics. This makes the hyperspectral image saliency map have advantages such as high accuracy and anti-interference. However, it is difficult to obtain hyperspectral data, and the characteristics of large data redundancy and high dimensionality increase the difficulty of data processing. In addition, training a neural network model requires pixel-level annotation, and the annotation cost is high. A semi-supervised neural network only needs to use a small amount of fully supervised data with pixel-level annotation, while most of the remaining data does not need to be annotated or only weakly annotated, which will effectively reduce the annotation cost. Therefore, applying a semi-supervised neural network to the generation of hyperspectral image saliency maps has important significance and value. Summary of the Invention

[0005] The present invention aims to provide a method for generating a hyperspectral image saliency map based on a semi-supervised neural network, which only uses a small amount of fully supervised data with pixel-level annotation and a large amount of weakly supervised data with weak annotation to generate a high-quality hyperspectral image saliency map, thereby solving the problems of difficult hyperspectral data processing and high pixel-level annotation cost of the neural network model.

[0006] A method for generating a hyperspectral image saliency map based on a semi-supervised neural network provided by the present invention includes the following steps:

[0007] Construct a semi-supervised neural network; the semi-supervised neural network includes a twin prediction module and an attention assistance module;

[0008] Pre-train the Siamese prediction module and the attention assistance module using a fully supervised dataset;

[0009] Use the pre-trained attention assistance module to generate a pixel-level mask for the weakly supervised dataset;

[0010] Use the pixel-level mask and the weakly supervised dataset to perform fully supervised training on the pre-trained Siamese prediction module;

[0011] Use the fully supervised trained Siamese prediction module to generate a saliency map for the input hyperspectral image.

[0012] In some embodiments, the attention assistance module includes a superpixel sampling network, a global average pooling layer, a fully connected layer, a 1×1 convolutional layer, a Sigmoid function layer, and a U 2 Net network; the processing process of the attention assistance module is as follows:

[0013] The superpixel sampling network extracts superpixel local features from the feature map in the input weakly supervised dataset and inputs them into the fully connected layer;

[0014] The global average pooling layer extracts global features from the feature map in the input weakly supervised dataset, upsamples the global features to the original size, and then inputs them into the fully connected layer;

[0015] The fully connected layer combines the local features and the global features into a fused feature and then inputs it into the convolutional layer;

[0016] The 1×1 convolutional layer generates attention weights based on the fused feature and outputs them through the Sigmoid function layer;

[0017] Calculate the Hadamard product of the attention weights and the feature map in the input weakly supervised dataset and input it into the U 2 Net network;

[0018] The U 2 Net network generates a pixel-level mask for the weakly supervised dataset according to the Hadamard product.

[0019] Wherein, the method for making the feature map is:

[0020] Take the saliency bounding box as a kind of weak label, and calculate the Hadamard product of the saliency bounding box and the original hyperspectral image to obtain the feature map.

[0021] In some embodiments, the Siamese prediction module is a Siamese network structure based on the U 2 Net network, including two prediction networks sharing weight parameters; the processing process of the Siamese prediction module is as follows:

[0022] Input the original hyperspectral image into the first prediction network to obtain a first predicted saliency map;

[0023] Input the boundary hyperspectral image into the second prediction network to obtain the second predicted saliency map;

[0024] Use the mean square error to represent the difference between the first predicted saliency map and the second predicted saliency map, and train by minimizing this difference.

[0025] Among them, the method for producing the boundary hyperspectral image is as follows:

[0026] Use the saliency bounding box as a kind of weak label, and use the saliency bounding box to crop the original hyperspectral image to obtain the boundary hyperspectral image.

[0027] Furthermore, the loss function used when training the twin prediction module is a hybrid loss function; the hybrid loss function includes binary cross-entropy, structural similarity index and detection box intersection over union

[0028] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:

[0029] 1. The present invention uses a semi-supervised neural network to extract hyperspectral image features and directly generate a predicted saliency map, which greatly saves the pixel-level annotation cost, improves the robustness and detection accuracy of the model, and can generate a saliency map with higher quality.

[0030] 2. The present invention constructs a twin prediction module, which constructs a twin network structure including two prediction networks based on the U 2 Net network to capture the regional-level implicit constraint ability between the weak label of the saliency bounding box and each pixel.

[0031] 3. The present invention constructs an attention assistance module, introduces a superpixel-level channel attention mechanism, makes full use of the unique spectral and channel information of the hyperspectral image, and can generate a high-quality pixel-level mask for the weakly supervised dataset to optimize the twin prediction module. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0033] Figure 1 It is a flowchart of the method for generating a saliency map of a hyperspectral image based on a semi-supervised neural network in an embodiment of the present invention.

[0034] Figure 2This is a flowchart of the processing procedure of the attention assistance module in the embodiments of the present invention. Detailed implementation manners

[0035] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0036] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0037] Embodiment

[0038] As Figure 1 shown, this embodiment proposes a method for generating a saliency map of a hyperspectral image based on a semi-supervised neural network, including the following steps:

[0039] Construct a semi-supervised neural network; the semi-supervised neural network includes a twin prediction module and an attention assistance module;

[0040] Pre-train the twin prediction module and the attention assistance module using a fully supervised dataset;

[0041] Use the pre-trained attention assistance module to generate a pixel-level mask for the weakly supervised dataset;

[0042] Use the pixel-level mask and the weakly supervised dataset to perform fully supervised training on the pre-trained twin prediction module;

[0043] Use the fully supervised trained twin prediction module to generate a saliency map for the input hyperspectral image.

[0044] 1. Attention assistance module

[0045] As Figure 2 shown, the attention assistance module includes a superpixel sampling network, a global average pooling layer, a fully connected layer, a 1×1 convolutional layer, a Sigmoid function layer, and a U 2 Net network; the processing procedure of the attention assistance module is as follows:

[0046] The superpixel sampling network extracts superpixel local features from the feature map in the input weakly supervised dataset and inputs them into the fully connected layer; the superpixel sampling network is a prior art and will not be elaborated here;

[0047] The global average pooling layer extracts global features from the feature maps in the input weakly supervised dataset, upsamples the global features to the original size, and then inputs them into the fully connected layer;

[0048] The fully connected layer combines the local features and the global features into fused features and then inputs them into the convolutional layer;

[0049] The 1×1 convolutional layer generates attention weights based on the fused features and outputs them through the Sigmoid function layer;

[0050] Calculate the Hadamard product of the attention weights and the feature maps in the input weakly supervised dataset and input it into the U 2 Net network;

[0051] The U 2 Net network generates a pixel-level mask for the weakly supervised dataset based on the Hadamard product.

[0052] Among them, the method for producing the feature maps is as follows: taking the saliency bounding box as a kind of weak label, calculating the Hadamard product of the saliency bounding box and the original hyperspectral image to obtain the feature maps. The saliency bounding box, as a kind of weak label, can be regarded as a pixel-level label with partial noise. The pixels outside the bounding box do not contain noise and are regarded as the background. Under the interference of noise, the background points inside the bounding box are misrecognized as foreground points. This error will cause the network to converge in the wrong direction during the gradient backpropagation process. Using the spectral curves of each pixel for superpixel clustering can distinguish the background and the foreground. In addition, the correlation between channels can also be used as a reliable basis for accurately distinguishing different types of objects. Therefore, the attention assistance module in this embodiment uses the bounding box as a weak label to enhance the salient region of the original hyperspectral image, then suppresses the background points through the superpixel-level sampling network, and finally generates a high-quality pixel-level mask for the weakly supervised set data to optimize the twin prediction module.

[0053] 2. Twin prediction module

[0054] The twin prediction module is a twin network structure based on the U 2 Net network, including two prediction networks sharing weight parameters; the processing process of the twin prediction module is as follows:

[0055] Input the original hyperspectral image into the first prediction network to obtain the first predicted saliency map;

[0056] Input the boundary hyperspectral image into the second prediction network to obtain the second predicted saliency map;

[0057] Use the mean square error to represent the difference between the first predicted saliency map and the second predicted saliency map, and train by minimizing this difference.

[0058] Among them, the method for producing the boundary hyperspectral image is as follows: using the saliency bounding box as a weak label, and cropping the original hyperspectral image with the saliency bounding box to obtain the boundary hyperspectral image. The difference between weakly supervised training and fully supervised training is that the weak label loses the constraint on each pixel. When using the bounding box as the weak label, this constraint is converted into a region-level implicit constraint. In this embodiment, a Siamese network structure based on the U 2 Net network with two prediction networks is adopted, and the implicit constraint between the weak label and the pixels can be learned by minimizing the difference between the two prediction networks.

[0059] An example of the training process of the semi-supervised neural network:

[0060] (1) Loss function: A hybrid loss function is adopted, including binary cross-entropy, structural similarity index, and intersection over union of detection boxes. The loss between each ground-truth saliency map and the predicted saliency map is calculated through the hybrid loss function, and the parameters of the neural network are optimized by the backpropagation algorithm.

[0061] (2) Data processing: Randomly select 30% of the dataset as the fully supervised dataset, 60% as the weakly supervised dataset, and the remaining 10% as the test set. Data augmentation is performed before training, and operations such as flipping, rotating 90°, and rotating 270° are performed on each image.

[0062] (3) Training settings: The semi-supervised neural network is trained for 25 epochs on the fully supervised dataset, 35 epochs on the weakly supervised dataset, the batch size is set to 8, the learning rate is 0.001, and the optimization method is the Adam algorithm.

[0063] Through the above training process, the Siamese prediction module in the semi-supervised neural network can be obtained, which can directly generate a saliency map for the input hyperspectral image.

[0064] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for generating a saliency map of a hyperspectral image based on a semi-supervised neural network, characterized in that, It includes the following steps: Construct a semi-supervised neural network; the semi-supervised neural network includes a twin prediction module and an attention assistance module; Pre-train the twin prediction module and the attention assistance module using a fully supervised dataset; Use the pre-trained attention assistance module to generate a pixel-level mask for the weakly supervised dataset; Use the pixel-level mask and the weakly supervised dataset to perform fully supervised training on the pre-trained twin prediction module; Use the fully supervised trained twin prediction module to generate a saliency map for the input hyperspectral image; The attention assistance module includes a superpixel sampling network, a global average pooling layer, a fully connected layer, a 1×1 convolutional layer, a Sigmoid function layer, and a U 2 Net network; the processing procedure of the attention assistance module is as follows: The superpixel sampling network extracts superpixel local features from the feature map in the input weakly supervised dataset and inputs them into a fully connected layer; The global average pooling layer extracts global features from the feature map in the input weakly supervised dataset, upsamples the global features to the original size, and then inputs them into a fully connected layer; The fully connected layer combines the local features and the global features into a fused feature and then inputs it into a convolutional layer; The 1×1 convolutional layer generates attention weights based on the fused feature and outputs them through a Sigmoid function layer; Calculate the Hadamard product of the attention weights and the feature maps in the input weakly supervised dataset and input it into U 2 Net network; The said U 2 Net network generates a pixel-level mask for the weakly supervised dataset according to the Hadamard product; the Siamese prediction module is a Siamese network structure based on the U 2 Net network, including two prediction networks sharing weight parameters; the processing process of the Siamese prediction module is as follows: Input the original hyperspectral image into the first prediction network to obtain a first predicted saliency map; Input the boundary hyperspectral image into the second prediction network to obtain a second predicted saliency map; Use the mean square error to represent the difference between the first predicted saliency map and the second predicted saliency map, and train by minimizing this difference.

2. The method for generating a hyperspectral image saliency map based on a semi-supervised neural network according to claim 1, characterized in that, The method for making the feature map is as follows: Use the saliency bounding box as a kind of weak label, and calculate the Hadamard product of the saliency bounding box and the original hyperspectral image to obtain the feature map.

3. The method for generating a hyperspectral image saliency map based on a semi-supervised neural network according to claim 1, wherein The method for making the boundary hyperspectral image is as follows: Use the saliency bounding box as a kind of weak label, and use the saliency bounding box to crop the original hyperspectral image to obtain the boundary hyperspectral image.

4. The method for generating a hyperspectral image saliency map based on a semi-supervised neural network according to claim 1, characterized in that, The loss function used when training the twin prediction module is a mixed loss function; the mixed loss function includes binary cross-entropy, structural similarity index, and detection box intersection over union.

Citation Information

Patent Citations

  • Weak supervision RGBD image saliency detection method and system based on image classification

    CN112861880A

  • Feature point tracking training and tracking methods, apparatus, electronic device, and storage medium

    WO2021253686A1