A hyperspectral remote sensing image methane point source detection method based on deep learning

By employing a deep learning-based method for methane point source detection in hyperspectral remote sensing images, and utilizing the separation of background and target subspaces and a multi-scale Transformer module, the method addresses the issues of low accuracy and poor interpretability in existing methane point source detection models, achieving higher accuracy in methane point source detection.

CN119672537BActive Publication Date: 2025-12-12SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411852384.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-12-12
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

Existing methane point source detection models are not very accurate and have poor interpretability. Traditional methods have significant limitations in combining hyperspectral data, and deep learning has a limited application scope in methane emission detection.

Method used

A method for detecting methane point sources in hyperspectral remote sensing images based on deep learning is proposed. By acquiring the background subspace and target subspace of the methane hyperspectral image, a deep learning network is used for model optimization. A total loss function is constructed to improve the detection accuracy. A multi-scale Transformer module and a constrained energy minimization loss function are used to suppress background information, enhance feature representation, and model remote feature interaction.

Benefits of technology

It improves the accuracy and stability of methane point source detection, enhances the interpretability of the model, and enables more accurate detection of methane point source locations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672537B_ABST
    Figure CN119672537B_ABST
Patent Text Reader

Abstract

The application provides a hyperspectral remote sensing image methane point source detection method based on deep learning, which comprises the following steps: obtaining an original methane hyperspectral image, separating a methane point source background subspace and a methane point source target subspace from the original methane hyperspectral image; inputting the original methane hyperspectral image, the methane point source background subspace and the methane point source target subspace into a constructed methane point source detection model for hyperspectral remote sensing images based on deep learning to obtain a reconstructed methane hyperspectral image and a methane point source detection position; establishing a total loss function based on the original methane hyperspectral image, the reconstructed methane hyperspectral image and the methane point source detection position, optimizing the detection model, obtaining a trained detection model when the total loss function value reaches the minimum, obtaining a to-be-detected methane hyperspectral image, inputting the to-be-detected methane hyperspectral image into the trained detection model, and obtaining a corresponding methane point source detection position. The application can improve the precision of the methane point source detection model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of atmosphere, more particularly, to a hyperspectral remote sensing image methane point source detection method based on deep learning. BACKGROUND

[0002] Shortwave infrared sensors have higher sensitivity to near-surface methane concentration changes than thermal infrared sensors, and have been widely used in multi-scale methane monitoring. Among them, the data volume and time resolution advantage of multispectral data makes it applicable to long-term point source monitoring, but its precision and detection limit of methane point source are quite different from those of hyperspectral data, and it is often used in combination with hyperspectral data. Hyperspectral remote sensing images have high potential in methane point source detection and concentration inversion due to their fine and nearly continuous spectral resolution advantage. Traditional estimation methods mainly include physical algorithms, CO2 proxy methods and matching filter methods. Physical algorithms are based on radiative transfer equations and are greatly affected by weather and surface illumination. CO2 proxy method is a kind of Proxy algorithm, as CO2 and methane have absorption characteristics at 1.6 μm, CO2 can be used to invert the concentration in the common absorption band, and CO2 is used as a reference to eliminate the light path change caused by scattering and correct the methane concentration. The matching filter method directly obtains the methane enhancement, rather than the methane concentration of the entire area, and has higher calculation efficiency than the above methods, and is suitable for point source detection tasks. Deep learning has been widely used in the field of environment, including solar power prediction, wind power prediction, etc., but its application in methane emission detection is relatively lacking. Current methane detection deep networks mostly use CNN and other basic architectures. For example, some scholars use UNet to first do binary classification on the image to determine the plume area, and then predict the methane column enhancement in the plume mask area to realize the positioning and quantification of methane point source emission. MethaNET uses a group of plumes of different sizes and shapes under different weather conditions (wind speed, etc.) and background noise, and learns features from the plume under a specific background wind speed, which can be used to predict plumes driven by different wind speeds.

[0003] In summary, the development trend of detection sensors in methane monitoring shows high spatial, high spectral, high temporal resolution and multi-source remote sensing combination research, providing a relatively complete data source for multi-scale methane emission monitoring. Traditional methods such as physical algorithms, CO2 proxy methods and matching filter methods are still widely used, and each method has certain differences in specific applications, and needs to be flexibly selected according to data and targets. Methane detection deep network mostly uses simple architecture of CNN, and complex model is relatively less, and the precision is improved compared with traditional method, which has certain potential in methane emission detection task but the current application range is relatively small. SUMMARY

[0004] The application provides a hyperspectral remote sensing image methane point source detection method based on deep learning to overcome the defects of low precision and poor interpretability of current methane point source detection models.

[0005] To solve the above technical problems, the technical scheme of the application is as follows:

[0006] The application provides a hyperspectral remote sensing image methane point source detection method based on deep learning, which comprises the following steps:

[0007] An original methane hyperspectral image is obtained, and a methane point source background subspace and a methane point source target subspace are separated from the original methane hyperspectral image;

[0008] The original methane hyperspectral image, the methane point source background subspace and the methane point source target subspace are input into a constructed methane point source detection model for hyperspectral remote sensing images based on deep learning to obtain a reconstructed methane hyperspectral image and a methane point source detection position;

[0009] A total loss function is established based on the original methane hyperspectral image, the reconstructed methane hyperspectral image and the methane point source detection position, and the constructed methane point source detection model for hyperspectral remote sensing images based on deep learning is optimized, and when the total loss function value reaches the minimum, a trained methane point source detection model for hyperspectral remote sensing images based on deep learning is obtained;

[0010] A to-be-detected methane hyperspectral image is input into the trained methane point source detection model for hyperspectral remote sensing images based on deep learning to obtain a corresponding methane point source detection position.

[0011] Preferably, the interpretable methane point source detection deep network model comprises a ReLU activation function layer, a splicing layer, a first matrix multiplication layer, a second matrix multiplication layer, a third matrix multiplication layer, a first reshape layer, a second reshape layer, a third reshape layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a Leaky ReLU activation function layer, a multi-scale Transformer module and a first softmax layer.

[0012] The input end of the ReLU activation function layer and the input end of the first convolutional layer are respectively used as the first input end and the third input end of the interpretable methane point source detection deep network model, and the second input end of the splicing layer and the first input end of the second matrix multiplication layer are both used as the second input end of the interpretable methane point source detection deep network model.

[0013] The methane point source background subspace is connected with an input end of a ReLU activation function layer, an output end of the ReLU activation function layer is connected with a first input end of a splicing layer, an output end of the splicing layer is connected with a first input end of a first matrix multiplication layer, a second input end of the first matrix multiplication layer is connected with a first output end of a second Reshape layer, an output end of the first matrix multiplication layer is connected with an input end of a first Reshape layer;

[0014] The original methane hyperspectral image is connected with an input end of a first convolutional layer, the first convolutional layer, a Leaky ReLU activation function layer, a multi-scale Transformer module, a second convolutional layer, a first Softmax layer and a second Reshape layer are connected in sequence;

[0015] The methane point source target subspace is connected with a second input end of the splicing layer and a first input end of a second matrix multiplication layer respectively, a second output end of the second Reshape layer is connected with a second input end of the second matrix multiplication layer, an output end of the second matrix multiplication layer is connected with an input end of a third Reshape layer, an output end of the third Reshape layer is connected with input ends of a third convolutional layer and a fourth convolutional layer respectively, output ends of the third convolutional layer and the fourth convolutional layer are connected with first and second input ends of a third matrix multiplication layer respectively;

[0016] An output end of the first Reshape layer and an output end of the third matrix multiplication layer are respectively a first output end and a second output end of the interpretable methane point source detection deep network model.

[0017] Preferably, the multi-scale Transformer module comprises a first Pixel Unshuffle layer, a second Pixel Unshuffle layer, a first Pixel shuffle layer, a second Pixel shuffle layer, a first Transformer subunit, a second Transformer subunit, a third Transformer subunit, a first Skip connection layer, a first element addition layer, a second element addition layer and a third element addition layer.

[0018] An input end of the first Skip connection layer is a first input end of the multi-scale Transformer module, and an output end of the first Skip connection layer is connected with a second input end of the third element addition layer.

[0019] The input end of the first Pixel Unshuffle layer is the second input end of the multi-scale Transformer module, the output end of the first Pixel Unshuffle layer is connected with the input end of the first Transformer subunit, the output end of the first Transformer subunit is connected with the first input end of the first element addition layer, the output end of the first element addition layer is connected with the input end of the first Pixel shuffle layer, the output end of the first Pixel shuffle layer is connected with the second input end of the second element addition layer, and the output end of the second element addition layer is connected with the first input end of the third element addition layer;

[0020] The input end of the second Transformer subunit is the third input end of the multi-scale Transformer module, and the output end of the second Transformer subunit is connected with the first input end of the second element addition layer;

[0021] The input end of the second Pixel Unshuffle layer is the fourth input end of the multi-scale Transformer module, the second Pixel Unshuffle layer, the third Transformer subunit and the second Pixel shuffle layer are sequentially connected, and the output end of the second Pixel shuffle layer is connected with the second input end of the first element addition layer;

[0022] The output end of the third element addition layer is the output end of the multi-scale Transformer module.

[0023] Preferably, the first Transformer subunit, the second Transformer subunit and the third Transformer subunit are the same in structure and comprise a layer normalization layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a fourth matrix multiplication layer, a fifth matrix multiplication layer, a second softmax layer, a fourth Reshape layer, a fourth element addition layer, a multilayer perceptron and a second Skip connection layer.

[0024] The input end of the layer normalization layer is the input end of the Transformer subunit, and the input end of the layer normalization layer is connected with the input end of the second Skip connection layer.

[0025] The output end of the second Skip connection layer is connected with the second input end of the fourth element addition layer.

[0026] The output ends of the layer normalization layers are connected with the input ends of the fifth convolutional layer, the sixth convolutional layer and the seventh convolutional layer respectively, the output ends of the fifth convolutional layer and the sixth convolutional layer are connected with the first and second input ends of the fourth matrix multiplication layer respectively, and the output end of the fourth matrix multiplication layer is connected with the input end of the second softmax layer;

[0027] The output end of the seventh convolutional layer is connected with the first and second input ends of the fifth matrix multiplication layer respectively, the output end of the fifth matrix multiplication layer is connected with the input end of the fourth Reshape layer, the output end of the fourth Reshape layer is connected with the input end of the eighth convolutional layer, the output end of the eighth convolutional layer is connected with the first input end of the fourth element addition layer, and the output end of the fourth element addition layer is connected with the input end of the multilayer perceptron;

[0028] The output end of the multilayer perceptron is the output end of the Transformer subunit.

[0029] Preferably, the convolution kernel of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fifth convolutional layer, the sixth convolutional layer, the seventh convolutional layer and the eighth convolutional layer is 1×1, and the convolution kernel of the fourth convolutional layer is 3×3.

[0030] Preferably, the determination method of the total loss function comprises:

[0031] L=L rec +ηL spa +μL CEM

[0032] Wherein, L represents the total loss function, L rec , L spa and L CEM respectively represent the reconstruction loss, the sparse loss and the constraint energy minimization loss function, and η and μ are two trade-off parameters.

[0033] Preferably, the determination method of the reconstruction loss comprises:

[0034]

[0035] Wherein, χ and respectively represent the original methane hyperspectral image and the reconstructed methane hyperspectral image, and ||·|| represents the L1 norm of the tensor, that is, the absolute sum of all tensor elements. 1,1,1

[0036] Preferably, the determination method of the sparse loss comprises:

[0037]

[0038] In the formula, C​g C represents the background coefficient of the original methane hyperspectral image o N represents the number of pixels.

[0039] Preferably, the determination method of the constraint energy minimization loss function comprises:

[0040]

[0041] wherein f(x) is a methane point source detection position map, f(o) is a detection response obtained from a methane point source reference position spectrum according to the original methane hyperspectral image, and v is a weighting parameter balancing the two terms.

[0042] The application further provides a hyperspectral remote sensing image methane point source detection system based on deep learning, which is used for implementing the above method and comprises:

[0043] An image acquisition module is configured to acquire an original methane hyperspectral image, and separate a methane point source background subspace and a methane point source target subspace from the original methane hyperspectral image.

[0044] A preprocessing module is configured to input the original methane hyperspectral image, the methane point source background subspace and the methane point source target subspace into a constructed hyperspectral remote sensing image methane point source detection model based on deep learning, and obtain a reconstructed methane hyperspectral image and a methane point source detection position.

[0045] A model training module is configured to establish a total loss function based on the original methane hyperspectral image, the reconstructed methane hyperspectral image and the methane point source detection position, optimize the constructed hyperspectral remote sensing image methane point source detection model based on deep learning, and obtain a trained hyperspectral remote sensing image methane point source detection model based on deep learning when the total loss function value reaches the minimum.

[0046] A model testing module is configured to obtain a to-be-detected methane hyperspectral image, input the to-be-detected methane hyperspectral image into the trained hyperspectral remote sensing image methane point source detection model based on deep learning, and obtain a corresponding methane point source detection position.

[0047] Compared with the prior art, the technical scheme of the application has the beneficial effects that:

[0048] The application provides a hyperspectral remote sensing image methane point source detection method based on deep learning, which comprises the following steps: acquiring an original methane hyperspectral image, separating a methane point source background subspace and a methane point source target subspace from the original methane hyperspectral image; inputting the original methane hyperspectral image, the methane point source background subspace and the methane point source target subspace into a constructed methane point source detection model for hyperspectral remote sensing image based on deep learning, to obtain a reconstructed methane hyperspectral image and a methane point source detection position; establishing a total loss function based on the original methane hyperspectral image, the reconstructed methane hyperspectral image and the methane point source detection position, and optimizing the detection model; when the total loss function value reaches the minimum, a trained detection model is obtained; obtaining a to-be-detected methane hyperspectral image, inputting the to-be-detected methane hyperspectral image into the trained detection model, and obtaining a corresponding methane point source detection position. The application can improve the precision of the methane point source detection model. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 A flowchart of the methane point source detection method based on deep learning for hyperspectral remote sensing image described in embodiment 1 is shown in the figure.

[0050] Figure 2 A structure diagram of the interpretable methane point source detection deep network model described in embodiment 2 is shown in the figure.

[0051] Figure 3 A structure diagram of the multi-scale Transformer module described in embodiment 2 is shown in the figure.

[0052] Figure 4 A structure diagram of the Transformer subunit described in embodiment 2 is shown in the figure.

[0053] Figure 5 An original methane hyperspectral image described in embodiment 2 is shown in the figure.

[0054] Figure 6 A methane point source detection position described in embodiment 2 is shown in the figure.

[0055] Figure 7 An enlarged view of the methane point source detection position described in embodiment 2 is shown in the figure.

[0056] Figure 8 A structure diagram of a methane point source detection system described in embodiment 3 is shown in the figure. DETAILED DESCRIPTION

[0057] The drawings are only used for illustrative description and cannot be understood as a limitation on the patent;

[0058] In order to better illustrate the embodiments, some components in the drawings may be omitted, enlarged or reduced, and do not represent the actual size of the product;

[0059] It will be appreciated by those skilled in the art that certain well-known structures and their descriptions can be omitted from the drawings.

[0060] The technical solutions of the present application will be further described below in combination with the drawings and examples.

[0061] Example 1

[0062] The present embodiment provides a hyperspectral remote sensing image methane point source detection method based on deep learning, as shown in the following steps: Figure 1

[0063] Obtain the original methane hyperspectral image, separate the original methane hyperspectral image to obtain the methane point source background subspace and the methane point source target subspace;

[0064] Input the original methane hyperspectral image, the methane point source background subspace and the methane point source target subspace into the constructed methane point source detection model for hyperspectral remote sensing image based on deep learning, to obtain the reconstructed methane hyperspectral image and the methane point source detection position;

[0065] Based on the original methane hyperspectral image, the reconstructed methane hyperspectral image and the methane point source detection position, a total loss function is established to optimize the constructed methane point source detection model for hyperspectral remote sensing image based on deep learning, and when the total loss function value reaches the minimum, a trained methane point source detection model for hyperspectral remote sensing image based on deep learning is obtained;

[0066] Obtain the to-be-detected methane hyperspectral image, input it into the trained methane point source detection model for hyperspectral remote sensing image based on deep learning, and obtain the corresponding methane point source detection position.

[0067] In the specific implementation process, first, the original methane hyperspectral image is obtained, and the methane point source background subspace and the methane point source target subspace are separated; the original methane hyperspectral image, the methane point source background subspace and the methane point source target subspace are input into the constructed methane point source detection model for hyperspectral remote sensing image based on deep learning, to obtain the reconstructed methane hyperspectral image and the methane point source detection position; then a total loss function is established to optimize the constructed detection model, and when the total loss function value reaches the minimum, the to-be-detected methane hyperspectral image is input into the trained methane point source detection model for hyperspectral remote sensing image based on deep learning, to obtain the corresponding methane point source detection position.

[0068] Example 2

[0069] The present application provides a hyperspectral remote sensing image methane point source detection method based on deep learning, comprising the following steps:

[0070] ​Obtaining a raw methane hyperspectral image, separating a methane point source background subspace and a methane point source target subspace from the raw methane hyperspectral image;

[0071] Inputting the raw methane hyperspectral image, the methane point source background subspace and the methane point source target subspace into the constructed deep learning-based hyperspectral remote sensing image methane point source detection model to obtain a reconstructed methane hyperspectral image and a methane point source detection position;

[0072] Based on the raw methane hyperspectral image, the reconstructed methane hyperspectral image and the methane point source detection position, a total loss function is established, and the constructed deep learning-based hyperspectral remote sensing image methane point source detection model is optimized, and when the total loss function value reaches the minimum, a trained deep learning-based hyperspectral remote sensing image methane point source detection model is obtained;

[0073] Obtaining a to-be-detected methane hyperspectral image, inputting the to-be-detected methane hyperspectral image into the trained deep learning-based hyperspectral remote sensing image methane point source detection model to obtain a corresponding methane point source detection position.

[0074] (1) Model background

[0075] The representation-based detection method shows significant and interpretable performance in the task of hyperspectral target detection. Among them, the subspace representation can decompose the image matrix into two low-dimensional submatrices, which can effectively reduce information redundancy and reduce computational complexity. The embodiment uses an optimization model based on image matrix subspace representation to decompose the original hyperspectral image into:

[0076] X=G+O+N (1)

[0077] In the formula, G, O and N are background components, target components and noise, respectively. In order to accurately and effectively represent the background and target components, the subspace representation of G and O is introduced; therefore, formula (1) can be re-expressed as:

[0078] X=S g C g +oC o +N (2)

[0079] In the formula, S g ,C g ,C o are background subspace, background coefficient and target coefficient, respectively. In order to simplify the process, the known target spectrum o is usually used as the target subspace. Based on formula (2), the detection problem can be converted into an optimization model combined with some common manual priori estimation of background and target coefficients, which is expressed as:

[0080]

[0081] where the first term is the fidelity term, and the other two terms, φ1(C g ) and φ2(C o ) are the prior terms applied on the background and target coefficients respectively, and α and β are the trade-off parameters, and 1 M and 1 N are M-dimensional and N-dimensional column vectors with 1.

[0082] To characterize the target and background coefficients, different priors can be used, but for formula (3), the background subspace usually needs to be known in advance, which is difficult to estimate and maintain its integrity and purity. In addition, the hand-crafted prior is not enough to accurately describe the properties of separating the background and target, and the nonlinear features are difficult to extract. In this embodiment, a deep learning network is used to represent the model, which can maintain physical interpretability, support nonlinear representation, adaptively learn the background subspace, and effectively solve the above problems through gradient descent.

[0083] (2) Constrained energy minimization loss function

[0084] The loss function in the deep learning network is a constrained energy minimization loss (CEM), which aims to design a finite impulse filter to minimize the background output energy and maximize the target response, that is, the filtering result of the target spectrum is 1, which can be represented as the following optimization problem:

[0085]

[0086] where x i represents the i-th pixel, and M represents the filtering matrix. Using the Lagrange multiplier method, it can be re-expressed as:

[0087]

[0088] where μ represents the Lagrange multiplier, represents the autocorrelation matrix. Taking the derivative of μ and M respectively, the solution is:

[0089]

[0090] Therefore, the detection result of CEM is M T X, which can be reshaped as a detection map. Inspired by the CEM idea, this embodiment uses the constrained energy minimization loss function to highlight the target information and suppress the background information.

[0091] (3) Deep learning-based hyperspectral remote sensing image methane point source detection model

[0092] In order to efficiently detect the position of the methane point source in the hyperspectral image, a deep learning network is introduced in the embodiment. First, the input HSI is extracted through the convolutional layer to obtain the bottom layer features. After that, the low-level features are enhanced through the multi-scale Transformer (MST) to enhance the feature representation and model the long-range feature interaction. Then, the enhanced features are projected into the joint coefficient space through the convolutional layer, and the non-negative constraint and the constraint of sum being one are realized through the Softmax function. Finally, the joint coefficient is decoded, and the input hyperspectral image is restored by multiplying the joint subspace (i.e. the cascade of the methane point source background subspace and the target subspace). The background subspace is defined as a learnable variable, and then a rectified linear unit (ReLU) activation is used to obtain the non-negative constraint. The detection block is parallel to the encoder-decoder structure, which separates the target coefficient from the joint coefficient and multiplies the target spectrum to synthesize the target component. Then, the target component is processed through the multi-scale convolutional layer to generate the detection map F.

[0093] As shown in Figure 2 , the interpretable methane point source detection deep network model comprises a ReLU activation function layer, a splicing layer, a first matrix multiplication layer, a second matrix multiplication layer, a third matrix multiplication layer, a first reshape layer, a second reshape layer, a third reshape layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a Leaky ReLU activation function layer, a multi-scale Transformer module and a first softmax layer;

[0094] The input end of the ReLU activation function layer and the input end of the first convolutional layer are respectively the first input end and the third input end of the interpretable methane point source detection deep network model, and the second input end of the splicing layer and the first input end of the second matrix multiplication layer are both the second input end of the interpretable methane point source detection deep network model;

[0095] The methane point source background subspace is connected with the input end of the ReLU activation function layer, the output end of the ReLU activation function layer is connected with the first input end of the splicing layer, the output end of the splicing layer is connected with the first input end of the first matrix multiplication layer, the second input end of the first matrix multiplication layer is connected with the first output end of the second reshape layer, and the output end of the first matrix multiplication layer is connected with the input end of the first reshape layer;

[0096] The original methane hyperspectral image is connected with the input end of the first convolutional layer, and the first convolutional layer, the Leaky ReLU activation function layer, the multi-scale Transformer module, the second convolutional layer, the first Softmax layer and the second reshape layer are connected in sequence;

[0097] The methane point source target subspace is connected with the second input end of the splicing layer and the first input end of the second matrix multiplication layer respectively, the second output end of the second Reshape layer is connected with the second input end of the second matrix multiplication layer, the output end of the second matrix multiplication layer is connected with the input end of the third Reshape layer, the output end of the third Reshape layer is connected with the input end of the third convolution layer and the fourth convolution layer respectively, and the output end of the third convolution layer and the fourth convolution layer is connected with the first input end and the second input end of the third matrix multiplication layer respectively;

[0098] The output end of the first Reshape layer and the output end of the third matrix multiplication layer are respectively taken as the first output end and the second output end of the interpretable methane point source detection deep network model.

[0099] 1) Encoder-decoder structure

[0100] The model of formula (3) is converted into a deep learning network in the embodiment, the fidelity term is converted into an encoder-decoder structure, and the prior term is retained in the loss function. Nonlinear representation and physical interpretability are compatible, and the optimization model can be effectively solved by gradient descent, and the solving process is as follows:

[0101] a. Extracting coefficients with an encoder: the encoder learns the nonlinear mapping between the hyperspectral image and the joint coefficients:

[0102]

[0103] In the formula, f represents that the encoder estimates the background coefficient and the target coefficient from the HSI, Θ represents the parameters involved in the encoder, and [·] represents a splicing operation. In the embodiment, a 1*1 convolution is first applied to extract low-level features, and a Leaky ReLU activation is used to improve nonlinearity.

[0104] Y0=f lk (W1 0 X) (9)

[0105] In the formula, Y0 is a bottom feature with C channels, W1 0 represents a k*k convolution numbered i, f lk is a Leaky ReLU function. In addition, a lightweight multi-scale Transformer module is used in the embodiment for multi-scale analysis and feature enhancement:

[0106] Y en =MST(Y0) (10)

[0107] In the formula, Y enTo enhance the features. The enhanced features are projected into the coefficient space by convolution and Softmax function (corresponding to the Softmax layer), and the required joint coefficients are obtained as:

[0108]

[0109] Here, the Softmax function is applied along the channel dimension to ensure the non-negative constraint and the sum-to-one constraint of the joint coefficients, and the R operation will expand the tensor into a matrix form.

[0110] b. Recover the image with the decoder: the decoder recovers the HSI from the joint coefficients by multiplying the joint subspace, which can be expressed as:

[0111]

[0112] In the formula, equation (24) can be equivalent to:

[0113]

[0114] In the formula, S g is an adaptive and learnable variable, and the ReLU function f is used to ensure the non-negativity of the background subspace, and the reshape operation here is to fold the matrix into a tensor form.

[0115] 2) Multi-scale Transformer

[0116] The multi-scale Transformer used in this embodiment can model long-range feature dependencies, utilize multi-scale features, and reduce computational requirements compared to single-scale networks. Thus, enhanced features are obtained, and the required coefficients are more accurately extracted.

[0117] As shown in Figure 3 , the multi-scale Transformer module includes a first Pixel Unshuffle layer, a second Pixel Unshuffle layer, a first Pixel shuffle layer, a second Pixel shuffle layer, a first Transformer subunit, a second Transformer subunit, a third Transformer subunit, a first Skip connection layer, a first element addition layer, a second element addition layer, and a third element addition layer.

[0118] The input end of the first Skip connection layer serves as the first input end of the multi-scale Transformer module, and the output end of the first Skip connection layer is connected with the second input end of the third element addition layer.

[0119] The input end of the first Pixel Unshuffle layer is the second input end of the multi-scale Transformer module, the output end of the first Pixel Unshuffle layer is connected with the input end of the first Transformer subunit, the output end of the first Transformer subunit is connected with the first input end of the first element addition layer, the output end of the first element addition layer is connected with the input end of the first Pixel shuffle layer, the output end of the first Pixel shuffle layer is connected with the second input end of the second element addition layer, and the output end of the second element addition layer is connected with the first input end of the third element addition layer;

[0120] The input end of the second Transformer subunit is the third input end of the multi-scale Transformer module, and the output end of the second Transformer subunit is connected with the first input end of the second element addition layer;

[0121] The input end of the second Pixel Unshuffle layer is the fourth input end of the multi-scale Transformer module, the second Pixel Unshuffle layer, the third Transformer subunit and the second Pixel shuffle layer are sequentially connected, and the output end of the second Pixel shuffle layer is connected with the second input end of the first element addition layer;

[0122] The output end of the third element addition layer is the output end of the multi-scale Transformer module.

[0123] The first Transformer subunit, the second Transformer subunit and the third Transformer subunit are of the same structure, as shown in Figure 4 The first Transformer subunit, the second Transformer subunit and the third Transformer subunit are of the same structure, as shown in

[0124] The input end of the layer normalization layer is the input end of the Transformer subunit, and the input end of the layer normalization layer is connected with the input end of the second Skip connection layer;

[0125] The output end of the second Skip connection layer is connected with the second input end of the fourth element addition layer;

[0126] The output ends of the layer normalization layers are connected with the input ends of the fifth convolutional layer, the sixth convolutional layer and the seventh convolutional layer respectively, the output ends of the fifth convolutional layer and the sixth convolutional layer are connected with the first and second input ends of the fourth matrix multiplication layer respectively, and the output end of the fourth matrix multiplication layer is connected with the input end of the second softmax layer;

[0127] The output end of the seventh convolutional layer and the output end of the second softmax layer are connected with the first and second input ends of the fifth matrix multiplication layer respectively, the output end of the fifth matrix multiplication layer is connected with the input end of the fourth Reshape layer, the output end of the fourth Reshape layer is connected with the input end of the eighth convolutional layer, the output end of the eighth convolutional layer is connected with the first input end of the fourth element addition layer, and the output end of the fourth element addition layer is connected with the input end of the multilayer perceptron;

[0128] The output end of the multilayer perceptron is the output end of the Transformer subunit.

[0129] a.The structure of the multi-scale Transformer module adopts a hierarchical design, which is composed of three layers of feature encoding at the pyramid spatial resolution, and each level contains a Transformer encoder to capture the context relationship. From the first level to the last level, the Pixel Unshuffle operation is applied to obtain the multi-scale spatial information without loss of information, and conversely, the Pixel Shuffle operation is the recovery process of the multi-scale spatial information to the original size spatial information. In order to reduce the information loss in the recovery process, the structure integrates the encoding features of the Transformer encoder with the features reconstructed from the lower level through element addition, which helps to maintain a lightweight structure compared with the commonly used connection method. At the same time, the output of the first layer is added to the input to obtain enhanced features through global residual connection, which can further retain the main information and ensure that the multi-scale Transformer module learns detailed information.

[0130] b.Transformer encoder: The self-attention mechanism is the core idea of the transformer. In the classic self-attention algorithm, the computational complexity is quadratic with the spatial resolution of the input, which makes it difficult to perform object detection tasks with the whole image as input. Therefore, this embodiment adopts a lightweight variant with linear complexity to solve this problem. Specifically, the context relationship between channels is modeled without targeting the original patch or pixel. In addition to the lightweight consideration, this design can also preserve spatial information, so that each pixel can be independently represented by a joint subspace. Otherwise, spatial-spectral confusion can lead to substandard coefficients, ultimately limiting detection performance. The structure of the multi-head self-attention encoder (MHCA) is as follows: Given a tensor Z, after layer normalization, three linear projections are calculated by 1x1 convolution, namely Query (Q), Key (K), and Value (V). The size of the three projections is hwxC. Next, the transpose similarity between Q and K is calculated, and the attention map A is generated using the Softmax function, and then multiplied by V to obtain the self-attention feature. The process is represented as:

[0131] A=softmax(Q T K×α) (14)

[0132] Attention(Z)=R(VA) (15)

[0133] In the formula, a is a learnable scale factor that adjusts the size of the dot product and ensures the stability of the Softmax output. To further capture the relationship between different channels, the Transformer model applies a multi-head mechanism. Specifically, the model divides the number of channels into multiple heads and calculates individual self-attention in parallel. For this purpose, the three projections Q, K, and V are reshaped into new projections Q', K', and V'. The calculation formula of multi-head self-attention is:

[0134] A'=softmax(Q' T K'×α') (16)

[0135] Attention(Z)=R(V'A') (17)

[0136] In the formula, A' is the multi-head attention map, and a' is a learnable vector that controls the dot product size of each head. Finally, the self-attention is linearly transformed (1x1 convolution) to obtain the residual result, and then added to the input of the Transformer encoder to generate the feature of interest Z':

[0137] Z'=W1 5 Attention(Z)+Z (18)

[0138] To perform nonlinear transformation on the features of interest, the output of MHCA is further passed through a multilayer perceptron module and generates the final enhanced features Z":

[0139] Z" = W1 7 f g (W1 6 Z') (19)

[0140] where f g represents the Gaussian error linear unit (GELU) activation. The adopted MLP module consists of two 1x1 convolutions and a hidden layer of nonlinearities.

[0141] 3) Methane point source detection module

[0142] With the learned target coefficients, the desired methane point source target component τ can be synthesized:

[0143] τ = R(oC o ) (20)

[0144] To convert the target component into the final detection map, L2 norm, max value method, nonlinear suppression function, etc. are common strategies, but these strategies ignore the dynamic spatial relationship. To solve this problem, the present embodiment uses a multi-scale convolution block to construct a dynamic mapping, which can effectively utilize the spatial correlation of the adjacent region:

[0145] F = R(W1 8 τ + W3 0 τ) (21)

[0146] where the outputs of the 1x1 convolution layer and the 3x3 convolution layer are combined by element-wise addition to generate the final detection map F. This fusion strategy can effectively utilize the pixel-level spectral information and block-level spatial information to determine the existence of pixel-level or sub-pixel-level targets.

[0147] 4) Loss function

[0148] As described above, the present embodiment constructs a methane point source detection network based on the methane point source detection module. To train the network, both the fidelity and the prior term are converted into the corresponding terms of the loss function. Among them, the fidelity term is the reconstruction loss, and the uncertain prior term is the sparse loss under the joint coefficient constraint. At the same time, the constrained energy minimization loss function is taken as the core term to accurately separate the background and the target. Therefore, the composite loss function is clipped as:

[0149] L = L rec + ηL spa + μL CEM (22)

[0150] where Lrec , L spa and L CEM respectively represent reconstruction loss, sparsity loss and constrained energy minimization loss function, which measure reconstruction effect, sparsity of joint coefficients and background target recognition degree respectively, and η and μ are two trade-off parameters to balance the three terms, and reconstruction loss is used to measure the difference between the input and output of the proposed SRN. Considering that L1 norm can better preserve texture and edge information than Frobenius norm, L1 norm is used for reconstruction in the embodiment:

[0151]

[0152] where χ and are the input original methane hyperspectral image and the reconstructed methane hyperspectral image respectively, and ||·||1 represents the L1 norm of a tensor, that is, the absolute sum of all tensor elements. 1,1,1

[0153] The sparsity loss is used to represent the sparsity of the coefficients responsible for representing each pixel. The sparsity constraint can make the model more concise and enhance the interpretability of the model. The classical L1 norm is used in the embodiment to constrain the sparsity of joint coefficients:

[0154]

[0155] where ||·||1 represents the L1 norm of a vector, that is, the absolute sum of all vector elements, and C g represents the background coefficient of the original methane hyperspectral image, and C o represents the target coefficient of the original methane hyperspectral image, and N represents the number of pixels.

[0156] Based on the core idea of CEM, that is, matching target features by maximizing the response of known target spectrum and limiting the output energy of the background, the constrained energy minimization loss function term is used to highlight the target and suppress the background. In the embodiment, the deep detector is defined as a mapping The mapping can be approximated as a finite impulse filter W, so the constrained energy minimization loss function can be expressed as:

[0157]

[0158] ​In the formula, f(χ) represents the methane point source detection location spectrum, and f(o) represents the detection response of the methane point source reference location spectrum obtained from the original methane hyperspectral image. Absolute value operation |·| ensures the distance is positive, and v is the trade-off parameter for the two balancing terms. The first term corresponds to the CEM optimization term, aiming to minimize the energy of the filtered result; the second term corresponds to the CEM condition, which maximizes the response of the target feature by forcing the filtered result of the target spectrum to 1. This highlights the target information while suppressing background information. However, this energy minimization term processes all pixels the same way, limiting detection performance. To address this issue, this embodiment improves the performance by introducing a weight matrix that considers the difference between each pixel and the target spectrum. Furthermore, since the depth detector described in this embodiment is based on tensor input, the target spectrum is unfolded into a tensor cube for unified calculation.

[0159]

[0160] In the formula, 1 H×W A matrix of 1s, ||·|| 1,1 L1 norm represents the L1 norm of a matrix, which is the absolute sum of all elements in the matrix.

[0161]

[0162] F=f(χ) (28)

[0163] G=f(o H×W×L (29)

[0164] In the formula, ||·||2 is the L2 norm of the vector, and o H×W×L Let be a tensor cube expanded by 'o'. Since the weight matrix D measures the difference between each pixel and the target spectrum, the constrained energy minimization loss function can represent the affinity of all pixels to the target spectrum. That is, the closer a pixel is to the target spectrum, the less energy it loses, and vice versa. This avoids the loss of a large amount of target information, thus preserving target information while more effectively suppressing background information.

[0165] In this embodiment, the convolutional kernels of the first, second, third, fifth, sixth, seventh, and eighth convolutional layers are 1×1, and the convolutional kernel of the fourth convolutional layer is 3×3. For example... Figure 5 The image shown is the original hyperspectral image of methane. A1-A7 in the image represent the selected detection regions. Figure 5 The output of the methane point source detection model is input into the model. Figure 6 ,like Figure 6 As shown, this is the methane point source detection location output by the methane point source detection model described in this embodiment. Figure 7Figure 3 is a magnified view of a methane point source detection position in the A1-A7 detection area, wherein the bright area is a methane point source and the dark area is a background.

[0166] The embodiment has the following three technical advantages: first, to overcome the bottleneck of traditional matching filtering method in precision and performance, a deep learning network with physical interpretability is innovatively introduced to improve the accuracy and stability of methane point source detection. Second, to cope with the challenge of high memory consumption of hyperspectral image data, a lightweight Multi Scale Transformer (MST) is designed, which adopts a hierarchical three-channel structure to effectively balance the model complexity and information retention capability. Finally, through the multi-head self-attention encoder, the context relationship between channels is modeled, which not only preserves spatial information but also enables each pixel to be expressed independently through joint subspace representation, thereby significantly enhancing the data feature extraction efficiency and the expression ability of high-dimensional data, and more accurately capturing complex spectral information.

[0167] Embodiment 3

[0168] The embodiment provides a deep learning-based hyperspectral remote sensing image methane point source detection system for implementing the method of embodiments 1 or 2, as shown in Figure 8 The embodiment provides a deep learning-based hyperspectral remote sensing image methane point source detection system for implementing the method of embodiments 1 or 2, as shown in

[0169] An image acquisition module is configured to acquire an original methane hyperspectral image, and separate a methane point source background subspace and a methane point source target subspace from the original methane hyperspectral image;

[0170] A preprocessing module is configured to input the original methane hyperspectral image, the methane point source background subspace, and the methane point source target subspace into a constructed deep learning-based hyperspectral remote sensing image methane point source detection model to obtain a reconstructed methane hyperspectral image and a methane point source detection position;

[0171] A model training module is configured to establish a total loss function based on the original methane hyperspectral image, the reconstructed methane hyperspectral image, and the methane point source detection position, optimize the constructed deep learning-based hyperspectral remote sensing image methane point source detection model, and obtain a trained deep learning-based hyperspectral remote sensing image methane point source detection model when the total loss function value reaches a minimum;

[0172] A model testing module is configured to obtain a to-be-detected methane hyperspectral image, input the to-be-detected methane hyperspectral image into the trained deep learning-based hyperspectral remote sensing image methane point source detection model, and obtain a corresponding methane point source detection position.

[0173] The same or similar reference signs correspond to the same or similar components;

[0174] The terms used to describe the positional relationship in the drawings are only used for exemplary illustration and cannot be understood as a limitation on the patent;

[0175] Obviously, the above-mentioned embodiments of the present application are only examples for clearly illustrating the present application, but are not intended to limit the implementation modes of the present application. Based on the above description, other different forms of changes or variations can also be made by those skilled in the art. Here, it is not necessary and also impossible to enumerate all the implementation modes. Any modification, equivalent replacement and improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the claims of the present application.

Claims

1. A method for detecting methane point sources in hyperspectral remote sensing images based on deep learning, characterized in that, Includes the following steps: Acquire the original methane hyperspectral image, and separate the methane point source background subspace and the methane point source target subspace from the original methane hyperspectral image; The original methane hyperspectral image, the methane point source background subspace, and the methane point source target subspace are input into the constructed deep learning-based hyperspectral remote sensing image methane point source detection model to obtain the reconstructed methane hyperspectral image and the methane point source detection location. Based on the original methane hyperspectral image, the reconstructed methane hyperspectral image, and the methane point source detection location, a total loss function is established. The constructed deep learning-based hyperspectral remote sensing image methane point source detection model is then optimized. When the total loss function value reaches its minimum, the trained deep learning-based hyperspectral remote sensing image methane point source detection model is obtained. Obtain the hyperspectral image of the methane to be detected, input it into the trained deep learning-based hyperspectral remote sensing image methane point source detection model, and obtain the corresponding methane point source detection location. The specific structure of the deep learning-based hyperspectral remote sensing image methane point source detection model is as follows: The inputs of the ReLU activation function layer and the first convolutional layer are respectively used as the first and third inputs of the deep learning-based hyperspectral remote sensing image methane point source detection model. The second input of the stitching layer and the first input of the second matrix multiplication layer are both used as the second input of the deep learning-based hyperspectral remote sensing image methane point source detection model. The methane point source background subspace is connected to the input of the ReLU activation function layer. The output of the ReLU activation function layer is connected to the first input of the splicing layer. The output of the splicing layer is connected to the first input of the first matrix multiplication layer. The second input of the first matrix multiplication layer is connected to the first output of the second reshape layer. The output of the first matrix multiplication layer is connected to the input of the first reshape layer. The original methane hyperspectral image is connected to the input of the first convolutional layer, and the first convolutional layer, the Leaky ReLU activation function layer, the multi-scale Transformer module, the second convolutional layer, the first Softmax layer, and the second Reshape layer are connected in sequence. The methane point source target subspace is connected to the second input of the splicing layer and the first input of the second matrix multiplication layer, respectively. The second output of the second reshape layer is connected to the second input of the second matrix multiplication layer, the output of the second matrix multiplication layer is connected to the input of the third reshape layer, the output of the third reshape layer is connected to the input of the third convolutional layer and the fourth convolutional layer, respectively, and the outputs of the third convolutional layer and the fourth convolutional layer are connected to the first and second inputs of the third matrix multiplication layer, respectively. The output of the first Reshape layer and the output of the third matrix multiplication layer serve as the first and second outputs of the deep learning-based hyperspectral remote sensing image methane point source detection model, respectively.

2. The method for detecting methane point sources in hyperspectral remote sensing images based on deep learning according to claim 1, characterized in that, The multi-scale Transformer module includes a first pixel unshuffle layer, a second pixel unshuffle layer, a first pixel shuffle layer, a second pixel shuffle layer, a first Transformer subunit, a second Transformer subunit, a third Transformer subunit, a first skip connection layer, a first element-wise addition layer, a second element-wise addition layer, and a third element-wise addition layer; The input of the first Skip Connection layer serves as the first input of the multi-scale Transformer module, and the output of the first Skip Connection layer is connected to the second input of the third element-adding layer. The input of the first pixel unshuffle layer serves as the second input of the multi-scale Transformer module. The output of the first pixel unshuffle layer is connected to the input of the first Transformer subunit. The output of the first Transformer subunit is connected to the first input of the first element-adding layer. The output of the first element-adding layer is connected to the input of the first pixel shuffle layer. The output of the first pixel shuffle layer is connected to the second input of the second element-adding layer. The output of the second element-adding layer is connected to the first input of the third element-adding layer. The input of the second Transformer subunit serves as the third input of the multi-scale Transformer module, and the output of the second Transformer subunit is connected to the first input of the second element-adding layer. The input of the second pixel unshuffle layer serves as the fourth input of the multi-scale Transformer module. The second pixel unshuffle layer, the third Transformer sub-unit, and the second pixel shuffle layer are connected in sequence. The output of the second pixel shuffle layer is connected to the second input of the first element-adding layer. The output of the third element addition layer serves as the output of the multi-scale Transformer module.

3. The method for detecting methane point sources in hyperspectral remote sensing images based on deep learning according to claim 2, characterized in that, The first Transformer subunit, the second Transformer subunit, and the third Transformer subunit have the same structure, including a layer normalization layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a fourth matrix multiplication layer, a fifth matrix multiplication layer, a second softmax layer, a fourth reshape layer, a fourth element-wise addition layer, a multilayer perceptron, and a second skip connection layer. The input of the layer normalization layer serves as the input of the Transformer subunit, and the input of the layer normalization layer is connected to the input of the second Skip connection layer. The output of the second Skip Connection layer is connected to the second input of the fourth element addition layer; The output of the normalization layer is connected to the input of the fifth, sixth and seventh convolutional layers, respectively. The outputs of the fifth and sixth convolutional layers are connected to the first and second inputs of the fourth matrix multiplication layer, respectively. The output of the fourth matrix multiplication layer is connected to the input of the second softmax layer. The output of the seventh convolutional layer and the output of the second softmax layer are connected to the first and second inputs of the fifth matrix multiplication layer, respectively. The output of the fifth matrix multiplication layer is connected to the input of the fourth reshape layer. The output of the fourth reshape layer is connected to the input of the eighth convolutional layer. The output of the eighth convolutional layer is connected to the first input of the fourth element-wise addition layer. The output of the fourth element-wise addition layer is connected to the input of the multilayer perceptron. The output of the multilayer perceptron serves as the output of the Transformer subunit.

4. The method for detecting methane point sources in hyperspectral remote sensing images based on deep learning according to claim 3, characterized in that, The first, second, third, fifth, sixth, seventh, and eighth convolutional layers have 1×1 convolutional kernels, while the fourth convolutional layer has a 3×3 convolutional kernel.

5. The method for detecting methane point sources in hyperspectral remote sensing images based on deep learning according to claim 1, characterized in that, The method for determining the total loss function includes: L=L rec +ηL spa +μL CEM Where L represents the total loss function, L rec L spa and L CEM Let represent the reconstruction loss, sparse loss, and constrained energy minimization loss functions, respectively, and η and μ are two trade-off parameters.

6. The method for detecting methane point sources in hyperspectral remote sensing images based on deep learning according to claim 5, characterized in that, The method for determining the reconstruction loss includes: Where, χ and Representing the original methane hyperspectral image and the reconstructed methane hyperspectral image, respectively, ‖·‖ 1,1,1 It represents the L1 norm of a tensor, which is the absolute sum of all tensor elements.

7. The method for detecting methane point sources in hyperspectral remote sensing images based on deep learning according to claim 5, characterized in that, The method for determining the sparse loss includes: In the formula, C g C represents the background coefficient of the original methane hyperspectral image. o This represents the target coefficient of the original methane hyperspectral image, where N represents the number of pixels.

8. The method for detecting methane point sources in hyperspectral remote sensing images based on deep learning according to claim 5, characterized in that, The method for determining the constraint energy minimization loss function includes: In the formula, f(χ) is the methane point source detection location spectrum, f(o) is the detection response of the methane point source reference location spectrum obtained from the original methane hyperspectral image, and v is the trade-off parameter for the two balancing terms.

9. A deep learning-based hyperspectral remote sensing image methane point source detection system, used to implement the method described in any one of claims 1-8, characterized in that, include: The image acquisition module is used to acquire the original methane hyperspectral image and separate the methane point source background subspace and the methane point source target subspace from the original methane hyperspectral image. The preprocessing module is used to input the original methane hyperspectral image, the methane point source background subspace, and the methane point source target subspace into the constructed deep learning-based hyperspectral remote sensing image methane point source detection model to obtain the reconstructed methane hyperspectral image and the methane point source detection location. The model training module is used to establish a total loss function based on the original methane hyperspectral image, the reconstructed methane hyperspectral image, and the methane point source detection location. This function is used to optimize the constructed deep learning-based hyperspectral remote sensing image methane point source detection model. When the total loss function value reaches its minimum, the trained deep learning-based hyperspectral remote sensing image methane point source detection model is obtained. The model testing module is used to obtain the hyperspectral image of the methane to be detected, input it into the trained deep learning-based hyperspectral remote sensing image methane point source detection model, and obtain the corresponding methane point source detection location.

Citation Information

Patent Citations

  • Method for quickly identifying and extracting methane point source based on hyperspectral imager

    CN113869143A

  • Regional atmospheric methane background value prediction method and system based on remote sensing data

    CN118658547A