A method and equipment for seamless spatial distribution reconstruction of methane based on low-coverage hyperspectral satellite data

Through the deep partial convolutional U-net network and satellite CH4 distribution mask, the problem of low coverage of satellite methane concentration data is solved, high-precision methane concentration data filling and continuous monitoring are achieved, and the temporal and spatial coverage of the data and the application effect are improved.

CN119312019BActive Publication Date: 2025-09-12UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411371538.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-09-12
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

Existing satellite remote sensing technology has a low coverage problem in methane concentration data monitoring, resulting in sparse and discontinuous data, which cannot meet the precise monitoring needs on a global scale. Traditional data filling methods cannot effectively solve this problem.

Method used

A deep partially convolutional U-net network combined with satellite CH4 distribution mask is used to construct a filling model. Through sample construction, model training and data filling application, seamless spatial distribution reconstruction of low-coverage methane concentration data is achieved.

Benefits of technology

By effectively extracting and utilizing information from the input data and simulating real conditions, the filling accuracy and continuity of low-coverage satellite CH4 concentration data can be improved, the dependence on atmospheric chemistry model data for the same period can be avoided, and the application scope of the filling model can be expanded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312019B_ABST
    Figure CN119312019B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and equipment for reconstructing the seamless spatial distribution of methane based on low-coverage hyperspectral satellite data, comprising: using atmospheric chemical model methane concentration distribution data and auxiliary environmental data as sample data, screening and mapping the measured methane concentration distribution data, constructing a reference mask, and combining the reference mask with the sample data to form a training sample; constructing a filling model, the filling model adopting a U-net structure including a deep partial convolution layer, a spatial and channel attention module, an upsampling layer, and a partial convolution layer; supervising the filling model for a methane concentration filling task based on an overall loss function constructed based on an unobserved area evaluation loss function, an observed area evaluation loss function, and a regional overall change loss function; obtaining the measured methane concentration distribution data and auxiliary environmental data as application data, using the filling model to perform data filling on the application data, and efficiently and quickly filling the low-coverage methane concentration data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of environmental observation, and in particular relates to a method and equipment for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low-coverage data. Background Art

[0002] The primary driver of the rise in global average surface temperature is the increase in atmospheric concentrations of greenhouse gases, including methane (CH4). To mitigate global warming, researchers want to understand the generation, distribution, and transport of atmospheric CH4. This requires precise and continuous global monitoring of atmospheric CH4.

[0003] Satellite remote sensing technology is currently widely used in various CH4 observation and research projects, becoming the mainstream technology for obtaining long-term global CH4 spatiotemporal distribution data. However, due to limitations in satellite orbits, meteorological conditions, satellite spectral quality, and the accuracy of inversion algorithms, the CH4 spatiotemporal distribution data obtained from satellite inversion is always sparse and does not meet application requirements.

[0004] Tropomi satellites, which provide a large amount of current CH4 concentration data, can only provide effective global CH4 distribution pixels for a single day, even under optimal conditions, covering less than 30% of the global land area. The average global land coverage rate is only around 15%. At certain times, when some areas are completely obscured by cloud, there may even be no CH4 concentration data for that day, completely failing to meet application requirements.

[0005] At the same time, due to the different number of observations in different regions, it is very likely to obtain erroneous analysis results when conducting long-term analysis such as monthly analysis. Therefore, it is necessary to obtain CH4 concentration data with high temporal and spatial continuity through filling, which is conducive to the analysis and monitoring of CH4 over a large area.

[0006] Current data imputation methods, such as machine learning, used for gases like NO2 are not suitable for imputing CH4 concentration data. This is because NO2 concentration coverage is significantly higher than CH4, typically exceeding 50%. Furthermore, some locations lack CH4 satellite observation data year-round, unlike NO2, which has full annual data coverage. Therefore, methods that impute NO2 concentrations are ineffective when applied to CH4. The low coverage of satellite CH4 concentration data can lead to distortion in these models in real-world applications.

[0007] In response to the above technical problems, a filling solution for low-coverage methane concentration data is urgently needed. Summary of the Invention

[0008] In view of the above, the purpose of the present invention is to provide a method and equipment for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low-coverage data, develop a deep partial convolutional U-net network, and combine it with the use of satellite CH4 distribution mask to complete the filling of low-coverage satellite CH4 concentration data.

[0009] To achieve the above-mentioned object of the invention, an embodiment provides a method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low-coverage data, comprising the following steps:

[0010] Sample construction: The atmospheric chemical model methane concentration distribution data and auxiliary environmental data are used as sample data. The measured methane concentration distribution data are screened and mapped to construct a reference mask for the sample data. The reference mask is used to mask the sample data to form a training sample.

[0011] Padding model construction: The padding model adopts a U-net structure consisting of an encoding part and a decoding part. The encoding part consists of multiple encoding modules of decreasing size, each of which consists of a depth-wise partial convolutional layer and a spatial and channel attention module. The decoding part consists of multiple decoding modules of increasing size, each of which consists of an upsampling layer and a partial convolutional layer connected in sequence, which establishes connections between depth-wise partial convolutional layers and partial convolutional layers of the same size.

[0012] Filling model training: Based on the evaluation loss function for the unobserved area, the evaluation loss function for the observed area, and the overall regional change loss function, an overall loss function is constructed. The filling model is then supervised and trained for the methane concentration filling task based on the training samples and the overall loss function.

[0013] Data filling application: obtain measured methane concentration distribution data and auxiliary environmental data as application data, construct a mask reference mask for the application data based on the measured methane concentration distribution data, use the reference mask to mask the application data to form an application sample, and use the trained filling model to perform data filling on the application sample to obtain full coverage methane concentration distribution data.

[0014] Preferably, the obtained measured methane concentration distribution data is screened, including:

[0015] If the data provider provides a QA value screening standard, the measured methane concentration distribution data will be screened according to the standard. Specifically, the methane concentration distribution data with a QA value of 1 will be retained, and the remaining data will be removed. If the QA value is not provided, the corresponding methane concentration data will be removed when the standard deviation between the inverted simulated spectrum and the actual collected spectrum is greater than the set threshold. The retained measured methane concentration distribution data will be used to construct the reference mask.

[0016] Preferably, mapping the screened and retained measured methane concentration distribution data includes:

[0017] A preset two-dimensional equally spaced grid mesh with a structure of H×W, where the difference in longitude or latitude between each grid point and its adjacent grid points is equal. The measured methane concentration distribution data is processed according to the following formula:

[0018]

[0019] The above formula averages the gas concentration of the pixels in each longitude-latitude grid point to obtain the methane concentration value at each longitude-latitude grid point. Among them, XCH4 s represents the CH4 gas concentration at the s-th pixel to be processed on the original pixel plane after linear interpolation. lon s , lat s are the longitude and latitude of this pixel, XCH4 <​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​A reference mask is constructed for the measured methane concentration distribution data in the application data. Specifically, the reference mask is constructed by assigning the valueless points in the measured CH4 concentration distribution data to 0 and the valued points to 1.

[0025] A reference mask is constructed for the auxiliary environment data in the application data, and all values ​​in the specific reference mask are 1.

[0026] Preferably, the sample data D and the reference mask M included in the training sample have the same data structure, and when performing inference transfer in the depth partial convolution layer and the partial convolution layer, for the convolution kernel K with a weight of F and a bias of b, after the input data D is slidingly convolved by the convolution kernel K, the calculated value d at the corresponding position after one convolution window and the updated mask m are expressed as:

[0027]

[0028] Among them, D′ and M′ are the input data and mask data of the corresponding convolution window, N is the number of data in the convolution window, and M′ is the n is the mask data of the corresponding position n in the convolution window, ⊙ is element-by-element calculation, the superscript T represents transposition, the current calculated value d and the updated mask m are used as the input data D′ and mask data M′ of the next layer;

[0029] For the partial convolution layer, its single convolution kernel will convolve all layers, that is, the structure of a single convolution kernel is C×H′×W′. For the deep partial convolution layer, its single convolution kernel will only convolve a single layer, that is, the structure of a single convolution kernel is 1×H′×W′, where H′×W′ is the convolution window size, and C and 1 both represent the number of channels.

[0030] Preferably, the loss function for evaluating the unobserved area is expressed as loss hole , the loss function of the observation area evaluation is expressed as loss valid :

[0031]

[0032] Among them, k is the data point index, K is the total amount of data, y k Represents the label data of data point k, y′ k Represents the predicted output data of the filling model for data point k, m k Represents the mask data of data point k.

[0033] Preferably, the overall regional change loss function is expressed as loss tv :

[0034]

[0035] Among them, || ||1 represents the L1 norm, (i, j)∈P represents that the position (i, j) belongs to the region P, and y com,i,j Represents the synthetic data at position (i, j), synthesized as follows:

[0036] y com,i,j =m i,j ×y i,j +(1-m i,j )×y′ i,j

[0037] Among them, m i,j Represents the mask data at position (i, j), y i,j Represents the label data of position (i, j), y′ i,j Represents the model prediction data at position (i, j);

[0038] The overall loss function is expressed as loss total :

[0039] loss total =α×loss valid +β×loss hole +γ×loss tv

[0040] Among them, α, β, and γ are the loss functions for evaluating the observed area. valid , loss function for evaluating the unobserved area loss hole And the overall regional change loss function loss tv The weight of .

[0041] Preferably, before the training samples are input into the filling model, a random mirror operation is performed on the atmospheric chemical model methane concentration distribution data. Specifically, according to a given probability, the training samples are mirrored in the longitude dimension or the latitude dimension, that is, the data is reversed.

[0042] Preferably, the auxiliary environmental data includes meteorological parameter data, geographical parameter data, and CH4 surface emission data.

[0043] To achieve the above-mentioned purpose, an embodiment of the present invention further provides a device for filling low-coverage methane concentration data, comprising:

[0044] A sample construction module is used to use atmospheric chemical model methane concentration distribution data and auxiliary environmental data as sample data, filter and map the measured methane concentration distribution data, construct a reference mask for the sample data, and use the reference mask to mask the sample data to form a training sample;

[0045] A padding model construction module, wherein the padding model adopts a U-net structure including an encoding part and a decoding part, wherein the encoding part includes multiple encoding modules of successively decreasing sizes, each encoding module includes a depth-wise partial convolutional layer and a spatial and channel attention module, and the decoding part includes multiple decoding modules of successively increasing sizes, each decoding module includes an upsampling layer and a partial convolutional layer connected in sequence, which establishes connections between depth-wise partial convolutional layers and partial convolutional layers of the same size;

[0046] The filling model training module is used to construct an overall loss function based on the evaluation loss function of the unobserved area, the evaluation loss function of the observed area, and the overall regional change loss function, and to perform supervised training of the filling model for the methane concentration filling task based on the training samples and the overall loss function;

[0047] The data filling application module is used to obtain measured methane concentration distribution data and auxiliary environmental data as application data, construct a mask reference mask for the application data based on the measured methane concentration distribution data, use the reference mask to mask the application data to form an application sample, and use the trained filling model to fill the application sample with data to obtain full coverage methane concentration distribution data.

[0048] To achieve the above-mentioned purpose of the invention, an embodiment also provides a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low-coverage data.

[0049] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the above-mentioned method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low coverage data is implemented.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] The filling model constructed by the present invention can effectively extract and utilize the information of each layer in the input data, and distinguish between valuable points and valueless points, thereby increasing the information ratio of valuable points in the network, and helping to extract effective information from low-coverage satellite CH4 concentration data. At the same time, using the CH4 distribution data mask measured by satellites, etc., the model can simulate the real situation, learn the boundary conditions of valuable points and valueless points that are closer to reality, and then infer a more realistic CH4 distribution situation through valuable points. In addition, the present invention only uses CH4 atmospheric chemical model data during training, avoiding the situation of relying on atmospheric chemical model CH4 data of the same period during subsequent filling, thereby improving the application scope of the filling model. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0053] Figure 1 This is a flow chart of a method for reconstructing seamless spatial distribution of methane based on hyperspectral satellite low coverage data provided by an embodiment;

[0054] Figure 2 is a structural diagram of a filling model provided in an embodiment;

[0055] Figure 3 : This is the measured methane concentration distribution data of a satellite in a certain area on a certain day provided in the embodiment. Dark blue represents a zero-value area.

[0056] Figure 4 Schematic diagram of the structure of the filling device for low coverage methane concentration data provided in the embodiment. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0058] The inventive concept of the present invention is: to address the technical problems that the coverage of CH4 concentration data measured by satellites and other sources is too low, the use of traditional data filling methods cannot achieve the expected results, and the filling results depend on atmospheric chemical model data. The embodiment of the present invention provides a method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low-coverage data. By constructing a filling model and combining it with the use of satellite and other sources' measured CH4 concentration distribution data masks, the filling of low-coverage satellite CH4 concentration data is completed.

[0059] like Figure 1 As shown, the embodiment provides a method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low coverage data, including the following steps:

[0060] S1, the atmospheric chemical model methane concentration distribution data and auxiliary environmental data are used as sample data, the measured methane concentration distribution data are screened and mapped to construct a reference mask for the sample data, and the reference mask is used to mask the sample data to form a training sample.

[0061] In an embodiment, a daily-scale CH4 concentration data set and auxiliary environmental data for training are obtained. The CH4 concentration data set for training includes atmospheric chemistry model CH4 concentration distribution data for training and satellite and other measured CH4 concentration distribution data for filling. For the atmospheric chemistry model CH4 concentration distribution data, its time period can be relatively short, such as 2 years, and does not need to overlap with the time of the CH4 concentration data that needs to be filled. Therefore, the atmospheric chemistry model CH4 concentration distribution data is only used during training, and the atmospheric chemistry model CH4 concentration data is not used during filling, thereby getting rid of the strong dependence on the atmospheric chemistry model CH4 concentration data. The auxiliary environmental data include meteorological parameter data, geographic parameter data, and CH4 surface emission data, wherein the parameter categories mainly included in the meteorological parameters and geographic parameters are shown in Table 1.

[0062] Table 1

[0063] Serial number Geographical parameters and meteorological parameters 1 Mean sea level pressure 2 2m temperature 3 2m condensation temperature 4 Boundary layer height 5 Total column concentration of water vapor 6 Zenithal Solar Radiation 7 100m longitudinal wind speed 8 100m zonal wind speed

[0064] For CH4 surface emission data, CH4 surface emission data related to industrial and human activities were obtained based on public information, and wetland CH4 surface emission data were generated through models such as LPJ-wsl, and the data were added together to form the total CH4 surface emission data.

[0065] In an embodiment, the measured CH4 concentration distribution data from satellites and the like are screened, including: using the QA value provided by the official source of the measured concentration data set to screen the measured CH4 concentration distribution data, specifically retaining the CH4 concentration distribution data with a QA value of 1, and removing the remaining data; when the QA value for screening is not provided, the CH4 concentration data corresponding to the standard deviation between the inverted simulated spectrum and the actual collected spectrum is greater than a set threshold is removed. It should be noted that the screening criteria are not fixed and can be empirically determined based on the inverted gas and the actual situation. In addition, other physical restrictions can be added, including but not limited to methods such as the cloud cover at the center of the data point being greater than a certain value. The screened and retained measured methane concentration distribution data are used to construct a reference mask.

[0066] In the embodiment, the measured CH4 concentration distribution data that has been screened and retained are mapped to plane grid points, that is, all data points are gridded to the same grid, including: a preset two-dimensional equally spaced grid with a structure of H×W, where the longitude or latitude difference between each grid point and the adjacent grid point is equal in degrees, and the measured methane concentration distribution data is processed according to the following formula:

[0067]

[0068] The above formula averages the gas concentration of the pixels in each longitude and latitude grid point to obtain the methane concentration value at each longitude and latitude grid point, where XCH4 sDenote the CH4 gas concentration at the s-th pixel to be processed on the original pixel plane after linear interpolation, lon s , lat s are the longitude and latitude of this pixel, XCH4 lon,lat represents the CH4 gas concentration at the corresponding longitude and latitude grid points on the equal longitude and latitude grid. lon and lat are the longitude and latitude of this pixel, grid is the difference in degrees between longitude and latitude grid points, and the symbol & represents the meaning of "and". Denote the sum of XCH4 for all points that satisfy the screening condition (lon s - lon) < grid / 2 & (lat s - lat) < grid / 2. s Sum. Denote the number of points that satisfy the screening condition (lon s - lon) < grid / 2 & (lat s - lat) < grid / 2. 1 is the cumulative value of the number of each point.

[0069] For the CH4 concentration data of the atmospheric chemistry model, geographical parameters, and meteorological parameter data, since these data are initially on an equally spaced grid, but the range and spacing of this grid may be different from the preset size of the present invention, so part of the data can be interpolated from the original grid to the preset equally spaced grid according to the linear interpolation rule.

[0070] In the embodiment, a reference mask is constructed based on the mapped measured CH4 concentration distribution data as the methane concentration distribution data of the atmospheric chemistry model, including: randomly selecting the satellite measured CH4 concentration distribution data of one day within 15 days before and after the corresponding day to construct the reference mask. The mapped measured CH4 concentration data is equal to the grid of the reference mask. When there is a measured methane concentration data value at the mapped grid point, the corresponding element value in the reference mask is set to 1. When there is no value at the mapped grid point, the corresponding element value in the reference mask is set to 0. The reference mask obtained in this way comes from the real measured CH4 concentration data.

[0071] For traditional mask coverage, generally, a pre-set irregular mask or a regular pattern mask dug at random positions is used. Since the boundary between the valid points and invalid points of the actual satellite measured CH4 concentration distribution data is often not fixed and has a large number of sides. The regular pattern mask dug at random positions can only generate masks with fewer fixed boundary patterns, and for example Figure 3For isolated satellite observations such as these, this type of masking method doesn't allow the model to effectively learn the impact of these points on their surroundings. For pre-set irregular masks, the boundaries between the valued and non-valued areas during training are also relatively fixed, which can also cause problems when filling in the measured CH4 concentration distribution data from satellites whose boundaries differ significantly from the training mask boundaries. Therefore, the present invention directly uses the area actually observed by the satellite as a mask. Furthermore, due to the influence of parameters such as the zenith angle, CH4 observation coverage varies in different seasons, such as winter and summer. Satellite CH4 concentration distribution data from one of the 15 days before and after the corresponding day is randomly selected as a reference mask to ensure that the concentration data for that day matches a mask that is likely to appear in the corresponding season, rather than a mask that is unlikely to appear in the actual season. Furthermore, by superimposing different masks on the same concentration data, during subsequent training, even though the labeled data is the same, it can actually be considered different training data. This method also expands the training set, facilitating subsequent training to obtain a more versatile model.

[0072] In the embodiment, a reference mask is also constructed for the auxiliary environmental data in the sample data. Specifically, all the values ​​in the reference mask are 1. Although the reference mask with all values ​​1 does not have a masking effect on the auxiliary environmental data, it can maintain the overall reference mask size of the sample data, achieve placeholders, and facilitate subsequent model training.

[0073] After obtaining the reference mask of the sample data, the sample data is masked using the reference mask to form a training sample. Specifically, the reference mask is multiplied by the sample data to achieve masking.

[0074] S2, builds a filling model based on the deep partial convolutional U-net network.

[0075] In the embodiment, Figure 2 As shown, the constructed filling model adopts a U-net structure including an encoding part and a decoding part, wherein the encoding part includes a plurality of encoding modules with successively decreasing sizes, each encoding module includes a depth partial convolution layer and a spatial and channel attention module, and the decoding part includes a plurality of decoding modules with successively increasing sizes, each decoding module includes an upsampling layer and a partial convolution layer connected in sequence, which establishes a connection between the depth partial convolution layer and the partial convolution layer of the same size, and the input data is encoded and decoded for data filling to obtain the filled prediction data.

[0076] The sample data D and the reference mask M contained in the training sample have the same data structure, both of which are C×H×W. When performing inference transmission in the depth-wise partial convolution layer and the partial convolution layer, for the convolution kernel K with weight F and bias b, after the input data D is sliding convolved by the convolution kernel K, the calculated value d at the corresponding position after one convolution window and the updated mask m are expressed as:

[0077]

[0078]

[0079] Among them, D′ and M′ are the input data and mask data of the corresponding convolution window, N is the number of data in the convolution window, and M′ is the n is the mask data of the corresponding position n in the convolution window, ⊙ is element-by-element calculation, the superscript T represents transposition, the current calculated value d and the updated mask m are used as the input data D′ and mask data M′ of the next layer; the valued points in the reference mask M of the present invention are actual specific values, and the valueless points are 0; the valued points in other updated input masks m are 1, and the valueless points are 0.

[0080] For the partial convolution layer, its single convolution kernel will convolve all layers, that is, the structure of a single convolution kernel is C×H′×W′. For the deep partial convolution layer, its single convolution kernel will only convolve a single layer, that is, the structure of a single convolution kernel is 1×H′×W′, where H′×W′ is the convolution window size, and C and 1 both represent the number of channels.

[0081] Since the actual CH4 concentration data observed by satellites and other methods are too low, most of the filling methods currently used only use data with valuable points, such as the machine learning filling model of the NO2 model. When the data coverage of this model is too low, the training data will be greatly reduced, and the filling ability in some areas without observations all year round will be lost, and the generalization ability will be weakened; or filling models such as FNO will also regard valueless points as valuable points and screen them out in the subsequent loss function. This type of method will also cause the model to have too many invalid signals in the signal conversion process because the valueless points account for too large a proportion in the model training, and it will not be able to reach the number of valid signals required by the Nyquist sampling theorem, resulting in deviations in the model fitting function and reducing the accuracy of the model filling. The depth partial convolution and partial convolution currently used in the present invention can separate the valued points from the valueless points when the data is input, and at the same time can improve the contribution of the valuable points to the overall model training. For example, for a 3×3 data, if there is only one valuable point, the traditional convolution will multiply all points and then average them, so the contribution of the valuable point will be reduced to 1 / 9. The partial convolution only multiplies the valuable points and divides them by the number of valuable points, avoiding the reduction of the contribution of the valuable points.

[0082] S3, constructs an overall loss function based on the evaluation loss function of the unobserved area, the evaluation loss function of the observed area, and the regional overall change loss function, and performs supervised training on the filling model for the methane concentration filling task based on the training samples and the overall loss function.

[0083] In the embodiment, the loss function for evaluating the unobserved area is expressed as loss hole , the loss function of the observation area evaluation is expressed as loss valid :

[0084]

[0085] Among them, k is the data point index, K is the total amount of data, y k Represents the label data of data point k, y′ k Represents the predicted output data of the filling model for data point k, m k Represents the mask data of data point k, mask data m k The value is 1 when the corresponding position is an observed area, and 0 when it is not an observed area.

[0086] The overall regional change loss function is expressed as loss tv :

[0087]

[0088] Among them, || ||1 represents the L1 norm, (i, j)∈P represents that the position (i, j) belongs to the region P, and y com,i,j Represents the synthetic data at position (i, j), synthesized as follows:

[0089] y com,i,j =m i,j ×y i,j +(1-m i,j )×y′ i,j

[0090] Among them, m i,j Represents the mask data at position (i, j), y i,j Represents the label data of position (i, j), y′ i,j Represents the model prediction data at position (i, j);

[0091] The overall loss function is expressed as loss total :

[0092] loss total =α×loss valid +β×loss hole +γ×loss tv

[0093] Among them, α, β, and γ are the loss functions for evaluating the observed area. valid , loss function for evaluating the unobserved area loss hole And the overall regional change loss function loss tv The weight of .

[0094] In the embodiment, based on the above-mentioned overall loss function, the filling model is supervised and trained using training samples for the methane concentration filling task. The specific process is as follows:

[0095] The daily distribution data was sorted chronologically. The initial mask data constructed based on the measured CH4 emission distribution data, the atmospheric chemistry model CH4 concentration distribution data, and the auxiliary environmental data for the same day were superimposed in terms of data dimensions to form a training sample dataset with a structure of C × H × W for training and evaluation. At the same time, the atmospheric chemistry model CH4 concentration data was dimensionally increased to form a training labeled dataset with a structure of 1 × H × W.

[0096] The training and evaluation datasets were divided into a training input dataset and an evaluation input dataset in a 4:1 ratio, in chronological order. The training and label datasets for the same day were randomly mirrored. Specifically, with a 50% probability, the training and label datasets for the same day were mirrored in either the longitude or latitude dimensions, flipping the data head-to-tail. Furthermore, data with initially reversed values, such as wind direction, were inverted. These operations prevented underfitting of the model in specific locations, such as those with long-term high CH4 concentrations, and improved model generalization.

[0097] During training, the number of training rounds and learning rate are set according to the settings. In each training round, a certain number of sample data are randomly extracted without replacement and input into the model for training. The parameters within the model are tuned, including but not limited to the Adam optimizer. The parameters within the model are adjusted according to a gradient-based algorithm to minimize the loss function until the data in the entire training input dataset is exhausted and a new training round begins. The learning rate of each training round varies according to a certain pattern. After completing all rounds, a trained filling model is obtained. Specifically, the number of training rounds is set to 70 and the learning rate is set to 0.005. The learning rate of the initial round is 0.005, and the learning rate of the next round of the final round is 0. A cosine function relationship is thus fitted, and the learning rates of the intermediate rounds are taken according to the corresponding points on the fitted cosine function relationship.

[0098] S4, obtain the measured methane concentration distribution data and auxiliary environmental data as application data, construct a mask reference mask for the application data based on the measured methane concentration distribution data, use the reference mask to mask the application data to form an application sample, use the trained filling model to fill the application sample with data, and obtain full coverage methane concentration distribution data.

[0099] In an embodiment, the measured CH4 concentration distribution data and auxiliary environmental data of a certain day with low coverage to be filled are obtained, and these data are superimposed in dimensions, and a mask reference mask is constructed for the application data at the same time, specifically including: constructing a reference mask for the measured methane concentration distribution data in the application data, specifically: assigning a valueless point in the measured CH4 concentration distribution data of the same day to 0, and modifying a valued point to 1 to construct a reference mask; constructing a reference mask for the auxiliary environmental data in the application data, specifically, all values ​​in the reference mask are 1, and the application data is masked using the reference mask to form an application sample with a structure of C×H×W, and the application sample is input into the trained filling model, and the predicted data is obtained through forward reasoning as the full coverage CH4 concentration distribution data after filling.

[0100] In order to verify the effect of the above filling method, an experimental example is also provided. The specific satellite CH4 dataset to be filled is the CH4 concentration dataset of TROPOMI, which is limited to a certain area, such as Figure 3 As shown in the figure, the preset grid size is 640*512. The data structure of the training sample is 10×640×512. In the padding model, the structure of a single partial convolution kernel is C×H′×W′, and its specific structure can be 10×7×7, 10×5×5, or 10×3×3 in different layers. For the deep partial convolution kernel, the single convolution kernel is 1×H′×W′, and its specific structure can be 1×7×7, 1×5×5, or 1×3×3 in different layers. In the constructed overall loss function, α, β, and γ are set to 1, 6, and 0.1, respectively.

[0101] like Figure 4As shown, an embodiment provides a filling device for low-coverage methane concentration data, including a sample construction module, a filling model construction module, a filling model training module, and a data filling application module, wherein the sample construction module is used to use atmospheric chemical model methane concentration distribution data and auxiliary environmental data as sample data, screen and map the measured methane concentration distribution data, and then construct a reference mask for the sample data, and use the reference mask to mask the sample data to form a training sample; the filling model constructed in the filling model construction module adopts a U-net structure including an encoding part and a decoding part, wherein the encoding part includes a plurality of encoding modules with successively decreasing sizes, each encoding module includes a depth part convolution layer and a space and channel attention module, and the decoding part includes a plurality of decoding modules with successively increasing sizes. Block, each decoding module contains an upsampling layer and a partial convolution layer connected in sequence, which establishes a connection between the partial convolution layer and the partial convolution layer of the same depth; the filling model training module is used to construct an overall loss function based on the unobserved area evaluation loss function, the observed area evaluation loss function, and the regional overall change loss function, and supervise the filling model for the methane concentration filling task based on the training samples and the overall loss function; the data filling application module is used to obtain the measured methane concentration distribution data and auxiliary environmental data as application data, construct a mask reference mask for the application data based on the measured methane concentration distribution data, use the reference mask to mask the application data to form an application sample, and use the trained filling model to perform data filling on the application sample to obtain full coverage methane concentration distribution data.

[0102] It should be noted that the low-coverage methane concentration data filling device provided in the above embodiment should be illustrated by the division of the above-mentioned functional modules when filling data. The above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the low-coverage methane concentration data filling device provided in the above embodiment and the low-coverage methane concentration data filling construction method embodiment are based on the same concept. The specific implementation process is detailed in the embodiment of the methane seamless spatial distribution reconstruction method based on hyperspectral satellite low-coverage data, which will not be repeated here.

[0103] Based on the same inventive concept, an embodiment further provides a computing device including a memory and one or more processors. The memory stores executable code. When the one or more processors execute the executable code, the computing device is used to implement the above-mentioned method for seamless spatial distribution reconstruction of methane based on hyperspectral satellite low-coverage data. The method specifically includes the following steps:

[0104] S1, using atmospheric chemical model methane concentration distribution data and auxiliary environmental data as sample data, screening and mapping the measured methane concentration distribution data to construct a reference mask for the sample data, and using the reference mask to mask the sample data to form a training sample;

[0105] S2, builds a filling model based on a deep partially convolutional U-net network;

[0106] S3, constructs an overall loss function based on the evaluation loss function of the unobserved area, the evaluation loss function of the observed area, and the regional overall change loss function, and performs supervised training of the filling model for the methane concentration filling task based on the training samples and the overall loss function;

[0107] S4, obtain the measured methane concentration distribution data and auxiliary environmental data as application data, construct a mask reference mask for the application data based on the measured methane concentration distribution data, use the reference mask to mask the application data to form an application sample, use the trained filling model to fill the application sample with data, and obtain full coverage methane concentration distribution data.

[0108] The computing device provided in the embodiment, in addition to the processor and memory, also includes hardware required for other services such as internal bus, network interface, memory, etc. at the hardware level. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low coverage data described in S1-S4. Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0109] Based on the same inventive concept, an embodiment further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low coverage data is implemented, specifically comprising the following steps:

[0110] S1, using atmospheric chemical model methane concentration distribution data and auxiliary environmental data as sample data, screening and mapping the measured methane concentration distribution data to construct a reference mask for the sample data, and using the reference mask to mask the sample data to form a training sample;

[0111] S2, builds a filling model based on a deep partially convolutional U-net network;

[0112] S3, constructs an overall loss function based on the evaluation loss function of the unobserved area, the evaluation loss function of the observed area, and the regional overall change loss function, and performs supervised training of the filling model for the methane concentration filling task based on the training samples and the overall loss function;

[0113] S4, obtain the measured methane concentration distribution data and auxiliary environmental data as application data, construct a mask reference mask for the application data based on the measured methane concentration distribution data, use the reference mask to mask the application data to form an application sample, use the trained filling model to fill the application sample with data, and obtain full coverage methane concentration distribution data.

[0114] In the embodiment, computer-readable media includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.

[0115] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for seamless spatial distribution reconstruction of methane based on low-coverage hyperspectral satellite data, characterized in that: The following steps are involved: Sample construction: The atmospheric chemical model methane concentration distribution data and auxiliary environmental data are used as sample data. The measured methane concentration distribution data are screened and mapped to construct a reference mask for the sample data. The reference mask is used to mask the sample data to form a training sample. Among them, the measured methane concentration distribution data that are screened and retained are mapped, including: A two-dimensional grid with equal spacing is preset, with a structure of H×W. The difference in longitude or latitude between each grid point and the adjacent grid points is equal. The measured methane concentration distribution data is processed according to the following formula: The above formula averages the gas concentrations of the pixels in each longitude-latitude grid cell to obtain the methane concentration values at each longitude-latitude grid cell, where XCH4 s represents the CH4 gas concentration at the s-th pixel to be processed on the original pixel plane after linear interpolation, lon s , lat s being the longitude and latitude of this pixel, XCH4 lon,lat representing the CH4 gas concentration at the corresponding longitude-latitude grid point on the longitude-latitude grid, lon and lat being the longitude and latitude of this pixel, grid being the difference in degrees between longitude-latitude grid cells, and the symbol & indicating the meaning of "and", indicating the summation of XCH4 s for all points that satisfy the screening condition (lon s - lon) < grid / 2 & (lat s - lat) < grid / 2, indicating the calculation of the number of all points that satisfy the screening condition (lon s - lon) < grid / 2 & (lat s - lat) < grid / 2, and 1 being the cumulative value of the number of each point; Padding model construction: The padding model adopts a U-net structure consisting of an encoding part and a decoding part. The encoding part consists of multiple encoding modules of decreasing size, each of which consists of a depth-wise partial convolutional layer and a spatial and channel attention module. The decoding part consists of multiple decoding modules of increasing size, each of which consists of an upsampling layer and a partial convolutional layer connected in sequence, which establishes connections between depth-wise partial convolutional layers and partial convolutional layers of the same size. Filling model training: Based on the evaluation loss function for the unobserved area, the evaluation loss function for the observed area, and the overall regional change loss function, an overall loss function is constructed. The filling model is then supervised and trained for the methane concentration filling task based on the training samples and the overall loss function. Data filling application: Obtain measured methane concentration distribution data and auxiliary environmental data as application data, construct a mask reference mask for the application data based on the measured methane concentration distribution data, use the reference mask to mask the application data to form an application sample, and use the trained filling model to fill the application sample with data to obtain full coverage methane concentration distribution data; The data structure of the sample data D and the reference mask M contained in the training sample is the same. When the inference is passed in the depthwise partial convolution layer and the partial convolution layer, for the convolution kernel K with weight F and bias b, after the input data D is sliding convolved by the convolution kernel K, the calculated value d at the corresponding position after one convolution window and the updated mask m are expressed as: Among them, D′ and M′ are the input data and mask data of the corresponding convolution window, N is the number of data in the convolution window, and M′ is the n is the mask data of the corresponding position n in the convolution window, ⊙ is element-by-element calculation, the superscript T represents transposition, the current calculated value d and the updated mask m are used as the input data D′ and mask data M′ of the next layer; For the partial convolution layer, its single convolution kernel will convolve all layers, that is, the structure of a single convolution kernel is C×H′×W′. For the deep partial convolution layer, its single convolution kernel will only convolve a single layer, that is, the structure of a single convolution kernel is 1×H′×W′, where H′×W′ is the convolution window size, and C and 1 both represent the number of channels.

2. The method for seamless spatial distribution reconstruction of methane based on hyperspectral satellite low coverage data according to claim 1 is characterized in that: Screening of the measured methane concentration distribution data, including: If the data provider provides a QA value screening standard, the measured methane concentration distribution data will be screened according to the standard. Specifically, the methane concentration distribution data with a QA value of 1 will be retained, and the remaining data will be removed. If the QA value is not provided, the corresponding methane concentration data will be removed when the standard deviation between the inverted simulated spectrum and the actual collected spectrum is greater than the set threshold. The retained measured methane concentration distribution data will be used to construct the reference mask.

3. The method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low coverage data according to claim 1 is characterized in that: Construct a reference mask for the sample data, including: A reference mask is constructed for the atmospheric chemistry model methane concentration distribution data in the sample data. Specifically, the measured methane concentration distribution data of one day within a fixed number of days before and after the corresponding day in the measured methane concentration data is randomly selected to construct a reference mask for the atmospheric chemistry model methane concentration distribution data. The grids of the mapped measured methane concentration data and the reference mask are determined to be equal. When there is a measured methane concentration data value at the mapped grid point, the corresponding element value in the reference mask is set to 1. When there is no value at the mapped grid point, the corresponding element value in the reference mask is set to 0. Construct a reference mask for the auxiliary environment data in the sample data, where all values ​​in the reference mask are 1; Construct a mask reference mask for the application data, including: A reference mask is constructed for the measured methane concentration distribution data in the application data. Specifically, the reference mask is constructed by assigning the valueless points in the measured CH4 concentration distribution data to 0 and the valued points to 1. A reference mask is constructed for the auxiliary environment data in the application data, and all values ​​in the specific reference mask are 1.

4. The method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low coverage data according to claim 1 is characterized in that: The loss function for evaluating the unobserved area is expressed as loss hole , the loss function of the observation area evaluation is expressed as loss valid : Among them, k is the data point index, K is the total amount of data, y k Represents the label data of data point k, y′ k Represents the predicted output data of the filling model for data point k, m k Represents the mask data of data point k.

5. The method for reconstructing seamless spatial distribution of methane based on hyperspectral satellite low coverage data according to claim 1, characterized in that: The overall regional change loss function is expressed as loss tv : Among them, ‖‖1 represents the L1 norm, (i, j)∈P means that the position (i, j) belongs to the region P, and y com,i,j Represents the synthetic data at position (i, j), synthesized as follows: y com,i,j =m i,j ×y i,j +(1-m i,j )×y′ i,j Among them, m i,j Represents the mask data at position (i, j), y i,j Represents the label data of position (i, j), y′ i,j Represents the model prediction data at position (i, j); The overall loss function is expressed as loss total : loss total =α×loss valid +β×loss hole +γ×loss tv Among them, α, β, and γ are the loss functions for evaluating the observed area. valid , loss function for evaluating the unobserved area loss hole And the overall regional change loss function loss tv The weight of .

6. The method for reconstructing seamless spatial distribution of methane based on hyperspectral satellite low coverage data according to claim 1, characterized in that: Before the training samples are fed into the model, a random mirroring operation is performed on the atmospheric chemistry model methane concentration distribution data. Specifically, the training samples are mirrored in the longitude or latitude dimension according to a given probability, that is, the data is reversed. The auxiliary environmental data includes meteorological parameter data, geographical parameter data, and CH4 surface emission data.

7. A methane seamless spatial distribution reconstruction device based on hyperspectral satellite low coverage data, characterized by: include: A sample construction module is used to use atmospheric chemical model methane concentration distribution data and auxiliary environmental data as sample data, filter and map the measured methane concentration distribution data, construct a reference mask for the sample data, and use the reference mask to mask the sample data to form a training sample; Among them, the measured methane concentration distribution data that are screened and retained are mapped, including: A two-dimensional grid with equal spacing is preset, with a structure of H×W. The difference in longitude or latitude between each grid point and the adjacent grid points is equal. The measured methane concentration distribution data is processed according to the following formula: The above formula averages the gas concentrations of the pixels in each longitude-latitude grid point to obtain the methane concentration values at each longitude-latitude grid point, where XCH4 s represents the CH4 gas concentration at the s-th pixel to be processed on the original pixel plane after linear interpolation, lon s , lat s are the longitude and latitude of this pixel, XCH4 lon,lat represents the CH4 gas concentration at the corresponding longitude-latitude grid point on the longitude-latitude grid, lon and lat are the longitude and latitude of this pixel, grid is the difference in degrees between longitude-latitude grid points, and the symbol & means and, represents the sum of XCH4 s for all points that satisfy the screening condition (lon s - lon) < grid / 2 & (lat s - lat) < grid / 2, represents the calculation of the number of all points that satisfy the screening condition (lon s - lon) < grid / 2 & (lat s - lat) < grid / 2, and 1 is the cumulative value of the number of each point; A padding model construction module, wherein the padding model adopts a U-net structure including an encoding part and a decoding part, wherein the encoding part includes multiple encoding modules of successively decreasing sizes, each encoding module includes a depth-wise partial convolutional layer and a spatial and channel attention module, and the decoding part includes multiple decoding modules of successively increasing sizes, each decoding module includes an upsampling layer and a partial convolutional layer connected in sequence, which establishes connections between depth-wise partial convolutional layers and partial convolutional layers of the same size; The filling model training module is used to construct an overall loss function based on the evaluation loss function of the unobserved area, the evaluation loss function of the observed area, and the overall regional change loss function, and to perform supervised training of the filling model for the methane concentration filling task based on the training samples and the overall loss function; A data filling application module is used to obtain measured methane concentration distribution data and auxiliary environmental data as application data, construct a mask reference mask for the application data based on the measured methane concentration distribution data, use the reference mask to mask the application data to form an application sample, and use the trained filling model to fill the application sample with data to obtain full coverage methane concentration distribution data; The data structure of the sample data D and the reference mask M contained in the training sample is the same. When the inference is passed in the depthwise partial convolution layer and the partial convolution layer, for the convolution kernel K with weight F and bias b, after the input data D is sliding convolved by the convolution kernel K, the calculated value d at the corresponding position after one convolution window and the updated mask m are expressed as: Among them, D′ and M′ are the input data and mask data of the corresponding convolution window, N is the number of data in the convolution window, and M′ is the n is the mask data of the corresponding position n in the convolution window, ⊙ is element-by-element calculation, the superscript T represents transposition, the current calculated value d and the updated mask m are used as the input data D′ and mask data M′ of the next layer; For the partial convolution layer, its single convolution kernel will convolve all layers, that is, the structure of a single convolution kernel is C×H′×W′. For the deep partial convolution layer, its single convolution kernel will only convolve a single layer, that is, the structure of a single convolution kernel is 1×H′×W′, where H′×W′ is the convolution window size, and C and 1 both represent the number of channels.

8. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the method for reconstructing the seamless spatial distribution of methane based on hyperspectral satellite low-coverage data according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Full-coverage atmospheric methane concentration data production method and system coupled with physical mechanism

    CN116486931A

  • Methane concentration data complementation method and device for machine learning

    CN117009750A