A method for inverting VOCS from satellite remote sensing column concentration using machine learning
Through machine learning methods, multi-source data fusion and deep learning modeling of satellite remote sensing data is solved, and the problems of insufficient data integration and single inversion model in the existing technology are achieved, high-resolution VOCs component inversion is achieved, and the accuracy and reliability of inversion are improved.
Patent Information
- Application Number
- CN202510387281.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing satellite remote sensing VOCs inversion technology faces many challenges, including insufficient data integration, unconsidered nonlinear relationships, single inversion model, lack of uncertainty analysis and lack of dynamic correction mechanisms, resulting in insufficient spatial and temporal resolution of the inversion results, which is difficult to meet the needs of refined pollution monitoring and control.
Using machine learning method, multi-source data preprocessing is carried out on satellite remote sensing column concentration data, ground VOCs component monitoring data, related pollutant data, meteorological data and model simulation data, multi-scale feature correlation diagram is constructed, dynamic feature embedding and nonlinear relationship modeling is carried out, multi-task deep learning inversion model is constructed, and adaptive weight optimization is carried out. Finally, spatial and temporal interpolation and uncertainty analysis are performed on satellite remote sensing column concentration data, multi-scale verification and dynamic correction are carried out to obtain high-resolution VOCs component inversion results.
It significantly improves the accuracy, reliability and practicality of satellite remote sensing VOCs inversion, improves the spatial and temporal resolution of the inversion results, enhances the ability to portray complex relationships of VOCs components, improves the credibility of the inversion results, and meets the needs of atmospheric environment management at different scales.
Smart Images

Figure CN119903362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of air pollution monitoring technology, and more specifically, to a method for realizing satellite remote sensing column concentration inversion VOCS by using machine learning. Background Art
[0002] Volatile organic compounds (VOCs) are important precursors of air pollution and play a key role in the formation of ozone and fine particulate matter (PM2.5), which seriously affects air quality and human health. Accurate monitoring and quantification of the spatiotemporal distribution of VOCs are essential for understanding atmospheric chemical processes and formulating effective pollution control strategies. Although traditional ground monitoring networks can provide precise point data, they are limited by spatial coverage and temporal continuity. In recent years, the rapid development of satellite remote sensing technology has provided new possibilities for large-scale VOCs monitoring.
[0003] However, the existing satellite remote sensing VOCs inversion technology faces many challenges and limitations. First, traditional methods have obvious deficiencies in processing multi-source heterogeneous data, and it is difficult to effectively integrate multi-dimensional information such as satellite remote sensing, ground monitoring, meteorological observations and model simulations, resulting in the inversion results failing to fully utilize the advantages of various types of data. Secondly, the existing technology does not take into account the complex nonlinear relationship between VOCs components, and cannot accurately characterize the interaction and transformation process between pollutants, which affects the accuracy of the inversion. Furthermore, a single inversion model is difficult to take into account the inversion accuracy of the total amount of VOCs and the main components at the same time, which restricts the application of the results in environmental management at different scales. In addition, the existing methods generally lack quantitative analysis and dynamic correction mechanisms for the uncertainty of the inversion results, which reduces the reliability and practicality of the results. In practical applications, these problems lead to insufficient spatiotemporal resolution of the inversion results, which is difficult to meet the needs of refined pollution monitoring and control. For example, in complex terrain or rapidly urbanizing areas, existing technologies often cannot accurately capture the local distribution characteristics and short-term change trends of pollutants, affecting the pertinence and effectiveness of pollution prevention and control measures. At the same time, in the case of severe regional transmission and photochemical pollution, it is difficult for existing technologies to provide reliable decision support for cross-regional joint prevention and control due to the inability to accurately distinguish between local emissions and external contributions. These problems seriously restrict the application potential of satellite remote sensing in air pollution monitoring and management.
[0004] In view of this, the present invention proposes a method for inverting VOCS from satellite remote sensing column concentration using machine learning to solve the above problems. Summary of the invention
[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a method for realizing satellite remote sensing column concentration inversion VOCS by machine learning, comprising:
[0006] Multi-source data preprocessing was performed on satellite remote sensing column concentration data, ground VOCs component monitoring data, related pollutant data, meteorological data, and model simulation data to obtain a standardized multi-source feature data set;
[0007] Performing feature hierarchical extraction and cross-modal association analysis based on the standardized multi-source feature data set to obtain a multi-scale feature association map;
[0008] Perform dynamic feature embedding and nonlinear relationship modeling based on the multi-scale feature association graph to obtain a pollutant component feature space;
[0009] Based on the pollutant component feature space, a multi-task deep learning inversion model is constructed, and adaptive weight optimization is performed to obtain an optimized inversion model;
[0010] Based on the optimized inversion model, the satellite remote sensing column concentration data is subjected to spatiotemporal interpolation and uncertainty analysis to obtain a high-resolution VOCs component inversion map;
[0011] Multi-scale verification and dynamic correction are performed based on the high-resolution VOCs component inversion map and ground VOCs component monitoring data to obtain the final VOCs component inversion result.
[0012] Preferably, the multi-source data preprocessing of satellite remote sensing column concentration data, ground VOCs component monitoring data, relevant pollutant data, meteorological data and model simulation data to obtain a standardized multi-source feature data set includes:
[0013] Normalizing the spatial resolution and filling missing values of the satellite remote sensing column concentration data to obtain a preprocessed remote sensing data matrix;
[0014] Performing time series alignment and outlier removal on the ground VOCs component monitoring data and related pollutant data to obtain a preprocessed ground monitoring data matrix;
[0015] Downscaling and spatiotemporal interpolation are performed on the meteorological data to obtain a preprocessed meteorological data matrix;
[0016] Performing grid resampling and error correction on the model simulation data to obtain a preprocessed model simulation data matrix;
[0017] The pre-processed remote sensing data matrix, ground monitoring data matrix, meteorological data matrix and model simulation data matrix are subjected to feature normalization and batch standardization processing to obtain a standardized multi-source feature data set.
[0018] Preferably, performing feature hierarchical extraction and cross-modal association analysis based on the standardized multi-source feature data set to obtain a multi-scale feature association graph includes:
[0019] Performing principal component analysis on the standardized multi-source feature data set to obtain a primary feature matrix, and performing feature hierarchical clustering based on the primary feature matrix to obtain a multi-scale feature set;
[0020] Performing modal decomposition on the multi-scale feature set to obtain a remote sensing modal feature subset, a ground monitoring modal feature subset, a meteorological modal feature subset, and a model simulation modal feature subset;
[0021] Performing mutual information calculation on the remote sensing modal feature subset, the ground monitoring modal feature subset, the meteorological modal feature subset, and the model simulation modal feature subset to obtain a cross-modal feature association matrix;
[0022] Performing graph structure modeling on the cross-modal feature association matrix to obtain an initial feature association graph, and performing edge weight optimization based on the initial feature association graph to obtain a weighted feature association graph;
[0023] Community detection and hierarchical clustering are performed on the weighted feature association graph to obtain a multi-scale feature association graph.
[0024] Preferably, the dynamic feature embedding and nonlinear relationship modeling are performed based on the multi-scale feature association graph to obtain the pollutant component feature space, including:
[0025] Performing graph convolution processing on the multi-scale feature association graph to obtain a graph embedding feature vector, and performing dynamic feature update according to the graph embedding feature vector to obtain a dynamic embedding feature matrix;
[0026] Performing a self-attention mechanism on the dynamic embedding feature matrix to obtain a weighted feature matrix, and performing feature dimensionality reduction according to the weighted feature matrix to obtain a reduced-dimensional feature matrix;
[0027] Performing kernel density estimation on the dimension-reduced feature matrix to obtain feature distribution probability density, and performing nonlinear relationship modeling based on the feature distribution probability density to obtain a nonlinear feature mapping function;
[0028] The nonlinear characteristic mapping function is regularized to obtain a stable characteristic mapping model, and the VOCs component characteristics are spatially projected according to the stable characteristic mapping model to obtain a pollutant component characteristic space.
[0029] Preferably, based on the pollutant component feature space, a multi-task deep learning inversion model is constructed, and adaptive weight optimization is performed to obtain an optimized inversion model, including:
[0030] Performing task decomposition on the pollutant component feature space to obtain a VOCs total amount inversion task subset and a VOCs main component inversion task subset;
[0031] Based on the VOCs total amount inversion task subset and the VOCs main component inversion task subset, a multi-task deep learning network is constructed, and a shared layer and a task-specific layer are designed for the multi-task deep learning network to obtain an initial inversion model;
[0032] Designing a loss function for the initial inversion model to obtain a multi-task joint loss function, and performing gradient balance optimization according to the multi-task joint loss function to obtain a preliminary optimized inversion model;
[0033] The preliminary optimized inversion model is adaptively weighted to obtain a task weight vector, and the model parameters are updated according to the task weight vector to obtain an optimized inversion model.
[0034] Preferably, based on the optimized inversion model, the satellite remote sensing column concentration data is subjected to spatiotemporal interpolation and uncertainty analysis to obtain a high-resolution VOCs component inversion map, including:
[0035] Performing spatiotemporal grid division on the satellite remote sensing column concentration data to obtain spatiotemporal data blocks, and performing feature extraction based on the spatiotemporal data blocks to obtain a spatiotemporal feature sequence;
[0036] Based on the optimized inversion model, the temporal and spatial characteristic sequence is predicted to obtain a primary inversion result, and temporal and spatial interpolation is performed according to the primary inversion result to obtain a high-resolution inversion result;
[0037] Performing Monte Carlo simulation on the high-resolution inversion result to obtain an uncertainty distribution of the inversion result, and performing confidence interval estimation based on the uncertainty distribution of the inversion result to obtain an uncertainty analysis result;
[0038] The high-resolution inversion results and uncertainty analysis results are layered to obtain a high-resolution VOCs component inversion map.
[0039] Preferably, the multi-scale verification and dynamic correction are performed based on the high-resolution VOCs component inversion map and the ground VOCs component monitoring data to obtain the final VOCs component inversion result, including:
[0040] Performing spatial scale segmentation on the high-resolution VOCs component inversion map to obtain a multi-scale inversion sub-map, and extracting an inversion feature vector based on the multi-scale inversion sub-map;
[0041] Decomposing the ground VOCs component monitoring data in time scale to obtain a multi-scale monitoring sequence, and extracting a monitoring feature vector according to the multi-scale monitoring sequence;
[0042] Performing correlation analysis on the inversion feature vector and the monitoring feature vector to obtain a multi-scale verification index set, and performing error distribution modeling based on the multi-scale verification index set to obtain an error distribution model;
[0043] Based on the error distribution model, the high-resolution VOCs component inversion map is dynamically corrected to obtain a corrected inversion map, and the results are fused according to the corrected inversion map to obtain the final VOCs component inversion result.
[0044] Preferably, the loss function design is performed on the initial inversion model to obtain a multi-task joint loss function, and gradient balance optimization is performed according to the multi-task joint loss function to obtain a preliminary optimized inversion model, including:
[0045] Defining subtask loss functions for the VOCs total amount inversion task subset and the VOCs main component inversion task subset respectively to obtain a subtask loss function set, and performing weighted combination according to the subtask loss function set to obtain an initial multi-task joint loss function;
[0046] Performing gradient conflict detection on the initial multi-task joint loss function to obtain a gradient conflict matrix, and performing gradient balancing adjustment according to the gradient conflict matrix to obtain a balanced multi-task joint loss function;
[0047] Based on the balanced multi-task joint loss function, the initial inversion model is back-propagated and trained to obtain a preliminary optimized inversion model.
[0048] Preferably, based on the error distribution model, the high-resolution VOCs component inversion map is dynamically corrected to obtain a corrected inversion map, and the results are fused according to the corrected inversion map to obtain the final VOCs component inversion result, including:
[0049] Performing Bayesian inference on the error distribution model to obtain an error posterior distribution, and calculating error correction coefficients based on the error posterior distribution to obtain a dynamic correction coefficient matrix;
[0050] Based on the dynamic correction coefficient matrix, the high-resolution VOCs component inversion map is corrected pixel by pixel to obtain a corrected inversion map;
[0051] Performing multi-scale feature decomposition on the corrected inversion map to obtain a multi-scale inversion feature set, and performing feature weighted fusion according to the multi-scale inversion feature set to obtain a fusion feature matrix;
[0052] Performing probability density estimation on the fusion feature matrix to obtain a probability distribution of an inversion result, and performing maximum a posteriori estimation based on the probability distribution of the inversion result to obtain a preliminary fusion result;
[0053] The preliminary fusion result is subjected to time series smoothing processing to obtain a smooth fusion result, and a consistency check is performed based on the smooth fusion result and the ground VOCs component monitoring data to obtain a final VOCs component inversion result.
[0054] A device for implementing satellite remote sensing column concentration inversion VOCS using machine learning, the device for implementing satellite remote sensing column concentration inversion VOCS using machine learning includes a memory and a processor, the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the method for implementing satellite remote sensing column concentration inversion VOCS using machine learning.
[0055] The technical effects and advantages of the method for realizing satellite remote sensing column concentration inversion VOCS by machine learning in the present invention are as follows:
[0056] The present invention significantly improves the accuracy, reliability and practicality of satellite remote sensing VOCs inversion. Through the deep fusion of multi-source data and the application of advanced algorithms, the spatiotemporal resolution of the inversion results is greatly improved, making it possible to finely characterize the distribution of pollutant concentrations. At the same time, the ability to characterize the complex relationship between VOCs components is significantly enhanced, providing important support for a deep understanding of atmospheric chemical processes. The introduced uncertainty analysis and dynamic correction mechanism greatly improve the credibility of the inversion results and provide a more reliable data basis for atmospheric pollution monitoring and decision-making. In addition, the multi-task learning framework realizes the collaborative inversion of the total amount and main components of VOCs, meeting the needs of atmospheric environmental management at different scales. Overall, it provides a more accurate and comprehensive scientific basis for fields such as atmospheric pollution monitoring, source analysis, pollution prevention and control, and environmental policy formulation, which helps to improve the level of refinement of regional air quality management, and has important practical application value for improving air quality and protecting public health. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a schematic diagram of the steps of the method for inverting VOCs concentration from satellite remote sensing columns using machine learning in the present invention;
[0058] Figure 2 This is a schematic diagram of the system for realizing satellite remote sensing column concentration inversion VOCS using machine learning in the present invention. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] This application example provides a method and system. The execution subjects of the method and system include but are not limited to: mechanical equipment, data processing platform, cloud server node, network upload device, etc. equipped with the system can be regarded as the general computing node of this application, and the data processing platform includes but is not limited to: remote sensing data processing system, environmental monitoring system, and atmospheric pollution model system.
[0061] The present invention provides a method for realizing VOCs inversion from satellite remote sensing column concentration by using machine learning, comprising the following steps:
[0062] Step S1: preprocess satellite remote sensing column concentration data, ground VOCs component monitoring data, relevant pollutant data, meteorological data and model simulation data to obtain a standardized multi-source feature data set;
[0063] Step S2: Perform feature hierarchical extraction and cross-modal correlation analysis based on the standardized multi-source feature data set to obtain a multi-scale feature correlation map;
[0064] Step S3: dynamic feature embedding and nonlinear relationship modeling are performed based on the multi-scale feature association graph to obtain the pollutant component feature space;
[0065] Step S4: Based on the pollutant component feature space, a multi-task deep learning inversion model is constructed, and adaptive weight optimization is performed to obtain an optimized inversion model;
[0066] Step S5: Based on the optimized inversion model, perform spatiotemporal interpolation and uncertainty analysis on the satellite remote sensing column concentration data to obtain a high-resolution VOCs component inversion map;
[0067] Step S6: Perform multi-scale verification and dynamic correction based on the high-resolution VOCs component inversion map and ground VOCs component monitoring data to obtain the final VOCs component inversion result.
[0068] The present invention improves data quality by preprocessing multi-source data, thereby providing a reliable data basis for subsequent analysis; deeply explores the potential associations between different data sources through feature hierarchical extraction and cross-modal association analysis, and constructs a multi-scale feature association graph; accurately expresses complex pollutant component feature space by using dynamic feature embedding and nonlinear relationship modeling; constructs a multi-task deep learning inversion model based on the pollutant component feature space, and performs adaptive weight optimization to effectively improve the accuracy and generalization ability of the model; performs spatiotemporal interpolation and uncertainty analysis on satellite remote sensing column concentration data to generate a high-resolution VOCs component inversion graph, thereby improving spatial resolution and time continuity; and accurately corrects the inversion results through multi-scale verification and dynamic correction in combination with ground VOCs component monitoring data, thereby finally obtaining a VOCs component inversion result with high accuracy and high reliability.
[0069] In the embodiment of the present invention, refer to Figure 1 , is a schematic flow chart of the steps of the method for inverting VOCs from satellite remote sensing column concentration using machine learning in the present invention. In this example, the steps of the method for inverting VOCs from satellite remote sensing column concentration using machine learning include:
[0070] Step S1: preprocess satellite remote sensing column concentration data, ground VOCs component monitoring data, relevant pollutant data, meteorological data and model simulation data to obtain a standardized multi-source feature data set;
[0071] In this embodiment, multi-source data are first collected, including satellite remote sensing column concentration data (such as remote sensing observation data from satellites such as TROPOMI and OMI), ground VOCs component monitoring data (such as VOCs concentration monitoring data at ground stations), relevant pollutant data (such as NOx, SO2, O3 and other pollutant observation data), meteorological data (such as temperature, humidity, wind speed, wind direction, etc.) and model simulation data (such as simulation results of models such as CMAQ and WRF-Chem). These multi-source heterogeneous data are preprocessed, including spatial resolution normalization and missing value filling, time series alignment and outlier removal, downscaling and spatiotemporal interpolation, grid resampling and error correction, etc., and finally feature normalization and batch standardization are performed to obtain a standardized multi-source feature data set with a unified format and scale, providing a high-quality data foundation for subsequent analysis and modeling.
[0072] Step S2: Perform feature hierarchical extraction and cross-modal correlation analysis based on the standardized multi-source feature data set to obtain a multi-scale feature correlation map;
[0073] In this embodiment, the standardized multi-source feature data set is first subjected to principal component analysis to extract the primary feature matrix, and then the feature hierarchical clustering is performed based on the matrix to obtain a multi-scale feature set. The multi-scale feature set is then subjected to modal decomposition to separate the remote sensing modal feature subset, the ground monitoring modal feature subset, the meteorological modal feature subset, and the model simulation modal feature subset. The mutual information between these different modal feature subsets is calculated, a cross-modal feature association matrix is constructed, and then the graph structure modeling is performed based on the association matrix to obtain an initial feature association graph. The initial feature association graph is edge weight optimized to obtain a weighted feature association graph, and finally a multi-scale feature association graph is constructed through community detection and hierarchical clustering, which reflects the complex association relationship between features of different scales and modalities.
[0074] Step S3: dynamic feature embedding and nonlinear relationship modeling are performed based on the multi-scale feature association graph to obtain the pollutant component feature space;
[0075] In this embodiment, the multi-scale feature association graph is subjected to graph convolution processing to extract the graph embedding feature vector, and the dynamic feature update is performed based on the vector to obtain a dynamic embedding feature matrix. Then the self-attention mechanism is applied to the matrix to obtain a weighted feature matrix, and then the feature dimension reduction is performed to obtain a reduced dimension feature matrix. The reduced dimension feature matrix is subjected to kernel density estimation to obtain the feature distribution probability density, and a nonlinear feature mapping function is constructed based on this. The mapping function is regularized to obtain a stable feature mapping model, and finally the model is used to perform spatial projection of the VOCs component characteristics to construct a pollutant component feature space, which accurately expresses the complex nonlinear relationship between the VOCs components and various influencing factors.
[0076] Step S4: Based on the pollutant component feature space, a multi-task deep learning inversion model is constructed, and adaptive weight optimization is performed to obtain an optimized inversion model;
[0077] In this embodiment, the pollutant component feature space is task-decomposed and divided into a VOCs total amount inversion task subset and a VOCs main component inversion task subset. Based on these task subsets, a multi-task deep learning network is constructed, and a shared layer and a task-specific layer are designed to obtain an initial inversion model. A multi-task joint loss function is designed for the initial model, and a preliminary optimized inversion model is obtained through gradient balance optimization. The model is further adaptively weighted, the task weight vector is calculated, and the model parameters are updated based on the vector, and finally an optimized inversion model is obtained, which can simultaneously invert the total amount of VOCs and the concentration of the main components with high precision.
[0078] Step S5: Based on the optimized inversion model, perform spatiotemporal interpolation and uncertainty analysis on the satellite remote sensing column concentration data to obtain a high-resolution VOCs component inversion map;
[0079] In this embodiment, the satellite remote sensing column concentration data is divided into time and space grids to obtain time and space data blocks, and the time and space feature sequences are extracted. The optimized inversion model is used to predict the time and space feature sequence to obtain the primary inversion result, and then time and space interpolation is performed to obtain the high-resolution inversion result. Monte Carlo simulation is performed on the high-resolution inversion result, the uncertainty distribution of the inversion result is analyzed, and the confidence interval is estimated to obtain the uncertainty analysis result. Finally, the high-resolution inversion result and the uncertainty analysis result are layered to generate a high-resolution VOCs component inversion map, which not only shows the spatial distribution of the VOCs component, but also quantifies the uncertainty of the inversion result.
[0080] Step S6: Perform multi-scale verification and dynamic correction based on the high-resolution VOCs component inversion map and ground VOCs component monitoring data to obtain the final VOCs component inversion result.
[0081] In this embodiment, the high-resolution VOCs component inversion map is segmented by spatial scale to obtain a multi-scale inversion sub-map, and the inversion feature vector is extracted. At the same time, the ground VOCs component monitoring data is decomposed by time scale to obtain a multi-scale monitoring sequence, and the monitoring feature vector is extracted. The correlation between the inversion feature vector and the monitoring feature vector is analyzed to obtain a multi-scale verification index set, and an error distribution model is established based on the index set. The error distribution model is used to dynamically correct the high-resolution VOCs component inversion map to obtain a corrected inversion map, and finally the final VOCs component inversion result is generated by result fusion, achieving high-precision and high-reliability VOCs component column concentration inversion.
[0082] In this embodiment, the detailed implementation steps of step S1 include:
[0083] The satellite remote sensing column concentration data is normalized in spatial resolution and filled with missing values to obtain the preprocessed remote sensing data matrix;
[0084] The ground VOCs component monitoring data and related pollutant data were time-series aligned and outliers were eliminated to obtain the preprocessed ground monitoring data matrix;
[0085] Downscaling and spatiotemporal interpolation are performed on meteorological data to obtain a preprocessed meteorological data matrix;
[0086] Performing grid resampling and error correction on the model simulation data to obtain a preprocessed model simulation data matrix;
[0087] The preprocessed remote sensing data matrix, ground monitoring data matrix, meteorological data matrix and model simulation data matrix are subjected to feature normalization and batch standardization to obtain a standardized multi-source feature data set.
[0088] In this embodiment, remote sensing column concentration data of multiple satellites (such as TROPOMI, OMI, GOME-2, etc.) are collected. These data contain column concentration information of various atmospheric pollutants (such as NO2, HCHO, CO, etc.) with different spatial resolutions and coverage. The spatial resolution of these remote sensing data is normalized to unify data of different resolutions into the same spatial grid. The specific method adopts techniques such as bilinear interpolation or convolution kernel resampling. The problem of missing values in satellite remote sensing data is handled, which is usually caused by factors such as cloud cover and instrument failure. The missing values are filled by spatiotemporal interpolation methods (such as Kriging interpolation, machine learning interpolation, etc.). The normalized and missing value-filled remote sensing data are organized into a structured matrix form as a preprocessed remote sensing data matrix.
[0089] Collect VOCs component monitoring data from ground monitoring sites, including concentration data of different types of VOCs (such as alkanes, alkenes, aromatic hydrocarbons, etc.). At the same time, collect monitoring data of related pollutants (such as NOx, SO2, O3, etc.) as auxiliary features. Perform time series alignment processing on these ground monitoring data from different sources and different sampling frequencies and unify them to the same time scale. Use statistical methods (such as the 3σ criterion, box plot method, etc.) to identify and eliminate outliers in the ground monitoring data to improve data quality. Integrate the processed ground VOCs component monitoring data and related pollutant data into a structured matrix form as the preprocessed ground monitoring data matrix.
[0090] Collect meteorological data in the study area, including parameters such as temperature, humidity, air pressure, wind speed, wind direction, and boundary layer height. These meteorological data usually come from meteorological station observations or meteorological reanalysis data sets (such as ERA5, NCEP, etc.). Downscale the high-resolution meteorological reanalysis data to match its spatial resolution with the scale of the study area. Perform spatiotemporal interpolation on the discrete observation data of the meteorological station to generate continuous meteorological field data. The processed meteorological data is organized into a structured matrix form as the preprocessed meteorological data matrix.
[0091] Collect simulation data from chemical transport models (such as CMAQ, WRF-Chem, GEOS-Chem, etc.), which contain the three-dimensional concentration distribution of VOCs components. Resample the grid of the model simulation data to make it consistent with the spatial resolution of the study area. Perform statistical error correction on the model simulation data based on historical observation data to reduce the systematic bias of the model. Arrange the resampled and error-corrected model simulation data into a structured matrix form as the preprocessed model simulation data matrix.
[0092] Each feature in the above four pre-processed data matrices (remote sensing data matrix, ground monitoring data matrix, meteorological data matrix and model simulation data matrix) is normalized. Commonly used methods include Min-Max normalization and Z-score normalization. Batch normalization technology is used to further process the features to reduce the distribution differences between features. The four data matrices after normalization and batch normalization are integrated to form a unified standardized multi-source feature data set, which provides a basis for subsequent analysis.
[0093] In this embodiment, the detailed implementation steps of step S2 include:
[0094] Perform principal component analysis on the standardized multi-source feature data set to obtain a primary feature matrix, and perform hierarchical clustering based on the primary feature matrix to obtain a multi-scale feature set;
[0095] Perform modal decomposition on the multi-scale feature set to obtain remote sensing modal feature subsets, ground monitoring modal feature subsets, meteorological modal feature subsets, and model simulation modal feature subsets;
[0096] Mutual information is calculated for remote sensing modal feature subsets, ground monitoring modal feature subsets, meteorological modal feature subsets, and model simulation modal feature subsets to obtain a cross-modal feature correlation matrix;
[0097] Model the graph structure of the cross-modal feature association matrix to obtain an initial feature association graph, and optimize the edge weights based on the initial feature association graph to obtain a weighted feature association graph;
[0098] Community detection and hierarchical clustering are performed on the weighted feature association graph to obtain a multi-scale feature association graph.
[0099] In this embodiment, the principal component analysis (PCA) algorithm is applied to the standardized multi-source feature data set to reduce the data dimension and extract the main features. The mathematical formula of PCA is expressed as: ; Where X is the original data matrix, W is the eigenvector matrix, is the feature matrix after dimensionality reduction. The primary feature matrix is obtained by retaining the principal components whose explained variance ratio is greater than a certain threshold (such as 95%). A hierarchical clustering algorithm, such as DBSCAN or a hierarchical clustering algorithm, is applied to the primary feature matrix to divide the features into clusters of different scales. Based on the clustering results, a multi-scale feature set is constructed, which contains feature representations at different scales.
[0100] Perform modal decomposition on the multi-scale feature set according to the data source and divide the features into different modal subsets. Extract features from satellite remote sensing data from the multi-scale feature set to form a remote sensing modal feature subset. Extract features from ground monitoring data from the multi-scale feature set to form a ground monitoring modal feature subset. Extract features from meteorological data from the multi-scale feature set to form a meteorological modal feature subset. Extract features from model simulation data from the multi-scale feature set to form a model simulation modal feature subset.
[0101] It should be noted that “features” refer to numerical indicators extracted from raw data that can characterize the distribution and changes of VOCs. For example, the remote sensing modality feature subset may include column concentration data, ultraviolet absorption index, HCHO / NO 2 ratio, seasonal change slope (seasonal HCHO concentration change rate (molec / cm 2 / month) etc.; the ground monitoring modal feature subset may include the concentration of main components (such as benzene (1.2ppb), toluene (2.8ppb), ethylene (4.1ppb)), the difference between daytime and nighttime concentrations, the autocorrelation coefficient lagged by 24 hours (0.76), the correlation between ozone and specific VOCs components etc.; the meteorological modal feature subset may include basic meteorological parameters, atmospheric stability index (such as Pasquill-Gifford stability grade (AF)), photochemical effective radiation (ultraviolet radiation intensity), cloud coverage etc.; the model simulation modal feature subset may include VOCs concentrations at different altitudes, contribution rates of different emission sources (%) (such as traffic sources (35%), industrial sources (42%)), characteristic parameters of the daily variation curve of emission intensity, photochemical aging degree (ratio that characterizes the degree of oxidation of VOCs components), etc.
[0102] The mutual information between different modal feature subsets is calculated to quantify the nonlinear correlation between different modal features. The mathematical expression of mutual information is: ; Where p(x, y) is the joint probability distribution of features X and Y in different modal feature subsets, and p(x) and p(y) are the marginal probability distributions of features X and Y in different modal feature subsets, respectively. The mutual information is calculated between all pairs of modal feature subsets, and an n×n mutual information matrix is constructed (n is the total number of modal feature subsets). The mutual information matrix is normalized to obtain the cross-modal feature association matrix, which describes the association strength between different modal features.
[0103] The cross-modal feature association matrix is converted into a graph structure, where nodes represent features and edges represent the association between features. The weight of the edge is determined by the value of the corresponding element in the association matrix. The edges in the graph are screened, and the edges with weights greater than a certain threshold are retained to form an initial feature association graph. The graph learning algorithm is applied to optimize the edge weights in the initial feature association graph to improve the expressiveness of the graph. The optimization method can adopt a graph structure learning algorithm based on gradient descent, such as the graph attention network (GAT). After optimization, a weighted feature association graph is obtained, which more accurately reflects the strength of association between features.
[0104] Apply community detection algorithms, such as Louvain algorithm or label propagation algorithm, to the weighted feature association graph to identify the community structure in the graph. The goal of community detection is to maximize modularity. ;in, is the weight between nodes i and j, and are the degrees of nodes i and j respectively, m is the sum of the weights of all edges in the graph, is an indicator function, which is 1 when nodes i and j belong to the same community, otherwise it is 0. Hierarchical clustering is performed on the identified communities to construct a hierarchical feature association structure. Based on the results of community detection and hierarchical clustering, a multi-scale feature association graph is constructed, which contains the association relationship between features at different scales.
[0105] In this embodiment, the detailed implementation steps of step S3 include:
[0106] Perform graph convolution processing on the multi-scale feature association graph to obtain a graph embedding feature vector, and perform dynamic feature update based on the graph embedding feature vector to obtain a dynamic embedding feature matrix;
[0107] The dynamic embedding feature matrix is processed by the self-attention mechanism to obtain a weighted feature matrix, and feature dimension reduction is performed based on the weighted feature matrix to obtain a reduced dimension feature matrix;
[0108] Perform kernel density estimation on the reduced dimension feature matrix to obtain the feature distribution probability density, and perform nonlinear relationship modeling based on the feature distribution probability density to obtain the nonlinear feature mapping function;
[0109] The nonlinear feature mapping function is regularized to obtain a stable feature mapping model, and the VOCs component characteristics are spatially projected according to the stable feature mapping model to obtain the pollutant component feature space.
[0110] In this embodiment, a graph convolutional network (GCN) is applied to the multi-scale feature association graph to extract the embedded representation of the node. The expression of the graph convolutional network is: ;in, Is to join the self-loop The adjacency matrix after (identity matrix) ), the elements in the adjacency matrix reflect the strength of the association between features; D is the degree matrix, is the node representation of the lth layer, is the weight matrix of the lth layer, is a nonlinear activation function. Through multiple layers of GCN, we get the embedding vectors of each node (feature) in the graph, which capture the local structural information of the node. For each time step, the graph embedding feature vector is updated according to the new observation data. The update process can be implemented using recurrent neural networks such as gated recurrent units (GRU) or long short-term memory networks (LSTM).
[0111] Apply the self-attention mechanism to the dynamic embedding feature matrix to calculate the correlation weights between features. The expression of the self-attention mechanism is: ; Among them, Attention is a self-attention mechanism, Q, K, and V are query (representing the target features we are concerned about, such as the concentration distribution characteristics of a specific VOCs component), key (representing the "identity" of each feature, used to evaluate its relevance to the query feature) and value (containing the actual information of each feature, which will be selectively aggregated) matrices, is the transpose of K, and d_k is the dimension of the key vector. Through the self-attention mechanism, a weighted feature matrix (representing the strength of association between each pair of features) is obtained, which emphasizes the representation of important features. Dimensionality reduction techniques such as t-SNE or UMAP are applied to the weighted feature matrix to further reduce the feature dimension. The reduced-dimensional features are organized into a matrix form to obtain a reduced-dimensional feature matrix.
[0112] Apply kernel density estimation to each feature dimension in the reduced feature matrix to estimate the feature The probability distribution of . The expression is: ; where K is the kernel function (such as Gaussian kernel), h is the bandwidth parameter, is the Jth sample point. Through KDE, we get the probability density function of the feature dimension, which describes the distribution characteristics of the feature. Based on the probability density function of the feature, we build a nonlinear relationship model between the features. We can use methods such as Gaussian process regression or radial basis function network. By constructing a nonlinear mapping function between features, the function can capture complex nonlinear relationships.
[0113] Add regularization terms to the nonlinear feature mapping function to control model complexity and prevent overfitting. The optimization objective after regularization is expressed as: ; Wherein, Loss(θ) is the original loss function, Ω(θ) is the regularization term, and λ is the regularization coefficient. The optimal regularization parameter is determined by cross-validation to obtain a stable feature mapping model. Using the stable feature mapping model, the VOCs component characteristics are spatially projected and converted to a new feature space. In this new feature space, the complex relationship between the VOCs components and various influencing factors is better expressed. The projected feature space is named the pollutant component feature space, which provides a good feature representation for the subsequent inversion model.
[0114] In this embodiment, step S4 includes the following steps:
[0115] The pollutant component feature space is decomposed into tasks to obtain the VOCs total amount inversion task subset and the VOCs main component inversion task subset;
[0116] Based on the VOCs total amount inversion task subset and the VOCs main component inversion task subset, a multi-task deep learning network is constructed, and the shared layer and task-specific layer of the multi-task deep learning network are designed to obtain the initial inversion model;
[0117] The loss function of the initial inversion model is designed to obtain a multi-task joint loss function, and gradient balance optimization is performed based on the multi-task joint loss function to obtain a preliminary optimized inversion model;
[0118] The preliminary optimized inversion model is adaptively weighted to obtain a task weight vector, and the model parameters are updated according to the task weight vector to obtain an optimized inversion model.
[0119] In this embodiment, the pollutant component feature space is analyzed to identify different types of inversion tasks. The inversion tasks are divided into two main subsets: the VOCs total amount inversion task subset, which is responsible for predicting the overall concentration of VOCs in the area; the VOCs main component inversion task subset, which is responsible for predicting the concentrations of each main VOCs component (such as alkanes, alkenes, aromatic hydrocarbons, etc.). For each task subset, the corresponding input features and target outputs are defined. A multi-task deep learning network architecture is designed, which can simultaneously handle the two tasks of VOCs total amount inversion and VOCs main component inversion. The network design includes: a shared layer for learning the underlying feature representation common to all tasks; a total amount inversion specific layer for VOCs total amount inversion task specific feature extraction and mapping; a component inversion specific layer for VOCs main component inversion task specific feature extraction and mapping. The designed network structure is instantiated to form an initial inversion model.
[0120] Define an appropriate loss function for each subtask. For the VOCs total amount inversion task, the mean square error (MSE) or mean absolute error (MAE) can be used as the loss function; for the VOCs main component inversion task, the sum of the mean square errors of the concentrations of each component or the weighted mean square error can be used as the loss function. Combine the loss functions of each subtask into a joint loss function in the following form: ; Among them, L_to is the loss of the VOCs total amount inversion task, L_com is the loss of the VOCs main component inversion task, and is the weight coefficient. During the training process, monitor the gradients of different tasks and detect gradient conflicts. When the gradient directions of different tasks are inconsistent, gradient conflicts will occur, affecting the convergence of the model. Use gradient balancing techniques, such as gradient normalization and gradient projection, to alleviate the gradient conflict problem. The adjusted gradient calculation method is: ;in, is the original gradient, is the normalized gradient. Use the balanced gradient to update the model parameters and obtain a preliminary optimized inversion model.
[0121] The performance of the initially optimized inversion model is evaluated, and the validation set performance index of each task is calculated. According to the task performance index, the task weight coefficient is dynamically adjusted. The uncertainty-based task weight adjustment method can be used, and the calculation formula of the task weight is: ;in, It's a task The prediction uncertainty is (using Monte Carlo Dropout sampling, keeping the dropout activation during inference, performing multiple sampling, and taking the variance of multiple predictions as the prediction uncertainty). For the task The calculated task weights are organized into a vector form as the task weight vector.
[0122] Assuming that the VOCs inversion system has three main tasks, the prediction uncertainty and task weights are explained in Table 1:
[0123] Table 1
[0124] Based on the task weight vector, the weight coefficients in the multi-task joint loss function are updated, the model is retrained and the parameters are updated. After multiple rounds of iterative optimization, the final optimized inversion model is obtained, which can complete the simultaneous inversion of the total amount of VOCs and the main components with high accuracy.
[0125] In this embodiment, the loss function is designed for the initial inversion model to obtain a multi-task joint loss function, and gradient balance optimization is performed according to the multi-task joint loss function to obtain the specific steps of the preliminary optimized inversion model:
[0126] Subtask loss functions are defined for the VOCs total amount inversion task subset and the VOCs main component inversion task subset respectively to obtain a subtask loss function set, and a weighted combination is performed based on the subtask loss function set to obtain an initial multi-task joint loss function;
[0127] Perform gradient conflict detection on the initial multi-task joint loss function to obtain a gradient conflict matrix, and perform gradient balancing adjustment based on the gradient conflict matrix to obtain a balanced multi-task joint loss function;
[0128] Based on the balanced multi-task joint loss function, the initial inversion model is back-propagated and trained to obtain a preliminary optimized inversion model.
[0129] In this embodiment, the loss function L_to is defined for the VOCs total amount inversion task, and regression loss functions such as mean square error (MSE) or mean absolute error (MAE) are selected. The loss function L_com is defined for the VOCs main component inversion task, and the sum of the mean square errors of the concentrations of each component is selected: ;in, is the actual concentration of the sth VOCs component, is the concentration of the sth VOCs component predicted by the model. is the total number of components. The above subtask loss functions are combined into an initial multi-task joint loss function. During the training process, the gradient direction of each task on each sample is calculated to form a gradient vector. The cosine similarity between the gradient vectors of different tasks is calculated, and the gradient conflict matrix C is constructed. The elements in the gradient conflict matrix C are ;in, It's a task The gradient vector of It's a task The cosine similarity range is between [-1, 1], and the smaller the value, the more serious the gradient conflict. According to the gradient conflict matrix, a gradient balancing strategy is designed. For task pairs with serious conflicts, methods such as gradient orthogonalization or gradient scaling are used. The formula for gradient orthogonalization is: ;in, and It's a task and tasks The gradient of It is the adjusted task Gradient. Apply the balanced adjusted gradient to the multi-task joint loss function to obtain the balanced multi-task joint loss function: ;in, and is the weight coefficient after equilibrium.
[0130] The initial inversion model is trained by back propagation algorithm using balanced multi-task joint loss function. After multiple rounds of iterative training, a preliminary optimized inversion model is obtained, which has good performance in each subtask.
[0131] In this embodiment, step S5 includes the following steps:
[0132] The satellite remote sensing column concentration data is divided into time and space grids to obtain time and space data blocks, and feature extraction is performed based on the time and space data blocks to obtain time and space feature sequences;
[0133] Based on the optimized inversion model, the temporal and spatial characteristic sequence is predicted to obtain the primary inversion result, and the temporal and spatial interpolation is performed according to the primary inversion result to obtain the high-resolution inversion result;
[0134] Monte Carlo simulation is performed on the high-resolution inversion results to obtain the uncertainty distribution of the inversion results, and confidence interval estimation is performed based on the uncertainty distribution of the inversion results to obtain the uncertainty analysis results;
[0135] The high-resolution inversion results and uncertainty analysis results are layered to obtain a high-resolution VOCs component inversion map.
[0136] In this embodiment, the satellite remote sensing column concentration data is divided into a spatiotemporal grid, and the data is divided into several data blocks in time and space. In space, the grid division can be performed according to latitude and longitude or administrative divisions; in time, it can be divided according to time granularity such as hours, days, and weeks. For each spatiotemporal data block, relevant features are extracted, including remote sensing features (such as NO2, HCHO column concentrations), meteorological features (such as temperature, wind speed, etc.), and spatiotemporal features (such as longitude and latitude, time, etc.). The extracted features are organized into a sequence form to form a spatiotemporal feature sequence as the input of the model.
[0137] The spatiotemporal feature sequence is input into the optimized inversion model to obtain the primary inversion results of the total amount and main components of VOCs. The primary inversion results reflect the distribution of VOCs at the original resolution of the satellite data. Spatiotemporal interpolation methods are applied to the primary inversion results to improve the spatial resolution of the inversion results. Methods such as Kriging interpolation, bilinear interpolation, or deep learning interpolation can be used. The spatiotemporal interpolation process takes into account spatial correlation and temporal continuity to generate continuous high-resolution inversion results. High-resolution inversion results provide more detailed information on the spatial distribution of VOCs, which is conducive to refined analysis and management.
[0138] Uncertainty analysis is performed on the high-resolution inversion results to quantify the reliability of the inversion results. The Monte Carlo simulation method is used to randomly perturb the model parameters and input features to generate multiple sets of inversion results. The specific steps include: adding random noise to the model parameters ,like ,in Follows Gaussian distribution N_σ; are the original model parameters, To add random noise Add random noise to the input features ,like ,in Following the Gaussian distribution N_σ, is the original input feature, The input features after adding random noise; using the disturbed model and features, generate multiple sets of inversion results; statistically analyze the distribution characteristics of multiple sets of inversion results to obtain the uncertainty distribution of the inversion results. According to the uncertainty distribution of the inversion results, calculate the confidence interval. The commonly used confidence interval is the 95% confidence interval, that is, [μ-1.96σ1, μ+1.96σ1], where μ is the mean of the inversion result and σ1 is the standard deviation. The confidence interval provides the uncertainty range of the inversion result, which helps to evaluate the reliability of the result.
[0139] The high-resolution inversion results and uncertainty analysis results are designed as different layers and superimposed through geographic information system (GIS) tools. The inversion result layer shows the spatial distribution of the total amount and main components of VOCs, usually using color gradients to represent the concentration level; the uncertainty layer shows the reliability of the inversion results, usually using transparency or specific patterns to represent the size of uncertainty. By superimposing the layers, a high-resolution VOCs component inversion map is generated, which not only shows the spatial distribution of VOCs, but also indicates the credibility of the results through uncertainty information, providing comprehensive information support for scientific research and decision-making.
[0140] In this embodiment, step S6 includes the following steps:
[0141] The high-resolution VOCs component inversion map is spatially segmented to obtain multi-scale inversion sub-maps, and the inversion feature vectors are extracted based on the multi-scale inversion sub-maps.
[0142] The ground VOCs component monitoring data is decomposed into time scales to obtain a multi-scale monitoring sequence, and the monitoring feature vector is extracted based on the multi-scale monitoring sequence;
[0143] Perform correlation analysis on the inversion feature vector and the monitoring feature vector to obtain a multi-scale verification index set, and perform error distribution modeling based on the multi-scale verification index set to obtain an error distribution model;
[0144] Based on the error distribution model, the high-resolution VOCs component inversion map is dynamically corrected to obtain the corrected inversion map, and the results are fused according to the corrected inversion map to obtain the final VOCs component inversion result.
[0145] In this embodiment, the high-resolution VOCs component inversion map is segmented by spatial scale and divided into sub-maps of different spatial scales. The segmentation method can be based on spatial resolution (such as 1km, 5km, 10km, etc.), administrative divisions (such as streets, counties, cities, etc.) or natural geographical units (such as river basins, plains, mountainous areas, etc.). For the inversion sub-map at each spatial scale, feature vectors are extracted, including statistical features (such as mean, variance, quantile, etc.), spatial features (such as spatial autocorrelation index, hot spot distribution, etc.) and component features (such as component proportion, principal component, etc.). The extracted feature vectors are organized into a matrix form as inversion feature vectors for subsequent verification analysis.
[0146] The ground VOCs component monitoring data is decomposed by time scale and divided into sequences of different time scales. The decomposition can be based on time frequency (such as hourly scale, daily scale, weekly scale, monthly scale, etc.) or characteristic time (such as high pollution period, low pollution period, seasonal changes, etc.). For the monitoring sequence at each time scale, feature vectors are extracted, including time characteristics (such as trend, cycle, anomaly, etc.), statistical characteristics (such as mean, variance, quantile, etc.) and component characteristics (such as component proportion, correlation, etc.). The extracted feature vectors are organized into a matrix form as monitoring feature vectors for subsequent verification analysis.
[0147] The inversion feature vector and the monitoring feature vector are paired according to the time-space correspondence, and the correlation index between them is calculated. The correlation indexes include Pearson correlation coefficient, Spearman rank correlation coefficient, mutual information, etc. These indexes measure the consistency between the inversion results and the monitoring data from different angles. At different spatial scales and time scales, the correlation indexes are calculated respectively to form a multi-scale verification index set. The multi-scale verification index set reflects the accuracy and reliability of the inversion results at different scales. According to the multi-scale verification index set, the distribution model of the inversion error is established. The error distribution model can adopt a parametric model (such as Gaussian distribution, t distribution, etc.) or a non-parametric model (such as kernel density estimation, histogram, etc.). The model parameters are determined by methods such as maximum likelihood estimation or moment estimation. The error distribution model describes the deviation distribution between the inversion result and the true value, providing a basis for subsequent correction.
[0148] In this embodiment, based on the error distribution model, the high-resolution VOCs component inversion map is dynamically corrected to obtain a corrected inversion map, and the results are fused according to the corrected inversion map to obtain the final VOCs component inversion result. The specific steps are as follows:
[0149] Perform Bayesian inference on the error distribution model to obtain the error posterior distribution, and calculate the error correction coefficient based on the error posterior distribution to obtain a dynamic correction coefficient matrix;
[0150] Based on the dynamic correction coefficient matrix, the high-resolution VOCs component inversion map is corrected pixel by pixel to obtain the corrected inversion map;
[0151] Perform multi-scale feature decomposition on the corrected inversion map to obtain a multi-scale inversion feature set, and perform feature weighted fusion based on the multi-scale inversion feature set to obtain a fusion feature matrix;
[0152] Perform probability density estimation on the fusion feature matrix to obtain the probability distribution of the inversion result, and perform maximum a posteriori estimation based on the probability distribution of the inversion result to obtain the preliminary fusion result;
[0153] The preliminary fusion results are smoothed in time series to obtain smooth fusion results, and a consistency check is performed between the smooth fusion results and the ground VOCs component monitoring data to obtain the final VOCs component inversion results.
[0154] In this embodiment, based on the error distribution model, the Bayesian inference method is applied to update the error distribution estimate. The formula is: ;in, is the posterior distribution of the parameter θ considering the spatial position e, time t and environmental factors E, is a likelihood function that takes into account the spatiotemporal characteristics (using the Gaussian spatiotemporal process model and dynamically constructing it through variogram analysis to capture the error correlation structure in different regions and time periods to obtain the likelihood function), represents the observation dataset of the rth data source used for VOCs inversion, The weights of multi-source data fusion (spatial position and the weight coefficient of the rth data source at time t), is a prior distribution with spatiotemporal variability (a prior construction driven by historical data), is a regularization term that controls the model complexity. is the regularization coefficient, which can be optimized by cross-validation. It is the environmental factor adjustment function, capturing the impact of environmental variables on parameter distribution.
[0155] ;in, is the quality indicator (such as signal-to-noise ratio, data integrity) of the data source r at location e and time t, is the sensitivity parameter (manual setting). The environmental factor adjustment function is a nonlinear function between environmental factors such as temperature, humidity, and light.
[0156] Through Bayesian inference, the posterior distribution of the error is obtained, which combines the information of prior knowledge and observed data. According to the posterior distribution of the error, the error correction coefficient is calculated. The calculation method of the correction coefficient is:
[0157] Error correction factor ; Among them, y_true is the true value, y_pred is the predicted value of the model, For expectations, Represents the VOCs concentration value (predicted value) obtained by the initial inversion of the model. It should be explained that in the expression, y_pred is a random variable, which represents the concept of "predicted value". When the predicted value is equal to the specific value x, the correction coefficients are calculated for different locations, different times, and different VOCs components to form a dynamic correction coefficient matrix.
[0158] The dynamic correction coefficient matrix is used to perform pixel-by-pixel correction on the high-resolution VOCs component inversion map. The correction formula is: ; where y_or is the original inversion value, foll is the correction function, It is the inversion value after pixel-by-pixel correction; it is usually a linear function or a nonlinear function. The corrected inversion map retains the spatial structural characteristics of the original inversion map and improves the accuracy of the numerical value. The corrected inversion map is subjected to multi-scale wavelet transform or other multi-scale decomposition methods to separate the features of different scales. The expression of multi-scale decomposition is: ;in, is the detail coefficient of the pth scale, is the approximate coefficient of the N3th scale, is the corrected inversion map. Different weights are assigned to features of different scales, and weighted fusion is performed to obtain a fusion feature matrix. The weight can be determined according to the importance or reliability of features at each scale.
[0159] It should be noted that the formulas of the present invention are all dimensionless and numerically calculated. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula for the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.
[0160] For each element in the fusion feature matrix, estimate its probability density function to obtain the probability distribution of the inversion result. Probability density estimation can use parametric methods (such as assuming normal distribution) or non-parametric methods (such as kernel density estimation). According to the probability distribution of the inversion result, calculate the maximum a posteriori estimate (MAP) as the preliminary fusion result. Apply time series smoothing to the preliminary fusion result to reduce abnormal fluctuations in time. Smoothing methods can use moving average, exponential smoothing or Kalman filtering. After smoothing, a continuous and stable inversion result time series is obtained. Compare the smoothed fusion results with the ground VOCs component monitoring data, and calculate consistency indicators such as correlation coefficient, mean square error, consistency index, etc. According to the consistency test results, make final adjustments to the inversion results to obtain the final VOCs component inversion results, which combines the advantages of satellite remote sensing and ground monitoring and has high accuracy and reliability.
[0161] This application significantly improves the accuracy, reliability and practicality of satellite remote sensing VOCs inversion. Through the deep fusion of multi-source data and the application of advanced algorithms, the spatiotemporal resolution of the inversion results has been greatly improved, making it possible to fine-tune the distribution of pollutant concentrations. At the same time, the ability to characterize the complex relationships between VOCs components has been significantly enhanced, providing important support for a deep understanding of atmospheric chemical processes. The introduced uncertainty analysis and dynamic correction mechanisms have greatly improved the credibility of the inversion results and provided a more reliable data basis for atmospheric pollution monitoring and decision-making. In addition, the multi-task learning framework realizes the collaborative inversion of the total amount and main components of VOCs, meeting the needs of atmospheric environmental management at different scales. Overall, it provides a more accurate and comprehensive scientific basis for fields such as atmospheric pollution monitoring, source analysis, pollution prevention and control, and environmental policy formulation, which helps to improve the level of refinement of regional air quality management, and has important practical application value for improving air quality and protecting public health.
[0162] The above describes the method for realizing satellite remote sensing column concentration inversion VOCS using machine learning in the embodiment of the present application. The following describes the system for realizing satellite remote sensing column concentration inversion VOCS using machine learning in the embodiment of the present application. Figure 2 In the embodiment of the present application, one embodiment of the system for realizing satellite remote sensing column concentration inversion VOCS using machine learning includes:
[0163] The data synthesis module is used to preprocess satellite remote sensing column concentration data, ground VOCs component monitoring data, related pollutant data, meteorological data and model simulation data to obtain a standardized multi-source feature data set;
[0164] The association graph construction module performs feature hierarchical extraction and cross-modal association analysis based on the standardized multi-source feature data set to obtain a multi-scale feature association graph;
[0165] The feature space construction module performs dynamic feature embedding and nonlinear relationship modeling based on multi-scale feature association graphs to obtain the pollutant component feature space;
[0166] The model building module builds a multi-task deep learning inversion model based on the pollutant component feature space, and performs adaptive weight optimization to obtain the optimized inversion model;
[0167] The inversion map construction module performs spatiotemporal interpolation and uncertainty analysis on satellite remote sensing column concentration data based on the optimized inversion model to obtain a high-resolution VOCs component inversion map;
[0168] The comprehensive result module performs multi-scale verification and dynamic correction based on the high-resolution VOCs component inversion map and ground VOCs component monitoring data to obtain the final VOCs component inversion result; each module is connected by wired and / or wireless means to realize data transmission between modules.
[0169] The present application also provides a device for implementing satellite remote sensing column concentration inversion VOCS using machine learning. The device for implementing satellite remote sensing column concentration inversion VOCS using machine learning includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the method for implementing satellite remote sensing column concentration inversion VOCS using machine learning in the above-mentioned embodiments.
[0170] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A method for inverting VOCS from satellite remote sensing column concentration using machine learning, characterized in that: include: Step S1: preprocess satellite remote sensing column concentration data, ground VOCs component monitoring data, relevant pollutant data, meteorological data and model simulation data to obtain a standardized multi-source feature data set; Step S2: performing feature hierarchical extraction and cross-modal association analysis based on the standardized multi-source feature data set to obtain a multi-scale feature association graph; Step S3: performing dynamic feature embedding and nonlinear relationship modeling based on the multi-scale feature association graph to obtain a pollutant component feature space; Step S4: constructing a multi-task deep learning inversion model based on the pollutant component feature space, and performing adaptive weight optimization to obtain an optimized inversion model; Step S5: Based on the optimized inversion model, performing spatiotemporal interpolation and uncertainty analysis on the satellite remote sensing column concentration data to obtain a high-resolution VOCs component inversion map; Step S6: performing multi-scale verification and dynamic correction based on the high-resolution VOCs component inversion map and the ground VOCs component monitoring data to obtain the final VOCs component inversion result; Based on the pollutant component feature space, a multi-task deep learning inversion model is constructed, and adaptive weight optimization is performed to obtain an optimized inversion model, including: Performing task decomposition on the pollutant component feature space to obtain a VOCs total amount inversion task subset and a VOCs main component inversion task subset; Based on the VOCs total amount inversion task subset and the VOCs main component inversion task subset, a multi-task deep learning network is constructed, and a shared layer and a task-specific layer are designed for the multi-task deep learning network to obtain an initial inversion model; Designing a loss function for the initial inversion model to obtain a multi-task joint loss function, and performing gradient balance optimization according to the multi-task joint loss function to obtain a preliminary optimized inversion model; The preliminary optimized inversion model is adaptively weighted to obtain a task weight vector, and the model parameters are updated according to the task weight vector to obtain an optimized inversion model.
2. The method for realizing satellite remote sensing column concentration inversion VOCS using machine learning according to claim 1, characterized in that: The multi-source data preprocessing is performed on the satellite remote sensing column concentration data, the ground VOCs component monitoring data, the relevant pollutant data, the meteorological data and the model simulation data to obtain a standardized multi-source feature data set, including: Normalizing the spatial resolution and filling missing values of the satellite remote sensing column concentration data to obtain a preprocessed remote sensing data matrix; Performing time series alignment and outlier removal on the ground VOCs component monitoring data and related pollutant data to obtain a preprocessed ground monitoring data matrix; Downscaling and spatiotemporal interpolation are performed on the meteorological data to obtain a preprocessed meteorological data matrix; Performing grid resampling and error correction on the model simulation data to obtain a preprocessed model simulation data matrix; The pre-processed remote sensing data matrix, ground monitoring data matrix, meteorological data matrix and model simulation data matrix are subjected to feature normalization and batch standardization processing to obtain a standardized multi-source feature data set.
3. The method for realizing satellite remote sensing column concentration inversion VOCS using machine learning according to claim 1, characterized in that: The step of performing feature hierarchical extraction and cross-modal association analysis based on the standardized multi-source feature data set to obtain a multi-scale feature association graph includes: Performing principal component analysis on the standardized multi-source feature data set to obtain a primary feature matrix, and performing feature hierarchical clustering based on the primary feature matrix to obtain a multi-scale feature set; Performing modal decomposition on the multi-scale feature set to obtain a remote sensing modal feature subset, a ground monitoring modal feature subset, a meteorological modal feature subset, and a model simulation modal feature subset; Performing mutual information calculation on the remote sensing modal feature subset, the ground monitoring modal feature subset, the meteorological modal feature subset, and the model simulation modal feature subset to obtain a cross-modal feature association matrix; Performing graph structure modeling on the cross-modal feature association matrix to obtain an initial feature association graph, and performing edge weight optimization based on the initial feature association graph to obtain a weighted feature association graph; Community detection and hierarchical clustering are performed on the weighted feature association graph to obtain a multi-scale feature association graph.
4. The method for realizing satellite remote sensing column concentration inversion VOCS using machine learning according to claim 1, characterized in that: The dynamic feature embedding and nonlinear relationship modeling are performed based on the multi-scale feature association graph to obtain the pollutant component feature space, including: Performing graph convolution processing on the multi-scale feature association graph to obtain a graph embedding feature vector, and performing dynamic feature update according to the graph embedding feature vector to obtain a dynamic embedding feature matrix; Performing a self-attention mechanism on the dynamic embedding feature matrix to obtain a weighted feature matrix, and performing feature dimensionality reduction according to the weighted feature matrix to obtain a reduced-dimensional feature matrix; Performing kernel density estimation on the dimension-reduced feature matrix to obtain feature distribution probability density, and performing nonlinear relationship modeling based on the feature distribution probability density to obtain a nonlinear feature mapping function; The nonlinear characteristic mapping function is regularized to obtain a stable characteristic mapping model, and the VOCs component characteristics are spatially projected according to the stable characteristic mapping model to obtain a pollutant component characteristic space.
5. The method for realizing satellite remote sensing column concentration inversion VOCS using machine learning according to claim 1, characterized in that: Based on the optimized inversion model, the satellite remote sensing column concentration data is subjected to spatiotemporal interpolation and uncertainty analysis to obtain a high-resolution VOCs component inversion map, including: Performing spatiotemporal grid division on the satellite remote sensing column concentration data to obtain spatiotemporal data blocks, and performing feature extraction based on the spatiotemporal data blocks to obtain a spatiotemporal feature sequence; Based on the optimized inversion model, the temporal and spatial characteristic sequence is predicted to obtain a primary inversion result, and temporal and spatial interpolation is performed according to the primary inversion result to obtain a high-resolution inversion result; Performing Monte Carlo simulation on the high-resolution inversion result to obtain an uncertainty distribution of the inversion result, and performing confidence interval estimation based on the uncertainty distribution of the inversion result to obtain an uncertainty analysis result; The high-resolution inversion results and uncertainty analysis results are layered to obtain a high-resolution VOCs component inversion map.
6. The method for realizing satellite remote sensing column concentration inversion VOCS using machine learning according to claim 1, characterized in that: The multi-scale verification and dynamic correction are performed based on the high-resolution VOCs component inversion map and the ground VOCs component monitoring data to obtain the final VOCs component inversion result, including: Performing spatial scale segmentation on the high-resolution VOCs component inversion map to obtain a multi-scale inversion sub-map, and extracting an inversion feature vector based on the multi-scale inversion sub-map; Decomposing the ground VOCs component monitoring data in time scale to obtain a multi-scale monitoring sequence, and extracting a monitoring feature vector according to the multi-scale monitoring sequence; Performing correlation analysis on the inversion feature vector and the monitoring feature vector to obtain a multi-scale verification index set, and performing error distribution modeling based on the multi-scale verification index set to obtain an error distribution model; Based on the error distribution model, the high-resolution VOCs component inversion map is dynamically corrected to obtain a corrected inversion map, and the results are fused according to the corrected inversion map to obtain the final VOCs component inversion result.
7. The method for realizing satellite remote sensing column concentration inversion VOCS using machine learning according to claim 5, characterized in that: The loss function is designed for the initial inversion model to obtain a multi-task joint loss function, and gradient balance optimization is performed according to the multi-task joint loss function to obtain a preliminary optimized inversion model, including: Defining subtask loss functions for the VOCs total amount inversion task subset and the VOCs main component inversion task subset respectively to obtain a subtask loss function set, and performing weighted combination according to the subtask loss function set to obtain an initial multi-task joint loss function; Performing gradient conflict detection on the initial multi-task joint loss function to obtain a gradient conflict matrix, and performing gradient balancing adjustment according to the gradient conflict matrix to obtain a balanced multi-task joint loss function; Based on the balanced multi-task joint loss function, the initial inversion model is back-propagated and trained to obtain a preliminary optimized inversion model.
8. The method for realizing satellite remote sensing column concentration inversion VOCS using machine learning according to claim 6, characterized in that: Based on the error distribution model, the high-resolution VOCs component inversion map is dynamically corrected to obtain a corrected inversion map, and the results are fused according to the corrected inversion map to obtain a final VOCs component inversion result, including: Performing Bayesian inference on the error distribution model to obtain an error posterior distribution, and calculating error correction coefficients based on the error posterior distribution to obtain a dynamic correction coefficient matrix; Based on the dynamic correction coefficient matrix, the high-resolution VOCs component inversion map is corrected pixel by pixel to obtain a corrected inversion map; Performing multi-scale feature decomposition on the corrected inversion map to obtain a multi-scale inversion feature set, and performing feature weighted fusion according to the multi-scale inversion feature set to obtain a fusion feature matrix; Performing probability density estimation on the fusion feature matrix to obtain a probability distribution of an inversion result, and performing maximum a posteriori estimation based on the probability distribution of the inversion result to obtain a preliminary fusion result; The preliminary fusion result is subjected to time series smoothing processing to obtain a smooth fusion result, and a consistency check is performed based on the smooth fusion result and the ground VOCs component monitoring data to obtain a final VOCs component inversion result.
9. A device for implementing satellite remote sensing column concentration inversion VOCS using machine learning, the device comprising a memory and a processor, the memory storing computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the method for implementing satellite remote sensing column concentration inversion VOCS using machine learning as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Ozone profile and sulfur dioxide column concentration collaborative inversion method
CN111929263A
Remote sensing-based detection system and method for gaseous pollutant from diesel vehicle exhaust
US20210131964A1