A method for analyzing the sources of atmospheric particulate matter based on deep learning
By constructing a large-scale preprocessing data set and combining a deep learning model, the problem of insufficient analysis of atmospheric particulate matter sources in traditional methods is solved, and accurate analysis and efficient prediction of atmospheric particulate matter sources are achieved.
Patent Information
- Application Number
- CN202411263515.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-09-10
AI Technical Summary
The traditional atmospheric particulate source analysis methods have problems such as high time and resource consumption and insufficient results. Especially when combining the original data with atmospheric particulate concentration data, the application of deep learning technology has not been fully explored.
By building a large-scale preprocessed data set, combining convolutional neural networks and recurrent neural networks, a deep learning model is established, and the model hyperparameters are optimized using integrated learning algorithms and Bayesian optimization to achieve accurate analysis of the source of atmospheric particulate matter.
The accurate analysis of the source of atmospheric particulate matter is achieved, the model's ability to identify the nonlinear relationship between particulate matter and pollution source is improved, the system's robustness and prediction capabilities are enhanced, and the analysis results are ensured with high accuracy and stability.
Smart Images

Figure CN119202597B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of atmospheric particulate matter analysis, and specifically to a method for analyzing the sources of atmospheric particulate matter based on deep learning. Background Art
[0002] Atmospheric particulate matter is one of the main components of air pollution and has significant hazards to the environment and human health. Particulate matter mainly comes from industrial emissions, transportation, construction, agricultural activities and natural processes. Among them, inhalable particulate matter with a diameter of less than 10 microns (PM10) and fine particulate matter with a diameter of less than 2.5 microns (PM2.5) are particularly noteworthy because they can penetrate deep into the respiratory tract and lungs, posing a direct threat to human health. In recent years, with the acceleration of urbanization and the increase in industrial activities, the problem of atmospheric particulate matter pollution has become increasingly serious. Especially in developing countries and emerging economies, particulate matter pollution has become one of the environmental problems that urgently need to be addressed. However, due to the complex sources and diverse components of particulate matter, traditional methods for analyzing the source of particulate matter face many challenges.
[0003] Traditional methods for analyzing the sources of atmospheric particulate matter mainly include chemical composition analysis, physical property analysis, and model analysis. These methods usually rely on laboratory analysis of sampled particulate matter or reverse deduction of pollution sources based on atmospheric chemical models. However, these methods have significant limitations: laboratory analysis requires a lot of time and resources, and it is difficult to analyze the sources of atmospheric particulate matter in real time; atmospheric chemical models usually require a lot of prior knowledge and assumptions, and the calculations are complex and the results are not accurate enough. In the field of atmospheric particulate matter source analysis, the application of deep learning technology is still in the exploratory stage. Although some studies have tried to use deep learning models to predict atmospheric particulate matter concentrations, there is still a lack of systematic research on the analysis of atmospheric particulate matter sources. In particular, how to effectively combine raw data (such as pollution source emission data, historical meteorological data, etc.) with atmospheric particulate matter concentration data, and use deep learning models to accurately analyze the specific sources of particulate matter, there is still a large technical gap in this field. Summary of the invention
[0004] In order to solve the above technical problems, the present invention is implemented by the following technical solutions: a method for analyzing the sources of atmospheric particulate matter based on deep learning, comprising:
[0005] Construct a large-scale preprocessing data set, including atmospheric particulate matter concentration data from different regions, seasons and meteorological conditions, and collect the corresponding potential pollution source emission data, historical meteorological data and other raw data, integrate the raw data to form a multi-dimensional preprocessing data set;
[0006] A new deep learning model is established using convolutional neural network (CNN) and recurrent neural network (RNN). The deep learning model of atmospheric particulate matter data is constructed by using the feature extraction capability of convolutional neural network and the time series data processing capability of recurrent neural network. The deep learning model is used to automatically learn the complex nonlinear relationship between the characteristics of atmospheric particulate matter and various pollution sources.
[0007] Introduce an integrated learning algorithm to form a prediction model, use Bayesian optimization to optimize the hyperparameters of the prediction model, process the characteristics of atmospheric particulate matter through sparse coding, extract representative key features, and enhance the robustness and performance of the model; apply the trained prediction model to real-time atmospheric monitoring data, and analyze the contribution ratio of various pollution sources to particulate matter concentration in real time; the prediction model dynamically adjusts parameters according to real-time data to ensure high accuracy of the analysis results; the prediction model is automatically updated and adjusted by adding new data, and the integrated learning algorithm of the prediction model of the source of atmospheric particulate matter is updated in real time;
[0008] During the prediction model training process, cross-validation is used to evaluate the generalization ability of the prediction model to ensure the consistent performance of the prediction model on different preprocessed data sets, and the hyperparameters of the prediction model are continuously adjusted through Bayesian optimization, so that the prediction model can maintain efficient and stable performance under different environments and conditions. This technical solution achieves accurate analysis of the sources of atmospheric particulate matter through the integration of raw data and the combination of deep learning models. The convolutional neural network can effectively extract particulate matter characteristics, while the recurrent neural network can process complex time series data.
[0009] Preferably, the construction of a large-scale preprocessed data set is not limited to atmospheric particulate matter concentration data and pollution source emission data, but also includes social economic activity data and traffic flow data related to atmospheric particulate matter concentration change factors. The preprocessed data set covers urban, suburban and rural areas in the spatial dimension, and covers daily data and extreme weather event data for multiple years in the temporal dimension;
[0010] In the process of constructing the preprocessing dataset, the spatial interpolation algorithm is used to interpolate the missing spatial data, and the time series analysis method is used to restore the missing time data, thus forming a complete and continuous spatiotemporal dataset, providing richer and more comprehensive input data for model training. This technical solution expands the dimension of the preprocessing dataset, enabling the prediction model to consider more factors that affect the concentration of atmospheric particulate matter, and improves the accuracy of particulate matter source analysis.
[0011] Preferably, the ensemble learning algorithm uses random forest (RF) and gradient boosted decision tree (GBDT) as the ensemble learning algorithm of the prediction model, uses the multi-decision tree ensemble algorithm of random forest to reduce overfitting, and the gradient boosted decision tree improves the prediction ability of the prediction model through iterative optimization, and is used for the analysis of nonlinear features;
[0012] On this basis, the Bayesian optimization algorithm is used to automatically tune the hyperparameters of the prediction model, including the number of trees, depth, learning rate, etc., and finally generate the optimal prediction model, so that it can maintain efficient prediction performance under diverse input data conditions. The random forest and gradient boosting decision tree are combined to utilize the advantages of both for nonlinear feature analysis, effectively reducing the overfitting phenomenon of the prediction model.
[0013] Preferably, the ensemble learning algorithm comprises the following steps:
[0014] Data preprocessing: Collect raw data, including atmospheric particulate matter concentration data, pollution source emission data and historical meteorological data, to form a preprocessing database, and perform preprocessing operations such as cleaning, missing value processing and normalization on the raw data to eliminate dimensional differences between the data;
[0015] Feature extraction: Extract key features that affect the concentration of atmospheric particulate matter from the original data, such as environmental variables: particulate matter concentration, wind speed, wind direction, temperature, and humidity. Use feature selection algorithms to screen out feature variables that affect the source analysis of atmospheric particulate matter, and calculate the proportion of each feature variable to obtain the proportion of each feature variable.
[0016] Prediction model training: The processed raw data is divided into a training set and a test set. The training set corresponds to the training data, and the test set corresponds to the test data. The prediction model learns the rules through the training data and verifies its generalization ability on the test data.
[0017] Use the random forest multi-decision tree ensemble algorithm to gradually build multiple decision trees in an iterative manner. Each tree is trained based on the residual of the previous tree and learns different rules in the training set. Cross-validation and hyperparameter adjustment are used to avoid overfitting, and the depth of the tree, learning rate, and number of iterations are adjusted to improve the generalization ability of the ensemble algorithm.
[0018] Prediction model verification: Use test data to verify the trained prediction model, evaluate the accuracy and stability of the prediction model, and evaluate the accuracy, mean square error (MSE), and R 2 Finally, the optimal prediction model is applied to the prediction and analysis of actual data;
[0019] Use training data to evaluate the prediction performance of the prediction model, calculate the accuracy and mean square error indicators, judge the analytical ability of the prediction model, and further verify the robustness of the prediction model through cross-validation and hyperparameter adjustment to prevent overfitting;
[0020] The trained prediction model is used to predict the actual observed data, analyze the contribution ratio of various pollution sources to the particulate matter concentration, and identify the main pollution sources as output results based on the analysis results.
[0021] The steps of the ensemble learning algorithm are described in detail, including data preprocessing, feature extraction, prediction model training, validation and application, emphasizing the role of random forests and gradient boosting decision trees in preventing overfitting and improving generalization ability.
[0022] Preferably, the integrated learning algorithm is provided with an optimization module to regularly retrain and optimize the prediction model, form a comparison chart based on the collected data and historical data, regularly output the historical comparison data at a set time, and automatically generate an analysis report based on actual needs to assist decision makers in formulating environmental governance strategies.
[0023] The optimization module can regularly output the comparison results of historical data and generate data reports according to preset time periods (such as weekly, monthly, quarterly, and annually). By comparing the data changes in different time periods, the changing trends of atmospheric particulate matter concentrations and sources are analyzed to assist environmental management departments in formulating more accurate pollution prevention and control strategies.
[0024] Preferably, the optimization module compares the new data with the data in the historical database, that is, compares the historical data with the new data, generates comparison results for each week, month, quarter and year, visualizes the comparison results through log transformation, and generates a log graph to show the magnitude and trend of data changes.
[0025] The log transformation can amplify smaller differences and reduce the impact of larger differences, thereby more intuitively displaying the changes in data at different time scales. It is particularly important when analyzing sudden changes in atmospheric particulate matter concentrations, ensuring sensitivity to data changes and analysis accuracy.
[0026] Preferably, the optimization module uses the contrast difference to calculate the difference between the contrast data, and the formula is as follows:
[0027] Difference(t)=New Data(t)-Historical Data(t)
[0028] Where Difference(t) is the data difference at time point t; NewData(t) is the new data collected or observed at time point t; HistoricalData(t) is the historical data or baseline data before time point t;
[0029] Among them, the log transformation formula in the optimization module is:
[0030] Log Transformed Data(t)=log(1+|Difference(t)|)×sign(Difference(t))
[0031] Among them, Log Transformed Data(t) represents the result of logarithmic transformation of the original data difference at time point t;
[0032] log is the natural logarithm, the logarithm with base e, |Difference(t)| is the absolute value of the data difference at time point t, Difference(t) is the data difference at time point t, and sign(Difference(t)) is the sign function used to return the sign of the data difference.
[0033] Among them, the prediction model fusion formula in the optimization module is:
[0034]
[0035] Where Ensemble Output(t) is the ensemble output at time point t, ∑ represents the sum of the outputs of all models, i is the index of the model, from 1 to N, N is the total number of models included in the ensemble, and w i is the weight of the i-th model, which is a non-negative real number. i (t) is the output of the i-th model at time point t.
[0036] The difference between historical data and new data is calculated by contrast difference, and the weighted prediction model is used to fuse the analytical results of multiple prediction models;
[0037] The contrast difference formula is used to quantify the magnitude of data changes between different time points. Through the weighted prediction model, the output results of multiple prediction models are weighted and fused to make the final analysis result more stable and reliable.
[0038] The weights of the prediction model are dynamically adjusted through a Bayesian optimization algorithm to ensure that under different environmental conditions and data inputs, the prediction model can adaptively optimize the weight distribution to improve the accuracy of atmospheric particulate matter source analysis.
[0039] Preferably, in the data preprocessing stage, that is, when integrating the original data, the high-dimensional data is reduced in dimension by the principal component analysis (PCA) algorithm to reduce the redundancy of the data and improve the computational efficiency and prediction accuracy of the prediction model;
[0040] The calculation formula of the principal component analysis algorithm is:
[0041] Z=XW
[0042] Among them, X is the input data matrix; W is the feature vector matrix; Z is the data matrix after dimension reduction. The PCA algorithm is used to extract the feature dimension that best explains the data variance for subsequent prediction model training.
[0043] The PCA algorithm performs eigenvalue decomposition on the data matrix, selects the first few principal components with the largest contribution rate, retains the main information of the data, and reduces the dimension of the input variables, alleviating the computational burden of the prediction model, enabling it to converge faster and improve prediction accuracy.
[0044] Preferably, the hyperparameters of the prediction model are optimized by an improved Bayesian optimization algorithm, in which Gaussian process regression (GPR) is introduced as a proxy model to model the optimization process of each iteration;
[0045] The core formula of the Bayesian optimization algorithm is:
[0046] x t+1 = argmax x α(x|D 1:t )
[0047] Among them, α(x|D 1:t ) is the acquisition function, D 1:t For the data set after iterating t times, find the next optimal hyperparameter by maximizing the acquisition function.
[0048] An improved Bayesian optimization algorithm is used to optimize the hyperparameters of the prediction model. Gaussian process regression is introduced as a proxy model in the algorithm to model the optimization process of each iteration.
[0049] Preferably, during the training and optimization process of the model, the Adam optimization algorithm in the adaptive learning rate optimization algorithm is used to update the parameters. The adaptive learning rate optimization algorithm dynamically adjusts the learning step size according to the change of the gradient to improve the convergence speed and generalization ability of the model on complex data sets;
[0050] The core calculation formula of the Adam optimization algorithm is:
[0051]
[0052] Among them, θ tis the parameter of the model at time t, α is the learning rate, is an unbiased estimate of the first-order momentum, is an unbiased estimate of the second-order momentum, To prevent division by zero,
[0053] In practical applications, the system continuously adjusts the update step size of parameters through an adaptive learning rate optimization algorithm, so that the model can adapt to different data distributions and improve the parsing accuracy. Especially when faced with large-scale and diversified data, the model can converge quickly and maintain stable prediction performance.
[0054] The Adam optimization algorithm combines the advantages of the momentum and RMSprop algorithms. It takes into account both the current gradient and the direction of the historical gradient when updating parameters, thereby achieving more stable and rapid model convergence. Through the adaptive learning rate optimization algorithm, the model can find the optimal solution more quickly in a complex and dynamic data environment and avoid falling into the local optimum, thereby improving the accuracy of atmospheric particulate matter source analysis and the breadth of model application.
[0055] The present invention provides a method for analyzing the sources of atmospheric particulate matter based on deep learning, which has the following beneficial effects:
[0056] 1. This deep learning-based atmospheric particulate matter source analysis method, through an adaptive learning rate optimization algorithm, can find the optimal solution more quickly in a complex and dynamic data environment and avoid falling into local optimality, thereby improving the accuracy of atmospheric particulate matter source analysis and the breadth of model application; through precise analysis of the sources of atmospheric particulate matter, it effectively guides the formulation of urban air pollution control measures, reduces the risk of residents being exposed to high concentrations of particulate matter, and improves the scientificity and effectiveness of air quality management.
[0057] 2. This deep learning-based method for analyzing the sources of atmospheric particulate matter is suitable for various complex environmental conditions and pollution sources, and can be widely used in urban air pollution monitoring, industrial emission control, traffic pollution control and other fields. It provides strong technical support for governments at all levels and environmental protection agencies, and helps to achieve more scientific and accurate environmental management and pollution prevention. At the same time, the present invention has shown significant advantages in solving problems such as the complexity of the sources of atmospheric particulate matter, real-time analysis and model robustness, and provides important technical means and support for air pollution control and environmental protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is a flowchart of a method for analyzing the source of atmospheric particulate matter based on deep learning in the present invention;
[0059] Figure 2 A flowchart of generating a log graph for the present invention;
[0060] Figure 3 Flowchart of the integrated learning algorithm of the present invention. DETAILED DESCRIPTION
[0061] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention are provided for the purpose of illustration and description, and are not intended to be exhaustive or to limit the present invention to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present invention, and to enable those of ordinary skill in the art to understand the present invention and thereby design various embodiments with various modifications suitable for specific uses.
[0062] like Figures 1 to 3 As shown, the present invention provides a technical solution: a method for analyzing the source of atmospheric particulate matter based on deep learning, comprising: constructing a large-scale preprocessing data set, including atmospheric particulate matter concentration data from different regions, seasons and meteorological conditions, and collecting corresponding potential pollution source emission data, historical meteorological data and other raw data, integrating the raw data to form a multi-dimensional preprocessing data set;
[0063] A new deep learning model is established using convolutional neural network (CNN) and recurrent neural network (RNN). The deep learning model of atmospheric particulate matter data is constructed by using the feature extraction capability of convolutional neural network and the time series data processing capability of recurrent neural network. The deep learning model is used to automatically learn the complex nonlinear relationship between the characteristics of atmospheric particulate matter and various pollution sources.
[0064] Introduce an integrated learning algorithm to form a prediction model, use Bayesian optimization to optimize the hyperparameters of the prediction model, process the characteristics of atmospheric particulate matter through sparse coding, extract representative key features, and enhance the robustness and performance of the model; apply the trained prediction model to real-time atmospheric monitoring data, and analyze the contribution ratio of various pollution sources to particulate matter concentration in real time; the prediction model dynamically adjusts parameters according to real-time data to ensure the high accuracy of the analysis results; the prediction model is automatically updated and adjusted by adding new data, and the integrated learning algorithm of the prediction model of the source of atmospheric particulate matter is updated in real time to ensure the high accuracy and timeliness of the analysis results, and provide continuous and reliable analysis results;
[0065] During the prediction model training process, cross-validation is used to evaluate the generalization ability of the prediction model to ensure the consistent performance of the prediction model on different preprocessed data sets, and the hyperparameters of the prediction model are continuously adjusted through Bayesian optimization, so that the prediction model can maintain efficient and stable performance in different environments and conditions.
[0066] This technical solution achieves accurate analysis of the sources of atmospheric particulate matter by integrating raw data with deep learning models. Convolutional neural networks can effectively extract particulate matter features, while recurrent neural networks can process complex time series data. This combination significantly improves the model's ability to identify the nonlinear relationship between particulate matter and pollution sources. In addition, the integrated learning algorithm enhances the robustness and predictive ability of the system through the fusion of multiple models, ensuring the high accuracy and stability of the analysis results.
[0067] The construction of the preprocessing data set specifically includes the following steps:
[0068] Data collection: Through the national air quality monitoring network, PM2.5 and PM10 concentration data were collected over the past three years, covering different areas such as cities, suburbs and rural areas; emission data of major pollution sources were obtained from the industrial sector and environmental protection departments, including particulate matter data emitted by factories, transportation, construction sites, etc.;
[0069] Historical meteorological data such as temperature, humidity, wind speed, wind direction, precipitation, etc. were obtained through meteorological stations; social and economic activity data such as population density, economic activity intensity, energy consumption, etc. were collected; traffic flow data were obtained, including road traffic volume, traffic congestion index, etc.
[0070] Data preprocessing: Clean the collected data, including processing missing values and outliers, and perform normalization to eliminate dimensional differences between different data sources; use spatial interpolation algorithms to interpolate missing spatial data to ensure the spatial continuity of the data; use time series analysis methods to restore missing time data to ensure the temporal integrity of the data; use principal component analysis (PCA) algorithms to reduce the dimensionality of high-dimensional data, extract main features, reduce data redundancy, and improve the computational efficiency of the model.
[0071] By constructing a multi-dimensional pre-processed data set covering different regions, seasons, and meteorological conditions, and combining it with a deep learning model of convolutional neural network (CNN) and recurrent neural network (RNN), the spatial distribution characteristics and temporal variation patterns of particulate matter can be effectively extracted. The model can automatically learn the complex nonlinear relationship between the characteristics of atmospheric particulate matter and various pollution sources, significantly improving the accuracy of atmospheric particulate matter source analysis.
[0072] The present invention applies the trained deep learning model to real-time atmospheric monitoring data to analyze in real time the contribution ratio of various pollution sources to the atmospheric particulate matter concentration. The model can dynamically adjust parameters according to real-time data to ensure the high accuracy of the analysis results, and automatically update and adjust the model through the continuous input of new data, and update the integrated learning algorithm of the prediction model of the atmospheric particulate matter source in real time. This dynamic adjustment and real-time analysis capability is particularly important in a rapidly changing environment.
[0073] The construction of a large-scale preprocessed dataset is not limited to atmospheric particulate matter concentration data and pollution source emission data, but also includes socioeconomic activity data and traffic flow data related to atmospheric particulate matter concentration change factors. The preprocessed dataset covers urban, suburban and rural areas in the spatial dimension, and daily data and extreme weather event data for multiple years in the temporal dimension.
[0074] In the process of constructing the preprocessed data set, the spatial interpolation algorithm is used to interpolate the missing spatial data, and the time series analysis method is used to restore the missing time data, thus forming a complete and continuous spatiotemporal data set, providing richer and more comprehensive input data for model training.
[0075] This technical solution expands the dimension of the preprocessed data set, enabling the prediction model to consider more factors that affect the concentration of atmospheric particulate matter, improving the accuracy of particulate matter source analysis, and completing missing data through spatial interpolation and time series analysis methods to ensure the integrity and continuity of the preprocessed data set, thereby improving the training effect and prediction ability of the model, enabling it to adapt to diverse environmental conditions and complex data distribution.
[0076] The ensemble learning algorithm uses random forest (RF) and gradient boosted decision tree (GBDT) as the ensemble learning algorithm of the prediction model. The multi-decision tree ensemble algorithm of random forest is used to reduce overfitting. The gradient boosted decision tree is used to improve the prediction ability of the prediction model through iterative optimization, which is used for the analysis of nonlinear features.
[0077] On this basis, the Bayesian optimization algorithm is used to automatically tune the hyperparameters of the prediction model, including the number of trees, depth, learning rate, etc., and finally generate the optimal prediction model to maintain efficient prediction performance under diverse input data conditions.
[0078] The combination of random forest and gradient boosting decision tree utilizes the advantages of both to perform nonlinear feature analysis, effectively reducing the overfitting phenomenon of the prediction model. The hyperparameters of the prediction model are tuned through Bayesian optimization to further improve the prediction performance of the prediction model, so that the final prediction model has stronger generalization ability and higher parsing accuracy under different conditions, especially when processing complex and nonlinear data.
[0079] The ensemble learning algorithm consists of the following steps:
[0080] Step 1: Collect raw data, including atmospheric particulate matter concentration data, pollution source emission data and historical meteorological data, to form a preprocessing database, and perform preprocessing operations such as cleaning, missing value processing and normalization on the raw data to eliminate dimensional differences between the data;
[0081] Step 2: Extract key features that affect the concentration of atmospheric particulate matter from the original data, such as extracting environmental variables: particulate matter concentration, wind speed, wind direction, temperature, and humidity. Use feature selection algorithms to screen out feature variables that affect the source analysis of atmospheric particulate matter, and calculate the proportion of each feature variable to obtain the proportion of each feature variable.
[0082] Step 3: Divide the processed original data into a training set and a test set. The training set corresponds to the training data, and the test set corresponds to the test data. The prediction model learns the rules through the training data and verifies its generalization ability on the test data. Use the random forest multi-decision tree ensemble algorithm to gradually build multiple decision trees in an iterative manner. Each tree is trained based on the residual of the previous tree and learns different rules in the training set. Avoid overfitting through cross-validation and hyperparameter adjustment, adjust the tree depth, learning rate and number of iterations to improve the generalization ability of the ensemble algorithm.
[0083] Step 4: Use the test data to verify the trained prediction model, evaluate the accuracy and stability of the prediction model, and evaluate the accuracy, mean square error (MSE), and R 2 Finally, the optimal prediction model is applied to the prediction and analysis of actual data;
[0084] Use the training data to evaluate the prediction performance of the prediction model, calculate the accuracy and mean square error indicators, judge the analytical ability of the prediction model, and further verify the robustness of the prediction model through cross-validation and hyperparameter adjustment to prevent overfitting; use the trained prediction model to predict the actual observation data, analyze the contribution ratio of various pollution sources to particulate matter concentration, and identify the main pollution sources as output results based on the analysis results.
[0085] The steps of the ensemble learning algorithm are described in detail, including data preprocessing, feature extraction, prediction model training, validation and application, emphasizing the role of random forests and gradient boosting decision trees in preventing overfitting and improving generalization ability.
[0086] Through a systematic integrated learning algorithm, the prediction model can be more stable and accurate when processing actual observation data. The data preprocessing step ensures the quality of the input data, the feature extraction step ensures that the prediction model focuses on key environmental variables, and cross-validation and hyperparameter adjustment effectively prevent overfitting. The final prediction model not only has high accuracy, but also has good robustness and stability.
[0087] The ensemble learning algorithm is used to reduce the limitations of a single model through the combination of random forest (RF) and gradient boosted decision tree (GBDT), and the hyperparameters of the model are automatically tuned using Bayesian optimization to further enhance the robustness of the model. The weighted fusion of the prediction model ensures the stability of the final analysis results, and can still provide stable and reliable prediction results even in complex and changing environments.
[0088] The integrated learning algorithm is equipped with an optimization module to regularly retrain and optimize the prediction model, compare the collected data with historical data to form a comparison chart, output the historical comparison data regularly at the set time, and automatically generate analysis reports based on actual needs to assist decision makers in formulating environmental governance strategies.
[0089] The optimization module can regularly output the comparison results of historical data and generate data reports according to the preset time period. By comparing the data changes in different time periods, it can analyze the changing trends of atmospheric particulate matter concentrations and sources, and assist environmental management departments in formulating more accurate pollution prevention and control strategies.
[0090] The introduction of the optimization module enables the prediction model to continuously adapt to new data and environmental changes, ensuring the long-term accuracy of the analysis results. By regularly generating comparison charts and analysis reports, the system provides decision makers with strong data support, facilitating the tracking of pollution trends and the formulation of targeted governance measures. This automated optimization process reduces the need for human intervention and improves the system's operating efficiency and application value.
[0091] The optimization module regularly retrains and optimizes the model, combines the comparative analysis of new data with historical data, and automatically generates comparative charts to help environmental management departments understand the changing trend of atmospheric particulate matter concentration in real time. The log transformation is used to visualize the data, amplify the smaller differences and reduce the impact of larger differences, so that the system can more keenly capture the slight changes in the data, improving the sensitivity and accuracy of data analysis.
[0092] The optimization module compares the new data with the data in the historical database, that is, compares the historical data with the new data, and generates comparison results for each week, month, quarter, and year. The comparison results are visualized through log transformation and a log graph is generated to show the magnitude and trend of data changes.
[0093] The log transformation can amplify smaller differences and reduce the impact of larger differences, thereby more intuitively displaying the changes in data at different time scales. It is particularly important when analyzing sudden changes in atmospheric particulate matter concentrations, ensuring sensitivity to data changes and analysis accuracy.
[0094] The introduction of log transformation enhances the sensitivity of data comparison and can more intuitively display the amplitude and trend of changes in atmospheric particulate matter concentration. Through comparative analysis and visual display, the system can quickly identify abnormal changes and long-term trends, providing more accurate data support for the formulation of pollution control and prevention measures.
[0095] The optimization module uses the contrast difference to calculate the difference between the contrast data. The formula is as follows:
[0096] Difference(t)=New Data(t)-Historical Data(t)
[0097] Where Difference(t) is the data difference at time point t; NewData(t) is the new data collected or observed at time point t; HistoricalData(t) is the historical data or baseline data before time point t;
[0098] Among them, the log transformation formula in the optimization module is:
[0099] Log Transformed Data(t)=log(1+|Difference(t)|)×sign(Difference(t))
[0100] Among them, Log Transformed Data(t) represents the result of logarithmic transformation of the original data difference at time point t;
[0101] log is the natural logarithm, the logarithm with base e, |Difference(t)| is the absolute value of the data difference at time point t, Difference(t) is the data difference at time point t, and sign(Difference(t)) is the sign function used to return the sign of the data difference.
[0102] Among them, the prediction model fusion formula in the optimization module is:
[0103]
[0104] Where Ensemble Output(t) is the ensemble output at time point t, ∑ represents the sum of the outputs of all models, i is the index of the model, from 1 to N, N is the total number of models included in the ensemble, and w i is the weight of the i-th model, which is a non-negative real number. i (t) is the output of the i-th model at time point t.
[0105] The difference between historical data and new data is calculated by contrast difference, and the weighted prediction model is used to fuse the analytical results of multiple prediction models;
[0106] The contrast difference formula is used to quantify the magnitude of data changes between different time points. Through the weighted prediction model, the output results of multiple prediction models are weighted and fused to make the final analysis result more stable and reliable.
[0107] The weights of the prediction model are dynamically adjusted through the Bayesian optimization algorithm to ensure that under different environmental conditions and data inputs, the prediction model can adaptively optimize the weight distribution to improve the accuracy of atmospheric particulate matter source analysis.
[0108] By comparing the differences between historical data and new data, and integrating the analysis results with a weighted prediction model, the comparative difference calculation can accurately quantify the changes in atmospheric particulate matter concentrations. Combined with the weighted prediction model, the system can still provide accurate and reliable analysis results when processing complex and changeable data. This solution further improves the adaptability of the prediction model in uncertain environments and ensures stability and accuracy during data analysis.
[0109] In the data preprocessing stage, that is, when integrating the original data, the principal component analysis (PCA) algorithm is used to reduce the dimensionality of high-dimensional data, reduce the redundancy of the data, and improve the calculation efficiency and prediction accuracy of the prediction model;
[0110] The calculation formula of the principal component analysis algorithm is:
[0111] Z=XW
[0112] Among them, X is the input data matrix; W is the feature vector matrix; Z is the data matrix after dimension reduction. The PCA algorithm is used to extract the feature dimension that best explains the data variance for subsequent prediction model training.
[0113] The PCA algorithm performs eigenvalue decomposition on the data matrix, selects the first few principal components with the largest contribution rate, retains the main information of the data, and reduces the dimension of the input variables, alleviating the computational burden of the prediction model, enabling it to converge faster and improve prediction accuracy.
[0114] Through the PCA algorithm, the system can effectively reduce the dimension of the data and reduce the amount of calculation while retaining the most representative features, thereby improving the training speed and prediction accuracy of the prediction model. This not only improves the overall performance of the prediction model, but also enables it to process large-scale data sets more quickly and adapt to the needs of real-time analysis.
[0115] The hyperparameters of the prediction model are optimized by using an improved Bayesian optimization algorithm. Gaussian process regression (GPR) is introduced as a proxy model in the algorithm to model the optimization process of each iteration.
[0116] The core formula of the Bayesian optimization algorithm is:
[0117] x t+1 = argmax x α(x|D 1:t )
[0118] Among them, α(x|D 1:t ) is the acquisition function, D 1:t For the data set after iterating t times, find the next optimal hyperparameter by maximizing the acquisition function.
[0119] An improved Bayesian optimization algorithm is used to optimize the hyperparameters of the prediction model. Gaussian process regression is introduced as a proxy model in the algorithm to model the optimization process of each iteration.
[0120] Through Bayesian optimization, the hyperparameters of the prediction model, including the learning rate, the depth and number of decision trees, are dynamically adjusted, so that the prediction model can learn with the optimal parameter combination during the training process, thereby improving the prediction accuracy and computational efficiency of the prediction model, especially showing superior optimization effects in high-dimensional data.
[0121] The combination of the improved Bayesian optimization algorithm and the GPR model makes the hyperparameter optimization process more efficient and accurate. Gaussian process regression provides a more accurate prediction model for the optimization process and can find the optimal parameter combination more quickly, thereby improving the overall performance and prediction accuracy of the prediction model, especially showing stronger analytical capabilities in complex data environments.
[0122] During the training and optimization process of the model, the Adam optimization algorithm in the adaptive learning rate optimization algorithm is used to update parameters. The adaptive learning rate optimization algorithm dynamically adjusts the learning step size according to the change of the gradient to improve the convergence speed and generalization ability of the model on complex data sets.
[0123] The core calculation formula of the Adam optimization algorithm is:
[0124]
[0125] Among them, θ t is the parameter of the model at time t, α is the learning rate, is an unbiased estimate of the first-order momentum, is an unbiased estimate of the second-order momentum, To prevent division by zero,
[0126] In practical applications, the system continuously adjusts the update step size of parameters through an adaptive learning rate optimization algorithm, so that the model can adapt to different data distributions and improve the parsing accuracy. Especially when faced with large-scale and diversified data, the model can converge quickly and maintain stable prediction performance.
[0127] The Adam optimization algorithm combines the advantages of the momentum and RMSprop algorithms. When updating parameters, it considers both the size of the current gradient and the direction of the historical gradient, thereby achieving more stable and faster model convergence.
[0128] Through the adaptive learning rate optimization algorithm, the model can find the optimal solution more quickly in a complex and dynamic data environment and avoid falling into the local optimum, thereby improving the accuracy of atmospheric particulate matter source analysis and the breadth of model application.
[0129] The introduction of the adaptive learning rate optimization algorithm Adam enables the model to converge faster in complex and dynamic environments and improves the generalization ability of the model. By dynamically adjusting the learning rate, the Adam optimization algorithm can achieve robust parameter updates under different data distributions, avoiding the model from falling into local optimality, and ultimately improving the accuracy of analysis and the breadth of application of the model.
[0130] By accurately analyzing the sources of atmospheric particulate matter, it helps environmental management departments identify major pollution sources and formulate targeted governance measures, reducing the harm of air pollution to public health. The high accuracy and real-time nature of the analysis results provide strong support for scientific decision-making and significantly improve the effectiveness of air quality management.
[0131] The present invention is applicable to various complex environmental conditions and pollution sources, and can be widely used in urban air pollution monitoring, industrial emission control, traffic pollution control and other fields, providing strong technical support for governments at all levels and environmental protection agencies, and helping to achieve more scientific and accurate environmental management and pollution prevention; at the same time, the present invention has shown significant advantages in solving problems such as the complexity of the sources of atmospheric particulate matter, real-time analysis and model robustness, and provides important technical means and support for air pollution control and environmental protection.
[0132] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without creative work should fall within the scope of protection of the present invention. The structures, devices and operating methods not specifically described and explained in the present invention are implemented according to the conventional means in the field unless otherwise specified and limited.
Claims
1. A method for analyzing the source of atmospheric particulate matter based on deep learning, characterized in that: include: Construct a large-scale preprocessed data set, including atmospheric particulate matter concentration data from different regions, seasons and meteorological conditions, and collect the corresponding raw data, integrate the raw data to form a multi-dimensional preprocessed data set; A deep learning model is established using convolutional neural networks and recurrent neural networks. The feature extraction capability of convolutional neural networks and the time series data processing capability of recurrent neural networks are used to build a deep learning model for atmospheric particulate matter data. The deep learning model is used to automatically learn the complex nonlinear relationship between atmospheric particulate matter characteristics and various pollution sources. Introduce an integrated learning algorithm to form a prediction model, use Bayesian optimization to optimize the hyperparameters of the prediction model, process the characteristics of atmospheric particulate matter through sparse coding, and extract representative key features; apply the trained prediction model to real-time atmospheric monitoring data, and analyze the contribution ratio of various pollution sources to particulate matter concentration in real time; the prediction model dynamically adjusts parameters according to real-time data; the prediction model is automatically updated and adjusted by adding new data, and the integrated learning algorithm of the prediction model of atmospheric particulate matter sources is updated in real time; During the prediction model training process, cross-validation is used to evaluate the generalization ability of the prediction model, and the hyperparameters of the prediction model are continuously adjusted through Bayesian optimization; The integrated learning algorithm is provided with an optimization module to regularly retrain and optimize the prediction model, form a comparison chart based on the collected data and historical data, regularly output the historical comparison data at a set time, and automatically generate an analysis report based on actual needs to assist decision makers in formulating environmental governance strategies; The optimization module uses the contrast difference to calculate the difference between the contrast data, and the formula is as follows: Difference(t)=New Data(t)-Historical Data(t); Where Difference(t) is the data difference at time point t; New Data(t) is the new data collected or observed at time point t; Historical Data(t) is the historical data or baseline data before time point t.
2. The method for analyzing the source of atmospheric particulate matter based on deep learning according to claim 1, characterized in that: The construction of the large-scale preprocessed data set is not limited to atmospheric particulate matter concentration data and pollution source emission data, but also includes socioeconomic activity data and traffic flow data related to atmospheric particulate matter concentration change factors. The preprocessed data set covers urban, suburban and rural areas in the spatial dimension, and covers daily data and extreme weather event data for multiple years in the temporal dimension. In the process of constructing the preprocessed data set, the spatial interpolation algorithm is used to interpolate the missing spatial data, and the time series analysis method is used to restore the missing time data, thus forming a complete and continuous spatiotemporal data set.
3. The method for analyzing the source of atmospheric particulate matter based on deep learning according to claim 2, characterized in that: The ensemble learning algorithm adopts random forest (RF) and gradient boosted decision tree (GBDT) as the ensemble learning algorithm of the prediction model, uses the multi-decision tree ensemble algorithm of random forest to reduce the overfitting phenomenon, and the gradient boosted decision tree improves the prediction ability of the prediction model through iterative optimization, which is used for the analysis of nonlinear features; On this basis, the Bayesian optimization algorithm is used to automatically tune the hyperparameters of the prediction model, including the number, depth, and learning rate of trees, and finally generate the optimal prediction model.
4. The method for analyzing the source of atmospheric particulate matter based on deep learning according to claim 3 is characterized by: The ensemble learning algorithm comprises the following steps: Data preprocessing: Collect raw data, including atmospheric particulate matter concentration data, pollution source emission data and historical meteorological data, to form a preprocessing database, and perform preprocessing operations such as cleaning, missing value processing and normalization on the raw data; Feature extraction: Extract key features from the original data, use feature selection algorithm to screen out characteristic variables that affect the source analysis of atmospheric particulate matter, and calculate the proportion of each characteristic variable to obtain the proportion of each characteristic variable; Prediction model training: The processed raw data is divided into a training set and a test set. The training set corresponds to the training data, and the test set corresponds to the test data. The prediction model learns the rules through the training data and verifies its generalization ability on the test data. Use the random forest multi-decision tree ensemble algorithm to gradually build multiple decision trees in an iterative manner. Each tree is trained based on the residual of the previous tree and learns different rules in the training set. Cross-validation and hyperparameter adjustment are used to avoid overfitting, and the depth, learning rate, and number of iterations of the tree are adjusted. Prediction model verification: Use test data to verify the trained prediction model, evaluate the accuracy and stability of the prediction model, and finally apply the optimal prediction model to the prediction and analysis of actual data; Use training data to evaluate the prediction performance of the prediction model, calculate the accuracy and mean square error indicators, judge the analytical ability of the prediction model, and further verify the robustness of the prediction model through cross-validation and hyperparameter adjustment; The trained prediction model is used to predict the actual observed data, analyze the contribution ratio of various pollution sources to the particulate matter concentration, and identify the main pollution sources as output results based on the analysis results.
5. The method for analyzing the source of atmospheric particulate matter based on deep learning according to claim 4, characterized in that: The optimization module compares the new data with the data in the historical database, that is, compares the historical data with the new data, generates comparison results for each week, each month, each quarter, and each year, and visualizes the comparison results through log transformation to generate a log graph.
6. The method for analyzing the source of atmospheric particulate matter based on deep learning according to claim 5, characterized in that: The log transformation formula in the optimization module is: Log Transformed Data(t)=log(1+|Difference(t)|)×sign(Difference(t)) Among them, Log Transformed Data(t) represents the result of logarithmic transformation of the original data difference at time point t; log is the natural logarithm with base e as the logarithm, |Difference(t)| is the absolute value of the data difference at time point t, Difference(t) is the data difference at time point t, sign(Difference(t)) is a sign function, which is used to return the sign of the data difference; Among them, the prediction model fusion formula in the optimization module is: Where Ensemble Output(t) is the ensemble output at time point t, ∑ represents the sum of the outputs of all models, i is the index of the model, ranging from 1 to N, N is the total number of models included in the ensemble, and w i is the weight of the i-th model, which is a non-negative real number. i (t) is the output of the i-th model at time point t; The difference between historical data and new data is calculated by contrast difference, and the weighted prediction model is used to fuse the analytical results of multiple prediction models; The contrast difference formula is used to quantify the magnitude of data changes between different time points. The output results of multiple prediction models are weighted and fused through the weighted prediction model. The weights of the prediction model are dynamically adjusted through a Bayesian optimization algorithm to ensure that under different environmental conditions and data inputs, the prediction model can adaptively optimize the weight distribution to improve the accuracy of atmospheric particulate matter source analysis.
7. The method for analyzing the source of atmospheric particulate matter based on deep learning according to claim 6, characterized in that: In the data preprocessing stage, that is, when integrating the original data, the principal component analysis (PCA) algorithm is used to reduce the dimensionality of high-dimensional data to reduce the redundancy of the data; The calculation formula of the principal component analysis algorithm is: Z=XW Among them, X is the input data matrix; W is the feature vector matrix; Z is the data matrix after dimensionality reduction.
8. The method for analyzing the source of atmospheric particulate matter based on deep learning according to claim 7, characterized in that: The hyperparameters of the prediction model are optimized by using an improved Bayesian optimization algorithm. Gaussian process regression (GPR) is introduced as a proxy model in the algorithm to model the optimization process of each iteration. The core formula of the Bayesian optimization algorithm is: x t+1 =argmax x α(x|D 1:t ) Among them, α(x|D 1:t ) is the acquisition function, D 1:t is the data set after iteration t times.
9. The method for analyzing the source of atmospheric particulate matter based on deep learning according to claim 8, characterized in that: During the training and optimization process of the model, the Adam optimization algorithm in the adaptive learning rate optimization algorithm is used to update the parameters. The adaptive learning rate optimization algorithm dynamically adjusts the learning step size according to the change of the gradient. The core calculation formula of the Adam optimization algorithm is: Among them, θ t+1 represents the parameters of the model at time t+1, θ t is the parameter of the model at time t, α is the learning rate, is an unbiased estimate of the first-order momentum, is an unbiased estimate of the second-order momentum, and ò is a minimum value that prevents division by zero.
Citation Information
Patent Citations
Atmospheric pollutant tracing method based on coupled machine learning and correlation analysis
CN113610243A
Evolutionary ensemble learning-based atmospheric pollutant concentration space-time prediction method
CN114578457A