A method and device for dynamically estimating and quantifying uncertainty of volatile organic compound emissions from a gas station
By constructing a multi-model parallel prediction and integrated analysis method, and combining multi-source environmental data, the uncertainty problem in the estimation of volatile organic compound emissions from gas stations was solved, achieving accurate emission estimation and reliable confidence interval derivation, thus improving the prediction accuracy and robustness of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA UNIV OF MINING & TECH (BEIJING)
- Filing Date
- 2026-01-09
- Publication Date
- 2026-06-05
Smart Images

Figure CN122157867A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of environmental science and air pollution control technology, specifically to a method and apparatus for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions from gas stations. Background Technology
[0002] Gas stations are a crucial link in the storage, transportation, and sales of petroleum products. The volatile organic compounds (VOCs) generated during their operations are significant precursors to urban ozone and secondary organic aerosols, posing a significant threat to regional air quality. Accurately estimating their emissions and quantifying uncertainties is fundamental to formulating scientific emission reduction policies and achieving refined management.
[0003] Currently, the mainstream methods for estimating VOCs emissions from gas stations mainly include the emission factor method. The emission factor method calculates emissions using "activity level × emission factor," which is simple to calculate but relies on fixed empirical coefficients, making it difficult to reflect spatiotemporal variations such as meteorological conditions, equipment operating status, and human activities, leading to significant biases in the estimation results. Traditional emission factor methods generally suffer from low spatiotemporal resolution and weak dynamic characterization capabilities. They often use regional or annual averages, failing to capture daily or monthly dynamic changes and spatial heterogeneity. Furthermore, these methods typically treat emissions as a deterministic process, lacking a systematic quantification of input data errors and model structural uncertainties, making it difficult to provide reliable confidence intervals and limiting their application in risk assessment and decision support. Summary of the Invention
[0004] This invention provides a method and apparatus for dynamic estimation and uncertainty quantification of volatile organic compound emissions from gas stations, in order to solve the problem that existing technologies lack systematic quantification of input data errors and model structure uncertainties, making it difficult to provide reliable confidence intervals and limiting their application in risk assessment and decision support.
[0005] In a first aspect, the present invention provides a method for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions from gas stations, the method comprising: Multiple datasets were constructed based on volatile organic compound (VOC) emission data from multiple gas stations and corresponding multi-source environmental data. Each dataset is input into a corresponding preset initial volatile organic compound emission prediction model for training, generating multiple target volatile organic compound emission prediction models. The predicted value of volatile organic compound emissions for each gas station is calculated using the predicted volatile organic compound emission model for each of the aforementioned targets. Based on all predicted values of volatile organic compound (VOC) emissions for each of the gas stations, calculate the total VOC emissions and uncertainty range for all the gas stations.
[0006] This invention systematically quantifies input data errors and model structure uncertainties through parallel prediction and ensemble analysis using multiple models. Each target model outputs predicted emissions for a single station from different dimensions. By combining the statistical characteristics of multiple sets of prediction results, the total emissions and reliable confidence intervals are accurately derived, thus overcoming the shortcoming of existing technologies in quantifying uncertainties.
[0007] In one optional implementation, the step of constructing multiple datasets based on volatile organic compound (VOC) emission data from multiple gas stations and corresponding multi-source environmental data includes: Acquire volatile organic compound emission data from multiple gas stations, as well as meteorological data, air quality data, and energy and industrial product output data for each gas station; Based on the locations of each gas station, a first rectangular buffer zone is constructed, and the average value of meteorological data and the average value of air quality data within the first rectangular buffer zone are calculated using the spatial averaging method. Based on the spatial discretization method, the energy and industrial product output data are allocated to multiple grid cells, and weights are assigned according to the resident population density in each grid cell to construct a second rectangular buffer area. The average output of industrial products within the second rectangular buffer area is calculated using the spatial averaging method. Logarithmic transformation is performed on the volatile organic compound (VOC) emission data of the gas station to obtain new VOC emission data for the gas station; The average values of the meteorological data, the average values of the air quality data, the average values of the industrial product output, and the new volatile organic compound emission data of the gas station are standardized. The first input variable combination is constructed using the average value of standardized meteorological data and the average value of air quality data; The average value of the standardized meteorological data and the average value of industrial product output are used to construct a second combination of input variables; Based on the same time and space, the first input variable combination and the second input variable combination are respectively correlated with the standardized volatile organic compound emission data of gas stations to obtain the first dataset and the second dataset.
[0008] This invention achieves effective fusion of multi-source heterogeneous data through precise spatial matching and standardized data preprocessing. The construction of rectangular buffer regions and the application of spatial averaging ensure accurate correspondence between meteorological, air quality, and industrial-related data and gas station locations; logarithmic transformation and standardization eliminate data skewness and dimensional differences, improving data quality. The construction of two types of input variable combinations comprehensively covers key influencing dimensions.
[0009] In an optional implementation, before the step of inputting each of the datasets into a corresponding preset initial volatile organic compound (VOC) emission prediction model for training to generate multiple target VOC emission prediction models, the method further includes: The first dataset and the second dataset are sampled to generate multiple sample subsets of the first dataset and multiple sample subsets of the second dataset; Based on each of the aforementioned sample subsets, multiple initial decision trees are constructed respectively; The mean squared error loss function of the initial decision tree is replaced with the quantile regression loss function to obtain the updated decision tree; The test samples of the sample subset are aggregated into the training sample values that fall into the leaf nodes of the corresponding updated decision tree to construct the conditional empirical distribution function of the updated decision tree. The predicted value of any quantile is extracted from the conditional empirical distribution function, and the quantile of the quantile is calculated to obtain the prediction interval of the gas station volatile organic compound emissions of the updated decision tree. An initial volatile organic compound (VOC) emission prediction model for the first dataset is constructed using the updated decision tree, conditional empirical distribution function, and prediction interval of VOC emissions from gas stations corresponding to the first dataset. An initial volatile organic compound (VOC) emission prediction model for the second dataset is constructed using the updated decision tree, conditional empirical distribution function, and prediction interval of VOC emissions from gas stations corresponding to the second dataset.
[0010] This invention constructs a decision tree by generating multiple sample subsets through sampling, and replaces the traditional mean squared error loss function with a quantile regression loss function to accurately capture the nonlinear relationship between emissions and input variables. By leveraging a conditional empirical distribution function to extract quantile prediction values and calculate prediction intervals, the model possesses probabilistic prediction capabilities, quantifying the uncertainty of emissions estimation. Two types of datasets are modeled separately, comprehensively covering different influencing dimensions, thus improving the model's prediction accuracy and robustness.
[0011] In one optional implementation, the step of inputting each of the datasets into a corresponding preset initial volatile organic compound (VOC) emission prediction model for training to generate multiple target VOC emission prediction models includes: Each dataset is divided into a training set and a validation set according to a preset ratio; Each training set is input into the corresponding preset initial volatile organic compound emission prediction model for training, resulting in multiple updated volatile organic compound prediction models. Each of the aforementioned validation sets is input into the corresponding updated volatile organic compound (VOC) emission prediction model for validation, resulting in multiple target VOC emission prediction models.
[0012] This invention separates model training and validation by dividing the training and validation sets into a predetermined ratio. The training set provides sufficient data to support the model's learning of the relationship between input variables and emissions, while the validation set effectively tests the model's generalization ability and promptly identifies problems such as overfitting. Through iterative optimization of training and validation, the final target model exhibits significantly improved prediction accuracy and stability.
[0013] In one optional implementation, the step of calculating the predicted value of the volatile organic compound (VOC) emissions of each gas station using the respective target VOC emission prediction models includes: Based on the spatial location information of each gas station, the meteorological indicators and industrial emission proxy indicators of each gas station are spatially matched to generate the input variable values of each gas station. The input variable values of the gas station are respectively input into each of the target volatile organic compound emission prediction models; The predicted value of the gas station's volatile organic compound (VOC) emissions is calculated using the respective target VOC emission prediction models.
[0014] This invention ensures the spatiotemporal consistency and accuracy of input variable values by precisely matching the spatial location of gas stations with multi-source environmental data. The input variable values are then substituted into an optimized target model, leveraging the model's ability to capture complex relationships to efficiently output predicted volatile organic compound (VOC) emissions for a single gas station.
[0015] In one optional implementation, calculating the total volatile organic compound (VOC) emissions and uncertainty range of all gas stations based on the predicted VOC emissions of each of the gas stations includes: The predicted values of the preset quantiles of each gas station are extracted by the target volatile organic compound emission prediction model to obtain the distribution characteristics of the volatile organic compound emissions of each gas station. The distribution characteristics of volatile organic compound emissions from each of the gas stations were sampled using the Monte Carlo sampling method, generating multiple predicted emission values for each gas station. Calculate the mean of all predicted emissions for each gas station and the confidence interval for a preset percentage; Based on the average volatile organic compound (VOC) emissions of all the gas stations and a preset confidence interval, the total VOC emissions and uncertainty interval of all the gas stations are determined.
[0016] This invention clarifies the emission distribution characteristics of individual stations by extracting quantile prediction values, and generates multiple sets of prediction samples using Monte Carlo sampling to systematically quantify the uncertainty in emission estimation. By summarizing the mean and confidence interval of individual stations, the total emission amount and uncertainty interval are accurately calculated.
[0017] In one optional implementation, the method further includes: The input variable values of each target volatile organic compound emission prediction model are analyzed using Shapley's interpretation values to obtain the influence values of each variable characteristic of the input variable values on the corresponding volatile organic compound emissions of the gas station. Based on the mean absolute value of the Shapley explanatory values, the characteristics of each variable are ranked, and the key factors affecting the volatile organic compound emissions of the gas station are determined according to the ranking results.
[0018] This invention uses Shapley's explanatory value to accurately quantify the impact of each variable's characteristics on emissions, and efficiently identifies key influencing factors by sorting by the absolute value mean.
[0019] Secondly, the present invention provides a device for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions from gas stations, the device comprising: The module is used to build multiple datasets based on the volatile organic compound emission data of multiple gas stations and the corresponding multi-source environmental data; The training module is used to input each of the datasets into the corresponding preset initial volatile organic compound emission prediction model for training, and generate multiple target volatile organic compound emission prediction models. The prediction module is used to calculate the predicted value of the volatile organic compound emissions of each gas station using the target volatile organic compound emission prediction model. The calculation module is used to calculate the total volatile organic compound (VOC) emissions and uncertainty range of all the gas stations based on all predicted values of VOC emissions of each of the gas stations.
[0020] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the dynamic estimation and uncertainty quantification method for volatile organic compound emissions at gas stations as described in the first aspect or any corresponding embodiment.
[0021] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the dynamic estimation and uncertainty quantification method for volatile organic compound emissions at gas stations as described in the first aspect or any corresponding embodiment. Attached Figure Description
[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the first method for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions from gas stations according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the second process for the dynamic estimation and uncertainty quantification method of volatile organic compound emissions from gas stations according to an embodiment of the present invention; Figure 3 This is a structural block diagram of a device for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions at gas stations according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0026] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0027] This invention provides a method for dynamically estimating and quantifying the uncertainty of volatile organic compound (VOC) emissions from gas stations. Through parallel prediction and integrated analysis using multiple models, it systematically quantifies input data errors and model structural uncertainties. Each target model outputs predicted emissions for a single station from different dimensions. By combining the statistical characteristics of multiple prediction results, the total emissions and reliable confidence intervals are accurately derived, thus overcoming the limitation of existing technologies in quantifying uncertainty.
[0028] According to an embodiment of the present invention, a method for dynamic estimation and uncertainty quantification of volatile organic compound emissions from gas stations is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0029] This embodiment provides a method for dynamically estimating and quantifying the uncertainty of volatile organic compound (VOC) emissions from gas stations. Figure 1 This is a flowchart of a method for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions from gas stations according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Construct multiple datasets based on the volatile organic compound emission data of multiple gas stations and the corresponding multi-source environmental data.
[0030] It should be noted that the volatile organic compound (VOC) emission data of gas stations refers to the monthly volatile organic compound (VOC) emission records of gas stations with the industry category of "retail of motor vehicle fuel" in the pollutant discharge permit management information platform over many years.
[0031] Multi-source environmental data refers to a collection of heterogeneous data related to VOCs emissions from gas stations, including meteorological data, air quality data, and energy and industrial product output data.
[0032] A dataset refers to a dataset formed by establishing a spatiotemporal correspondence between the input variables of "meteorological parameters + air quality parameters" and standardized gas station VOCs emission data, or a dataset formed by establishing a spatiotemporal correspondence between the input variables of "meteorological parameters + industrial emission proxy indicators" and standardized gas station VOCs emission data.
[0033] In this embodiment of the invention, the monthly volatile organic compound (VOC) emission data of multiple gas stations over many years, as well as the meteorological data, air quality data, energy and industrial product output data corresponding to each gas station, are first logarithmically transformed to eliminate skewed distribution characteristics. Then, multi-source environmental data are accurately associated with gas station locations through spatial matching. After standardizing all data, spatiotemporal correspondences are established with the standardized emission data according to the combination logic of meteorological parameters + air quality parameters and meteorological parameters + industrial emission proxy indicators, respectively, to construct the first dataset and the second dataset.
[0034] Step S102: Input each dataset into the corresponding preset initial volatile organic compound emission prediction model for training, and generate multiple target volatile organic compound emission prediction models.
[0035] It should be noted that the preset initial volatile organic compound emission prediction model refers to the basic model framework based on quantile random forest in this invention, which has the basic model structure of sample sampling, decision tree construction, and quantile loss calculation.
[0036] A target volatile organic compound (VOC) emission prediction model refers to an optimal, mature model that, after training on the training set of the corresponding dataset and performing performance verification and parameter optimization on the validation set, possesses accurate VOC emission prediction capabilities and can output emission distribution characteristics.
[0037] In this embodiment of the invention, each dataset is divided into a corresponding training set and a validation set according to a preset ratio. Each training set is input into a corresponding preset initial volatile organic compound emission prediction model for iterative training. Then, each validation set is substituted into the trained model for generalization ability verification and model performance optimization. Through closed-loop iteration of training and validation, the parameters of the model are tuned and the accuracy is calibrated. Finally, multiple target volatile organic compound emission prediction models that are adapted to the corresponding dataset features and have stable prediction performance are generated.
[0038] Step S103: Calculate the predicted value of volatile organic compound emissions for each gas station using the volatile organic compound emission prediction model for each target.
[0039] It should be noted that the predicted value of volatile organic compound (VOC) emissions from gas stations refers to the estimated value of VOC emissions at each gas station's corresponding quantile, calculated and output by the target VOC emission prediction model.
[0040] In this embodiment of the invention, precise spatial matching of multi-source environmental data is completed according to the spatial location information of each gas station, generating standardized input variable values corresponding to each gas station. Each input variable value is then input into each target volatile organic compound emission prediction model. Relying on the model's ability to learn the nonlinear mapping relationship between input variables and emissions, after processing by the conditional empirical distribution function within the model, the predicted value of volatile organic compound emissions at each gas station at a preset quantile is calculated and output by each target volatile organic compound emission prediction model.
[0041] Step S104: Based on all predicted values of volatile organic compound (VOC) emissions from each gas station, calculate the total VOC emissions and uncertainty range for all gas stations.
[0042] It should be noted that the total volatile organic compound (VOC) emissions from gas stations refer to the overall calculated value of VOC emissions from all gas stations within a given time period, obtained by summing up the predicted average VOC emissions from all gas stations.
[0043] The uncertainty interval refers to the range of values within which the total volatile organic compound emissions of all gas stations are likely to fall under a preset confidence level, after the Monte Carlo sampling method is used to quantify the estimation error based on the distribution characteristics of emissions from each gas station.
[0044] In this embodiment of the invention, the predicted values of volatile organic compound (VOC) emissions from each gas station are extracted from preset quantiles to clarify the distribution characteristics of emissions. The Monte Carlo sampling method is used to repeatedly sample the distribution characteristics to generate multiple sets of predicted emission values for each gas station. The mean of the predicted values for a single station and the preset proportional confidence interval are calculated. Then, the prediction results of all gas stations are summarized to complete the total calculation. Furthermore, the total VOC emissions of all gas stations and the corresponding emission uncertainty interval are derived.
[0045] This embodiment provides a method for dynamically estimating and quantifying the uncertainty of volatile organic compound (VOC) emissions from gas stations. Figure 2 This is a flowchart of a method for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions from gas stations according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Construct multiple datasets based on the volatile organic compound emission data of multiple gas stations and the corresponding multi-source environmental data.
[0046] Specifically, step S201 includes: Step S2011: Obtain volatile organic compound emission data from multiple gas stations, as well as corresponding meteorological data, air quality data, and energy and industrial product output data for each gas station.
[0047] It should be noted that meteorological data refers to environmental data related to the meteorological conditions of the area where the gas station is located.
[0048] Air quality data refers to atmospheric environmental quality monitoring data for the area where the gas station is located.
[0049] Energy and industrial product output data refers to production data of energy and industrial products related to volatile organic compound emissions from gas stations.
[0050] In this embodiment of the invention, monthly volatile organic compound (VOC) emission data from hundreds of gas stations over several years within a certain range are collected. The data comes from gas stations whose industry category is "retail of motor vehicle fuel" in the pollutant discharge permit management information platform. The meteorological data comes from the real-time product dataset CLDAS-V2.0 of the meteorological bureau's land surface data assimilation system. The air quality data comes from the observation values of thousands of monitoring stations. The energy and industrial product output data come from the provincial monthly statistical data of the statistics bureau.
[0051] Specifically, meteorological data includes commonly used indicators such as surface temperature, rainfall, surface air pressure, specific humidity, soil relative humidity, air temperature, and wind speed; air quality data includes PM2.5 and PM2.5. 10 Conventional parameters such as SO2, NO2, CO, and O3; energy and industrial product output data including indicators such as gasoline, diesel, naphtha, liquefied petroleum gas, petroleum asphalt, power generation equipment, automobiles, steel, aluminum, cement, primary plastics, and flat glass.
[0052] Step S2012: Based on the location of each gas station, a first rectangular buffer area is constructed, and the average value of meteorological data and the average value of air quality data within the first rectangular buffer area are calculated using the spatial averaging method.
[0053] It should be noted that the location of a gas station refers to the latitude and longitude coordinates that represent the actual geographical location of each gas station, and is the core benchmark point for constructing the buffer zone.
[0054] The first rectangular buffer area refers to the rectangular space area constructed with the gas station location as the center and a preset first radius.
[0055] The spatial averaging method refers to a method of summarizing and calculating the average value of spatially distributed data of the same type within a defined spatial area.
[0056] The average value of meteorological data refers to the average value of various types of meteorological indicators within the first rectangular buffer area, calculated using the spatial averaging method.
[0057] The average value of air quality data refers to the average value of each type of air quality index within the second rectangular buffer area, calculated using the spatial averaging method.
[0058] In this embodiment of the invention, for all gas station locations, a rectangular spatial range is constructed with the gas station location as the center and a preset first radius. The average value of meteorological elements in this area is calculated using a spatial averaging method, which is used as the meteorological input index for the gas station. The average value of air quality parameters in this area is also calculated, which is used as the air quality input index for the gas station.
[0059] Step S2013: Based on the spatial discretization method, energy and industrial product output data are allocated to multiple grid cells, and weights are assigned according to the resident population density in each grid cell to construct a second rectangular buffer area.
[0060] It should be noted that spatial discretization refers to the method of dividing a continuous geographical space or a large administrative region into multiple regular grid units of equal area, thereby realizing the decomposition and allocation of macro-scale data into micro-scale spatial units.
[0061] A grid cell refers to a regular rectangular spatial unit formed by spatial discretization.
[0062] Resident population density refers to the number of resident people per unit area within each grid cell.
[0063] The second rectangular buffer zone refers to a rectangular spatial area centered on the gas station location, adapted to the radiation range of industrial activities.
[0064] In this embodiment of the invention, a second rectangular buffer area is constructed by using a spatial discretization method to allocate the monthly output of energy and major industrial products summarized by provincial administrative units to grid units. The allocation weight is based on the resident population density within the grid.
[0065] Step S2014: The spatial averaging method is used to calculate the average output of industrial products within the second rectangular buffer area.
[0066] It should be noted that the average output of industrial products refers to the weighted average value of industrial product output within the second rectangular buffer area, calculated using the spatial averaging method.
[0067] In this embodiment of the invention, the average output of industrial products within the buffer zone is calculated using a spatial averaging method.
[0068] Step S2015: Perform a logarithmic transformation on the volatile organic compound (VOC) emission data of the gas station to obtain new VOC emission data for the gas station.
[0069] It should be noted that logarithmic transformation refers to a commonly used data preprocessing method.
[0070] The new data on volatile organic compound (VOC) emissions from gas stations refers to the VOC emissions data from gas stations after logarithmic transformation.
[0071] In this embodiment of the invention, a logarithmic transformation is performed on the volatile organic compound (VOC) emission data of gas stations to eliminate the skewed distribution characteristics of the data, and finally a new VOC emission data of gas stations with a more uniform distribution and stronger stability is obtained.
[0072] Step S2016 involves standardizing the average values of meteorological data, air quality data, industrial product output, and new volatile organic compound emission data from gas stations.
[0073] In this embodiment of the invention, all data are standardized, including the average values of meteorological data, air quality data, industrial product output, and new volatile organic compound emissions data from gas stations.
[0074] Step S2017: Use the average value of the standardized meteorological data and the average value of the air quality data to construct the first input variable combination.
[0075] It should be noted that the first input variable combination refers to the feature combination formed by integrating the average value of the standardized meteorological data and the average value of the air quality data.
[0076] In this embodiment of the invention, two types of input variable combinations are designed. The first type is a combination of meteorological parameters and air quality parameters (i.e., the first input variable combination).
[0077] In step S2018, the average value of the standardized meteorological data and the average value of industrial product output are used to construct the second input variable combination.
[0078] It should be noted that the second input variable combination refers to the feature combination formed by integrating the average value of standardized meteorological data with the average value of industrial product output.
[0079] In this embodiment of the invention, the second type is a combination of meteorological parameters and industrial emission proxy indicators (i.e., the first input variable combination).
[0080] Step S2019: Based on the same time and space, establish a correspondence between the first input variable combination and the second input variable combination and the standardized volatile organic compound emission data of gas stations, respectively, to obtain the first dataset and the second dataset.
[0081] It should be noted that the correspondence refers to the one-to-one matching relationship between the combination of input variables (first or second) and the emission data under the same spatiotemporal conditions.
[0082] The first dataset refers to the sample set constructed by combining the first input variables with standardized emission data according to the same spatiotemporal correspondence.
[0083] The second dataset refers to the sample set constructed by combining the second input variables with the standardized emission data according to the same spatiotemporal correspondence.
[0084] In this embodiment of the invention, two types of input variable combinations are respectively correlated with volatile organic compound (VOC) emissions according to time and space to obtain a first dataset and a second dataset. The first type of input variable combination includes multiple meteorological indicators and multiple air quality parameters, which are used to analyze the comprehensive impact of meteorological driving and atmospheric chemical processes on VOC concentration changes. The second type of input variable combination includes multiple meteorological indicators and multiple energy and industrial major product output indicators, which are used to construct a proxy relationship chain of "production-emission".
[0085] In some optional implementations, prior to step S202, the method further includes: Step a1: Sample the first dataset and the second dataset to generate multiple sample subsets of the first dataset and multiple sample subsets of the second dataset.
[0086] It should be noted that a sample subset refers to a small set of data extracted from the original dataset (first or second dataset) through sampling. Each subset contains multiple sample pairs of input variables and corresponding emission data.
[0087] In this embodiment of the invention, sampling with replacement is performed on both the first dataset and the second dataset to generate multiple sample subsets of the first dataset and multiple sample subsets of the second dataset. Specifically, the multiple sample subsets of the first dataset are divided into a training set and a test set according to a preset ratio, and similarly, the multiple sample subsets of the second dataset are divided into a training set and a test set according to a preset ratio.
[0088] Step a2: Construct multiple initial decision trees based on each sample subset.
[0089] It should be noted that the initial decision tree refers to the basic prediction model built based on a single subset of samples, and is the core component of the quantile random forest ensemble model.
[0090] In this embodiment of the invention, a quantile random forest method is used to construct a volatile organic compound (VOC) emission prediction model. The core component of the quantile random forest ensemble model is the decision tree. Therefore, multiple decision trees are constructed for each sample subset. The basic structure of the quantile random forest is as follows:
[0091] In the formula, Y represents the emission of volatile organic compounds; X represents the input variable; Indicates the first i A decision tree; Indicates hyperparameters; This indicates the error term.
[0092] It is worth mentioning that volatile organic compound (VOC) emission prediction models were constructed for the first variable input combination and the second variable input combination, respectively. Therefore, the first dataset corresponds to one VOC emission prediction model, and the second dataset corresponds to another VOC emission prediction model.
[0093] Step a3: Replace the mean squared error loss function of the initial decision tree with the quantile regression loss function to obtain the updated decision tree.
[0094] It should be noted that the quantile regression loss function refers to a loss assessment function adapted to the needs of quantile prediction.
[0095] The mean squared error loss function refers to the loss assessment function commonly used in traditional decision tree construction.
[0096] Updating a decision tree refers to the decision tree obtained by replacing the mean squared error loss function of the initial decision tree with the quantile regression loss function.
[0097] In this embodiment of the invention, during the node splitting process of each decision tree, a quantile regression loss function is used instead of the traditional mean squared error loss function to optimize the prediction accuracy at a specific quantile. The quantile loss function is as follows:
[0098] In the formula, τ is the quantile level, and u is the residual. I This is an indicator function.
[0099] Step a4: Aggregate the test sample values of the sample subset that fall into the leaf nodes of the corresponding update decision tree, and construct the conditional empirical distribution function of the update decision tree.
[0100] It should be noted that test samples refer to independent samples that are separated from the sample subset and used to verify the model's performance.
[0101] A leaf node is the lowest level node in a decision tree, and it is the final unit to which a test sample is classified.
[0102] Training sample values refer to the sample data in the sample subset used to train and update the decision tree.
[0103] Aggregation refers to the process of summarizing and organizing all training sample values within the leaf nodes to which the test sample falls.
[0104] The conditional empirical distribution function refers to the distribution function constructed based on the aggregated training sample values under specific input conditions (the classification results of the leaf nodes corresponding to the test samples).
[0105] In this embodiment of the invention, during the prediction phase: the training sample values of the test sample falling into the leaf nodes of each tree are aggregated to construct a conditional empirical distribution function. Specifically, for the test sample... Calculate its in the first i The conditional empirical distribution function constructed from the training sample values falling into the leaf nodes of the tree is:
[0106] In the formula, K Indicates the number of decision trees; This represents the training sample value within the leaf node.
[0107] Step a5: Extract the predicted value of any quantile from the conditional empirical distribution function, calculate the quantile of the quantile, and obtain the prediction interval of the gas station volatile organic compound emissions for updating the decision tree.
[0108] It should be noted that quantiles refer to the numerical points that divide the sorted emissions dataset into multiple equally probable intervals.
[0109] Quantiles refer to statistical values directly associated with quantiles, that is, the specific emission prediction values corresponding to quantiles.
[0110] The prediction interval refers to the range of values formed by the leftmost (smallest) quantile among the extracted predicted values of multiple quantiles as the lower limit and the rightmost (largest) quantile as the upper limit.
[0111] In this embodiment of the invention, predicted values for arbitrary quantiles are extracted from the conditional empirical distribution function, and the corresponding confidence intervals are calculated. Specifically, the τ quantile is extracted from the conditional empirical distribution function:
[0112] By calculating the quantiles corresponding to multiple τ values, the prediction range of volatile organic compound emissions from gas stations can be obtained for updating the decision tree.
[0113] Step a6: Using the updated decision tree, conditional empirical distribution function, and prediction interval of volatile organic compound emissions from gas stations corresponding to the first dataset, construct the initial volatile organic compound emission prediction model for the first dataset.
[0114] In this embodiment of the invention, all updated decision trees corresponding to the first dataset are used as the core unit of integration. The conditional empirical distribution function corresponding to each updated decision tree and the prediction interval of volatile organic compound emissions from gas stations are integrated. The prediction results of multiple updated decision trees are fused and calibrated through an ensemble learning strategy, so that the model can combine the prediction advantages of multiple decision trees, while retaining the quantile prediction ability and uncertainty quantification ability. Finally, an initial volatile organic compound emission prediction model adapted to the characteristics of the first dataset is constructed.
[0115] Step a7: Using the updated decision tree, conditional empirical distribution function, and prediction interval of volatile organic compound emissions from gas stations corresponding to the second dataset, construct the initial volatile organic compound emission prediction model for the second dataset.
[0116] In this embodiment of the invention, all updated decision trees corresponding to the second dataset are used as the core unit of integration. The conditional empirical distribution function corresponding to each updated decision tree and the prediction interval of volatile organic compound emissions from gas stations are integrated. The prediction results of multiple updated decision trees are fused and calibrated through an ensemble learning strategy, so that the model can combine the prediction advantages of multiple decision trees, while retaining the quantile prediction ability and uncertainty quantification ability. Finally, an initial volatile organic compound emission prediction model adapted to the characteristics of the second dataset is constructed.
[0117] Step S202: Input each dataset into the corresponding preset initial volatile organic compound emission prediction model for training to generate multiple target volatile organic compound emission prediction models.
[0118] In some optional implementations, step S202 above includes: Step S2021: Divide each dataset into a training set and a validation set according to a preset ratio.
[0119] It should be noted that the preset ratio refers to K-1:1.
[0120] The training set refers to the subset of samples used for model training, which is divided into K-1 subsets from each dataset.
[0121] The validation set refers to a subset of samples from each dataset that is independent of the training set.
[0122] In this embodiment of the invention, K-fold cross-validation is used to train and evaluate the model. The source domain data is divided into K parts, of which K-1 parts are used for training and the remaining part is used for validation. Therefore, each dataset is divided into K-1 sample subsets as the training set and 1 sample subset as the validation set.
[0123] In step S2022, each training set is input into the corresponding preset initial volatile organic compound emission prediction model for training, resulting in multiple updated volatile organic compound prediction models.
[0124] It should be noted that updating the volatile organic compound prediction model refers to the optimized prediction model obtained after iterative optimization using the training set.
[0125] In this embodiment of the invention, the training set of the first dataset is input into the initial volatile organic compound emission prediction model corresponding to the first dataset for training. The training is carried out in K rounds, and the prediction error of the model on different data subsets is calculated. During the training, the hyperparameters of the quantile random forest, including the number of trees, maximum depth, feature sampling rate and number of quantile points, are adjusted by the grid search method to optimize the model performance, so as to obtain the updated volatile organic compound prediction model corresponding to the first dataset.
[0126] Similarly, the training set of the second dataset is input into the initial volatile organic compound (VOC) emission prediction model corresponding to the second dataset for training. This process is repeated K times, and the prediction error of the model on different data subsets is calculated. During training, the hyperparameters of the quantile random forest, including the number of trees, maximum depth, feature sampling rate, and number of quantile points, are adjusted using the grid search method to optimize the model performance, thereby obtaining the updated VOC prediction model corresponding to the second dataset.
[0127] In step S2023, each validation set is input into the corresponding updated volatile organic compound emission prediction model for validation, resulting in multiple target volatile organic compound emission prediction models.
[0128] In this embodiment of the invention, each validation set corresponds to an updated volatile organic compound (VOC) emission prediction model that is adapted to the input (the validation set of the first dataset corresponds to the first updated VOC emission prediction model, and the validation set of the second dataset corresponds to another updated VOC emission prediction model). The prediction accuracy, quantile loss value, and prediction interval coverage are the core validation indicators to evaluate the model's generalization ability and prediction reliability on the independent validation set. If the model validation result does not meet the preset standard, the model parameters are adjusted and retrained until the validation indicators meet the requirements. Then the final structure of the model is determined, and finally multiple target VOC emission prediction models with high prediction accuracy and stable uncertainty quantification capability are obtained.
[0129] Step S203: Calculate the predicted value of volatile organic compound emissions for each gas station using the volatile organic compound emission prediction model for each target.
[0130] Specifically, step S203 includes: Step S2031: Based on the spatial location information of each gas station, the meteorological indicators and industrial emission proxy indicators of each gas station are spatially matched to generate the input variable values of each gas station.
[0131] It should be noted that spatial location information refers to the core data that represents the specific geographical location of each gas station. The core data is latitude and longitude coordinates, and may also include auxiliary spatial information such as administrative division codes.
[0132] Spatial matching refers to the process of accurately associating multi-source environmental data (such as regional-scale meteorological data and grid-scale industrial data) with the corresponding gas station and its surrounding buffer zone based on the latitude and longitude coordinates of the gas station.
[0133] Industrial emission proxy indicators, including energy and major industrial product output indicators, represent the proxy chain of "production-emissions". Among them, gasoline and diesel output directly correspond to the regional oil supply scale, reflecting the base level of refueling volume and volatile matter; naphtha, liquefied petroleum gas and other petrochemical product outputs reflect crude oil consumption in the upstream refining process; automobile output is linked to fuel consumption demand from a life cycle perspective, forming an emission driver in the "manufacturing-use" chain; the output of indirect related products such as steel and cement represents the scale of infrastructure investment, which may affect emission intensity through the expansion of gas station construction.
[0134] Input variable values refer to standardized data obtained after spatial matching and data cleaning, covering core characteristics affecting emissions such as meteorological indicators, energy and industrial product output data.
[0135] In this embodiment of the invention, the spatial location information of all gas stations is spatially matched with meteorological indicators and industrial emission proxy indicators (including energy and major industrial product output indicators) to obtain the input variable value for each gas station.
[0136] Step S2032: Input the input variable values of the gas station into the volatile organic compound emission prediction model of each target.
[0137] In this embodiment of the invention, the input variable values of each gas station are respectively input into the target volatile organic compound emission prediction model corresponding to the first dataset and the target volatile organic compound emission prediction model corresponding to the second dataset.
[0138] Step S2033: Calculate the predicted value of the gas station's volatile organic compound emissions using the volatile organic compound emission prediction model for each target.
[0139] In this embodiment of the invention, the predicted value of the volatile organic compound (VOC) emissions of each gas station is calculated using two optimized target VOC emission prediction models.
[0140] Step S204: Based on all predicted values of volatile organic compound (VOC) emissions from each gas station, calculate the total VOC emissions and uncertainty range for all gas stations.
[0141] In some optional implementations, step S204 above includes: Step S2041: Extract the predicted values of the preset quantiles of each gas station through the volatile organic compound emission prediction model of each target gas station to obtain the distribution characteristics of the volatile organic compound emissions of each gas station.
[0142] It should be noted that the distribution characteristics of volatile organic compound (VOC) emissions from gas stations refer to the characteristic information that characterizes the overall distribution pattern of VOC emissions from gas stations, formed by integrating the predicted values of each preset quantile.
[0143] In this embodiment of the invention, the predicted values of the specified quantiles of each gas station are calculated by the target volatile organic compound emission prediction model, thereby obtaining the distribution characteristics of the volatile organic compound emissions of each gas station.
[0144] Step S2042: The Monte Carlo sampling method is used to sample the distribution characteristics of volatile organic compound emissions of each gas station and generate multiple predicted emission values for each gas station.
[0145] It should be noted that Monte Carlo sampling is a random sampling method based on probability and statistics theory.
[0146] Emissions projections refer to estimates of volatile organic compound (VOC) emissions from individual gas stations generated through Monte Carlo sampling.
[0147] In this embodiment of the invention, during the prediction process of the quantile random forest model (i.e., the target volatile organic compound emission prediction model), each gas station is sampled 200 times to generate 200 predicted emission values for each gas station.
[0148] Step S2043: Calculate the mean of all predicted emissions for each gas station and the confidence interval of the preset proportion.
[0149] It should be noted that the preset confidence interval refers to the 95% confidence interval.
[0150] In this embodiment of the invention, the mean, standard deviation, and 95% confidence interval (i.e. uncertainty interval) of 200 predicted emission values for each gas station are calculated.
[0151] The relative uncertainty is calculated by dividing the half-width of the confidence interval (the difference between the upper and lower limits) by the mean. The specific formula is as follows: .
[0152] Step S2044: Based on the average volatile organic compound (VOC) emissions of all gas stations and a preset confidence interval, determine the total VOC emissions and uncertainty interval of all gas stations.
[0153] In this embodiment of the invention, the prediction results of all gas stations are summarized to obtain the mean of the total volatile organic compound emissions of all gas stations and the 95% confidence interval (i.e., the uncertainty interval).
[0154] By summing up the volatile organic compound (VOC) emissions from all gas stations, we can obtain the total VOC emissions from all gas stations and the range within which they are uncertain.
[0155] Step S205: The Shapley interpretation values are used to analyze the input variable values of each target volatile organic compound emission prediction model to obtain the influence values of each variable characteristic of the input variable values on the corresponding volatile organic compound emissions of the gas station.
[0156] It should be noted that the Shapley Explanation Value (SHAP value) refers to the core indicator of the model explanation method based on cooperative game theory. The core principle is to fairly allocate the contribution of each variable feature to the model's prediction results. By considering the impact of all possible combinations of each variable feature with other variable features on the prediction results, the average marginal contribution of the variable is calculated, thereby quantifying the degree of influence of the variable on the prediction results.
[0157] The impact value refers to the quantitative impact of a single variable characteristic on the volatile organic compound emissions of a gas station, obtained through Shapley's interpretation.
[0158] Variable features refer to the specific feature dimensions contained in the input variable values, such as surface temperature and wind speed in meteorological features, and steel production in energy and industrial product output features.
[0159] In this embodiment of the invention, key factors affecting volatile organic compound (VOC) emissions from gas stations are identified based on Shapley additive explanatory value analysis (SHAP). Specifically, the input variable features (i.e., the specific features contained in the input variable values) are analyzed based on game theory SHAP values, and the impact value of each feature on the prediction result is calculated.
[0160] Step S206: Based on the absolute mean of Shapley's explanatory values, sort the characteristics of each variable, and determine the key factors affecting the volatile organic compound emissions of gas stations based on the sorting results.
[0161] It should be noted that the absolute mean of the Shapley explanatory values refers to the arithmetic mean calculated by taking the absolute values of the Shapley explanatory values for a single variable feature across all gas station samples.
[0162] The ranking result refers to the order in which the characteristics of each variable are arranged from largest to smallest according to the mean of the absolute value of the Shapley interpretation.
[0163] Key factors refer to the variable features that rank highest in terms of the mean absolute value of Shapley's explanatory values, selected from the ranking results.
[0164] In this embodiment of the invention, the variable characteristics are sorted according to the mean absolute value of SHAP, the key factors affecting the volatile organic compound emissions of gas stations are identified, and the nonlinear relationship between key factors such as automobile production, aluminum production, flat glass production, wind speed and surface air pressure and volatile organic compound emissions is analyzed.
[0165] In terms of data processing, this invention effectively solves the problems of single input data and low spatiotemporal resolution in traditional methods by spatiotemporal matching and fusion of multi-source heterogeneous data (meteorology, air quality, and industrial output), and by performing logarithmic transformation and standardization preprocessing on emission data. This provides a high-quality, high-dimensional data foundation for model construction. In terms of model construction, the invention introduces a "meteorological-industrial emission proxy index" as an input variable combination, linking macroeconomic socioeconomic activities with micro-emission processes, significantly improving the model's explanatory power and prediction accuracy. The use of the Quantile Random Forest (QRF) algorithm not only captures the complex nonlinear relationship between meteorological conditions, human activities, and VOCs emissions, but also possesses good robustness due to its ensemble learning characteristics. Most importantly, the QRF model naturally outputs the probability distribution of emissions. Combined with Monte Carlo simulation, it can achieve dynamic estimation of the total VOCs emissions from all gas stations and quantify the uncertainty of the confidence interval, resulting in more scientific and reliable results. Furthermore, SHAP value analysis can clearly reveal the nonlinear driving mechanisms of key factors such as automobile production, aluminum production, temperature, and air pressure, providing solid technical support for identifying emission hotspots and formulating differentiated control policies, and has significant application value in environmental management.
[0166] This embodiment also provides a device for dynamic estimation and uncertainty quantification of volatile organic compound emissions from gas stations. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0167] This embodiment provides a device for dynamically estimating and quantifying the uncertainty of volatile organic compound (VOC) emissions from gas stations, such as... Figure 3 As shown, it includes: Module 301 is used to construct multiple datasets based on the volatile organic compound emission data of multiple gas stations and the corresponding multi-source environmental data. Training module 302 is used to input each dataset into the corresponding preset initial volatile organic compound emission prediction model for training, and generate multiple target volatile organic compound emission prediction models. The prediction module 303 is used to calculate the predicted value of the volatile organic compound emissions of each gas station through the volatile organic compound emission prediction model of each target. The calculation module 304 is used to calculate the total volatile organic compound emissions and uncertainty range of all gas stations based on all predicted values of volatile organic compound emissions of each gas station.
[0168] In some alternative implementations, the construction module 301 includes: The acquisition unit is used to acquire volatile organic compound emission data from multiple gas stations, as well as meteorological data, air quality data, and energy and industrial product output data for each gas station. The construction unit is used to construct a first rectangular buffer area based on the location of each gas station, and to calculate the average value of meteorological data and the average value of air quality data within the first rectangular buffer area using the spatial averaging method. The allocation unit is used to distribute energy and industrial product output data to multiple grid cells based on the spatial discretization method, and to allocate weights according to the resident population density in each grid cell to construct a second rectangular buffer area. The mean calculation unit is used to calculate the average output of industrial products within the second rectangular buffer area using the spatial averaging method. The transformation unit is used to perform a logarithmic transformation on the volatile organic compound (VOC) emission data of gas stations to obtain new VOC emission data for gas stations. The processing unit is used to standardize the average values of meteorological data, air quality data, industrial product output, and new volatile organic compound emission data from gas stations. Construct a first combination unit to use the average value of standardized meteorological data and the average value of air quality data to construct the first combination of input variables; Construct a first combination unit to use the average value of standardized meteorological data and the average value of industrial product output to construct a second combination of input variables; A unit is established to establish a correspondence between the first combination of input variables and the second combination of input variables and the standardized volatile organic compound emission data of gas stations, based on the same time and space, to obtain the first dataset and the second dataset.
[0169] In some alternative implementations, prior to training module 302, the device further includes: The sampling unit is used to sample the first dataset and the second dataset to generate multiple sample subsets of the first dataset and multiple sample subsets of the second dataset. Construct decision tree sub-units to build multiple initial decision trees based on each sample subset; The replacement sub-unit is used to replace the mean squared error loss function of the initial decision tree with the quantile regression loss function to obtain the updated decision tree; The aggregation subunit is used to aggregate the test sample values of the sample subset that fall into the leaf node of the corresponding updated decision tree, and construct the conditional empirical distribution function of the updated decision tree. Extract sub-units to extract the predicted value of any quantile from the conditional empirical distribution function, calculate the quantile of the quantile, and obtain the prediction interval of the gas station volatile organic compound emissions for updating the decision tree; The first model subunit is constructed to build the initial volatile organic compound emission prediction model for the first dataset by using the updated decision tree, conditional empirical distribution function and prediction interval of volatile organic compound emissions from gas stations corresponding to the first dataset. A second model subunit is constructed to build an initial volatile organic compound emission prediction model for the second dataset by using the updated decision tree, conditional empirical distribution function, and prediction interval of volatile organic compound emissions from gas stations corresponding to the second dataset.
[0170] In some alternative implementations, training module 302 includes: The partitioning unit is used to divide each dataset into a training set and a validation set according to a preset ratio; The training unit is used to input each training set into the corresponding preset initial volatile organic compound emission prediction model for training, so as to obtain multiple updated volatile organic compound prediction models. The validation unit is used to input each validation set into the corresponding updated volatile organic compound emission prediction model for validation, thereby obtaining multiple target volatile organic compound emission prediction models.
[0171] In some alternative implementations, the prediction module 303 includes: The matching unit is used to spatially match the meteorological indicators and industrial emission proxy indicators of each gas station based on the spatial location information of each gas station, and generate the input variable values of each gas station. The input unit is used to input the input variable values of the gas station into the volatile organic compound emission prediction model of each target. The prediction unit is used to calculate the predicted value of the volatile organic compound (VOC) emissions of the gas station using the VOC emission prediction model for each target.
[0172] In some alternative implementations, the computing module 304 includes: The predicted value extraction unit is used to extract the predicted values of the preset quantiles of each gas station through the predicted model of each target volatile organic compound emission, so as to obtain the distribution characteristics of the volatile organic compound emissions of each gas station. The feature sampling unit is used to sample the distribution characteristics of volatile organic compound emissions from each gas station using the Monte Carlo sampling method, and generate multiple predicted emission values for each gas station. The calculation unit is used to calculate the mean of all predicted emissions for each gas station and the confidence interval of a preset proportion; The total emission calculation unit is used to determine the total volatile organic compound (VOC) emissions and uncertainty range of all gas stations based on the average VOC emissions of all gas stations and a preset confidence interval.
[0173] In some alternative embodiments, the device further includes: The analysis unit is used to analyze the input variable values of each target volatile organic compound emission prediction model using Shapley's interpretation values, and to obtain the influence values of each variable characteristic of the input variable values on the corresponding volatile organic compound emissions of the gas station. The sorting unit is used to sort the characteristics of each variable based on the absolute mean of the Shapley explanatory values, and to determine the key factors affecting the volatile organic compound emissions of gas stations based on the sorting results.
[0174] The device for dynamic estimation and uncertainty quantification of volatile organic compound (VOC) emissions from gas stations provided in this embodiment of the invention can execute the method for dynamic estimation and uncertainty quantification of VOC emissions from gas stations provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0175] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0176] The following is a detailed reference. Figure 4 This diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 401, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 402 or a program loaded from memory 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of the electronic device. The processor 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0177] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0178] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a memory 408, or installed from a ROM 402. When the computer program is executed by the processor 401, it performs the functions defined in the dynamic estimation and uncertainty quantification method for volatile organic compound emissions at gas stations according to embodiments of the present invention.
[0179] Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0180] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the method for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions from gas stations shown in the above embodiments is implemented.
[0181] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for dynamically estimating and quantifying the uncertainty of volatile organic compound (VOC) emissions from gas stations, characterized in that... The method includes: Multiple datasets were constructed based on volatile organic compound (VOC) emission data from multiple gas stations and corresponding multi-source environmental data. Each dataset is input into a corresponding preset initial volatile organic compound emission prediction model for training, generating multiple target volatile organic compound emission prediction models. The predicted value of volatile organic compound emissions for each gas station is calculated using the predicted volatile organic compound emission model for each of the aforementioned targets. Based on all predicted values of volatile organic compound (VOC) emissions for each of the gas stations, calculate the total VOC emissions and uncertainty range for all the gas stations.
2. The method according to claim 1, characterized in that, The process involves constructing multiple datasets based on volatile organic compound (VOC) emission data from multiple gas stations and corresponding multi-source environmental data, including: Acquire volatile organic compound emission data from multiple gas stations, as well as meteorological data, air quality data, and energy and industrial product output data for each gas station; Based on the locations of each gas station, a first rectangular buffer zone is constructed, and the average value of meteorological data and the average value of air quality data within the first rectangular buffer zone are calculated using the spatial averaging method. Based on the spatial discretization method, the energy and industrial product output data are allocated to multiple grid cells, and weights are assigned according to the resident population density in each grid cell to construct a second rectangular buffer area. The average output of industrial products within the second rectangular buffer area is calculated using the spatial averaging method. Logarithmic transformation is performed on the volatile organic compound (VOC) emission data of the gas station to obtain new VOC emission data for the gas station; The average values of the meteorological data, the average values of the air quality data, the average values of the industrial product output, and the new volatile organic compound emission data of the gas station are standardized. The first input variable combination is constructed using the average value of standardized meteorological data and the average value of air quality data; The average value of the standardized meteorological data and the average value of industrial product output are used to construct a second combination of input variables; Based on the same time and space, the first input variable combination and the second input variable combination are respectively correlated with the standardized volatile organic compound emission data of gas stations to obtain the first dataset and the second dataset.
3. The method according to claim 2, characterized in that, Before the step of inputting each of the datasets into a corresponding preset initial volatile organic compound emission prediction model for training to generate multiple target volatile organic compound emission prediction models, the method further includes: The first dataset and the second dataset are sampled to generate multiple sample subsets of the first dataset and multiple sample subsets of the second dataset; Based on each of the aforementioned sample subsets, multiple initial decision trees are constructed respectively; The mean squared error loss function of the initial decision tree is replaced with the quantile regression loss function to obtain the updated decision tree; The test samples of the sample subset are aggregated into the training sample values that fall into the leaf nodes of the corresponding updated decision tree to construct the conditional empirical distribution function of the updated decision tree. The predicted value of any quantile is extracted from the conditional empirical distribution function, and the quantile of the quantile is calculated to obtain the prediction interval of the gas station volatile organic compound emissions of the updated decision tree. An initial volatile organic compound (VOC) emission prediction model for the first dataset is constructed using the updated decision tree, conditional empirical distribution function, and prediction interval of VOC emissions from gas stations corresponding to the first dataset. An initial volatile organic compound (VOC) emission prediction model for the second dataset is constructed using the updated decision tree, conditional empirical distribution function, and prediction interval of VOC emissions from gas stations corresponding to the second dataset.
4. The method according to claim 1, characterized in that, The step involves inputting each of the datasets into a corresponding preset initial volatile organic compound (VOC) emission prediction model for training, generating multiple target VOC emission prediction models, including: Each dataset is divided into a training set and a validation set according to a preset ratio; Each training set is input into the corresponding preset initial volatile organic compound emission prediction model for training, resulting in multiple updated volatile organic compound prediction models. Each of the aforementioned validation sets is input into the corresponding updated volatile organic compound (VOC) emission prediction model for validation, resulting in multiple target VOC emission prediction models.
5. The method according to claim 1, characterized in that, The step of calculating the predicted value of volatile organic compound (VOC) emissions for each gas station using the target VOC emission prediction model includes: Based on the spatial location information of each gas station, the meteorological indicators and industrial emission proxy indicators of each gas station are spatially matched to generate the input variable values of each gas station. The input variable values of the gas station are respectively input into each of the target volatile organic compound emission prediction models; The predicted value of the gas station's volatile organic compound (VOC) emissions is calculated using the respective target VOC emission prediction models.
6. The method according to claim 1, characterized in that, The calculation of the total volatile organic compound (VOC) emissions and uncertainty range of all gas stations, based on the predicted values of VOC emissions from each of the gas stations, includes: The predicted values of the preset quantiles of each gas station are extracted by the target volatile organic compound emission prediction model to obtain the distribution characteristics of the volatile organic compound emissions of each gas station. The distribution characteristics of volatile organic compound emissions from each of the gas stations were sampled using the Monte Carlo sampling method, generating multiple predicted emission values for each gas station. Calculate the mean of all predicted emissions for each gas station and the confidence interval for a preset percentage; Based on the average volatile organic compound (VOC) emissions of all the gas stations and a preset confidence interval, the total VOC emissions and uncertainty interval of all the gas stations are determined.
7. The method according to claim 1, characterized in that, The method further includes: The input variable values of each target volatile organic compound emission prediction model are analyzed using Shapley's interpretation values to obtain the influence values of each variable characteristic of the input variable values on the corresponding volatile organic compound emissions of the gas station. Based on the mean absolute value of the Shapley explanatory values, the characteristics of each variable are ranked, and the key factors affecting the volatile organic compound emissions of the gas station are determined according to the ranking results.
8. A device for dynamically estimating and quantifying the uncertainty of volatile organic compound emissions from gas stations, characterized in that, The device includes: The module is used to build multiple datasets based on the volatile organic compound emission data of multiple gas stations and the corresponding multi-source environmental data; The training module is used to input each of the datasets into the corresponding preset initial volatile organic compound emission prediction model for training, and generate multiple target volatile organic compound emission prediction models. The prediction module is used to calculate the predicted value of the volatile organic compound emissions of each gas station using the target volatile organic compound emission prediction model. The calculation module is used to calculate the total volatile organic compound (VOC) emissions and uncertainty range of all the gas stations based on all predicted values of VOC emissions of each of the gas stations.
9. An electronic device, characterized in that, include: The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes the computer instructions to perform the dynamic estimation and uncertainty quantification method for volatile organic compound emissions from gas stations as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the dynamic estimation and uncertainty quantification method for volatile organic compound emissions from gas stations as described in any one of claims 1 to 7.