Reverse prediction method and device for air pollutant emission, electronic equipment and medium
By training an air pollutant reverse prediction model using multi-source data collected by vehicle-mounted equipment in urban traffic, the problem of insufficient spatial resolution in existing technologies has been solved. This model enables accurate prediction of pollutant emission concentrations over historical periods, improves the reliability and generalization ability of source tracing, and assists in pollutant management and risk prevention.
Patent Information
- Application Number
- CN202511007827.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies for predicting air pollutants in urban traffic suffer from insufficient spatial resolution, which limits the reliability and generalization ability of dynamic source tracing of pollutant emissions.
By acquiring geographic data of the target area in the first time period, and inputting it into an air pollutant reverse prediction model trained by vehicle-mounted equipment during movement, the model outputs the pollutant emission concentration in the first time period. By using multi-source data and machine learning algorithms to construct a reverse prediction model, the model can accurately predict the pollutant emission concentration in historical time periods.
It enhances the reliability and generalization of dynamic source tracing of urban air pollutant emissions, enabling efficient location of pollutant hotspots, identification of emission change patterns, and provision of scientific evidence for pollutant control and prevention of residents' exposure risks.
Smart Images

Figure CN120875155A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental monitoring technology, and in particular to methods, devices, electronic equipment and media for reverse prediction of air pollutant emissions. Background Technology
[0002] Nitrogen oxides (NO2), carbon monoxide (CO), and inhalable particulate matter (PM2.5) emitted by urban traffic 10 PM 2.5 Pollutants such as [unspecified pollutants] have become key factors threatening residents' health. Therefore, achieving accurate spatiotemporal dynamic prediction of urban traffic air pollutants is of great significance for maintaining clean air development and protecting residents' health in my country. Current technologies for predicting urban traffic air pollutants suffer from insufficient spatial resolution, limiting the reliability and generalization ability of dynamic source tracing of urban air pollutant emissions. Summary of the Invention
[0003] This invention provides a method, device, electronic equipment, and medium for reverse prediction of air pollutant emissions, which can predict the concentration of pollutant emissions over historical periods and effectively improve the reliability and generalization ability of dynamic source tracing of urban air pollutant emissions.
[0004] According to one aspect of the present invention, a method for reverse prediction of air pollutant emissions is provided, the method comprising:
[0005] Acquire the first geographic data generated within the target area corresponding to the first time period; wherein, the geographic data is data information related to pollutant emission concentration;
[0006] The first geographic data is input into the target air pollutant reverse prediction model, and the emission concentration of the first pollutant corresponding to the first time period is output.
[0007] The target air pollutant reverse prediction model is trained by collecting mobile observation data generated in the target area during the second time period using vehicle-mounted equipment during movement; the mobile observation data consists of second geographical data and second pollutant emission concentrations; the first geographical data was collected earlier than the second geographical data.
[0008] According to another aspect of the present invention, an air pollutant emission reverse prediction device is provided, the device comprising:
[0009] A geographic data acquisition unit is used to acquire first geographic data generated within a target area corresponding to a first time period; wherein, the geographic data is data information related to pollutant emission concentration;
[0010] The pollutant emission concentration output unit is used to input the first geographical data into the target air pollutant reverse prediction model and output the first pollutant emission concentration corresponding to the first time period.
[0011] The target air pollutant reverse prediction model is trained by collecting mobile observation data generated in the target area during the second time period using vehicle-mounted equipment during movement; the mobile observation data consists of second geographical data and second pollutant emission concentrations; the first geographical data was collected earlier than the second geographical data.
[0012] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0013] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the air pollutant emission reverse prediction method according to any embodiment of the present invention.
[0014] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the air pollutant emission reverse prediction method according to any embodiment of the present invention.
[0015] The technical solution of this invention acquires first geographical data generated within a target area corresponding to a first time period, then inputs the first geographical data into a target air pollutant reverse prediction model, and outputs the emission concentration of the first pollutant corresponding to the first time period. This technical solution can predict pollutant emission concentrations for historical time periods, effectively improving the reliability and generalization ability of dynamic source tracing of urban air pollutant emissions. It solves the problem that existing technologies in the field of urban traffic air pollutant prediction suffer from insufficient spatial resolution, making it difficult to guarantee the reliability of dynamic source tracing of urban air pollutant emissions, and significantly limiting the generalization ability.
[0016] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the air pollutant emission reverse prediction method provided in Embodiment 1 of the present invention;
[0019] Figure 2 This is a schematic diagram of the reverse prediction process for air pollutant emissions provided in Embodiment 2 of the present invention;
[0020] Figure 3 This is a schematic diagram of the target air pollutant reverse prediction model provided in Embodiment 2 of this application;
[0021] Figure 4 This is a schematic diagram of the air pollutant emission reverse prediction method provided in Embodiment 3 of the present invention;
[0022] Figure 5 This is a schematic diagram of artificial intelligence pixel segmentation provided in Embodiment 3 of this application;
[0023] Figure 6 This is a schematic diagram of the visual artificial intelligence multimodal feature extraction and fusion process provided in Embodiment 3 of this application;
[0024] Figure 7 This is a schematic diagram of reverse prediction of air pollutant emissions provided in Embodiment 3 of this application;
[0025] Figure 8 This is a schematic diagram of the air pollutant emission reverse prediction device provided in Embodiment 4 of the present invention;
[0026] Figure 9 This is a schematic diagram of the structure of an electronic device that implements the air pollutant emission reverse prediction method of this invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Example 1
[0030] Figure 1 This is a flowchart of an air pollutant emission reverse prediction method according to Embodiment 1 of the present invention. This embodiment is applicable to situations where pollutant emission concentrations are predicted over historical time periods. The method can be executed by an air pollutant emission reverse prediction device, which can be implemented in hardware and / or software and can be configured within a device. For example, the device can be a backend server or other equipment with communication and computing capabilities. Figure 1 As shown, the method includes:
[0031] S110. Obtain the first geographic data generated within the target area corresponding to the first time period; wherein, the geographic data is data information related to pollutant emission concentration.
[0032] In this scheme, the first time period can be set based on the predicted pollutant emission concentration. For example, the first time period can be set to τ∈[τ start ,τ end The result can be obtained by splicing together n discrete time intervals τ = [τ1; τ2; ...; τ...]. n ].
[0033] The target area refers to an area with clearly defined geographical boundaries.
[0034] In this embodiment, the geographic data refers to data information associated with pollutant emission concentrations. The first geographic data includes meteorological information, street view images, and driving condition information. The meteorological information includes temperature, humidity, air pressure, wind speed, and wind direction. The driving condition information includes vehicle speed, direction, acceleration, gradient, and altitude.
[0035] In this solution, the first geographic data generated within the target area corresponding to the first time period can be obtained using the data acquisition equipment. Specifically, street view images of the first time period τ are obtained using internet map services, hourly average road driving condition information is obtained using traffic situation awareness services, and hourly meteorological information of the city at a 1-kilometer resolution is obtained based on open data interfaces.
[0036] Furthermore, using timestamps τ "YY-MM-DD h:m:s" The city geographic information database (GID) is formed by aligning the primary keys. Wherein, τ "YY-MM-DD h:m:s" The entire format represents a timestamp recorded in the format YY-MM-DD h:m:s, where YY represents the year, MM represents the month, DD represents the day, and h:m:s represents the hour, minute, and second, respectively.
[0037] The city geographic information database (GID) table is stored on a PostgreSQL database, and the image files are stored in the form of absolute path indexes.
[0038] S120. Input the first geographic data into the target air pollutant reverse prediction model and output the emission concentration of the first pollutant corresponding to the first time period; wherein, the target air pollutant reverse prediction model is trained by collecting mobile observation data generated in the target area corresponding to the second time period through vehicle-mounted equipment during movement; the mobile observation data consists of the second geographic data and the emission concentration of the second pollutant; the collection time of the first geographic data is earlier than the collection time of the second geographic data.
[0039] In this embodiment, the pollutant emission concentrations include nitrogen oxides (NO2), carbon monoxide (CO), and inhalable particulate matter (PM2.5). 10 PM 2.5 Emission concentration.
[0040] The second time period can be set based on the predicted pollutant emission concentration. The first time period is earlier than the second time period.
[0041] In this scheme, the target air pollutant reverse prediction model is trained using mobile observation data collected by vehicle-mounted equipment during a second time period within the target area. This mobile observation data consists of two parts: second geographical data and second pollutant emission concentrations. The target air pollutant reverse prediction model trained based on this mobile observation data is primarily used to perform correlation analysis between traffic flow and pollutant concentrations.
[0042] Furthermore, the first geographic data is input into the target air pollutant reverse prediction model. The model then outputs the emission concentration of the first pollutant corresponding to the first geographic data for the first time period. The core basis of the prediction process is to analyze mobile observation data within a second time period to clarify the correlation between traffic flow and the concentration of the second pollutant. Based on this correlation, an accurate prediction of the emission concentration of the first pollutant corresponding to the first geographic data for the first time period is made.
[0043] In this scheme, the reverse prediction of the emission concentration of the first pollutant can yield retrospective results. Where τ,(x,y) represent time and latitude / longitude coordinates, respectively. The matrix is saved to a PostgreSQL database using timestamps as the primary key. The results are then spatially aggregated into a 50m-cell road network for evolutionary distribution analysis and visualization.
[0044] Furthermore, a spatiotemporal evolution distribution visualization tool was developed using Python in conjunction with Plotly and Matplotlib. First, this tool was used to visualize the spatiotemporal distribution of four pollutants, intuitively presenting the concentration changes and spatial distribution characteristics of the pollutants. Next, the K-Means++ algorithm was applied to classify the predicted concentration values of these four pollutants into five levels, in ascending order of concentration: low, relatively low, medium, relatively high, and high. The K-Means++ algorithm, combined with sample data from the second pollutant's movement observation, automatically delineates the upper and lower bounds of the numerical intervals. This classification method helps to more clearly identify spatial differences in pollution levels. Subsequently, the emission evolution distribution of the inverse prediction results was visualized into an image format for easy display and analysis. Simultaneously, the classification layers were stored in ESRIShapefile, GeoJSON, Parquest, or Feather formats. These formats effectively preserve the spatial data's attribute information, providing strong support for subsequent reanalysis and identification of driving factors. The entire process fully leverages Python's data processing and visualization capabilities, as well as the clustering advantages of the K-Means++ algorithm, providing a complete analysis and visualization solution for air pollution research.
[0045] In this embodiment, the visualization of the first pollutant emission concentration can be expressed as dots, squares, honeycomb (area) shapes, or columnar shapes. Columnar shapes can visually represent the numerical gradient through differences in length or color. As a geographical process, the first pollutant emission concentration exhibits significant spatiotemporal differences, primarily driven by human activities, encompassing core scenarios such as daily consumption, industrial production, and transportation.
[0046] This solution can efficiently locate pollutant hotspots and accurately identify emission change patterns, helping to build or improve high spatiotemporal resolution pollutant concentration datasets. This provides a scientific basis for urban traffic pollutant emission control and resident exposure risk prevention. It is of vital importance to environmental governance efforts such as air pollutant formation mechanisms, secondary pollutant prevention, and traffic exhaust emission reduction.
[0047] The technical solution of this invention acquires first geographical data generated within a target area corresponding to a first time period, then inputs the first geographical data into a target air pollutant reverse prediction model, and outputs the emission concentration of the first pollutant corresponding to the first time period. By executing this technical solution, reverse prediction of pollutant emission concentrations for historical time periods can be achieved, effectively improving the reliability and generalization ability of dynamic source tracing of urban air pollutant emissions.
[0048] Example 2
[0049] Figure 2 This is a schematic diagram of the air pollutant emission reverse prediction process provided in Embodiment 2 of the present invention. The relationship between this embodiment and the above embodiments is a detailed description of the training process of the target air pollutant reverse prediction model. Figure 2 As shown, the method includes:
[0050] S210. Divide the mobile observation data into a first training set and a second training set.
[0051] The mobile observation data consists of secondary geographic data and secondary pollutant emission concentrations. Secondary geographic data includes meteorological information, street view images, and driving condition information. Meteorological information includes temperature, humidity, air pressure, wind speed, and wind direction. Driving condition information includes vehicle speed, direction, acceleration, gradient, and altitude.
[0052] In this scheme, mobile observation data generated within the target area corresponding to the second time period is collected by vehicle-mounted equipment during movement. Specifically, the concentration of the second pollutant emission is collected based on the air quality station in the vehicle-mounted equipment, meteorological information is collected based on the portable weather station in the vehicle-mounted equipment, street view images are collected based on the street view video camera, and driving condition information is collected based on the GPS sensor. Among the driving condition information, vehicle speed, direction, and altitude can be directly collected by the sensors; while acceleration needs to be calculated based on vehicle speed data, and slope can be derived from altitude information.
[0053] Furthermore, meteorological information, street view images, driving condition information, and the concentration of second pollutant emissions will be timestamped. "YY-MM-DD h:m:s" Primary key alignment forms a moving observation database (MOD). Wherein, τ "YY-MM-DD h:m:s"The entire format represents a timestamp recorded in the format YY-MM-DD h:m:s, where YY represents the year, MM represents the month, DD represents the day, and h:m:s represents the hour, minute, and second, respectively.
[0054] The database tables are stored on a PostgreSQL database, and the image files are stored using absolute path indexes.
[0055] In this embodiment, the mobile observation data is divided into a first training set and a second training set according to a preset division standard. For example, the mobile observation data can be divided into the first training set and the second training set in a 7:3 ratio. Alternatively, the mobile observation data can be divided into the first training set and the second training set in a 9:1 ratio. The data division ratio can be flexibly adjusted, with common ratios being 7:3, 8:2, or 9:1.
[0056] Furthermore, for the initially partitioned training set, 3-10 fold cross-validation is used for further splitting. Taking 3 fold as an example, the training set is divided into three equal parts. Each time, two parts are used for model training, and the remaining part is used as the validation set. Through three iterations, the model accuracy is checked in real time using the validation set during each training iteration, which helps to adjust parameters (such as the number of iterations, regularization coefficient, etc.). After the model has been trained and optimized through cross-validation, the global accuracy is evaluated using the isolated test set from the initial partition, thereby determining the model's generalization ability.
[0057] S220. Determine the parameter configuration of the base learner in the air pollutant inverse prediction model to be trained.
[0058] Among them, the air pollutant inverse prediction model to be trained is used to analyze the mobile observation data in the second time period to clarify the correlation between traffic flow and the concentration of the second pollutant.
[0059] In this embodiment, the base learner is a fundamental component in ensemble learning, and the rationality of its parameter configuration directly determines the final performance of the ensemble model. Due to differences in the algorithmic characteristics of base learners, their corresponding parameter systems also differ. For example, when using a decision tree as the base learner, the core parameters are mainly designed around the tree structure growth logic, including the maximum tree depth (to suppress model overfitting), the minimum number of samples required to split internal nodes (to avoid invalid splits due to insufficient data), the minimum number of samples required for leaf nodes (to ensure the statistical significance of leaf nodes), and splitting criteria (such as information gain, Gini index, etc., which determine the priority of feature selection). When using a support vector machine as the base learner, the parameters focus on the optimization logic of the classification boundary. Key parameters include: regularization parameters (to balance model complexity and classification error), kernel function type (such as linear kernel, RBF kernel, etc., which determine the mapping method of the feature space), and kernel coefficients (affecting the measurement scale of the kernel function on sample distance). These hyperparameters have a significant impact on model accuracy and error control, and all have been determined experimentally.
[0060] In this scheme, parameters suitable for the base learner in the air pollutant back prediction model to be trained can be initially selected based on the time series characteristics of pollutant emission concentration, sample size, and feature dimensions.
[0061] Specifically, the base learner loss function is set as follows:
[0062]
[0063] in, Let y be the loss function of the base learner. i Actual pollutant concentration, The predictions of the meta-learner, where λ is the L1 regularization coefficient; β j These are the weight coefficients of the meta-learner, where n represents the number of samples and p represents the number of parameters.
[0064] S230. Input the first training set into the base learner for training to obtain the base learner training result.
[0065] Specifically, the first training set is input into the base learner and training is carried out to obtain the training result of the base learner.
[0066] S240. Configure the parameters of the meta-learner in the air pollutant inverse prediction model to be trained.
[0067] The meta-learner is the top-level model in the stacked ensemble, taking the predictions from the base learners as input and outputting the final prediction. The parameters of the meta-learner include backbone model parameters, task sampling parameters, and optimizer parameters.
[0068] Specifically, parameters for the meta-learner in the air pollutant inverse prediction model can be initially selected based on the time series characteristics of pollutant emission concentrations, sample size, and feature dimensions.
[0069] In this scheme, the meta-learner loss function is:
[0070] The input expression for the meta-learner is:
[0071] in, Let y be the loss function of the meta-learner. i Actual pollutant concentration; The prediction value of the b-th base learner The prediction value of the meta-learner, β j h is the meta-learner weight coefficient. j (x i The output value of the j-th base learner, i.e., the meta-feature. β0 bias term.
[0072] Optionally, the base learner includes at least two of the following: random forest, extreme gradient boosting machine, extreme random forest, and lightweight gradient boosting machine.
[0073] The meta-learner includes regularized regression.
[0074] In this scheme, we use: Random Forest: Composed of multiple decision trees, it reduces overfitting through bootstrap sampling and random feature selection. Extreme Gradient Boosting Machine: An optimized version based on gradient boosting trees, it improves efficiency through regularization and parallel computation. Extreme Random Forest: Similar to Random Forest, but the split points of the decision trees are completely random (rather than based on information gain), resulting in faster training speed and lower variance. Lightweight Gradient Boosting Machine: Another efficient gradient boosting framework, employing histogram algorithms and leaf growth strategies.
[0075] Regularized regression is divided into L1 regularization (Lasso regression), L2 regularization (Ridge regression), and a combination of the two (Elastic Net).
[0076] In this embodiment, the air pollutant inverse prediction model to be trained is connected in parallel with at least two base learners for parallel regression. Then, Lasso regression is used as a meta-learner to connect the training results of the previous at least two base learners and the second training set for global training.
[0077] By constructing a reverse prediction model for air pollutants that integrates multi-source street view images, meteorological information, and driving condition information, this model can perform reverse analysis and dynamic backtracking of the emission concentrations of four key urban traffic pollutants. This overcomes the current technical bottleneck of lacking intelligent means to reverse predict the spatiotemporal dynamics of air pollutants.
[0078] Optionally, determine the parameter configuration of the base learner in the air pollutant inverse prediction model to be trained, including:
[0079] Determine the initial hyperparameters of the base learner in the air pollutant back prediction model to be trained; wherein, the initial hyperparameters refer to the parameters set before training the base learner;
[0080] Calculate the median error of the base learner under the initial hyperparameters;
[0081] Calculate the error gradient based on the median error;
[0082] The initial hyperparameters are adjusted based on the error gradient to obtain the adjusted hyperparameters, and the adjusted hyperparameters are used as the parameter configuration of the base learner.
[0083] In this approach, before training begins, a Bayesian hyperparameter optimizer is used to fine-tune the hyperparameters of each base learner, and the hyperparameters of the meta-learner are also optimized, resulting in the parameters of each model's hyperparameters. {RF,XGB,EF,LGB,Lasso} This optimizer uses the gradient descent algorithm to fine-tune parameters with the goal of minimizing the median error. The hyperparameters are optimized for 50 epochs, and after optimization, the optimal parameters for each model are saved as a pickle file. The number of hyperparameter optimization epochs can be flexibly adjusted based on the data size and learning rate. When the data size is large or the learning rate setting requires more refined validation, the number of optimization epochs can be increased to ensure parameter fit. An early training stopping mechanism is also introduced: if the model performance (such as validation set accuracy, loss value, and other core metrics) does not improve after more than 20 consecutive training epochs, training is automatically terminated. This mechanism effectively prevents the model from overfitting the training data in ineffective training, improving training efficiency while ensuring model performance.
[0084] Specifically, when constructing the air pollutant inverse prediction model to be trained, the initial hyperparameters of the base learner are determined in advance. Based on the determined initial hyperparameters, the base learner is configured, and the median error of the training results of the base learner under this configuration is calculated using the first training set, which is used as an evaluation metric for the initial performance of the model.
[0085] Furthermore, using the median error as the objective function, gradient calculation methods are employed to determine the gradient values of the error relative to each hyperparameter. These gradient values reflect the direction and extent of the influence of hyperparameter adjustments on error changes. Based on the calculated error gradients, an optimization algorithm is used to iteratively adjust the initial hyperparameters, gradually reducing the median error of the model. Finally, the adjusted hyperparameter combination is used as the optimal parameter configuration for the base learner.
[0086] The median error is inherently insensitive to outliers and anomalies, a characteristic that makes it highly advantageous when processing pollutant monitoring data. Since pollutant observation sensors are typically sensitive to environmental changes, numerical fluctuations (i.e., non-systematic small fluctuations) are prone to occur during monitoring. The median error effectively reduces the interference of these fluctuations on error assessment results, better meeting the needs of actual monitoring scenarios.
[0087] By constructing a reverse prediction model for air pollutants to be trained, which integrates multi-source street view images, meteorological information and driving condition information, the model can perform reverse analysis and dynamic backtracking of the emission concentrations of four key urban traffic pollutants.
[0088] S250. Using the training results of the base learner and the second training set as input, train the meta learner to obtain the target air pollutant reverse prediction model.
[0089] In this scheme, the training results of the base learner and the second training set are used as joint inputs to train the meta learner, and finally the reverse prediction model of the target air pollutant is obtained.
[0090] In this embodiment, Figure 3 This is a schematic diagram of the target air pollutant reverse prediction model provided in Embodiment 2 of this application, as shown below. Figure 3 As shown, NO2, CO, and PM were constructed respectively. 10 PM 2.5 (Units: ppb, ppb, μg / m³) 3 μg / m 3 The air pollutant inverse prediction model to be trained is obtained by connecting four base learners (random forest RF, extreme gradient booster XGB, extreme random forest EF, and lightweight gradient booster LGB) in parallel for parallel regression. Then, Lasso regression is used as the meta learner to connect the training results of the base learners of the previous four base learners and the second training set for global training to obtain the target air pollutant inverse prediction model.
[0091] Specifically, during model training, the Bayesian-optimized hyperparameter file is first read and used as the initial parameters for training. The order of the observed data is shuffled, and the dataset is split in a 9:1 ratio. A small portion of the dataset is used as the test set for model performance diagnostics, while the remainder is used for training. Next, 10-fold cross-validation is used for training (the training and validation sets are split in a 9:1 ratio in each round), and the optimal model parameters are automatically saved after each training round. With a maximum of 100 rounds, early stopping is initiated when the model loss (L) stops decreasing after more than 10 rounds to prevent overfitting. Target air pollutant back-prediction models for the emission concentrations of four pollutants are trained sequentially. Finally, the trained target air pollutant back-prediction model structure and hyperparameters are locally saved as Pickle binary files.
[0092] Furthermore, the number of folds in the cross-validation method can be flexibly adjusted according to specific circumstances such as data size and distribution characteristics; in addition, traditional OOB (out-of-bag) partitioning or proportion-based dataset partitioning can also be selected.
[0093] Furthermore, for each learner and the global model of the target air pollutant back prediction model, the retained test set is used as input, and the classic machine learning metrics R2, RMSE, and MAE are employed to comprehensively evaluate the performance of the learners / models, as shown in the following formulas:
[0094]
[0095] Among them, y i Actual pollutant concentration, To predict pollutant concentrations.
[0096] The performance validation results clearly demonstrate the performance of the target air pollutant inverse prediction model and each base learner (including RF, ET, XGB, and LGB) on the test set. In comparison, the target air pollutant inverse prediction model exhibits the best predictive ability for the concentrations of all four pollutants, with its coefficient of determination (R²) being the highest. 2 The accuracy reached 0.95-0.99. This data fully demonstrates that the model possesses excellent predictive performance in the test set scenario, accurately capturing the changing patterns of pollutant concentrations. During the machine learning training process, the test set was strictly isolated from the training and validation sets.
[0097] The technical solution of this invention involves dividing mobile observation data into a first training set and a second training set; determining the parameter configuration of the base learner in the air pollutant inverse prediction model to be trained; inputting the first training set into the base learner for training to obtain the training result of the base learner; constructing the parameter configuration of the meta learner in the air pollutant inverse prediction model to be trained; and using the training result of the base learner and the second training set as input to train the meta learner to obtain the target air pollutant inverse prediction model. By implementing this technical solution, key pollution sources such as high-emission vehicles, industrial pollution sources, and road dust are accurately identified, the spatiotemporal diffusion field of pollutants is reconstructed, and a target air pollutant inverse prediction model is constructed. This allows for the analysis of the correlation between traffic flow and pollutant concentration, quantifying the emission contribution of different vehicle types, and providing objective evidence for precise pollution control. This technical approach is of crucial significance for research on the formation mechanism of air pollutants, prevention of secondary pollutants, and environmental governance work such as traffic exhaust emission reduction. Conducting inverse prediction of air pollutant concentrations in major cities further enables the visualization of spatiotemporal changes in air pollutant concentrations and the analysis of influencing factors, providing a quantitative tool for air pollutant emission reduction. Meanwhile, this technology can reverse the prediction of pollutant emission concentrations over historical periods, effectively improving the reliability and generalization ability of dynamic source tracing of urban air pollutant emissions.
[0098] Example 3
[0099] Figure 4 This is a schematic diagram of the air pollutant emission reverse prediction method provided in Embodiment 3 of the present invention. The relationship between this embodiment and the above embodiments is a detailed description of the mobile observation data processing process. Figure 4 As shown, the method includes:
[0100] S410: Collect the emission concentration of the second pollutant generated in the target area corresponding to the second time period through the air quality module in the vehicle-mounted equipment; collect meteorological information generated in the target area corresponding to the second time period through the meteorological module in the vehicle-mounted equipment; collect street view images generated in the target area corresponding to the second time period through the street view device in the vehicle-mounted equipment; and collect driving condition information generated in the target area corresponding to the second time period through the positioning module in the vehicle-mounted equipment.
[0101] In this solution, the air quality module consists of an air quality station, the meteorological module consists of a portable weather station, the street view device consists of a street view video camera, and the positioning module consists of a GPS sensor.
[0102] Specifically, the system collects the concentration of the second pollutant emission based on the air quality station in the vehicle, collects meteorological information based on the portable weather station in the vehicle, collects street view images based on the street view video camera, and collects driving condition information based on the GPS sensor. Among the driving condition information, vehicle speed, direction, and altitude can be directly collected by the sensors; while acceleration needs to be calculated based on vehicle speed data, and gradient can be derived from altitude information.
[0103] Furthermore, meteorological information, street view images, driving condition information, and second pollutant emission concentrations are aligned using time strings as the primary key to form a mobile observation database.
[0104] S420. Extract features from the emission concentration of the second pollutant to obtain pollutant emission concentration features; extract features from the meteorological information to obtain meteorological features; extract features from the street view image to obtain street view features; and extract features from the driving condition information to obtain driving condition features; wherein the street view features are presented by the proportion of different types of pixels.
[0105] In this scheme, the emission concentration of the second pollutant is extracted based on artificial intelligence multimodal feature extraction technology to obtain pollutant emission concentration features; meteorological information is extracted to obtain meteorological features; street view images are extracted to obtain street view features; and driving condition information is extracted to obtain driving condition features.
[0106] Specifically, MOD feature matrix The columns include: MOD Street View Features (S MOD ), Driving condition characteristics (V) MOD ), meteorological characteristics (M) MOD It includes pollutant emission concentration characteristics (P0.05) measured by mobile sensors. MOD ), among which, PL MOD This represents the concentration of four air pollutants.
[0107] Optionally, feature extraction is performed on the street view image to obtain street view features, including:
[0108] Feature extraction is performed on the street view image to obtain the number of first-type pixels and the area of second-type pixels; wherein, the first-type pixels refer to pixels corresponding to ground features with discrete, countable features; and the second-type pixels refer to pixels corresponding to ground features with continuous, areal features.
[0109] Based on the number of first-type pixels, calculate the ratio of the first-type pixels to the total number of pixels in the street view image; and based on the area of the second-type pixels, calculate the ratio of the area of the second-type pixels to the total area of the street view image.
[0110] in, Figure 5 This is a schematic diagram of artificial intelligence instance segmentation provided in Embodiment 3 of this application, as shown below. Figure 5 As shown, the first type of pixel refers to the pixel corresponding to a ground feature with discrete, countable features, including pedestrians, riders, bicycles, motorcycles, cars, and trucks. In this embodiment, the second type of pixel refers to the pixel corresponding to a ground feature with continuous, areal features. The second type of pixel includes roads, buildings, walls, fences, bridges, trees, grasslands, bare land, water bodies, and the sky.
[0111] Specifically, feature extraction can be performed on street view images based on Mask2Former instance segmentation to obtain the number of first-type pixels and the area of second-type pixels.
[0112] Furthermore, based on the number of first-type pixels, their ratio in the total number of pixels in the street view image is calculated; simultaneously, based on the area of second-type pixels, their ratio in the total area of the street view image is calculated, thus obtaining the ratio of different types of pixels (SVR). C∈{Car,Truck,Bike,Person,Tree,...,Building} ).
[0113] By calculating the ratio of various types of pixels in street view images, a quantitative basis that is more in line with the characteristics of actual scenes is provided for the reverse prediction model of target air pollutants, thereby effectively improving the accuracy and reliability of prediction results.
[0114] Optionally, when the first type of pixel is a target vehicle, the ratio of the first type of pixel in the total number of pixels in the street view image is calculated based on the number of the first type of pixel, including:
[0115] Determine the number of vehicles of the first type in the street view image; wherein, the first type of vehicle refers to motor vehicles powered by battery energy storage.
[0116] The adjusted total number of vehicles is obtained by subtracting the number of the first type of vehicles from the total number of vehicles in the street view image.
[0117] The ratio of the remaining vehicle types is obtained by dividing the number of other types of vehicles by the adjusted total number of vehicles; wherein, the remaining vehicle types refer to the types of vehicles other than the first type of vehicles in the target vehicles.
[0118] In this scheme, the first category of vehicles refers to motor vehicles that use battery energy storage as their power source. For example, the first category of vehicles can be pure electric vehicles.
[0119] Furthermore, the pollutant emissions from Category I vehicles contribute relatively little during operation and can interfere with the final prediction accuracy; therefore, Category I vehicles need to be excluded to correct the ratio of other vehicle types. Thus, for the visual feature S in MOD... MOD The correction process involves the following steps: First, based on the Baidu Appollo open dataset and YOLOv8, YOLOv10, or YOLOv11 models, identify the emission types of cars and trucks on the road. Second, classify pure electric vehicles and non-electric vehicles according to license plate color. Finally, use pure electric vehicles as a mask and correct the ratio of non-electric vehicles (Cars) to trucks (Trucks).
[0120] Specifically, count the number of vehicles of the first type in the street view image; subtract the number of vehicles of the first type from the total number of vehicles in the street view image to calculate the adjusted total number of vehicles. Divide the number of vehicles of the other types by the adjusted total number of vehicles to obtain the ratio of the corresponding vehicle type.
[0121] By correcting the number of vehicles in the first category, the prediction accuracy and reliability of the target air pollutant reverse prediction model are effectively improved.
[0122] S430. Associate each time point within the second time period with the corresponding pollutant emission concentration characteristics, meteorological characteristics, street scene characteristics, and driving condition characteristics to construct mobile observation data pairs.
[0123] Specifically, MOD feature matrix The columns include: MOD Street View Features (S MOD ), Driving condition characteristics (V) MOD ), meteorological characteristics (M) MOD It includes pollutant emission concentration characteristics (P0.05) measured by mobile sensors. MOD The above-mentioned type of feature vectors in the database. The dependent variable matrix is constructed along the time axis. Finally, it is stored in a PostgreSQL database table. The following sections are used for training.
[0124] Furthermore, the MOD feature matrix establishes a mapping relationship between street view features, average information, meteorological information, and the emission concentrations of different pollutants. A large number of mobile observations are used to collect greenhouse gas emission samples, and outliers are removed to obtain an effective sample size. Through the mapping relationship between multimodal street view features, vehicle speed, meteorological information, and emission concentration levels, the powerful learning ability of artificial intelligence models on nonlinear correlations is used for regression fitting. The concentration values of different air pollutants are automatically predicted from the multimodal features.
[0125] In this scheme, after acquiring the first geographic data and completing the construction of the city geographic information database (GID), the following steps are also included: extracting features from meteorological information to obtain meteorological features; extracting features from street view images to obtain street view features; and extracting features from driving condition information to obtain driving condition features. Among these, street view features are presented through the proportion of different types of pixels. Each time point within the first time period is associated with the corresponding meteorological features, street view features, and driving condition features.
[0126] Specifically, visual artificial intelligence is used to extract GID features and form a multimodal feature database. GID Feature Matrix The columns include: GID Street View Features (S GID ), Driving condition characteristics (V) GID ), meteorological characteristics (M) GID ), the above-mentioned type of feature vectors in the database ( The dependent variable matrix is constructed along the time axis. Finally, it is stored in a PostgreSQL database table. It will be used for prediction later.
[0127] Finally, the GID feature matrix This is used to obtain the feature matrix for historical time periods, and compared with the MOD feature matrix. Maintaining consistency across the column dimensions of the data table (for reverse prediction of spatiotemporal variations in emissions of various air pollutants).
[0128] Furthermore, Figure 6 This is a schematic diagram of the visual artificial intelligence multimodal feature extraction and fusion process provided in Embodiment 3 of this application, as shown below. Figure 6 As shown, a city geographic information database (GID) and a mobile observation database (MOD) are constructed based on multi-source urban geographic data. Visual artificial intelligence is used to extract features from GID and MOD and form a multimodal feature database. GID feature matrix. The columns include: GID Street View Features (S GID ), Driving condition characteristics (V) GID ), meteorological characteristics (M) GID Three types of features. MOD feature matrix. The columns include: MOD Street View Features (S MOD ), Driving condition characteristics (V) MOD ), meteorological characteristics (M) MOD It includes pollutant emission concentration characteristics (P0.05) measured by mobile sensors. MOD The above-mentioned type of feature vectors in the database ( as well as The dependent variable matrix is constructed along the time axis. Finally, it is stored in a PostgreSQL database table. and These are then used for training and prediction, respectively.
[0129] The technical solution of this invention involves collecting the emission concentration of a second pollutant within a target area corresponding to a second time period through an air quality module in an onboard device; collecting meteorological information within the target area corresponding to a second time period through a meteorological module in the onboard device; collecting street view images within the target area corresponding to a second time period through a street view device in the onboard device; and collecting driving condition information within the target area corresponding to a second time period through a positioning module in the onboard device. Features are extracted from the second pollutant emission concentration to obtain pollutant emission concentration features; features are extracted from the meteorological information to obtain meteorological features; features are extracted from the street view images to obtain street view features; and features are extracted from the driving condition information to obtain driving condition features. Each time point within the second time period is associated with the corresponding pollutant emission concentration features, meteorological features, street view features, and driving condition features to construct mobile observation data pairs. By implementing this technical solution, key pollution sources such as high-emission vehicles, industrial pollution sources, and road dust are accurately identified, the spatiotemporal diffusion field of pollutants is reconstructed, and a reverse prediction model for target air pollutants is constructed. This model analyzes the correlation between traffic flow and pollutant concentration, quantifies the emission contribution of different vehicle types, and provides an objective basis for precise pollution control.
[0130] In this plan, Figure 7 This is a schematic diagram of the reverse prediction of air pollutant emissions provided in Embodiment 3 of this application, as shown below. Figure 7 As shown, a multi-source sensing database is constructed: a GID database is built through urban geographic data collection via the Internet, and an MOD database is built through mobile observation data collection of air pollutants. Visual artificial intelligence multimodal feature processing and fusion: feature extraction, feature registration, and feature fusion are performed on the GID and MOD databases. Spatiotemporal inverse prediction analysis and visualization of air pollutants: training and validation of the target air pollutant inverse prediction model, prediction and storage of the target air pollutant inverse prediction model, and spatiotemporal evolution analysis and visualization of pollutants.
[0131] Based on multi-source street view imagery, meteorological parameters, and vehicle speed conditions, a GeoELM (Geometric Elastic Array Model) for predicting target air pollutants was constructed to perform inverse analysis and dynamic backtracking of four key urban traffic pollutants. This aims to address the current technical bottleneck of lacking intelligent methods for predicting the spatiotemporal dynamics of continuous air pollutants. By integrating multi-source heterogeneous information such as vehicle-mounted mobile sensor data, massive street view imagery, and micro-meteorological data, and employing a Transformer-based visual artificial intelligence technology system (Mask2Former + GeoELM), accurate pollutant emission prediction (goodness-of-fit R² = 95-99%) and high spatial resolution (50m road segment) were achieved. This effectively inverted the spatiotemporal dynamic evolution of four typical air pollutants, providing decision support for optimizing traffic control strategies. This method can be used for historical observation data supplementation, emission reduction effectiveness assessment, and secondary pollution prevention. It helps to quickly locate pollutant hotspots and has significant application value for improving emission reduction policies for complex built-up environments.
[0132] Example 4
[0133] Figure 8 This is a schematic diagram of the air pollutant emission reverse prediction device provided in Embodiment 4 of the present invention. Figure 8 As shown, the device includes:
[0134] The geographic data acquisition unit 810 is used to acquire first geographic data generated within the target area corresponding to the first time period; wherein, the geographic data is data information related to pollutant emission concentration;
[0135] The pollutant emission concentration output unit 820 is used to input the first geographical data into the target air pollutant reverse prediction model and output the first pollutant emission concentration corresponding to the first time period.
[0136] The target air pollutant reverse prediction model is trained by collecting mobile observation data generated in the target area during the second time period using vehicle-mounted equipment during movement; the mobile observation data consists of second geographical data and second pollutant emission concentrations; the first geographical data was collected earlier than the second geographical data.
[0137] Optionally, the device includes:
[0138] A training set partitioning unit is used to divide the mobile observation data into a first training set and a second training set.
[0139] The parameter configuration determination unit for the base learner is used to determine the parameter configuration of the base learner in the air pollutant back prediction model to be trained.
[0140] The base learner training result obtaining unit is used to input the first training set into the base learner for training, and obtain the base learner training result;
[0141] The parameter configuration determination unit for the meta-learner is used to construct the parameter configuration of the meta-learner in the air pollutant inverse prediction model to be trained.
[0142] The target air pollutant inverse prediction model obtaining unit is used to train the meta-learner by taking the training results of the base learner and the second training set as input, so as to obtain the target air pollutant inverse prediction model.
[0143] Optionally, the base learner includes at least two of the following: random forest, extreme gradient boosting machine, extreme random forest, and lightweight gradient boosting machine.
[0144] The meta-learner includes regularized regression.
[0145] Optionally, the parameter configuration determination unit for the base learner is specifically used for:
[0146] Determine the initial hyperparameters of the base learner in the air pollutant back prediction model to be trained; wherein, the initial hyperparameters refer to the parameters set before training the base learner;
[0147] Calculate the median error of the base learner under the initial hyperparameters;
[0148] Calculate the error gradient based on the median error;
[0149] The initial hyperparameters are adjusted based on the error gradient to obtain the adjusted hyperparameters, and the adjusted hyperparameters are used as the parameter configuration of the base learner.
[0150] Optionally, the geographic data includes meteorological information, street view images, and driving condition information, and the device further includes:
[0151] The data acquisition unit is used to collect the emission concentration of the second pollutant generated in the target area corresponding to the second time period through the air quality module in the vehicle-mounted device; and to collect meteorological information generated in the target area corresponding to the second time period through the meteorological module in the vehicle-mounted device; and to collect street view images generated in the target area corresponding to the second time period through the street view device in the vehicle-mounted device; and to collect driving condition information generated in the target area corresponding to the second time period through the positioning module in the vehicle-mounted device.
[0152] Optionally, the device further includes:
[0153] The feature extraction unit is used to extract features from the emission concentration of the second pollutant to obtain pollutant emission concentration features; and to extract features from the meteorological information to obtain meteorological features; and to extract features from the street view image to obtain street view features; and to extract features from the driving condition information to obtain driving condition features; wherein the street view features are presented by the proportion of different types of pixels.
[0154] The mobile observation data pair construction unit is used to associate each time point within the second time period with the corresponding pollutant emission concentration characteristics, meteorological characteristics, street scene characteristics, and driving condition characteristics to construct mobile observation data pairs.
[0155] Optional, feature extraction unit, specifically used for:
[0156] Feature extraction is performed on the street view image to obtain the number of first-type pixels and the area of second-type pixels; wherein, the first-type pixels refer to pixels corresponding to ground features with discrete, countable features; and the second-type pixels refer to pixels corresponding to ground features with continuous, areal features.
[0157] Based on the number of first-type pixels, calculate the ratio of the first-type pixels to the total number of pixels in the street view image; and based on the area of the second-type pixels, calculate the ratio of the area of the second-type pixels to the total area of the street view image.
[0158] Optionally, the feature extraction unit is also used for:
[0159] Determine the number of vehicles of the first type in the street view image; wherein, the first type of vehicle refers to motor vehicles powered by battery energy storage.
[0160] The adjusted total number of vehicles is obtained by subtracting the number of the first type of vehicles from the total number of vehicles in the street view image.
[0161] The ratio of the remaining vehicle types is obtained by dividing the number of other types of vehicles by the adjusted total number of vehicles; wherein, the remaining vehicle types refer to the types of vehicles other than the first type of vehicles in the target vehicles.
[0162] The air pollutant emission reverse prediction device provided in the embodiments of the present invention can execute the air pollutant emission reverse prediction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0163] Example 5
[0164] Figure 9A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, capable of local processing via edge computing devices (such as Jetson, Raspberry Pi, Arduino, etc.) or microcontrollers; or capable of transmitting data to a server or workstation via local storage media for subsequent post-processing and analysis after observation tasks are completed. Examples include laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0165] like Figure 9 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0166] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0167] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as air pollutant emission reverse prediction methods.
[0168] In some embodiments, the air pollutant emission reverse prediction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the air pollutant emission reverse prediction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the air pollutant emission reverse prediction method by any other suitable means (e.g., by means of firmware).
[0169] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0170] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0171] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0173] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0174] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0175] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0176] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for reverse prediction of air pollutant emissions, characterized in that, include: Acquire the first geographic data generated within the target area corresponding to the first time period; wherein, the geographic data is data information related to pollutant emission concentration; The first geographic data is input into the target air pollutant reverse prediction model, and the emission concentration of the first pollutant corresponding to the first time period is output. The target air pollutant reverse prediction model is trained by collecting mobile observation data generated in the target area during the second time period using vehicle-mounted equipment during movement; the mobile observation data consists of second geographical data and second pollutant emission concentrations; the first geographical data was collected earlier than the second geographical data.
2. The method according to claim 1, characterized in that, The training process of the target air pollutant reverse prediction model includes: The mobile observation data is divided into a first training set and a second training set; Determine the parameter configuration of the base learner in the air pollutant back prediction model to be trained; The first training set is input into the base learner for training to obtain the base learner training result; Construct the parameter configuration of the meta-learner in the air pollutant inverse prediction model to be trained; The training results of the base learner and the second training set are used as input to train the meta learner, thereby obtaining the target air pollutant reverse prediction model.
3. The method according to claim 2, characterized in that, The base learner includes at least two of the following: random forest, extreme gradient boosting machine, extreme random forest, and lightweight gradient boosting machine. The meta-learner includes regularized regression.
4. The method according to claim 2, characterized in that, Determine the parameter configuration of the base learner in the air pollutant inverse prediction model to be trained, including: Determine the initial hyperparameters of the base learner in the air pollutant back prediction model to be trained; wherein, the initial hyperparameters refer to the parameters set before training the base learner; Calculate the median error of the base learner under the initial hyperparameters; Calculate the error gradient based on the median error; The initial hyperparameters are adjusted based on the error gradient to obtain the adjusted hyperparameters, and the adjusted hyperparameters are used as the parameter configuration of the base learner.
5. The method according to claim 2, characterized in that, The geographic data includes meteorological information, street view images, and driving condition information; the method further includes: The system collects the emission concentration of the second pollutant generated in the target area corresponding to the second time period through the air quality module in the vehicle-mounted device; and collects meteorological information generated in the target area corresponding to the second time period through the meteorological module in the vehicle-mounted device; and collects street view images generated in the target area corresponding to the second time period through the street view device in the vehicle-mounted device; and collects driving condition information generated in the target area corresponding to the second time period through the positioning module in the vehicle-mounted device.
6. The method according to claim 5, characterized in that, After collecting motion observation data generated within the target area corresponding to the second time period based on the vehicle-mounted equipment, the method further includes: The emission concentration of the second pollutant is subjected to feature extraction to obtain pollutant emission concentration features; the meteorological information is subjected to feature extraction to obtain meteorological features; the street view image is subjected to feature extraction to obtain street view features; and the driving condition information is subjected to feature extraction to obtain driving condition features; wherein, the street view features are presented by the proportion of different types of pixels. By associating each time point within the second time period with the corresponding pollutant emission concentration characteristics, meteorological characteristics, street scene characteristics, and driving condition characteristics, mobile observation data pairs are constructed.
7. The method according to claim 6, characterized in that, Feature extraction is performed on the street view image to obtain street view features, including: Feature extraction is performed on the street view image to obtain the number of first-type pixels and the area of second-type pixels; wherein, the first-type pixels refer to pixels corresponding to ground features with discrete, countable features; and the second-type pixels refer to pixels corresponding to ground features with continuous, areal features. Based on the number of first-type pixels, calculate the ratio of the first-type pixels to the total number of pixels in the street view image; and based on the area of the second-type pixels, calculate the ratio of the area of the second-type pixels to the total area of the street view image.
8. The method according to claim 7, characterized in that, When the first type of pixel is a target vehicle, the ratio of the first type of pixel in the total number of pixels in the street view image is calculated based on the number of the first type of pixel, including: Determine the number of vehicles of the first type in the street view image; wherein, the first type of vehicle refers to motor vehicles powered by battery energy storage. The adjusted total number of vehicles is obtained by subtracting the number of the first type of vehicles from the total number of vehicles in the street view image. The ratio of the remaining vehicle types is obtained by dividing the number of other types of vehicles by the adjusted total number of vehicles; wherein, the remaining vehicle types refer to the types of vehicles other than the first type of vehicles in the target vehicles.
9. An air pollutant emission reverse prediction device, characterized in that, include: A geographic data acquisition unit is used to acquire first geographic data generated within a target area corresponding to a first time period; wherein, the geographic data is data information related to pollutant emission concentration; The pollutant emission concentration output unit is used to input the first geographical data into the target air pollutant reverse prediction model and output the first pollutant emission concentration corresponding to the first time period. The target air pollutant reverse prediction model is trained by collecting mobile observation data generated in the target area during the second time period using vehicle-mounted equipment during movement; the mobile observation data consists of second geographical data and second pollutant emission concentrations; the first geographical data was collected earlier than the second geographical data.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the air pollutant emission reverse prediction method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the air pollutant emission reverse prediction method according to any one of claims 1-8.