Model construction method, air quality prediction method and electronic device
By training an air quality prediction model using the rate of change of an air quality numerical model, the problem of insufficient accuracy and interpretability of existing air quality prediction models is solved, achieving efficient and accurate prediction of air pollutant change processes and providing a real-time and interpretable solution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 3CLEAR SCI & TECH CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-26
AI Technical Summary
Existing air quality prediction models lack accuracy, reliability, and interpretability, and are computationally expensive, failing to meet the real-time forecasting needs of operational businesses.
The numerical model process rates of various changes in air quality output from the numerical air quality model are used as labeled data to train the target air quality prediction model. By training each process module in parallel, the accuracy of the basic capabilities of the modules is ensured, and the coordination between the modules is optimized to output the process rates of various changes in air pollutants.
It improves the accuracy and interpretability of air quality prediction model results, provides accurate basis for pollution source analysis, reduces computational costs, and meets the demand for real-time prediction.
Smart Images

Figure CN122087449A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence forecasting and early warning technology for air quality, specifically to a model building method, an air quality prediction method, and an electronic device. Background Technology
[0002] With the acceleration of industrialization and urbanization, air pollution has become one of the key challenges affecting public health and the ecological environment. Air quality forecasting is the core support for air quality management and atmospheric environmental management. Air quality prediction is the prediction of the concentration of air pollutants in the future. Currently, artificial intelligence network models are used to predict air quality; however, the accuracy, reliability, consistency, and interpretability of the prediction results output by these models still have shortcomings. Summary of the Invention
[0003] The purpose of this application is to provide a model building method, an air quality prediction method, and an electronic device to improve the accuracy of the output results of the trained target air quality prediction model, improve the model training efficiency, and enable the prediction of the process rate and total concentration of various changes in air pollutants through the target air quality prediction model, thereby improving the reliability and interpretability of the prediction results.
[0004] To achieve the above objectives, firstly, this application provides a model construction method, the method comprising: Obtain a training sample set, wherein the training samples in the training sample set include first information and numerical model operational data. The first information includes the air pollutant concentration of the first grid area at time T, the air pollutant emission information and meteorological information of the first grid area at each time from time T to time T+N. The numerical model operational data includes the numerical model process rate of various change processes. The numerical model process rate is the process rate of the change process of air pollutants in the first grid area at multiple first times, which is output by the air quality numerical model based on the first information. The multiple first times include each time from time T+1 to time T+N. For each of the aforementioned change processes, the process rate of the numerical model of the change process in the training samples is used as labeled data to train the process module corresponding to the change process, thereby obtaining the first process module corresponding to the change process. Using the numerical model process rates of the various change processes in the training samples as labeled data, the air quality prediction model is trained to obtain the target air quality prediction model, which is composed of first process modules corresponding to the various change processes.
[0005] Optionally, the step of using the numerical model process rate of the change process in the training samples as labeled data to train the process module corresponding to the change process includes: The first information is input to the process module corresponding to the change process to obtain the first process rate output by the process module. The first process rate includes the process rate of the change process of air pollutants in the first grid area at the multiple first moments. Based on the first process rate, the numerical model process rate of the change process, and the first preset loss function, a first loss value is determined, and the process module is trained based on the first loss value.
[0006] Optionally, the numerical model service data further includes numerical model concentration information, which is the concentration information of air pollutants in the first grid area output by the air quality numerical model at the multiple first moments; The step of training the air quality prediction model using the numerical model process rates of the various change processes in the training samples as labeled data includes: The first information is input into the air quality prediction model to obtain the second process rate output by multiple first process modules, and the first concentration information of air pollutants in the first grid area at multiple first moments output by the air quality prediction model. The process loss value is determined based on the second process rate, the numerical model process rate, and the second preset loss function; The concentration loss value is determined based on the first concentration information, the numerical model concentration information, and the third preset loss function; The target loss value is determined based on the process loss value and the concentration loss value; The air quality prediction model is trained based on the target loss value.
[0007] Optionally, determining the process loss value based on the second process rate, the numerical model process rate, and the second preset loss function includes: Determine the weights of each of the various change processes; For each of the aforementioned change processes, the loss value corresponding to the change process is determined based on the second process rate and the numerical model process rate corresponding to the change process, as well as the second preset loss function and the weight of the change process. The process loss value is determined based on the loss values corresponding to the various change processes.
[0008] Optionally, the various change processes include emission processes, horizontal advection processes, vertical advection processes, horizontal diffusion processes, vertical diffusion processes, chemical reaction processes, wet deposition processes, and dry deposition processes; the meteorological information includes wind speed, atmospheric stability, temperature, humidity, solar radiation intensity, boundary layer height, precipitation, and turbulence intensity; Determining the weights of each of the multiple change processes includes: Based on the wind speed, determine the respective weights of the horizontal advection process and the vertical advection process; The weights of the horizontal diffusion process and the vertical diffusion process are determined based on the atmospheric stability. The weights of the chemical reaction process are determined based on the temperature, humidity, and solar radiation intensity. The weights of the emission processes are determined based on the boundary layer height. The weights of the wet deposition process are determined based on the precipitation and humidity. The weights of the dry deposition process are determined based on the turbulence intensity, wind speed, atmospheric stability, humidity, underlying surface characteristics of the first grid region, and the properties of the air pollutants.
[0009] Optionally, determining the weights of each of the multiple change processes includes: For each of the aforementioned change processes, the confidence level of the output of the first process module corresponding to the change process is obtained, and the confidence level is used to characterize the credibility of the second process rate output by the first process module; the weight of the change process is determined based on the confidence level.
[0010] Optionally, the various processes include emission processes, sedimentation processes, and chemical reaction processes; The step of training the air quality prediction model using the numerical model process rates of the various change processes in the training samples as labeled data further includes: Based on the first concentration information, the second process rate output by the first process module corresponding to the emission process, the second process rate output by the first process module corresponding to the sedimentation process, the second process rate output by the first process module corresponding to the chemical process, and the fourth preset loss function and the fifth preset loss function, the physical loss value is determined. The fourth preset loss function is used to determine the mass loss value of the air pollutant, and the fifth preset loss function is used to determine the elemental loss value of the elements contained in the air pollutant. Determining the target loss value based on the process loss value and the concentration loss value includes: The target loss value is obtained by weighting the process loss value, the concentration loss value, and the physical loss value.
[0011] Secondly, this application provides an air quality prediction method, the method comprising: Obtain second information, which includes the air pollutant concentration, air pollutant emission information and meteorological information of the second grid area at time T', as well as the predicted air pollutant emission information and meteorological information of the second grid area at multiple second times, wherein the multiple second times include each time from time T'+1 to time T'+N. The second information is input into the target air quality prediction model to obtain the concentration of air pollutants in the second grid area at multiple second times as output by the target air quality prediction model, and the third process rate output by multiple process modules in the target air quality prediction model. The multiple process modules correspond one-to-one with multiple change processes of air pollutants. The third process rate includes the process rate of the change process of air pollutants in the second grid area at multiple second times. The target air quality prediction model is trained according to the model construction method provided in the first aspect of this application.
[0012] Optionally, the multiple change processes include many of the following: emission process, horizontal advection process, vertical advection process, horizontal diffusion process, vertical diffusion process, gas phase chemical reaction process, liquid phase chemical reaction process, dry deposition process, wet deposition process, inorganic aerosol thermodynamic equilibrium process, and secondary organic aerosol process.
[0013] Thirdly, this application provides an electronic device, comprising: A memory on which computer programs are stored; A processor is configured to execute the computer program in the memory to implement the steps of the model building method provided in the first aspect of this application.
[0014] Fourthly, this application provides an electronic device, comprising: A memory on which computer programs are stored; A processor is configured to execute the computer program in the memory to implement the steps of the air quality prediction method provided in the second aspect of this application.
[0015] By employing the above technical solution, the process rates of each change process output by the air quality numerical model are used as labeled data for training. This enables the target air quality prediction model to output the process rates of each change process of air pollutants during the application phase, while ensuring the accuracy of the output process rates. Compared with the method of outputting a single air pollutant concentration in related technologies, this can improve the interpretability and credibility of the prediction results. The process rates of each change process output by the target air quality prediction model can be used as source apportionment input data for pollution source analysis, thereby providing an accurate basis for air quality control decisions.
[0016] Furthermore, for each change process, the corresponding process module is trained, enabling parallel training of multiple process modules, improving training efficiency, and ensuring the accuracy of the basic simulation capabilities of each process module. After training each process module, the air quality prediction model composed of the first process modules corresponding to various change processes can be trained to obtain the target air quality prediction model. This optimizes the coordination between the various first process modules, ensuring the accuracy of the air pollutant concentration information output by the trained target air quality numerical model.
[0017] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the present application and form part of the specification. They are used together with the following detailed description to explain the present application, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a model construction method according to an exemplary embodiment.
[0019] Figure 2 This is a flowchart illustrating an exemplary method for training process modules corresponding to a change process.
[0020] Figure 3 This is a flowchart illustrating an exemplary method for training an air quality prediction model.
[0021] Figure 4 This is a flowchart illustrating an exemplary method for determining process loss values.
[0022] Figure 5 This is a schematic diagram illustrating model training using a time-parallel approach.
[0023] Figure 6 This is a flowchart illustrating an exemplary air quality prediction method.
[0024] Figure 7This is a schematic diagram illustrating an exemplary air quality prediction method.
[0025] Figure 8 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0026] The specific embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this application.
[0027] Numerical air quality models based on physicochemical mechanisms can provide relatively accurate air quality predictions. Examples of such models include CMAQ (Community Multiscale Air Quality Modeling System) and NAQPMS (Nested Air Quality Prediction Modeling System). However, while these models accurately simulate the lifecycle of pollutants in the atmosphere by solving complex systems of partial differential equations, they are computationally expensive and time-consuming, failing to meet the real-time forecasting requirements of operational applications.
[0028] In related technologies, artificial intelligence models can be used to predict air quality, thereby improving the efficiency of air quality prediction and shortening the required calculation time. However, the accuracy of the prediction results output by the models in these technologies is insufficient, and there are interpretability issues.
[0029] For example, when training artificial intelligence models in related technologies, their training paradigm relies on fitting the final air pollutant concentration field output by the numerical air quality model. Therefore, it cannot guarantee the authenticity and reliability of the simulation of internal air pollutant change processes, which limits the interpretability of model prediction results and their practical value in applications such as process contribution analysis.
[0030] Specifically, the artificial intelligence model in this technology designs a "basic simulator" whose structure follows the atmospheric chemical transport equation and is decomposed into five neural network modules, simulating five processes: emission (E), horizontal transport (T), diffusion (D), chemical reaction (C), and deposition (S). The air pollutant concentration at the next moment is determined by the sum of the current air pollutant concentration and the contributions of the five processes. However, the training objective of this artificial intelligence model is to match the final air pollutant concentration output by the model with the output of traditional numerical models. The parameters of the five neural network modules are adjusted through backpropagation by fitting the final air pollutant concentration. This may lead the model to learn non-mechanistic concentration changes in order to match the final air pollutant concentration. For example, the chemical reaction module may incorrectly "compensate" for concentration changes that should be handled by the horizontal transport module. In other words, the simulation of each process module within the model may deviate from physical reality, resulting in insufficient credibility and interpretability of its output predictions.
[0031] In addition, other network models are used in related technologies for air quality prediction, such as the improved deep belief network model, which uses a diffusion coefficient to calculate the theoretical diffusion contribution. However, this approach simplifies the complex physical process and differs from the actual diffusion process.
[0032] This application provides a model building method, an air quality prediction method, and an electronic device, which improve the accuracy of the output results of the trained target air quality prediction model, improve the model training efficiency, and enable the prediction of the process rate and total concentration of various changes in air pollutants through the target air quality prediction model, thereby improving the reliability and interpretability of the prediction results.
[0033] Figure 1 This is a flowchart illustrating a model building method according to an exemplary embodiment, which can be applied to a server, such as... Figure 1 As shown, the model construction method includes steps 11 to 13.
[0034] Step 11: Obtain the training sample set.
[0035] The training samples in the training sample set include first information and numerical model operational data. The first information includes the air pollutant concentration of the first grid area at time T, the air pollutant emission information and meteorological information of the first grid area at each time from time T to time T+N. The numerical model operational data includes the numerical model process rate of various change processes. The numerical model process rate is the process rate of the change process of air pollutants in the first grid area at multiple first times, which is output by the air quality numerical model based on the first information. The multiple first times include each time from time T+1 to time T+N.
[0036] There are many types of air pollutants, such as PM2.5 (particulate matter with a diameter of 2.5 micrometers or less), PM10 (particulate matter with a diameter of 10 micrometers or less), NO2 (nitrogen dioxide), SO2 (sulfur dioxide), CO (carbon monoxide), and O3 (ozone).
[0037] Air pollutant emission information for the first grid area at a given time may include the air pollutant emission concentration in the first grid area from the previous time to the present time. For example, with an interval of 1 hour between two adjacent times, the air pollutant emission information may include the air pollutant emission concentration in the first grid area within one hour. The air pollutant emission information may also include the emission concentration of precursors to air pollutants.
[0038] Meteorological information for the first grid area at a given moment may include wind speed, atmospheric stability, temperature, humidity, solar radiation intensity, boundary layer height, precipitation, precipitation intensity, turbulence diffusion coefficient, cloud water content, ion concentration, and rainwater pH value. For example, precipitation may be the amount of precipitation within one hour.
[0039] The training sample set includes multiple training samples, which may include data from different grid areas and at different times. For example, training sample 1 includes the air pollutant concentration of grid area 1 at 00:00 on XX / XX / XX / XX, and the air pollutant emission information and meteorological information for each time point from 00:00 on XX / XX / XX / XX to 00:00 on XX / XX / XX / XX. For instance, the interval between two adjacent times is 1 hour, and N can be 24. As another example, training sample 2 includes the air pollutant concentration of grid area 2 at 00:00 on YY / YY / XX, and the air pollutant emission information and meteorological information for each time point from 00:00 on YY / YY / XX to 00:00 on YY / YY / XX. XX and YY are omitted dates.
[0040] The first grid region mentioned above is used to distinguish it from the second grid region mentioned below, and is not intended to be limited to a fixed grid region. Similarly, the first moment mentioned above is used to distinguish it from the second moment mentioned below, and is not intended to be limited to a fixed moment. The first grid region can be a geographical area divided according to preset rules, and there is no limit to the size of the first grid region.
[0041] For example, various processes may include: emission processes, horizontal advection processes, vertical advection processes, horizontal diffusion processes, vertical diffusion processes, gas-phase chemical reaction processes, liquid-phase chemical reaction processes, dry deposition processes, wet deposition processes, inorganic aerosol thermodynamic equilibrium processes, and secondary organic aerosol processes. The principles of each process can be found in relevant technologies.
[0042] The process rate in a numerical model refers to the rate of change of air pollutants in the first grid region based on the initial information output of the air quality numerical model at multiple first time points. Here, process rate refers to the rate of change of a physical or chemical process per unit time; it can be understood as the amount of change in air pollutant concentration, or the contribution of the change process to the air pollutant concentration.
[0043] Since process rate refers to the amount of change, the process rate at a certain moment can be understood as the amount of change in pollutant concentration from the previous moment to the current moment. Taking horizontal advection as an example, it can be understood as the contribution of the horizontal advection process to the concentration of air pollutants within 1 hour, that is, the hourly contribution of the horizontal advection process to the concentration of air pollutants.
[0044] For example, the numerical model process rate for emission processes is denoted as ΔC_emission; the numerical model process rate for horizontal advection processes is denoted as ΔC_advhor; the numerical model process rate for vertical advection processes is denoted as ΔC_advvert; the numerical model process rate for horizontal diffusion processes is denoted as ΔC_difhor; the numerical model process rate for vertical diffusion processes is denoted as ΔC_difvert; the numerical model process rate for gas-phase chemical reaction processes is denoted as ΔC_gaschemistry; the numerical model process rate for dry deposition processes is denoted as ΔC_drydeposition; the numerical model process rate for wet deposition processes is denoted as ΔC_wetdeposition; the numerical model process rate for inorganic aerosol thermodynamic equilibrium processes is denoted as ΔC_isorr; the numerical model process rate for secondary organic aerosol processes is denoted as ΔC_soa; and the numerical model process rate for liquid-phase chemical reaction processes is denoted as ΔC_aqueous.
[0045] The numerical model operational data may also include numerical model concentration information, which is the concentration information of air pollutants in the first grid area output by the air quality numerical model at multiple first moments. In one embodiment, the numerical model concentration information may be the concentration change, that is, the change in air pollutant concentration from the previous moment to the current moment. In this embodiment, the change in air pollutant concentration may be the sum of the numerical model process rates of each change process.
[0046] In another embodiment, the concentration information in the numerical model can also be the concentration of air pollutants in the first grid region at a first time. The concentration of air pollutants at the first time can be obtained by adding the concentration of air pollutants at the previous time to the change in pollutant concentration between the previous time and the first time. As an example, the concentration of air pollutants at time T is used as the initial concentration. After obtaining the change in pollutant concentration from time T to time T+1, the concentration of air pollutants at time T+1 is added to the concentration of air pollutants at time T. The concentration of air pollutants at time T+2 is obtained similarly by adding the concentration of air pollutants at time T+1 to the concentration of air pollutants at time T+2.
[0047] In this process, the first information can be pre-input into the air quality numerical model to obtain the numerical model business data output by the air quality numerical model.
[0048] Numerical air quality models can be CMAQ or NAQPMS. These models accurately simulate the lifecycle of pollutants in the atmosphere by solving complex systems of partial differential equations. They also quantify the specific contribution of each process—transportation, chemical processes, and deposition—to pollutants using built-in process analysis tools (such as IPR (Integrated Process Rates)). In other words, numerical air quality models not only provide concentration results for air pollutants (numerical model concentration information) but also output the contribution of each process at various time points (numerical model process rate) through integrated process rate analysis tools.
[0049] Step 12: For each change process, the process rate of the numerical model of the change process in the training sample is used as the labeled data to train the process module corresponding to the change process, so as to obtain the first process module corresponding to the change process.
[0050] Step 13: Using the numerical model process rates of various change processes in the training samples as labeled data, train the air quality prediction model to obtain the target air quality prediction model. The air quality prediction model consists of the first process modules corresponding to various change processes.
[0051] The model training process in this application can be divided into two stages. In the first stage, the process module corresponding to each change process is trained independently and in parallel using the numerical model process rate output by the air quality numerical model, so as to ensure the accuracy of the basic simulation capability of each process module.
[0052] The second stage involves assembling the first-process modules corresponding to the various changes trained in the first stage to form an air quality prediction model. The air quality prediction model is then trained using the numerical model process rate and concentration information output by the numerical air quality model. This training process can be fine-tuned. The second stage optimizes the coordination between the first-process modules while ensuring the accuracy of the prediction results of each module, thereby ensuring the accuracy of the air pollutant concentrations output by the target air quality prediction model.
[0053] In the first stage, parallel training of the process modules corresponding to each change process decouples the individual process modules, resulting in high training efficiency and speed. In one embodiment, to further improve training speed, multiple process modules can be trained on different processing units, such as graphics processing units (GPUs). For example, the emission process module, horizontal advection process module, vertical advection process module, horizontal diffusion process module, and vertical diffusion process module, as emission / advection branches, can be trained on a single processing unit. The gas-phase chemical reaction process module, liquid-phase chemical reaction process module, inorganic aerosol thermodynamic equilibrium process module, and secondary organic aerosol process module, as chemical process branches, can be trained on a single processing unit. The dry sedimentation process module and wet sedimentation process module, as sedimentation process branches, can be trained on a single processing unit.
[0054] Therefore, the network models in related technologies do not utilize the contribution of air pollutant change processes output by the air quality numerical model for training, resulting in insufficient reliability and interpretability of the output prediction results. In this application, the process rates of each change process output by the air quality numerical model are used as labeled data for training. This enables the target air quality prediction model to output the process rates of each change process of air pollutants during the application phase, ensuring the accuracy of the output process rates. Compared to the method of outputting a single air pollutant concentration in related technologies, this improves the interpretability and reliability of the prediction results. The process rates of each change process output by the target air quality prediction model can be used as source apportionment input data for pollution source analysis, thereby providing accurate basis for air quality control decisions.
[0055] Furthermore, for each change process, the corresponding process module is trained, enabling parallel training of multiple process modules, improving training efficiency, and ensuring the accuracy of the basic simulation capabilities of each process module. After training each process module, the air quality prediction model composed of the first process modules corresponding to various change processes can be trained to obtain the target air quality prediction model. This optimizes the coordination between the various first process modules, ensuring the accuracy of the air pollutant concentration information output by the trained target air quality numerical model.
[0056] In this application, different machine learning sub-models can be used for each process module.
[0057] For example, the horizontal advection process module can employ 1D-CNN (One-Dimensional Convolutional Neural Network) or ConvLSTM (Convolutional Long Short-Term Memory). The CNN filter captures wind direction and speed patterns. The input to the horizontal advection process module includes initial meteorological information, specifically the U and V components of wind speed. The U component represents the vector magnitude of wind speed in the east-west horizontal direction, and the V component represents the vector magnitude of wind speed in the north-south horizontal direction. The output of the horizontal advection process module is the contribution of the horizontal advection process to air pollutants, which is also known as the process rate.
[0058] The vertical advection process module can employ either 1D-CNN or PINN (Physics-Informed Neural Networks). PINN, in particular, can directly embed physical formulas, offering strong interpretability. The input to the vertical advection process module includes first information, which consists of the W component of wind speed, representing the vector magnitude of wind speed in the vertical direction, and the air pollutant concentration in the first grid region at time T, including the concentration along the vertical profile. The output of the vertical advection process module includes the change in air pollutant concentration after being transported by the vertical advection process, i.e., the contribution of the vertical advection process to air pollutants.
[0059] This paper presents a schematic diagram of the physical formula embedded in the PINN in the vertical advection process module of this application. The governing equations for the advection part are shown in the following formula (1): (1) Among them, C iCi represents the concentration of the i-th air pollutant, in μg / m³. When calculating the change in pollutant concentration at a given moment, Ci can be the air pollutant concentration at the previous moment. (x, y, z) represent the three-dimensional coordinates of a point in the grid region, where x is the east-west coordinate, y is the north-south coordinate, and z is the vertical height coordinate. This represents the divergence operator. This represents the velocity vector of the fluid.
[0060] The physical formula for the vertical advection process is shown in equation (2): (2) The discrete scheme adopts a conformal upwind scheme with the following boundary conditions: ground: W=0; top layer: entrainment flux parameterization. The formula for the entrainment flux E (unit μg / m² / s) is shown in formula (3): (3) in, The value represents the rate of concentration change caused by vertical advection (in μg / m³ / s). C represents the concentration of air pollutants. W represents the vertical component of wind speed (in m / s), with positive values indicating upward motion and negative values indicating downward motion. It represents the vertical flux divergence, which represents the net inflow and outflow of pollutants per unit volume. This indicates the entrainment velocity (unit: m / s), which is the exchange velocity between the top of the boundary layer and the free atmosphere. This indicates the concentration in the free atmosphere above the boundary layer (unit: μg / m³). This represents the average concentration within the planetary boundary layer (unit: μg / m³). It represents the product.
[0061] The horizontal diffusion process module can employ a CNN or a diffusion model. The input to the horizontal diffusion process module includes first information, where the input meteorological information includes the turbulent diffusion coefficient, and the air pollutant concentration in the first grid region at time T includes the concentration gradient of the horizontal profile. The output of the horizontal diffusion process module includes the horizontal diffusion flux of air pollutants, i.e., the change in air pollutant concentration after being transported through the horizontal diffusion process.
[0062] The vertical diffusion module can employ either PINN or 1D-CNN. PINN can introduce boundary layer parameterization constraints, and the embedded physical equations in PINN ensure that the diffusion process conforms to physical laws such as Fick's law. The input to the vertical diffusion module includes initial information, such as boundary layer height, vertical turbulent diffusion coefficient, and temperature gradient in the vertical direction. The output of the vertical diffusion module includes the vertical diffusion flux of air pollutants, i.e., the change in air pollutant concentration after being transported through the vertical diffusion process.
[0063] This paper presents a schematic diagram of the physical formula for the PINN embedding in the vertical diffusion process module of this application. The physical formula for the vertical diffusion process is shown in the following formula (4): (4) Turbulence is parameterized as a stability function based on the Richardson number. Diffusion coefficient. It can be determined by formula (5): (5) in, This represents the rate of concentration change caused by vertical diffusion (unit: μg / m³ / s). The vertical turbulent diffusion coefficient (unit: m² / s) is a parameter that characterizes the ability of matter to diffuse in the vertical direction. This represents the turbulent diffusion term, which describes the diffusion transport caused by the concentration gradient. The reference diffusion coefficient or the diffusion coefficient under neutral conditions (unit: m² / s) is typically on the order of 0.1-10 m² / s within the boundary layer. f(Ri) represents the Richardson number stability function, used to describe the influence of atmospheric stability on turbulence. Ri is the Richardson number, a parameter used to measure atmospheric stability. The formula for Ri is shown in formula (6) below: (6) Where Ri>0 represents a stable stratification, Ri<0 represents an unstable stratification, and Ri≈0 represents a neutral stratification. g represents gravitational acceleration. The vector represents potential temperature, U represents the vector magnitude of wind speed in the horizontal east-west direction, and the V component represents the vector magnitude of wind speed in the horizontal north-south direction.
[0064] PBLH represents the planetary boundary layer height, measured in meters (m), indicating the thickness of the lower atmosphere that directly interacts with the Earth's surface. 'p' is an empirical exponent, typically between 1 and 2, describing how the diffusion coefficient varies with altitude. The temperature gradient in the vertical direction can be used to calculate... .
[0065] The dry deposition process module can employ either XGBoost (eXtreme Gradient Boosting) or a random forest model. XGBoost or random forest models can call their built-in functions to calculate the importance score of each feature. The feature importance ranking clearly shows the dominant role of land surface type or meteorological factors in deposition velocity. The input to the dry deposition process module can include first information, which may further include geographic feature information of a first grid area. This geographic feature information can be used to characterize the land surface type of the first grid area. The input meteorological information may include wind speed and humidity. The input to the dry deposition process module may also include the properties of air pollutants. The output of the dry deposition process module may include the dry deposition velocity of air pollutants, i.e., the change in the concentration of air pollutants after being transported through the dry deposition process.
[0066] The gas-phase chemical reaction process module can employ either a graph neural network (GNN) or a Transformer model. In a GNN model, the edge weights reflect the chemical reaction rate, while the attention mechanism in a Transformer model identifies key reaction pathways or precursors. The inputs to the gas-phase chemical reaction process module may include primary information, such as emission information (concentrations of air pollutant precursors) and meteorological information (temperature and solar radiation intensity). The outputs of the gas-phase chemical reaction process module may include the change in air pollutant concentration after the gas-phase chemical reaction process.
[0067] The liquid-phase chemical reaction process module can employ either a GNN or a LSTM, with the LSTM targeting the time dependence of the liquid-phase reaction. The inputs to the liquid-phase chemical reaction process module may include initial information, such as meteorological information including cloud water content, ion concentration, rainwater pH, and temperature. The outputs of the liquid-phase chemical reaction process module may include the change in air pollutant concentrations after the liquid-phase chemical reaction process.
[0068] The wet deposition process module can employ XGBoost or Random Forest models. These models can call their built-in functions to calculate the importance score of each feature; for example, feature importance can indicate that precipitation intensity is the most critical factor. The input to the wet deposition process module can include primary information, such as precipitation intensity and cloud shedding. The input can also include the properties of air pollutants, such as their solubility. The output of the wet deposition process module can include the wet deposition removal rate of air pollutants, i.e., the change in air pollutant concentration after the wet deposition process.
[0069] Both the inorganic aerosol thermodynamic equilibrium process module and the secondary organic aerosol process module can employ either a GNN or a Transformer model. GNNs can handle aerosol microphysical processes. Transformer models can handle complex nonlinear relationships; for example, attention weights can characterize the effects of temperature and humidity on the formation of secondary organic aerosols. The inputs to both modules can include primary information, where emission information may include the concentration of precursors to air pollutants, and meteorological information may include temperature, humidity, and solar radiation intensity. The output of the inorganic aerosol thermodynamic equilibrium process module can include the change in air pollutant concentration after the inorganic aerosol thermodynamic equilibrium process. The output of the secondary organic aerosol process module can include the change in air pollutant concentration after the secondary organic aerosol process.
[0070] In addition, the role of the emission process module is to reflect the emission information of air pollutants in the grid area at various times. The emission process module does not need to be trained, which means that the emission process module does not need to perform physical or chemical calculations.
[0071] Figure 2 This is a flowchart illustrating an exemplary method for training process modules corresponding to a changing process, such as... Figure 2 As shown, step 12 may include steps 121 and 122.
[0072] Step 121: Input the first information into the process module corresponding to the change process to obtain the first process rate output by the process module. The first process rate includes the process rate of the change process of air pollutants in the first grid area at multiple first moments.
[0073] Step 122: Determine the first loss value based on the first process rate, the numerical model process rate of the changing process, and the first preset loss function; and train the process module based on the first loss value.
[0074] In this application, model training can be performed in batches, and a training batch can be trained using multiple training samples. In this case, inputting the first information into the process module can be understood as inputting the first information of each of the multiple training samples in a batch into the process module, and the process module can output the first process rate corresponding to each training sample.
[0075] For example, the first preset loss function can be the mean squared error function, and the first loss value corresponding to the i-th change process is calculated by the following formula (7). : (7) in, The numerical model process rate represents the i-th change process. This represents the first process rate of the i-th change process. MSE() is the mean squared error function; the calculation of the mean squared error can be found in relevant techniques. If batch training is used, the mean of the process rate prediction errors corresponding to multiple training samples in a batch can be used as the first loss value to train the process module. Training methods can include, for example, gradient descent.
[0076] It should be noted that when calculating the loss value in the mean squared error function, the prediction error is calculated by comparing the numerical model process rate and the first process rate corresponding to the same first time point, the same first grid region, and the same air pollutant. In other words, when comparing the first process rate output by the process module with the numerical model process rate output by the air quality numerical model, the error calculation is performed on data from the same first grid region, the same first time point, and the same air pollutant. The calculation of the process loss value and concentration loss value mentioned below follows the same principle.
[0077] By using the above technical solution, the process modules are trained based on the first loss value corresponding to the change process, which can ensure the accuracy of the basic simulation capabilities of each process module.
[0078] It should also be noted that the training of the process module can refer to the training process of the network model in related technologies. It is trained through multiple iterations. In each iteration, the process module is updated accordingly. In the next training, the training is based on the latest process module obtained at the end of the previous iteration. Therefore, the process module corresponding to the change process mentioned above can be regarded as a module that is updated with each iteration during the training process.
[0079] Figure 3 This is a flowchart illustrating an exemplary method for training an air quality prediction model, such as... Figure 3 As shown, step 13 may include steps 131 to 135.
[0080] Step 131: Input the first information into the air quality prediction model to obtain the second process rate output by multiple first process modules, and the first concentration information of air pollutants in the first grid area at each first time point output by the air quality prediction model.
[0081] In one embodiment, the first concentration information may be the concentration change, that is, the change in the concentration of air pollutants from the previous moment to the first moment. In this embodiment, the concentration change of air pollutants may be the sum of the second process rates of each change process. In another embodiment, the first concentration information may also be the concentration of air pollutants in the first grid region at the first moment.
[0082] The first concentration information corresponds to the concentration information in the numerical model; that is, both are either used as concentration changes or as specific pollutant concentrations.
[0083] Step 132: Determine the process loss value based on the second process rate, the numerical model process rate, and the second preset loss function.
[0084] Step 133: Determine the concentration loss value based on the first concentration information, the numerical model concentration information, and the third preset loss function.
[0085] For example, the third preset loss function can be the mean squared error function, and the concentration loss value can be calculated using the following formula (8). .
[0086] (8) in, This represents the concentration information in the numerical model. This represents the initial concentration information. If batch training is used, the average of the prediction errors of the concentration information corresponding to multiple training samples in a batch can be used as the concentration loss value.
[0087] Step 134: Determine the target loss value based on the process loss value and the concentration loss value.
[0088] For example, the sum of the process loss value and the concentration loss value can be used as the target loss value. As another example, the weighted average of the process loss value and the concentration loss value can be used as the target loss value, where the weights of the process loss value and the concentration loss value can be preset.
[0089] Step 135: Train the air quality prediction model based on the target loss value.
[0090] Training methods can include gradient descent and similar approaches.
[0091] By using the above technical solutions, the air quality prediction model is trained to obtain the target air quality prediction model. This model can optimize the coordination between the various first process modules, taking into account process loss values and concentration loss values. This not only enables each process module to accurately learn the corresponding change process, but also ensures the accuracy of air pollutant concentration information.
[0092] It should be noted that in the second stage of training, each module of the first process can be regarded as a module that is updated with each iteration. That is, the air quality prediction model is continuously updated during the training process.
[0093] Figure 4 This is a flowchart illustrating an exemplary method for determining process loss values, such as... Figure 4As shown, step 132 may include steps 1321 to 1323.
[0094] Step 1321: Determine the weights of each of the various change processes.
[0095] First, we will introduce the first implementation method for determining the weights of various change processes.
[0096] Various processes are involved, including emission processes, horizontal advection processes, vertical advection processes, horizontal diffusion processes, vertical diffusion processes, chemical reaction processes, wet deposition processes, and dry deposition processes. Chemical reaction processes can include gas-phase chemical reaction processes, liquid-phase chemical reaction processes, inorganic aerosol thermodynamic equilibrium processes, and secondary organic aerosol processes. Meteorological information includes wind speed, atmospheric stability, temperature, humidity, solar radiation intensity, boundary layer height, precipitation, and turbulence intensity.
[0097] The implementation method of step 1321 can be as follows: The weights of the horizontal and vertical advection processes are determined based on wind speed. Based on atmospheric stability, determine the respective weights of the horizontal and vertical diffusion processes; The weights of chemical reaction processes are determined based on temperature, humidity, and solar radiation intensity. The weights of emission processes are determined based on the boundary layer height; The weights of the wet deposition process are determined based on precipitation and humidity. The weights of the dry deposition process are determined based on turbulence intensity, wind speed, atmospheric stability, humidity, underlying surface characteristics of the first grid region, and the properties of air pollutants.
[0098] For example, a correspondence can be established between wind speed and the weight of the horizontal advection process. The weight of the horizontal advection process can be positively correlated with wind speed, that is, the greater the wind speed, the greater the weight of the horizontal advection process. Similarly, a correspondence can be established between wind speed and the weight of the vertical advection process. The weight of the vertical advection process can also be positively correlated with wind speed, that is, the greater the wind speed, the greater the weight of the vertical advection process.
[0099] Establish the correspondence between atmospheric stability and the weights of horizontal diffusion processes, and between atmospheric stability and the weights of vertical diffusion processes. The weights of horizontal and vertical diffusion processes are negatively correlated with atmospheric stability; that is, the higher the atmospheric stability, the lower the weights of the horizontal and vertical diffusion processes.
[0100] Establish a correspondence between the first product of temperature, humidity, and solar radiation intensity and the weights of each chemical reaction process. The weights of chemical reaction processes are positively correlated with the first product; the larger the first product, the greater the weight of each chemical reaction process.
[0101] Establish a correspondence between boundary layer height and the weight of emission processes. The weight of emission processes is negatively correlated with boundary layer height; the greater the boundary layer height, the smaller the weight of emission processes.
[0102] Establish a correspondence between the second product of precipitation and humidity and the weight of the wet deposition process. The weight of the wet deposition process is positively correlated with the second product; the larger the second product, the greater the weight of the deposition process.
[0103] Establish the correspondence between meteorological and atmospheric dynamic factors, underlying surface characteristics, air pollutant properties, and the weights of dry deposition processes. Meteorological and atmospheric dynamic factors include turbulence intensity, wind speed, atmospheric stability, and humidity. Underlying surface characteristics of the grid area include surface roughness, vegetation type and physiological state, and the physicochemical properties of soil / water surfaces. Air pollutant properties include the pollutant's intrinsic attributes, such as its physical form (e.g., particle size distribution), chemical activity, and solubility.
[0104] It should be noted that the established correspondences can be in the form of functions. The weights obtained from the correspondences can be used as initial weights. After obtaining the initial weights of each change process, normalization can be performed to normalize the initial weights of each change process to between 0 and 1, and the sum of the weights of each change process is 1, thus obtaining the weights of each of the multiple change processes.
[0105] In this way, it can be ensured that under specific meteorological conditions, higher learning weights are given to key change processes, so that the model can develop the simulation capabilities of each process module in a balanced manner under different weather conditions.
[0106] A second implementation method for determining the weights of various change processes is introduced.
[0107] For each change process, the confidence level of the output of the first process module corresponding to the change process is obtained. The confidence level is used to characterize the credibility of the second process rate output by the first process module. The weight of the change process is determined based on the confidence level.
[0108] The first process module can output the confidence level of the second process rate along with the second process rate. The higher the confidence level, the more reliable the output second process rate is and the lower the uncertainty.
[0109] When a certain process of change is difficult to simulate under specific meteorological conditions (for example, chemical processes under stable weather conditions are very complex), its uncertainty increases. It will automatically increase, thereby reducing the weight of the change process and preventing the model from focusing too much on targets that are difficult to learn.
[0110] For example, a correspondence between confidence level and initial weight can be established. After obtaining the initial weight based on the confidence level, the weights of various change processes can be obtained through normalization.
[0111] Step 1322: For each change process, determine the loss value corresponding to the change process based on the second process rate and the numerical model process rate corresponding to the change process, as well as the second preset loss function and the weight of the change process.
[0112] In one embodiment, the second preset loss function can be as shown in formula (9). The second preset loss function is constructed based on the weights of the change process and the mean square error function. The loss value corresponding to the i-th change process can be determined by the following formula (9). : (9) 2 represents the rate of the second process in the i-th change process. This represents the weight of the i-th change process.
[0113] When calculating the loss value using formula (9), The weight of the i-th change process can be determined by either of the two implementation methods described above.
[0114] In one embodiment, the second preset loss function can be as shown in formula (10), wherein the second preset loss function is based on the weights of the change process, the mean square error function, and the uncertainty. The loss value corresponding to the i-th change process can be determined by the following formula (10). : (10) When calculating the loss value using formula (10), The weight of the i-th change process can be determined through the second implementation method described above. This represents the uncertainty of the rate of the second process corresponding to the i-th change process. Uncertainty is negatively correlated with confidence level; that is, the higher the confidence level, the greater the uncertainty. The larger.
[0115] Step 1323: Determine the process loss value based on the loss values corresponding to the various change processes.
[0116] As shown in formula (11), the sum of the loss values corresponding to the various change processes can be used as the process loss value. .
[0117] (11) Where I represents the number of various change processes.
[0118] By using the above technical solution, when calculating the process loss value, the contribution of each change process can be more accurately measured by using the weights of each change process, thereby accurately calculating the process loss value.
[0119] In other implementations, the weights of the various change processes may not be considered when calculating the process loss value.
[0120] In this application, various processes include emission processes, sedimentation processes, and chemical reaction processes.
[0121] In one embodiment, training an air quality prediction model using the numerical model process rates of various change processes in the training samples as labeled data further includes: Based on the first concentration information, the second process rate output by the first process module corresponding to the emission process, the second process rate output by the first process module corresponding to the sedimentation process, the second process rate output by the first process module corresponding to the chemical process, and the fourth preset loss function and the fifth preset loss function, the physical loss value is determined. The fourth preset loss function is used to determine the mass loss value of air pollutants, and the fifth preset loss function is used to determine the elemental loss value of the elements contained in the air pollutants. Based on the process loss value and concentration loss value, determine the target loss value, including: The target loss value is obtained by weighting the process loss value, concentration loss value, and physical loss value.
[0122] Based on the principle of mass conservation, the change in pollutant mass within each grid area is required to be equal to the net mass contributed by each change process. The pollutant mass can be obtained based on the pollutant concentration and the volume of the grid area.
[0123] For example, the fourth and fifth preset loss functions can be the mean squared error function.
[0124] In one embodiment, the mass loss value L_mass can be calculated using the following formula (12): L_mass=MSE(Observational quality change, Σ process contribution quality) (12) The observation quality of the j-th air pollutant in the first grid region can be obtained by multiplying the first concentration information of the j-th air pollutant in the first grid region with the regional volume of the first grid region. The process contribution quality can be obtained by multiplying the second process rate of the change process of the j-th air pollutant with the regional volume of the first grid region.
[0125] The observed mass change of the j-th air pollutant can be the change in the observed mass of the j-th air pollutant at the first moment relative to the observed mass at the previous moment. The mass contribution of the Σ process can be obtained from the following masses of the j-th air pollutant: the mass directly added by the emission process, the net flux through the grid boundary by the advection process (calculated using upwind differential), the net mass change caused by the chemical reaction process, the mass removed by the dry deposition process, and the mass removed by the wet deposition process.
[0126] For example, if there are J types of air pollutants, the mass loss value L_mass can be the average of the mass prediction errors corresponding to each of the J air pollutants.
[0127] In another embodiment, the mass loss value L_mass can be calculated using the following formula (13): L_mass=MSE(Input total mass, Output total mass + Sedimentation + Chemical conversion) (13) For the first grid region, the total input mass of the j-th air pollutant includes the emission contribution mass of the j-th air pollutant's emission process at the previous time step, and the total mass of the j-th air pollutant at the previous time step. The total output mass can be the total mass of the j-th air pollutant at the first time step. The deposition amount includes the dry deposition amount and wet deposition amount of the j-th air pollutant at the first time step, and the chemical transformation amount includes the transformation mass of the j-th air pollutant through gas-phase chemical reaction processes and liquid-phase chemical reaction processes at the first time step.
[0128] As an example, the J types of air pollutants may include nitrogen-containing and sulfur-containing pollutants. Based on the principle of element conservation, the total number of atoms of key chemical elements (such as nitrogen and sulfur) remains conserved during chemical reactions. Taking nitrogen as an example, tracking all nitrogen-containing air pollutants (NO2, NO, HNO3, etc.), the nitrogen loss value L_species = MSE(initial total number of nitrogen atoms, predicted total number of nitrogen atoms + nitrogen atoms lost due to sedimentation). The initial total number of nitrogen atoms can be calculated based on the total input mass of various nitrogen-containing pollutants, the predicted total number of nitrogen atoms can be calculated based on the total output mass of various nitrogen-containing pollutants, and the nitrogen atoms lost due to sedimentation can be calculated based on the dry and wet sedimentation amounts of various nitrogen-containing pollutants. The calculation of the loss values of other elements can refer to the calculation method for the nitrogen loss value.
[0129] If multiple chemical elements are tracked, such as nitrogen and sulfur, the average of the loss values corresponding to each element can be used as the elemental loss value.
[0130] For example, the sum of the mass loss value of air pollutants and the elemental loss value of the elements contained in the air pollutants can be used as the physical loss value.
[0131] The target loss value L_total can be determined by the following formula (14): 14
[0132] Indicates the process loss value. The weights representing the process loss values Indicates the concentration loss value. The weights representing the concentration loss values, Indicates the physical loss value. The weights represent the physical loss values. , and The value can be preset.
[0133] In addition, positive definite constraints can be imposed, such as requiring pollutant concentrations to be non-negative, or constraining the sign of the process contribution (emission contribution is non-negative, deposition contribution is non-positive).
[0134] The target air quality prediction model trained using the above technical solutions can output not only the concentration of air pollutants, but also the process rate of each change process. The process rate makes the prediction results interpretable.
[0135] In one embodiment, operational forecast data and process analysis data from the 2013-2023 air quality numerical model can be used for training. To improve training efficiency, such as... Figure 5 As shown, training can be performed in a time-parallel manner. Data loading refers to loading data from 2013 to 2023, and data sharding refers to dividing the data into different processes for processing. Specifically, data from 2013 to 2015 is processed in process 1, data from 2016 to 2018 in process 2, data from 2019 to 2021 in process 3, and data from 2022 to 2023 in process 4. Gradient aggregation can be understood as the aggregation of the difference information calculated by each process, and parameter updates are performed using the aggregated gradients for model training. Both the first stage of training the process module and the second stage of training the air quality prediction model can be performed in a time-parallel manner.
[0136] Figure 6 This is a flowchart illustrating an exemplary air quality prediction method, which can be applied to a server, such as... Figure 5 As shown, the air quality prediction method includes steps 61 and 62.
[0137] Step 61, obtain the second information.
[0138] The second information includes the air pollutant concentration, air pollutant emission information, and meteorological information of the second grid area at time T', as well as the predicted air pollutant emission information and meteorological information of the second grid area at multiple second times, including each time from time T'+1 to time T'+N.
[0139] Step 62: Input the second information into the target air quality prediction model to obtain the concentration of air pollutants in the second grid area at multiple second time points output by the target air quality prediction model, and the third process rate output by multiple process modules in the target air quality prediction model. These multiple process modules correspond one-to-one with various changes in air pollutants, and the third process rate includes the process rate of the changes in air pollutants in the second grid area at multiple second time points. The target air quality prediction model is trained according to the model construction method provided in this application.
[0140] The second information may include data from multiple second grid areas, and the types of air pollutants may also be multiple, meaning the target air quality prediction model can simultaneously predict the concentrations of multiple air pollutants in multiple grid areas. Time T' is the current time, and times T'+1 to T'+N can be future times. The air pollutant emission information and meteorological information at the second time can be predicted data.
[0141] It should be noted that the second grid region and the first grid region mentioned above are used to distinguish the grid regions in the model training stage and the model application stage. In some embodiments, the first grid region and the second grid region can also be the same. For example, data from the regions of grid region 1, grid region 2, etc. are used for model training. After obtaining the target air quality prediction model, the target air quality prediction model can be used to predict the air quality of grid region 1 and grid region 2 at future times.
[0142] Figure 7 This is a schematic diagram illustrating an exemplary air quality prediction method, such as... Figure 7 As shown, the target air quality prediction model may include emission process module, horizontal advection process module, vertical advection process module, horizontal diffusion process module, vertical diffusion process module, gas phase chemical reaction process module, liquid phase chemical reaction process module, dry deposition process module, wet deposition process module, inorganic aerosol thermodynamic equilibrium process module, and secondary organic aerosol process module.
[0143] Using the second information as input to the target air quality prediction model, the model can output the concentration of air pollutants in the second grid area at multiple second time points, as well as the third process rate of each change process in the second grid area at multiple second time points. The third process rate can be used as source apportionment input data. Using process analysis data, pollutant sources can be allocated based on different region IDs and industries, enabling pollution source analysis and providing accurate data for air quality control decisions.
[0144] The various process modules, their inputs, and the networks used in the target air quality prediction model have been described above. The network structures of modules such as one-dimensional convolutional neural networks, convolutional neural networks, physical information neural networks, graph neural networks, XGBoost models, ConvLSTM, and diffusion models can all refer to the network model structures in related technologies.
[0145] As an example, The third process rate represents the emission process. The rate of the third process in the horizontal advection process is represented. The third process rate represents the vertical advection process. The rate of the third process in the horizontal diffusion process is represented. The rate of the third process in the vertical diffusion process is represented. This indicates the rate of the third process in a gas-phase chemical reaction. This represents the rate of the third process in the dry sedimentation process. This represents the rate of the third process in the wet sedimentation process. The rate of the third process in the thermodynamic equilibrium process of inorganic aerosols is represented. This indicates the rate of the third process in the secondary organic aerosol process. This indicates the rate of the third process in a liquid-phase chemical reaction. (C) j,t-1 The concentration of the j-th air pollutant in the second grid region at time t-1 can be represented by C. The rates of each of the third processes mentioned above can represent the process rates of the change in the j-th air pollutant in the second grid region at time t-1. j,t This can represent the concentration of the j-th air pollutant in the second grid region at time t. Wherein, to The sum of, plus C j,t-1 It can be used as C j,t .
[0146] Figure 8 This is a block diagram illustrating an electronic device 1900 according to an exemplary embodiment. For example, the electronic device 1900 may be provided as a server. (Refer to...) Figure 8The electronic device 1900 includes a processor 1922, which may be one or more, and a memory 1932 for storing computer programs executable by the processor 1922. The computer program stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processor 1922 may be configured to execute the computer program to perform the aforementioned model building method or air quality prediction method.
[0147] Additionally, the electronic device 1900 may also include a power supply component 1926 and a communication component 1950. The power supply component 1926 can be configured to perform power management of the electronic device 1900, and the communication component 1950 can be configured to enable communication of the electronic device 1900, such as wired or wireless communication. Furthermore, the electronic device 1900 may also include an input / output (I / O) interface 1958. The electronic device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM Mac OS X TM Unix TM Linux TM etc.
[0148] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the model building method or air quality prediction method described above. For example, the computer-readable storage medium may be the memory 1932 including the program instructions described above, which may be executed by the processor 1922 of the electronic device 1900 to complete the model building method or air quality prediction method described above.
[0149] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a processor, which, when executed by the processor, implements the steps of the model building method or air quality prediction method described above.
[0150] The preferred embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solution of this application, and these simple modifications all fall within the protection scope of this application.
[0151] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this application will not describe the various possible combinations separately.
[0152] Furthermore, various different implementations of this application can be combined in any way, as long as they do not violate the spirit of this application, they should also be regarded as the content disclosed in this application.
Claims
1. A model construction method characterized by comprising: The method comprises: obtaining a training sample set, wherein each training sample in the training sample set comprises first information and numerical model business data, the first information comprises air pollutant concentration of a first grid area at T time, air pollutant emission information of the first grid area at each time from T time to T+N time, and meteorological information, and the numerical model business data comprises numerical model process rates of various change processes, wherein the numerical model process rate is a process rate of the change process of the air pollutant of the first grid area at a plurality of first times output by an air quality numerical model based on the first information, and the plurality of first times comprise each time from T+1 time to T+N time; for each change process, training a process module corresponding to the change process by taking the numerical model process rate of the change process in the training sample as labeled data, to obtain a first process module corresponding to the change process; training an air quality prediction model by taking the numerical model process rate of each change process in the training sample as labeled data, to obtain a target air quality prediction model, wherein the air quality prediction model is composed of first process modules corresponding to the various change processes respectively.
2. The method of claim 1, wherein, The method comprises: inputting the first information into the process module corresponding to the change process to obtain a first process rate output by the process module, wherein the first process rate comprises process rates of the change process of the air pollutant of the first grid area at the plurality of first times; determining a first loss value according to the first process rate, the numerical model process rate of the change process, and a first preset loss function, and training the process module according to the first loss value.
3. The method of claim 1, wherein, The numerical model business data further comprises numerical model concentration information, wherein the numerical model concentration information is concentration information of the air pollutant of the first grid area at the plurality of first times output by the air quality numerical model; The method comprises: inputting the first information into the air quality prediction model to obtain a second process rate output by each of the plurality of first process modules and first concentration information of the air pollutant of the first grid area at the plurality of first times output by the air quality prediction model; determining a process loss value according to the second process rate, the numerical model process rate, and a second preset loss function; determining a concentration loss value according to the first concentration information, the numerical model concentration information, and a third preset loss function; determining a target loss value according to the process loss value and the concentration loss value; training the air quality prediction model according to the target loss value.
4. The method of claim 3, wherein, The process loss value is determined according to the second process rate, the numerical model process rate, and a second preset loss function. The weight of each of the plurality of change processes is determined. For each of the plurality of change processes, a loss value corresponding to the change process is determined according to the second process rate and the numerical model process rate of the change process, and the weight of the change process. The process loss value is determined according to the loss value corresponding to each of the plurality of change processes.
5. The method of claim 4, wherein, The plurality of change processes include an emission process, a horizontal advection process, a vertical advection process, a horizontal diffusion process, a vertical diffusion process, a chemical reaction process, a wet deposition process, and a dry deposition process. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity.
6. The method of claim 4, wherein, The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity.
7. The method of claim 3, wherein, The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity.
8. An air quality prediction method characterized by, The weight of each of the plurality of change processes is determined according to the wind speed, the atmospheric stability, the temperature, the humidity, the solar radiation intensity, the boundary layer height, the precipitation, and the turbulence intensity. The weight of each of the plurality of change processes is obtained according to the confidence level of the first process module output corresponding to the change process. The plurality of change processes include an emission process, a deposition process, and a chemical reaction process. The training of the air quality prediction model by taking the numerical model process rate of each of the plurality of change processes in the training sample as the labeled data further includes: A physical loss value is determined according to the first concentration information, the second process rate output by the first process module corresponding to the emission process, the second process rate output by the first process module corresponding to the deposition process, the second process rate output by the first process module corresponding to the chemical process, and fourth and fifth preset loss functions. The target loss value is determined according to the process loss value and the concentration loss value. The target loss value is obtained according to the weighted values of the process loss value, the concentration loss value, and the physical loss value. The method includes: obtaining second information, the second information comprising air pollutant concentration, air pollutant emission information and meteorological information of a second grid area at a time T', and predicted air pollutant emission information and meteorological information of the second grid area at a plurality of second times, the plurality of second times comprising each of T'+1 to T'+N; inputting the second information into a target air quality prediction model to obtain air pollutant concentration of the second grid area at the plurality of second times output by the target air quality prediction model, and third process rates output by a plurality of process modules in the target air quality prediction model, the plurality of process modules corresponding to a plurality of change processes of air pollutants one by one, the third process rates comprising process rates of corresponding change processes of air pollutants in the second grid area at the plurality of second times, wherein the target air quality prediction model is trained according to the model construction method in any one of claims 1-7.
9. The method of claim 8, wherein, The plurality of change processes comprises one or more of the following: emission process, horizontal advection process, vertical advection process, horizontal diffusion process, vertical diffusion process, gas phase chemical reaction process, liquid phase chemical reaction process, dry deposition process, wet deposition process, inorganic aerosol thermodynamic equilibrium process, and secondary organic aerosol process.
10. An electronic device, comprising: comprising: a memory having a computer program stored thereon; a processor configured to execute the computer program in the memory to implement the steps of the method in any one of claims 1-7, or to implement the steps of the method in claim 8 or 9.